Agent skill

Backup And Disaster Recovery

by rampstackco in rampstackco/claude-skills

Plan and run backups, set recovery objectives, and run disaster recovery drills.

MITAuto-check passedDevOps & Cloud

Install Backup And Disaster Recovery

skills CLI
$ npx skills add rampstackco/claude-skills --skill backup-and-disaster-recovery -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rampstackco/claude-skills backup-and-disaster-recovery --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rampstackco/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/backup-and-disaster-recovery .claude/skills/backup-and-disaster-recovery && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
backup-and-disaster-recovery
GitHub stars
945
Token cost
~2.7k tokens
SKILL.md length
1,445 words
Files
3 (incl. references)
Skills in repo
103
Repo updated
First seen
Licence
MIT

At a glance

Plan and run backups, set recovery objectives, and run disaster recovery drills.

  • Works in 7 steps: Inventory state → Set RPO and RTO per tier → Verify or design backup architecture → …
  • Defining RPO/RTO targets
  • SKILL.md covers When to use, When NOT to use, Required inputs and The framework: 4 questions, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Backup And Disaster Recovery is an agent skill from rampstackco/claude-skills. Plan and run backups, set recovery objectives, and run disaster recovery drills. Use this skill when defining RPO/RTO targets, designing backup architecture, deciding what to back up and how often, planning for full-region or platform outages, or running a restoration drill. Triggers on backup, restore, RPO, RTO, disaster recovery, DR, business continuity, what if the database is gone, what if our hosting goes down, recovery drill, ransomware planning. Also triggers when an incident reveals a gap in restoration…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `README.md` and `references/restore-runbook-template.md`).

It sits in DevOps & Cloud, covering Backup and disaster recovery. The repository describes itself as: Stack-agnostic Claude Skills covering the full website lifecycle: brand, design, content, SEO, dev, ops, growth, and research. Build, ship, audit, optimize. The licence is MIT.

When your agent uses it

  • Defining RPO/RTO targets
  • Designing backup architecture
  • Deciding what to back up and how often
  • Planning for full-region

Example prompts

  • “/backup-and-disaster-recovery”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Inventory state
  2. Set RPO and RTO per tier
  3. Verify or design backup architecture
  4. Document the restore runbook
  5. Run a drill
  6. Document drill results
  7. Schedule the next drill

What it can do on your machine

Read from SKILL.md and the folder at commit 482c9bf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Backup And Disaster Recovery loads about 2.7k tokens when it runs, and up to ~4.1k if it reads all its reference files. Until then it costs about 139 tokens; SKILL.md has 1,445 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~139
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rampstackco/claude-skills at commit 482c9bf, republished under its MIT licence (© rampstackco). 1,445 words, ~2,718 tokens.

Download SKILL.mdSave it as .claude/skills/backup-and-disaster-recovery/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
backup-and-disaster-recovery
description
Plan and run backups, set recovery objectives, and run disaster recovery drills. Use this skill when defining RPO/RTO targets, designing backup architecture, deciding what to back up and how often, planning for full-region or platform outages, or running a restoration drill. Triggers on backup, restore, RPO, RTO, disaster recovery, DR, business continuity, what if the database is gone, what if our hosting goes down, recovery drill, ransomware planning. Also triggers when an incident reveals a gap in restoration capability.
category
operations
catalog_summary
RPO/RTO targets, backup strategy, restoration drills
display_order
6

Backup and Disaster Recovery

Plan for the worst case: the database is gone, the host is down for a week, the deploy was poisoned, ransomware encrypted everything. The skill is in advance preparation, not reaction.


When to use

  • Setting up backups for a new system
  • Reviewing and validating backup architecture
  • Defining RPO (recovery point objective) and RTO (recovery time objective)
  • Running a disaster recovery drill
  • Diagnosing gaps after an incident
  • Planning for ransomware, data corruption, or insider threats
  • Migrating to a new platform (DR planning belongs in the migration plan)

When NOT to use

  • Active incident response (use incident-response)
  • Routine deploy rollbacks (use launch-runbook)
  • Code or content versioning (covered by Git, CMS revision history)
  • Routine database snapshots (use this skill to set them up; routine review goes in monitoring)

Required inputs

  • The systems in scope (databases, file storage, code, configs, secrets)
  • The hosting platforms and providers
  • Existing backup tooling and what it covers
  • Tolerance for data loss (in time)
  • Tolerance for downtime (in time)
  • Compliance requirements (some regulations mandate specific backup standards)

The framework: 4 questions

Every disaster recovery plan answers four questions explicitly.

Question 1: What needs to be recoverable?

List every system that holds state. Categorize by criticality.

Tier 1: must recover. Without it, the business stops. (Customer database, transaction log, primary content store.)

Tier 2: should recover. Loss is painful but not fatal. (Analytics, logs, secondary services.)

Tier 3: nice to recover. Easy to rebuild. (Caches, derived data, temporary state.)

The tier drives RPO, RTO, backup frequency, and storage spend.

Question 2: How much data loss is acceptable? (RPO)

RPO is the maximum age of data that's acceptable to lose, measured in time.

  • RPO = 1 hour: hourly backups or continuous replication needed
  • RPO = 1 day: daily backups acceptable
  • RPO = 1 week: weekly backups acceptable

For most production data, RPO of 1 hour or less is the target. For critical financial systems, near-zero RPO (continuous replication).

For derived or rebuildable data, RPO of 1 day or longer is fine.

Question 3: How much downtime is acceptable? (RTO)

RTO is the maximum time to restore service after a disaster.

RTO targetImplies
< 5 minutesHot standby with automatic failover
< 1 hourWarm standby with manual failover or fast restore from recent snapshot
< 24 hoursCold backup with documented restore process
Days to weeksBest-effort, accept extended downtime

RTO drives architecture spend. Aggressive RTOs (< 1 hour) are expensive. Loose RTOs (days) are cheap.

Question 4: What's the disaster?

Plan for specific scenarios. Each has different implications.

Hardware failure. Disk dies. Standard backups solve this. Most modern hosts handle automatically.

Provider outage. Region or vendor goes down. Cross-region or cross-provider redundancy needed for low RTO.

Data corruption. Bad migration, bug, accidental delete. Point-in-time restore needed. The latest backup might be corrupted; you need history.

Ransomware or compromise. Attacker encrypts or deletes. Backups must be immutable or air-gapped, otherwise the attacker takes them too.

Account compromise. Attacker has admin credentials, deletes everything. Same defense as ransomware: immutable backups, separate access control.

Vendor lock-out. Account suspended, billing dispute, vendor disappears. Backups outside the vendor needed.

Insider threat. Disgruntled employee deletes or exfiltrates. Audit logs, separation of duties, immutable backups.

A backup strategy that handles only hardware failure isn't a strategy. It's the easiest case.


Workflow

Step 1: Inventory state

Every system that holds state goes on a list:

SystemData typeTierCurrent backupTested?

If you can't list it, you can't protect it. Often the inventory itself reveals gaps (the "we forgot about that database" moment).

Step 2: Set RPO and RTO per tier

For each tier, agree on RPO and RTO. Get sign-off from the people who'd be impacted by a disaster.

Push back on aspirational targets that aren't backed by infrastructure spend. RTO of 5 minutes for a system without a hot standby is not real.

Step 3: Verify or design backup architecture

For each system, ensure:

  • Frequency matches RPO.
  • Retention covers point-in-time recovery (typically 30+ days for production data).
  • Storage location is separate from the source. Same disk, same account, same region: not enough.
  • Immutability or write-once storage for at least some backup copies. Defends against ransomware.
  • Encryption at rest. Standard for compliance.
  • Tested restore procedure. Untested backups are not backups.

The "3-2-1 rule" is a useful starting point: 3 copies of data, 2 different storage types, 1 offsite (or off-account, off-platform).

Step 4: Document the restore runbook

For each system, write the runbook:

  1. How to detect the disaster (cross-reference monitoring)
  2. How to decide to restore (decision criteria, who authorizes)
  3. The exact restore steps (commands, screenshots, sequence)
  4. How to verify the restore worked
  5. How to switch traffic back
  6. Communication template (status page, customer notice)

The runbook is for the worst night of someone's career. Write it for tired, panicked you.

Step 5: Run a drill

The first restore should never be during a real disaster.

Drills can be:

  • Tabletop: walk through the runbook on paper. Useful for finding gaps in the plan.
  • Partial: restore to a non-production environment. Verify the data, validate the steps.
  • Full: simulate the disaster. Production failover or full restore. Maximum confidence, maximum risk.

For most teams: quarterly tabletop, annual partial drill, full drill before major launches or after major architecture changes.

Show full SKILL.md (573 more words)Show less
Step 6: Document drill results

After each drill, document:

  • What was tested
  • What worked
  • What broke
  • What the actual RPO and RTO were (vs. targets)
  • Action items

If the actual RTO was 6 hours when the target was 1 hour, the target is fiction. Either fix the gap or revise the target.

Step 7: Schedule the next drill

Calendar it. Assign an owner. Backups that aren't drilled drift toward useless.


Special topics

Database point-in-time recovery

Many managed databases offer point-in-time recovery (PITR) within a retention window (often 7-35 days). This typically achieves RPO of seconds to minutes.

For longer retention, schedule periodic exports to immutable storage.

PITR alone isn't enough. If the database service itself is compromised, PITR is gone too. Always have at least one backup outside the source service.

File storage backups

Object stores (S3, GCS, Azure Blob) usually offer:

  • Versioning (recover overwritten objects)
  • Replication (cross-region)
  • Object lock or immutability (defense against deletion)

Set all three for production-critical buckets. Don't rely on the storage provider's default retention.

Code and config backups

Code lives in Git. The Git host (GitHub, GitLab, etc.) is your backup, but a single host is a single point of failure.

For high-criticality code:

  • Mirror to a second host or your own server
  • Periodic offline exports

Configs and secrets need separate handling:

  • Infrastructure-as-code: in Git, mirrored
  • Runtime configs: backed up alongside the system
  • Secrets: in a secret manager with its own backup story
Backups of backups

The backup system itself can fail. Backup metadata, backup credentials, encryption keys: all must be backed up.

If your backup is encrypted with a key you've lost, the backup is useless.

Compliance backups

Some regulations require specific retention (e.g., 7 years for financial data). Comply with the highest applicable standard.

Don't conflate compliance retention with operational backup. Compliance often allows much slower restore (just need to be able to produce the data eventually).


Failure patterns

Untested backups. The single most common failure. Backups appear to work; restore fails. Test.

Backups in the same account or region as the source. Account compromise or region outage takes both.

No immutability. Ransomware encrypts the backups too. Use object lock or air-gapped storage.

RTO and RPO that aren't measured. Target says "1 hour" but no one has verified the actual RTO. Assume the actual is longer than the target until proven otherwise.

Restore runbook only in someone's head. Person leaves or is unavailable; runbook is gone. Document.

Backups but no DR plan. "We have backups" isn't a plan. The plan is the runbook plus the architecture plus the drilling.

Optimism bias. "It won't happen to us." It happens. Plan as if it will.

Backups too old or too new. Want point-in-time history (in case corruption isn't immediately discovered). Daily snapshots with 30+ day retention. Or continuous replication with separate periodic snapshots for history.

Skipping drills "because we're busy." Then you'll be busier during the disaster.

No communication plan. Restoring data is half the job. Telling customers, stakeholders, and internal teams what's happening is the other half.


Output format

A DR plan document includes:

  • Inventory: every stateful system
  • Tiering: criticality per system
  • Targets: RPO and RTO per tier
  • Architecture: backup tooling, frequency, storage, immutability
  • Runbooks: restore procedures per system
  • Drill schedule: what gets tested when
  • Drill log: results of past drills
  • Communication templates: what to say during a real DR event

Reference files

© rampstackco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/backup-and-disaster-recovery of rampstackco/claude-skills.

  • SKILL.md
  • README.md
  • references/restore-runbook-template.md

Open the folder on GitHubat commit 482c9bf

Compare with similar skills

Backup And Disaster Recovery next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Backup And Disaster Recovery compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Backup And Disaster Recovery this skillrampstackco/claude-skills945—~2.7kAutomated safety check: PassMIT
OmniRoute Backup and Sync CLIdiegosouzapw/OmniRoute75k—~948Automated safety check: PassMIT
Pymobiledevice3 Device Operatordoronz88/pymobiledevice32.9k—~1.8kAutomated safety check: NotesGPL-3.0
Myclaw BackupLeoYeAI/openclaw-backup659—~1.8kAutomated safety check: PassMIT
OmniRoute Database Backupsdiegosouzapw/OmniRoute75k—~395Automated safety check: PassMIT
Tempsgotempsh/temps831—~2kAutomated safety check: PassApache-2.0

Similar skills

  • OmniRoute Backup and Sync CLI

    diegosouzapw/OmniRoute

    Backup and restore OmniRoute data from the CLI. Trigger incremental snapshots, sync to cloud storage, manage backup schedules, and restore from archive files.

    75k GitHub stars~948 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Pymobiledevice3 Device Operator

    doronz88/pymobiledevice3

    Operate iOS and iPadOS devices with pymobiledevice3, from a local checkout or straight from PyPI via uvx on a fresh workstation.

    2.9k GitHub stars~1.8k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Myclaw Backup

    LeoYeAI/openclaw-backup

    Backup and restore all OpenClaw configuration, agent memory, skills, and workspace data.

    659 GitHub stars~1.8k tokensUpdated 7 mo ago
    DevOps & CloudAuto-check passed
  • OmniRoute Database Backups

    diegosouzapw/OmniRoute

    Trigger system backups, restore from backup files, and manage the SQLite database lifecycle. Supports export, import, and incremental snapshot strategies.

    75k GitHub stars~395 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Temps

    gotempsh/temps

    Manage, deploy, operate, and instrument applications with Temps.

    831 GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Hermes Offsite Backup

    OnlyTerp/hermes-optimization-guide

    Set up encrypted off-machine Hermes backups. An agent skill from OnlyTerp/hermes-optimization-guide.

    688 GitHub stars~772 tokensUpdated 16 days ago
    DevOps & CloudAuto-check passed

More from rampstackco/claude-skills

All 103 skills in this repo
  • After Action Report

    rampstackco/claude-skills

    Run a structured after-action review (postmortem, retrospective) on a launch, incident, or completed project to capture timeline, root cause analysis, contributing factors, and actionable lessons.

    945 GitHub stars~2.5k tokensUpdated 3 days ago
    Auto-check passed
  • Analytics Strategy

    rampstackco/claude-skills

    Design measurement frameworks including event taxonomy, KPI hierarchy, dashboard architecture, attribution models, and analytics implementation strategy.

    945 GitHub stars~2.4k tokensUpdated 3 days ago
    Auto-check passed
  • Brand Style Guide

    rampstackco/claude-skills

    Build or audit a comprehensive brand style guide that documents the full brand system including story, logo system, color, typography, imagery, voice, applications, and dos/don'ts.

    945 GitHub stars~2.1k tokensUpdated 3 days ago
    Auto-check passed
  • Brand Voice

    rampstackco/claude-skills

    Develop or document a complete brand voice and tone system covering voice attributes, tone shifts by context, vocabulary preferences, grammar rules, and copy examples.

    945 GitHub stars~2.2k tokensUpdated 3 days ago
    Auto-check passed
  • Content And Copy

    rampstackco/claude-skills

    Write or edit website copy, blog content, and editorial pieces with attention to voice, structure, and goal.

    945 GitHub stars~2.1k tokensUpdated 3 days ago
    Auto-check passed
  • Content Strategy

    rampstackco/claude-skills

    Develop a content strategy covering editorial positioning, content pillars, formats, calendar, governance, and topical authority planning.

    945 GitHub stars~2.6k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about Backup And Disaster Recovery

What does Backup And Disaster Recovery do?

Plan and run backups, set recovery objectives, and run disaster recovery drills. Backup And Disaster Recovery is an agent skill from rampstackco/claude-skills. Plan and run backups, set recovery objectives, and run disaster recovery drills.

When should I use Backup And Disaster Recovery?

Backup And Disaster Recovery fits situations like: defining RPO/RTO targets; designing backup architecture; deciding what to back up and how often; planning for full-region.

How do I install Backup And Disaster Recovery in Claude Code?

Run `npx skills add rampstackco/claude-skills --skill backup-and-disaster-recovery -a claude-code`. Or copy the skill folder (skills/backup-and-disaster-recovery in rampstackco/claude-skills) into .claude/skills/backup-and-disaster-recovery in your project. Claude Code loads it when a task matches its description.

How do I install Backup And Disaster Recovery in Codex?

Run `npx skills add rampstackco/claude-skills --skill backup-and-disaster-recovery -a codex`. Or copy the skill folder (skills/backup-and-disaster-recovery in rampstackco/claude-skills) into .agents/skills/backup-and-disaster-recovery in your project. Codex loads it when a task matches its description.

Can I use Backup And Disaster Recovery in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rampstackco/claude-skills --skill backup-and-disaster-recovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/backup-and-disaster-recovery, .gemini/skills/backup-and-disaster-recovery, .github/skills/backup-and-disaster-recovery and .opencode/skills/backup-and-disaster-recovery in your project.

What does Backup And Disaster Recovery need to run?

SKILL.md names no scripts, command-line tools or credentials: Backup And Disaster Recovery is instructions for the agent only.

Does Backup And Disaster Recovery access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Backup And Disaster Recovery safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Backup And Disaster Recovery use?

Backup And Disaster Recovery is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Backup And Disaster Recovery use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Backup And Disaster Recovery?

Skills that share tags, products or a category with Backup And Disaster Recovery: OmniRoute Backup and Sync CLI (diegosouzapw/OmniRoute, 75k stars), Pymobiledevice3 Device Operator (doronz88/pymobiledevice3, 2.9k stars), Myclaw Backup (LeoYeAI/openclaw-backup, 659 stars) and OmniRoute Database Backups (diegosouzapw/OmniRoute, 75k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Backup And Disaster Recovery?

rampstackco (a GitHub organization) maintains it in rampstackco/claude-skills, which has 945 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 7, 2026.

Source: rampstackco/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.