Convex Backup
openclaw/clawhub
Set up Convex backups and run a restore DRILL that proves recovery — snapshot, restore into a throwaway preview, assert the data came back — plus a schedule matched to your RPO and a gated recovery…
Write a disaster recovery plan for a service or system — covering RPO/RTO targets, failure scenario runbooks, backup and restore procedures, DR testing cadence, and communication templates.
$ npx skills add mohitagw15856/pm-claude-skills --skill disaster-recovery-plan -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mohitagw15856/pm-claude-skills disaster-recovery-plan --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/disaster-recovery-plan .claude/skills/disaster-recovery-plan && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "disaster-recovery-plan" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/disaster-recovery-plan into .claude/skills/disaster-recovery-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "disaster-recovery-plan", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/disaster-recovery-planType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mohitagw15856/pm-claude-skills --skill disaster-recovery-plan -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mohitagw15856/pm-claude-skills disaster-recovery-plan --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/disaster-recovery-plan .agents/skills/disaster-recovery-plan && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "disaster-recovery-plan" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/disaster-recovery-plan into .agents/skills/disaster-recovery-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "disaster-recovery-plan", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitagw15856/pm-claude-skills --skill disaster-recovery-plan -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mohitagw15856/pm-claude-skills disaster-recovery-plan --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/disaster-recovery-plan .cursor/skills/disaster-recovery-plan && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "disaster-recovery-plan" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/disaster-recovery-plan into .cursor/skills/disaster-recovery-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "disaster-recovery-plan", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mohitagw15856/pm-claude-skills.git --path skills/disaster-recovery-plan--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mohitagw15856/pm-claude-skills --skill disaster-recovery-plan -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mohitagw15856/pm-claude-skills disaster-recovery-plan --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/disaster-recovery-plan .gemini/skills/disaster-recovery-plan && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "disaster-recovery-plan" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/disaster-recovery-plan into .gemini/skills/disaster-recovery-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "disaster-recovery-plan", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mohitagw15856/pm-claude-skills disaster-recovery-planInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mohitagw15856/pm-claude-skills --skill disaster-recovery-plan -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/disaster-recovery-plan .github/skills/disaster-recovery-plan && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "disaster-recovery-plan" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/disaster-recovery-plan into .github/skills/disaster-recovery-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "disaster-recovery-plan", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mohitagw15856/pm-claude-skills --skill disaster-recovery-plan -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mohitagw15856/pm-claude-skills disaster-recovery-plan --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/disaster-recovery-plan .opencode/skills/disaster-recovery-plan && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "disaster-recovery-plan" agent skill from https://github.com/mohitagw15856/pm-claude-skills/tree/main/skills/disaster-recovery-plan into .opencode/skills/disaster-recovery-plan/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "disaster-recovery-plan", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
disaster-recovery-planWrite a disaster recovery plan for a service or system — covering RPO/RTO targets, failure scenario runbooks, backup and restore procedures, DR testing cadence, and communication templates.
Disaster Recovery Plan is an agent skill from mohitagw15856/pm-claude-skills. Write a disaster recovery plan for a service or system — covering RPO/RTO targets, failure scenario runbooks, backup and restore procedures, DR testing cadence, and communication templates. Use when asked to write a DR plan, document failover procedures, create recovery runbooks, define RTO/RPO targets, or prepare for a disaster recovery game day. Produces a full DR document with per-scenario recovery runbooks, backup validation procedures, testing schedule, and communication templates.
Its SKILL.md is about 5.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Backup and disaster recovery and Runbooks and postmortems. The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlawspsqlcurlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
health.aws.amazon.comstatus.cloud.google.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Disaster Recovery Plan loads about 5.8k tokens when it runs. Until then it costs about 129 tokens; SKILL.md has 1,854 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 1,854 words, ~5,785 tokens.
.claude/skills/disaster-recovery-plan/SKILL.md (or your agent's skills folder).Produce a complete disaster recovery plan for a service or system — giving engineers, SREs, and on-call responders everything they need to recover from a disaster scenario in the shortest possible time. A good DR plan is tested regularly, has exact commands (not vague instructions), and makes RTO/RPO targets measurable so the team knows whether recovery succeeded.
Ask for these if not already provided:
Team: [Team name] | Tech lead: [Name] Criticality tier: [Tier 1 / Tier 2 / Tier 3] | Last tested: [Date] Next DR test: [Date] | Document owner: [Name] Last updated: [Date] | Review cycle: Quarterly
Emergency? Skip to Section 3 — Failure Scenario Runbooks. Find the scenario that matches your situation and follow the steps exactly.
| Target | Value | Rationale |
|---|---|---|
| RPO (Recovery Point Objective) | [X minutes/hours] | [e.g. "Last committed transaction — database replication is synchronous"] |
| RTO (Recovery Time Objective) | [Y minutes/hours] | [e.g. "Revenue impact begins at 30 min; target recovery in 15 min"] |
| MTTR target (non-disaster) | [Z minutes] | [Operational incidents, not DR events] |
| Data retention (backups) | [N days/weeks] | [Compliance requirement or operational policy] |
| Backup frequency | [Every X hours] | [RPO-driven — backup interval must be ≤ RPO] |
What these mean in practice:
| Scenario | Likelihood | Impact | RTO target | RPO target | Runbook |
|---|---|---|---|---|---|
| Single availability zone failure | Medium | [Partial / Full outage] | [15 min] | [0 — no data loss] | Section 3.1 |
| Full region failure | Low | Full outage | [60 min] | [5 min] | Section 3.2 |
| Database corruption / data loss | Low | Full outage | [90 min] | [RPO value] | Section 3.3 |
| Critical dependency outage | High | [Partial degradation] | [30 min] | [N/A] | Section 3.4 |
| Security breach / ransomware | Very low | Full outage + investigation | [4 hours] | [Last clean backup] | Section 3.5 |
| Accidental bulk data deletion | Low | Partial or full data loss | [60 min] | [RPO value] | Section 3.6 |
Trigger: One AZ becomes unreachable — pods/instances in that zone stop responding.
Detection: PagerDuty alert [AlertName] fires, or cloud provider status page shows AZ degradation.
Expected RTO: [15 minutes] | Expected RPO: Zero (no data loss if multi-AZ replication is working)
Step 1 — Confirm the failure
# Check pod/instance health across zones
kubectl get pods -o wide -n [namespace] | grep -v Running
# Check which nodes are affected
kubectl get nodes -o wide | grep -v Ready
# Verify cloud provider AZ status
# AWS: https://health.aws.amazon.com/health/status
# GCP: https://status.cloud.google.comStep 2 — Assess whether auto-recovery has occurred
# If using auto-scaling, check if replacement instances launched
kubectl get pods -n [namespace] --watch
# Check deployment replica count
kubectl get deployment [service-name] -n [namespace]
# Verify load balancer health checks are passing
[cloud provider CLI command to check target group health]Step 3 — Force rescheduling if auto-recovery stalled
# Cordon the affected node so no new pods schedule on it
kubectl cordon [node-name]
# Drain the node — moves all pods to healthy nodes
kubectl drain [node-name] --ignore-daemonsets --delete-emptydir-data
# Verify pods have rescheduled successfully
kubectl get pods -o wide -n [namespace]Step 4 — Verify service health
# Smoke test key endpoints
curl -s -o /dev/null -w "%{http_code}" https://[service-url]/health
curl -s -o /dev/null -w "%{http_code}" https://[service-url]/[critical-endpoint]
# Check error rate in monitoring
[dashboard link or query]Recovery confirmed when: All pods are Running, health check returns 200, error rate is at baseline.
Trigger: The primary region is entirely unavailable. Detection: All service health checks failing, cloud provider status page confirms region-wide event. Expected RTO: [60 minutes] | Expected RPO: [5 minutes — based on cross-region replication lag]
Step 1 — Confirm regional failure (5 minutes)
# Confirm the primary region is unreachable
ping [primary-region-endpoint] || echo "Primary region unreachable"
# Check replication lag on standby region database
[command to check replica lag — e.g. for RDS: aws rds describe-db-instances --region [dr-region]]Step 2 — Declare DR event and notify (2 minutes)
Post to #incidents:
🔴 DR EVENT — [Service Name] — Region Failure
Primary region: [region] — UNREACHABLE
Activating failover to: [dr-region]
Incident commander: [Name]
Next update: 15 minutesPage [Engineering Manager] and [CTO/VP Eng] via PagerDuty.
Step 3 — Promote DR database (10 minutes)
# AWS RDS — promote read replica to primary
aws rds promote-read-replica \
--db-instance-identifier [dr-replica-identifier] \
--region [dr-region]
# Wait for promotion to complete
aws rds wait db-instance-available \
--db-instance-identifier [dr-replica-identifier] \
--region [dr-region]
# Record the new database endpoint
aws rds describe-db-instances \
--db-instance-identifier [dr-replica-identifier] \
--region [dr-region] \
--query 'DBInstances[0].Endpoint.Address'Step 4 — Deploy service in DR region (20 minutes)
# Update service configuration to point at DR database
kubectl set env deployment/[service-name] \
DATABASE_URL=[new-dr-database-url] \
-n [namespace] \
--context [dr-region-context]
# Scale up the DR deployment
kubectl scale deployment/[service-name] --replicas=[N] \
-n [namespace] \
--context [dr-region-context]
# Verify all pods are running
kubectl get pods -n [namespace] --context [dr-region-context]Step 5 — Cut over DNS / load balancer (5 minutes)
# Update DNS to point to DR region load balancer
# AWS Route 53:
aws route53 change-resource-record-sets \
--hosted-zone-id [zone-id] \
--change-batch file://dr-failover-dns.json
# Verify DNS propagation (may take up to [TTL] seconds)
dig [service-domain] @8.8.8.8Step 6 — Verify end-to-end
# Full smoke test against DR endpoint
curl -s https://[service-url]/health
[run automated smoke test suite if available]Recovery confirmed when: DNS resolves to DR region, smoke tests pass, error rate is at baseline.
Post-failover actions (not urgent — after service is stable):
Trigger: Data in the database is corrupted, deleted, or otherwise incorrect due to a software bug, operator error, or hardware fault. Detection: Application errors referencing missing/invalid data, monitoring alerts on query error rate, user reports. Expected RTO: [90 minutes] | Expected RPO: [Backup interval — e.g. 1 hour]
Step 1 — Stop the bleeding immediately
# Put the service into maintenance mode to prevent further writes to corrupted data
[command to enable maintenance mode — e.g. kubectl set env deployment/[name] MAINTENANCE_MODE=true]
# Or: scale down the service to zero to prevent writes
kubectl scale deployment/[service-name] --replicas=0 -n [namespace]Step 2 — Assess scope of corruption
# Identify which tables/records are affected
[SQL query to check data integrity — e.g.]
# psql $DATABASE_URL -c "SELECT COUNT(*) FROM [table] WHERE [integrity check condition]"
# Determine when corruption started (cross-reference with deploy times and error logs)
[log query to find earliest error — e.g. in Datadog:]
# service:[service-name] status:error "[corruption error message]" | sort by timestamp ascStep 3 — Identify the correct restore point
# List available backups
[command to list backups — e.g. for RDS:]
aws rds describe-db-snapshots \
--db-instance-identifier [db-identifier] \
--query 'DBSnapshots[*].[SnapshotCreateTime,DBSnapshotIdentifier]' \
--output table
# Choose the most recent backup BEFORE corruption started
# Record the chosen snapshot ID: [snapshot-id]Step 4 — Restore from backup
# Restore to a NEW database instance (never overwrite production directly)
aws rds restore-db-instance-from-db-snapshot \
--db-instance-identifier [service-name]-restored-[date] \
--db-snapshot-identifier [snapshot-id] \
--region [region]
# Wait for restore to complete
aws rds wait db-instance-available \
--db-instance-identifier [service-name]-restored-[date]
# Get the restored instance endpoint
aws rds describe-db-instances \
--db-instance-identifier [service-name]-restored-[date] \
--query 'DBInstances[0].Endpoint.Address'Step 5 — Validate restored data
# Connect to restored database and verify integrity
psql [restored-db-endpoint] -U [user] -d [database] -c "[data integrity query]"
# Confirm record counts match expectations
psql [restored-db-endpoint] -U [user] -d [database] -c "SELECT COUNT(*) FROM [critical-table]"Step 6 — Point service at restored database
kubectl set env deployment/[service-name] \
DATABASE_URL=postgres://[user]:[pass]@[restored-endpoint]/[db] \
-n [namespace]
kubectl scale deployment/[service-name] --replicas=[N] -n [namespace]Recovery confirmed when: Service is running against restored database, data integrity checks pass, error rate is at baseline.
Trigger: A service that [service name] depends on is unavailable or degraded. Detection: Increased error rate or latency on endpoints that call [dependency], alerts from dependency owner. Expected RTO: Depends on dependency — [30 minutes for mitigation, resolution depends on dependency owner]
Dependency map:
| Dependency | Criticality | Degraded behaviour | Mitigation |
|---|---|---|---|
| [Database] | Critical — all writes fail | Full outage | Activate DR database (Section 3.3) |
| [Cache — Redis] | High — latency increases | Performance degradation | Bypass cache, serve from DB |
| [Auth service] | Critical — auth fails | All authenticated endpoints fail | Return cached tokens (if implemented) |
| [Message queue] | Medium — async processing delays | Writes succeed, async jobs queue | Queue backlog — see on-call runbook |
| [External API — name] | Low — feature X unavailable | Graceful degradation | Feature flag to disable feature X |
Mitigation steps:
# Enable circuit breaker / fallback for [dependency] if implemented
kubectl set env deployment/[service-name] [DEPENDENCY]_CIRCUIT_BREAKER=open -n [namespace]
# Enable feature flag to disable [dependency-backed feature]
[feature flag CLI command or dashboard link]
# Check if dependency has a status page
# [Dependency status URL]Escalation: Contact [dependency] on-call via [PagerDuty / Slack #[channel]]. Share your service's error rate and the time dependency errors started.
Trigger: Evidence of unauthorized access, data exfiltration, or encryption of service data. Detection: Security tooling alert, unusual access patterns, user reports of data exposure. Expected RTO: [4+ hours — prioritise containment over speed] | Expected RPO: [Last verified clean backup]
Step 1 — Isolate immediately
# Take the service offline — do not attempt to recover while breach is active
kubectl scale deployment/[service-name] --replicas=0 -n [namespace]
# Revoke all API keys and service account credentials immediately
[command to rotate secrets — e.g. via Vault or cloud provider]
# Block all external access at network level
[firewall/security group command to deny all inbound traffic]Step 2 — Notify security team immediately Page [Security lead] via PagerDuty. Do NOT attempt to remediate without security team involvement.
Post to #security-incidents (private channel, not #incidents):
🔴 SECURITY INCIDENT — [Service Name]
Time detected: [Time]
Evidence: [One sentence — what was observed]
Actions taken: Service isolated, credentials revoked
Awaiting: Security team guidanceStep 3 — Preserve evidence
# Export current logs before any remediation
[log export command — preserve evidence for forensics]
# Snapshot the current state of all infrastructure
[snapshot/image command]Steps 4+ — Follow security team guidance. Do not restore from backup until security team confirms the attack vector is closed.
Trigger: An operator, script, or application bug has deleted records in bulk. Detection: Sudden drop in record counts, user reports of missing data, application errors. Expected RTO: [60 minutes] | Expected RPO: [Backup interval]
# Step 1 — Stop further writes immediately
kubectl scale deployment/[service-name] --replicas=0 -n [namespace]
# Step 2 — Determine what was deleted and when
psql $DATABASE_URL -c "
SELECT schemaname, tablename,
n_dead_tup, last_autovacuum
FROM pg_stat_user_tables
ORDER BY n_dead_tup DESC LIMIT 10;
"
# Step 3 — Check if deletion is recoverable via MVCC (PostgreSQL)
# Records may still be recoverable if VACUUM has not run
psql $DATABASE_URL -c "
SELECT * FROM [table]
WHERE xmax != 0 -- recently deleted rows
LIMIT 100;
"
# Step 4 — If not recoverable via MVCC, restore from backup
# Follow Section 3.3 (Database Corruption runbook) from Step 3 onward| Data store | Backup type | Frequency | Retention | Location |
|---|---|---|---|---|
| [Primary database] | Automated snapshots | Every [N] hours | [N] days | [S3 bucket / cloud storage path] |
| [Primary database] | Transaction log backups | Continuous | [N] days | [Location] |
| [Secondary store — e.g. Redis] | RDB dump | Daily | [N] days | [Location] |
| [Blob/object storage] | Cross-region replication | Continuous | [N] days | [DR region bucket] |
| [Config / secrets] | Terraform state + Vault backup | On change | Indefinite | [Location] |
# Test restore of latest database backup to a throwaway instance
aws rds restore-db-instance-from-db-snapshot \
--db-instance-identifier [service-name]-backup-test-$(date +%Y%m%d) \
--db-snapshot-identifier $(aws rds describe-db-snapshots \
--db-instance-identifier [db-id] \
--query 'sort_by(DBSnapshots, &SnapshotCreateTime)[-1].DBSnapshotIdentifier' \
--output text)
# Wait for restore, then run integrity checks
psql [test-instance-endpoint] -c "[integrity check query]"
# Confirm row counts match recent production values (allow ≤ RPO difference)
psql [test-instance-endpoint] -c "SELECT COUNT(*) FROM [critical-table]"
# Destroy the test instance
aws rds delete-db-instance \
--db-instance-identifier [service-name]-backup-test-$(date +%Y%m%d) \
--skip-final-snapshotRegular testing is mandatory. An untested DR plan is not a DR plan.
| Test type | Frequency | Who runs it | Pass criteria |
|---|---|---|---|
| Backup restore validation | Weekly (automated) | On-call rotation | Restore completes, integrity checks pass |
| Zone failover drill | Monthly | Engineering team | RTO target met, zero data loss |
| Region failover drill | Quarterly | Engineering + SRE | RTO/RPO targets met |
| Full DR game day | Annually | Engineering + stakeholders | All scenarios exercised, gaps documented |
| Chaos engineering (infra failures) | Weekly (automated) | Chaos engineering tooling | Service degrades gracefully, recovers automatically |
Incident commander responsibilities:
Notify these people at DR event start:
| Role | Name | Contact | When to notify |
|---|---|---|---|
| Engineering manager | [Name] | [Slack / Phone] | Immediately |
| CTO / VP Engineering | [Name] | [Phone] | Tier 1 services: immediately |
| Customer success lead | [Name] | [Slack] | If customer-facing impact |
| Security lead | [Name] | [Slack / PagerDuty] | If breach suspected |
| Legal / compliance | [Name] | [Email / Phone] | If data loss involves PII |
DR event declared:
🔴 DR EVENT — [Service Name]
Time: [HH:MM UTC]
Scenario: [Zone failure / Region failure / Data loss / etc.]
Impact: [Who is affected and how]
RTO target: [X minutes]
Incident commander: [Name]
War room: [Slack channel / call link]
Next update: [Time + 15 min]Status update (every 15 minutes):
🔴 DR UPDATE — [Service Name] — [HH:MM UTC]
Status: [Investigating / Executing recovery / Verifying]
Progress: [One sentence on current step]
Blockers: [Any — or "None"]
Updated RTO estimate: [Time]
Next update: [Time + 15 min]Recovery confirmed:
✅ DR RESOLVED — [Service Name] — [HH:MM UTC]
Total downtime: [X minutes]
Data loss: [None / X minutes of transactions]
RTO target: [X min] — Actual: [Y min] — [MET / MISSED]
RPO target: [X min] — Actual: [Y min] — [MET / MISSED]
Root cause: [One sentence]
Post-incident review: [Scheduled for / Link when created]Run this checklist quarterly and before any major infrastructure change:
Backups:
Failover infrastructure:
Runbooks:
Access:
Monitoring:
© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/disaster-recovery-plan of mohitagw15856/pm-claude-skills.
Open the folder on GitHubat commit 1cbf1f0
Disaster Recovery Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Disaster Recovery Plan this skillmohitagw15856/pm-claude-skills | 1.4k | — | ~5.8k | Automated safety check: Pass | MIT | |
| Convex Backupopenclaw/clawhub | 9.5k | — | ~1k | Automated safety check: Pass | MIT | |
| Oraclecloud Incident Runbookjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~2.7k | Automated safety check: Pass | MIT | |
| OmniRoute Backup and Sync CLIdiegosouzapw/OmniRoute | 75k | — | ~948 | Automated safety check: Pass | MIT | |
| Trader Memory Coretradermonty/claude-trading-skills | 3k | 2 repos | ~4.3k | Automated safety check: Pass | MIT | |
| Pymobiledevice3 Device Operatordoronz88/pymobiledevice3 | 2.9k | — | ~1.8k | Automated safety check: Notes | GPL-3.0 |
openclaw/clawhub
Set up Convex backups and run a restore DRILL that proves recovery — snapshot, restore into a throwaway preview, assert the data came back — plus a schedule matched to your RPO and a gated recovery…
jeremylongshore/tons-of-skills-marketplace
Self-service incident runbook for OCI outages — health probes, instance recovery, cross-AD/region failover.
diegosouzapw/OmniRoute
Backup and restore OmniRoute data from the CLI. Trigger incremental snapshots, sync to cloud storage, manage backup schedules, and restore from archive files.
tradermonty/claude-trading-skills
Track investment theses across their lifecycle — from screening idea to closed position with postmortem.
doronz88/pymobiledevice3
Operate iOS and iPadOS devices with pymobiledevice3, from a local checkout or straight from PyPI via uvx on a fresh workstation.
nrwl/nx
Author or scope a first-party Nx migration. An agent skill from nrwl/nx.
mohitagw15856/pm-claude-skills
Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…
mohitagw15856/pm-claude-skills
Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.
mohitagw15856/pm-claude-skills
Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.
mohitagw15856/pm-claude-skills
Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.
mohitagw15856/pm-claude-skills
Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.
mohitagw15856/pm-claude-skills
Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.
Categories
Write a disaster recovery plan for a service or system — covering RPO/RTO targets, failure scenario runbooks, backup and restore procedures, DR testing cadence, and communication templates. Disaster Recovery Plan is an agent skill from mohitagw15856/pm-claude-skills. Write a disaster recovery plan for a service or system — covering RPO/RTO targets, failure scenario runbooks, backup and restore procedures, DR testing cadence, and communication templates.
Disaster Recovery Plan fits situations like: asked to write a DR plan; document failover procedures; create recovery runbooks; define RTO/RPO targets.
Run `npx skills add mohitagw15856/pm-claude-skills --skill disaster-recovery-plan -a claude-code`. Or copy the skill folder (skills/disaster-recovery-plan in mohitagw15856/pm-claude-skills) into .claude/skills/disaster-recovery-plan in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mohitagw15856/pm-claude-skills --skill disaster-recovery-plan -a codex`. Or copy the skill folder (skills/disaster-recovery-plan in mohitagw15856/pm-claude-skills) into .agents/skills/disaster-recovery-plan in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill disaster-recovery-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/disaster-recovery-plan, .gemini/skills/disaster-recovery-plan, .github/skills/disaster-recovery-plan and .opencode/skills/disaster-recovery-plan in your project.
Going by SKILL.md and its folder, Disaster Recovery Plan needs the command-line tools its instructions call (kubectl, aws, psql and curl).
SKILL.md names 2 domains. In commands or code: health.aws.amazon.com and status.cloud.google.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Disaster Recovery Plan is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.8k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Disaster Recovery Plan: Convex Backup (openclaw/clawhub, 9.5k stars), Oraclecloud Incident Runbook (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), OmniRoute Backup and Sync CLI (diegosouzapw/OmniRoute, 75k stars) and Trader Memory Core (tradermonty/claude-trading-skills, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,434 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 9, 2026.
Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.