Agent skill

Oraclecloud Incident Runbook

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Self-service incident runbook for OCI outages — health probes, instance recovery, cross-AD/region failover.

MITAuto-check passedDevOps & Cloud

Install Oraclecloud Incident Runbook

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill oraclecloud-incident-runbook -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace oraclecloud-incident-runbook --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/oraclecloud-incident-runbook .claude/skills/oraclecloud-incident-runbook && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
oraclecloud-incident-runbook
GitHub stars
2.8k
Token cost
~2.7k tokens
SKILL.md length
559 words
Files
2 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Self-service incident runbook for OCI outages — health probes, instance recovery, cross-AD/region failover.

  • Works in 6 steps: Independent Health Probes → Severity Classification → Instance Recovery Actions → …
  • OCI instances go down
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 4 more sections
  • Calls pip

What it does

Oraclecloud Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Self-service incident runbook for OCI outages — health probes, instance recovery, cross-AD/region failover. Use when OCI instances go down, the status page is silent, or you need automated recovery without waiting for support. Trigger with "oraclecloud incident", "oci outage runbook", "oci failover", "oci instance recovery".

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/one-pager.md`). Compatibility notes: Designed for Claude Code

It sits in DevOps & Cloud, covering Runbooks and postmortems, Incident response and Backup and disaster recovery. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • OCI instances go down
  • The status page is silent
  • You need automated recovery without waiting for support
  • With oraclecloud incident

Example prompts

  • “oraclecloud incident”
  • “oci outage runbook”
  • “oci failover”
  • “/oraclecloud-incident-runbook”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(oci:*), Bash(python3:*), Grep

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Independent Health Probes
  2. Severity Classification
  3. Instance Recovery Actions
  4. Cross-AD Failover
  5. Cross-Region Failover
  6. Status Monitoring Script

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(oci:*)
    • Bash(python3:*)
    • Grep

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.oracle.com
    • ocistatus.oraclecloud.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Oraclecloud Incident Runbook loads about 2.7k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 559 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 559 words, ~2,661 tokens.

Download SKILL.mdSave it as .claude/skills/oraclecloud-incident-runbook/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
oraclecloud-incident-runbook
description
Self-service incident runbook for OCI outages — health probes, instance recovery, cross-AD/region failover. Use when OCI instances go down, the status page is silent, or you need automated recovery without waiting for support. Trigger with "oraclecloud incident", "oci outage runbook", "oci failover", "oci instance recovery".
allowed-tools
Read, Write, Edit, Bash(oci:*), Bash(python3:*), Grep
compatibility
Designed for Claude Code
version
1.8.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, oraclecloud, oci

Oracle Cloud Incident Runbook

Overview

Self-service runbook for when OCI instances go down and the status page stays green. OCI's status page has a history of not acknowledging outages in real time (London Jan 2026 — 502s and instances disappearing for 10 minutes with no status update). OCI Support response times average 4+ hours for Sev-1 tickets. This runbook gives you health probes, automated instance recovery, cross-AD failover, and cross-region failover — all executable without waiting on Oracle.

Purpose: Detect OCI service degradation independently, recover instances automatically, and fail over to alternate availability domains or regions when the primary is impacted.

Prerequisites

  • OCI CLI installed and configured — ~/.oci/config validated (see oraclecloud-install-auth)
  • Python 3.8+ with the OCI SDK — pip install oci
  • Pre-configured resources: at least one compute instance, a VCN with subnets in multiple ADs
  • IAM policies: manage instances, manage volumes, inspect work-requests in the target compartment
  • Boot volume backups enabled (recovery depends on having a recent backup)

Instructions

Step 1: Independent Health Probes

Do not trust the OCI status page alone. Run your own health checks against the OCI API:

python
import oci
import time

config = oci.config.from_file("~/.oci/config")

def probe_oci_health(config):
    """Probe OCI API endpoints independently of the status page."""
    results = {}

    # Probe 1: Identity service (lightest call)
    try:
        start = time.time()
        identity = oci.identity.IdentityClient(config)
        identity.list_regions()
        results["identity"] = {"status": "healthy", "latency_ms": int((time.time() - start) * 1000)}
    except oci.exceptions.ServiceError as e:
        results["identity"] = {"status": "degraded", "error": str(e.status)}

    # Probe 2: Compute service
    try:
        start = time.time()
        compute = oci.core.ComputeClient(config)
        compute.list_instances(compartment_id=config["tenancy"], limit=1)
        results["compute"] = {"status": "healthy", "latency_ms": int((time.time() - start) * 1000)}
    except oci.exceptions.ServiceError as e:
        results["compute"] = {"status": "degraded", "error": str(e.status)}

    # Probe 3: Networking service
    try:
        start = time.time()
        network = oci.core.VirtualNetworkClient(config)
        network.list_vcns(compartment_id=config["tenancy"], limit=1)
        results["networking"] = {"status": "healthy", "latency_ms": int((time.time() - start) * 1000)}
    except oci.exceptions.ServiceError as e:
        results["networking"] = {"status": "degraded", "error": str(e.status)}

    return results

health = probe_oci_health(config)
for service, status in health.items():
    print(f"  {service}: {status['status']} ({status.get('latency_ms', 'N/A')}ms)")
Step 2: Severity Classification
LevelConditionDetectionResponse Time
P1 — Total OutageAll API probes fail, instances unreachableAll 3 probes return degradedImmediate — trigger cross-region failover
P2 — Partial DegradationSome APIs slow (>2s latency), instance agent unresponsiveLatency >2000ms on any probe15 min — attempt instance recovery, prepare failover
P3 — IntermittentSporadic 429/500 errors, some requests succeedOccasional probe failures1 hour — enable retries, monitor trend
Step 3: Instance Recovery Actions
bash
# Check current instance state
oci compute instance get --instance-id "$INSTANCE_OCID" \
  --query 'data."lifecycle-state"' --raw-output

# Action: Reset (hard reboot) — fastest recovery for hung instances
oci compute instance action --instance-id "$INSTANCE_OCID" \
  --action RESET

# Action: Reboot (graceful) — for instances still responding to ACPI
oci compute instance action --instance-id "$INSTANCE_OCID" \
  --action SOFTRESET

# Action: Stop then Start — forces reallocation to new host hardware
oci compute instance action --instance-id "$INSTANCE_OCID" \
  --action STOP
# Wait for STOPPED state
oci compute instance action --instance-id "$INSTANCE_OCID" \
  --action START
Step 4: Cross-AD Failover

When an entire availability domain is impacted, launch a replacement instance in a different AD using the latest boot volume backup:

python
import oci

config = oci.config.from_file("~/.oci/config")
compute = oci.core.ComputeClient(config)
blockstorage = oci.core.BlockstorageClient(config)

# Find latest boot volume backup
backups = blockstorage.list_boot_volume_backups(
    compartment_id="COMPARTMENT_OCID",
    boot_volume_id="BOOT_VOLUME_OCID",
    sort_by="TIMECREATED",
    sort_order="DESC",
    limit=1,
).data

if not backups:
    raise RuntimeError("No boot volume backups found — cannot failover")

latest_backup = backups[0]
print(f"Using backup: {latest_backup.id} from {latest_backup.time_created}")

# Launch replacement instance in alternate AD
launch_details = oci.core.models.LaunchInstanceDetails(
    availability_domain="AD-2",  # Different from the impacted AD
    compartment_id="COMPARTMENT_OCID",
    shape="VM.Standard.E4.Flex",
    shape_config=oci.core.models.LaunchInstanceShapeConfigDetails(ocpus=2, memory_in_gbs=16),
    display_name="failover-instance",
    source_details=oci.core.models.InstanceSourceViaBootVolumeDetails(
        source_type="bootVolume",
        boot_volume_id=latest_backup.id,
    ),
    create_vnic_details=oci.core.models.CreateVnicDetails(
        subnet_id="SUBNET_OCID_IN_AD2",
    ),
)

response = compute.launch_instance(launch_details)
print(f"Failover instance launching: {response.data.id}")
Step 5: Cross-Region Failover

For region-wide outages, use a pre-replicated boot volume in the DR region:

bash
# List available boot volume replicas in DR region
oci bv boot-volume-backup list \
  --compartment-id "$COMPARTMENT_OCID" \
  --query 'data[0].id' --raw-output \
  --region us-phoenix-1

# Launch instance in DR region
oci compute instance launch \
  --compartment-id "$COMPARTMENT_OCID" \
  --availability-domain "AD-1" \
  --shape "VM.Standard.E4.Flex" \
  --shape-config '{"ocpus":2,"memoryInGBs":16}' \
  --display-name "dr-failover-instance" \
  --source-boot-volume-id "$DR_BACKUP_OCID" \
  --subnet-id "$DR_SUBNET_OCID" \
  --region us-phoenix-1
Step 6: Status Monitoring Script

Run this in a loop during an incident to detect recovery:

bash
#!/bin/bash
while true; do
  STATUS=$(oci compute instance get --instance-id "$INSTANCE_OCID" \
    --query 'data."lifecycle-state"' --raw-output 2>&1)
  TIMESTAMP=$(date -u +%H:%M:%S)
  echo "$TIMESTAMP | Instance: $STATUS"
  if [ "$STATUS" = "RUNNING" ]; then
    echo "$TIMESTAMP | Instance recovered."
    break
  fi
  sleep 30
done
Show full SKILL.md (252 more words)Show less

Output

Successful execution produces:

  • Independent health probe results for Identity, Compute, and Networking services
  • Severity classification (P1/P2/P3) based on probe results
  • Instance recovery action (reset, reboot, or stop/start) applied to the impacted instance
  • Cross-AD failover instance launched from the latest boot volume backup (if AD-level failure)
  • Cross-region failover instance launched in the DR region (if region-level failure)
  • Continuous monitoring log tracking instance lifecycle state until recovery

Error Handling

ErrorCodeCauseSolution
NotAuthenticated401API key expired during incidentUse oci session authenticate for token-based short-term auth
NotAuthorizedOrNotFound404IAM policy missing for recovery actionsPre-create policy: allow group sre to manage instances in compartment prod
TooManyRequests429Rate limited during mass recoveryOCI has no Retry-After header — back off 30 seconds between calls
InternalError500OCI service itself is degradedThis confirms the outage — proceed to cross-region failover
No boot volume backups—Backups not configuredCannot failover without backups — enable in Block Storage > Boot Volumes > Backup Policies
ServiceError status -1—API timeout (region unreachable)Switch to DR region immediately: --region us-phoenix-1

Examples

Quick health check (CLI one-liner):

bash
# Test if OCI API is responsive
oci iam region list --output table && echo "OCI API: OK" || echo "OCI API: UNREACHABLE"

Automated recovery wrapper:

bash
#!/bin/bash
set -euo pipefail
INSTANCE_OCID="${1:?Usage: recover.sh <instance-ocid>}"
STATE=$(oci compute instance get --instance-id "$INSTANCE_OCID" \
  --query 'data."lifecycle-state"' --raw-output)
case "$STATE" in
  RUNNING) echo "Instance is healthy" ;;
  STOPPED) oci compute instance action --instance-id "$INSTANCE_OCID" --action START ;;
  *)       oci compute instance action --instance-id "$INSTANCE_OCID" --action RESET ;;
esac

Resources

Next Steps

After stabilizing the incident, run oraclecloud-debug-bundle to collect forensic data for the postmortem, or review oraclecloud-prod-checklist to harden your environment against future outages.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/.curated/oraclecloud-incident-runbook of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/one-pager.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Oraclecloud Incident Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Oraclecloud Incident Runbook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Oraclecloud Incident Runbook this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2.7kAutomated safety check: PassMIT
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0
Activation Governance Chaos RolloutAli-Marandi/DataSense107—~1.9kAutomated safety check: PassMIT
Incident Response686f6c61/alfred-dev117—~1.1kAutomated safety check: PassMIT
Superset Incident Triagesuperset-sh/superset15k—~1kAutomated safety check: PassCustom licence
Post-Incident DebriefVeryGoodOpenSource/vgv-wingspan109—~1.9kAutomated safety check: PassMIT

Similar skills

  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Incident Response

    686f6c61/alfred-dev

    Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.

    117 GitHub stars~1.1k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Superset Incident Triage

    superset-sh/superset

    Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.

    15k GitHub stars~1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Post-Incident Debrief

    VeryGoodOpenSource/vgv-wingspan

    Produces a blameless post-incident debrief with timeline, root cause and follow-up actions after an outage, failed release or significant bug, while details are fresh.

    109 GitHub stars~1.9k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Oraclecloud Incident Runbook

What does Oraclecloud Incident Runbook do?

Self-service incident runbook for OCI outages — health probes, instance recovery, cross-AD/region failover. Oraclecloud Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Self-service incident runbook for OCI outages — health probes, instance recovery, cross-AD/region failover.

When should I use Oraclecloud Incident Runbook?

Oraclecloud Incident Runbook fits situations like: OCI instances go down; the status page is silent; you need automated recovery without waiting for support; with oraclecloud incident.

How do I install Oraclecloud Incident Runbook in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill oraclecloud-incident-runbook -a claude-code`. Or copy the skill folder (skills/.curated/oraclecloud-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/oraclecloud-incident-runbook in your project. Claude Code loads it when a task matches its description.

How do I install Oraclecloud Incident Runbook in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill oraclecloud-incident-runbook -a codex`. Or copy the skill folder (skills/.curated/oraclecloud-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/oraclecloud-incident-runbook in your project. Codex loads it when a task matches its description.

Can I use Oraclecloud Incident Runbook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill oraclecloud-incident-runbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/oraclecloud-incident-runbook, .gemini/skills/oraclecloud-incident-runbook, .github/skills/oraclecloud-incident-runbook and .opencode/skills/oraclecloud-incident-runbook in your project.

What does Oraclecloud Incident Runbook need to run?

Going by SKILL.md and its folder, Oraclecloud Incident Runbook needs the command-line tools its instructions call (pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(oci:*), Bash(python3:*), Grep. Compatibility (from SKILL.md): Designed for Claude Code.

Does Oraclecloud Incident Runbook access the network?

SKILL.md names 2 domains. As links in the text: docs.oracle.com and ocistatus.oraclecloud.com. This is read from the text; nothing was executed.

Is Oraclecloud Incident Runbook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Oraclecloud Incident Runbook use?

Oraclecloud Incident Runbook is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Oraclecloud Incident Runbook use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 436 tokens, read only when the agent opens those files.

What are the alternatives to Oraclecloud Incident Runbook?

Skills that share tags, products or a category with Oraclecloud Incident Runbook: Oncall (pigweed-project/pigweed, 548 stars), Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars), Incident Response (686f6c61/alfred-dev, 117 stars) and Superset Incident Triage (superset-sh/superset, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Oraclecloud Incident Runbook?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.