Agent skill

Vercel Incident Runbook

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Vercel incident response procedures with triage, instant rollback, and postmortem.

MITAuto-check passedDevOps & Cloud

Install Vercel Incident Runbook

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill vercel-incident-runbook -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace vercel-incident-runbook --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/vercel-incident-runbook .claude/skills/vercel-incident-runbook && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vercel-incident-runbook
GitHub stars
2.8k
Token cost
~2k tokens
SKILL.md length
313 words
Files
4 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Vercel incident response procedures with triage, instant rollback, and postmortem.

  • Works in 7 steps: Rapid Triage (First 5 Minutes) → Decision Tree → Instant Rollback (< 30 Seconds) → …
  • Responding to Vercel-related outages
  • SKILL.md covers Overview, Prerequisites, Instructions and Incident Severity Levels, plus 5 more sections
  • Calls vercel, jq and curl; reaches api.vercel.com and vercel-status.com; needs VERCEL_TOKEN

What it does

Vercel Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Vercel incident response procedures with triage, instant rollback, and postmortem. Use when responding to Vercel-related outages, investigating production errors, or running post-incident reviews for deployment failures. Trigger with phrases like "vercel incident", "vercel outage", "vercel down", "vercel on-call", "vercel emergency", "vercel broken".

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/errors.md`, `references/examples.md` and `references/immediate-actions-by-error-type.md`). Compatibility notes: Designed for Claude Code

It sits in DevOps & Cloud, covering Incident response and Runbooks and postmortems. It works with Vercel and Slack. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Responding to Vercel-related outages
  • Investigating production errors
  • Running post-incident reviews for deployment failures
  • With phrases like vercel incident

Example prompts

  • “vercel incident”
  • “vercel outage”
  • “vercel down”
  • “/vercel-incident-runbook”

Requirements

  • A credential in VERCEL_TOKEN
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Grep, Bash(vercel:*), Bash(curl:*)

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Rapid Triage (First 5 Minutes)
  2. Decision Tree
  3. Instant Rollback (< 30 Seconds)
  4. Investigate Root Cause
  5. Enable Maintenance Page (If Needed)
  6. Communication Templates
  7. Postmortem Template

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Bash(vercel:*)
    • Bash(curl:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • vercel
    • jq
    • curl
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.vercel.com
    • vercel-status.com

    Also links to:

    • vercel.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VERCEL_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Vercel Incident Runbook loads about 2k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 94 tokens; SKILL.md has 313 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 313 words, ~1,956 tokens.

Download SKILL.mdSave it as .claude/skills/vercel-incident-runbook/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
vercel-incident-runbook
description
Vercel incident response procedures with triage, instant rollback, and postmortem. Use when responding to Vercel-related outages, investigating production errors, or running post-incident reviews for deployment failures. Trigger with phrases like "vercel incident", "vercel outage", "vercel down", "vercel on-call", "vercel emergency", "vercel broken".
allowed-tools
Read, Grep, Bash(vercel:*), Bash(curl:*)
compatibility
Designed for Claude Code
version
1.18.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, vercel, incident-response, runbook

Vercel Incident Runbook

Overview

Step-by-step incident response for Vercel deployment failures, function errors, and platform outages. Covers rapid triage, instant rollback, communication templates, and postmortem procedures.

Prerequisites

  • Access to Vercel dashboard and CLI
  • Access to Vercel status page (vercel-status.com)
  • Communication channels (Slack, PagerDuty) configured
  • Log drain or runtime log access

Instructions

Step 1: Rapid Triage (First 5 Minutes)
bash
# 1. Check if it's a Vercel platform issue
curl -s "https://www.vercel-status.com/api/v2/summary.json" \
  | jq '.status.description, [.components[] | select(.status != "operational") | {name, status}]'

# 2. Check current production deployment status
vercel ls --prod
vercel inspect $(vercel ls --prod --json | jq -r '.[0].url')

# 3. Check recent deployments — did a deploy just happen?
curl -s -H "Authorization: Bearer $VERCEL_TOKEN" \
  "https://api.vercel.com/v6/deployments?target=production&limit=5&projectId=prj_xxx" \
  | jq '.deployments[] | {uid, state, createdAt: (.createdAt/1000 | todate), url}'

# 4. Check function logs for errors
vercel logs $(vercel ls --prod --json | jq -r '.[0].url') --level=error --limit=20
Step 2: Decision Tree
Is vercel-status.com showing an incident?
├── YES → Vercel platform issue
│   ├── Subscribe to updates on status page
│   ├── Post internal status: "Vercel platform incident — monitoring"
│   └── No action needed from us — wait for Vercel resolution
│
└── NO → Issue is in our deployment
    ├── Did a deployment happen in the last 30 minutes?
    │   ├── YES → Likely deployment regression
    │   │   └── ROLLBACK immediately (Step 3)
    │   └── NO → Application-level issue
    │       ├── Check function logs for new errors
    │       ├── Check external dependency status (DB, APIs)
    │       └── Investigate and hotfix (Step 4)
    │
    └── Is the issue region-specific?
        ├── YES → Check function regions, possible edge issue
        └── NO → Global issue, check code and env vars
Step 3: Instant Rollback (< 30 Seconds)
bash
# Option A: Rollback to previous production deployment (fastest)
vercel rollback
# This instantly swaps production traffic — no rebuild needed

# Option B: Rollback to a specific known-good deployment
vercel rollback dpl_xxxxxxxxxxxx

# Option C: Via API (for automation/PagerDuty integration)
curl -X POST "https://api.vercel.com/v9/projects/my-app/promote" \
  -H "Authorization: Bearer $VERCEL_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"deploymentId": "dpl_known_good_id"}'

# Verify rollback succeeded
vercel ls --prod
curl -s https://yourdomain.com/api/health | jq .
Step 4: Investigate Root Cause
bash
# Collect evidence while it's fresh
mkdir incident-$(date +%Y%m%d)
cd incident-$(date +%Y%m%d)

# Function logs around the incident time
vercel logs https://yourdomain.com --limit=200 > function-logs.txt

# Deployment diff — what changed?
curl -s -H "Authorization: Bearer $VERCEL_TOKEN" \
  "https://api.vercel.com/v13/deployments/dpl_broken" \
  | jq '.meta' > broken-deployment-meta.json

# Compare env vars between working and broken deployments
vercel env ls > env-vars.txt

# Check git diff between last good and broken commit
git log --oneline -10
git diff dpl_good_commit..dpl_broken_commit -- api/ src/
Step 5: Enable Maintenance Page (If Needed)
json
// vercel.json — temporary maintenance mode via rewrite
{
  "rewrites": [
    {
      "source": "/((?!_next|api/health).*)",
      "destination": "/maintenance.html"
    }
  ]
}
html
<!-- public/maintenance.html -->
<!DOCTYPE html>
<html>
<head><title>Maintenance</title></head>
<body>
  <h1>We'll be right back</h1>
  <p>We're performing scheduled maintenance. Please check back shortly.</p>
</body>
</html>
Step 6: Communication Templates

Internal — Slack (Incident Start)

:rotating_light: INCIDENT: [Project Name] production issue detected
Status: Investigating
Impact: [Description of user impact]
Start time: [UTC timestamp]
On-call: @[engineer]
Thread: replies here

Internal — Slack (Mitigation)

:white_check_mark: MITIGATED: [Project Name]
Action: Rolled back to deployment dpl_xxx
Impact duration: [X minutes]
Root cause: [Brief description]
Postmortem: [link] scheduled for [date]

External — Status Page

Title: Degraded performance on [service]
Body: We are investigating reports of [issue]. Some users may experience
[impact]. Our team is actively working on a resolution.
Update: The issue has been resolved. [Brief root cause].
Step 7: Postmortem Template
markdown
# Incident Postmortem: [Title]

## Summary
- Duration: [start] to [end] ([X minutes])
- Impact: [users/requests affected]
- Severity: [P1/P2/P3]

## Timeline (UTC)
- HH:MM — [event]
- HH:MM — Alert fired
- HH:MM — On-call acknowledged
- HH:MM — Root cause identified
- HH:MM — Rollback executed
- HH:MM — Service restored

## Root Cause
[What broke and why]

## Resolution
[What was done to fix it]

## Action Items
- [ ] [Preventive action] — Owner: @xxx — Due: [date]
- [ ] [Detection improvement] — Owner: @xxx — Due: [date]
- [ ] [Process improvement] — Owner: @xxx — Due: [date]

Incident Severity Levels

SeverityDefinitionResponse TimeRollback?
P1Production down, all users affected< 5 minImmediate
P2Degraded, some users affected< 15 minIf not fixable in 30 min
P3Minor issue, workaround exists< 1 hourNo
P4Cosmetic or non-urgentNext business dayNo

Output

  • Incident categorized and triaged within 5 minutes
  • Instant rollback executed if deployment regression detected
  • Communication sent to internal and external stakeholders
  • Postmortem scheduled with action items

Error Handling

ScenarioAction
Vercel status page shows incidentMonitor, communicate, no deployment changes
vercel rollback failsUse API promotion: POST to /v9/projects/.../promote
Rollback deployment also brokenDeploy from a known-good git tag
Cannot access Vercel dashboardUse CLI with saved VERCEL_TOKEN
Log retention expiredCheck external log drain provider

Examples

Contain a deployment regression during an incident

Declare the severity and incident lead, capture the affected deployment ID and baseline error rate, then use Vercel’s instant rollback to the previously verified deployment. Validate recovery with synthetic requests from the monitoring system before announcing restoration. Do not change unrelated configuration while containment is active; preserve logs and the rollback receipt for the post-incident review, then schedule remediation as a separately reviewed change.

Resources

Next Steps

For data handling and compliance, see vercel-data-handling.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/.curated/vercel-incident-runbook of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/errors.md
  • references/examples.md
  • references/immediate-actions-by-error-type.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Vercel Incident Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vercel Incident Runbook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vercel Incident Runbook this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2kAutomated safety check: PassMIT
Learningskortix-ai/suna20k—~1.1kAutomated safety check: PassCustom licence
Oncallpigweed-project/pigweed548—~963Automated safety check: PassApache-2.0
Axiom SRE Investigatoropenclaw/clawhub9.5k—~7.1kAutomated safety check: PassMIT
Activation Governance Chaos RolloutAli-Marandi/DataSense107—~1.9kAutomated safety check: PassMIT
Alerting Irmgrafana/skills2821 repos~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Learnings

    kortix-ai/suna

    The project's episodic memory: a timestamped ledger of rules paid for with real outages and near-misses, one entry per incident.

    20k GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Oncall

    pigweed-project/pigweed

    Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).

    548 GitHub stars~963 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Axiom SRE Investigator

    openclaw/clawhub

    Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.

    9.5k GitHub stars~7.1k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Design, validate, and govern fail-closed customer-activation automations that use an Outbox/worker pattern.

    107 GitHub stars~1.9k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Incident Response

    686f6c61/alfred-dev

    Protocolo de respuesta ante incidentes en produccion: triaje, mitigacion, causa raiz y postmortem.

    117 GitHub stars~1.1k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Vercel Incident Runbook

What does Vercel Incident Runbook do?

Vercel incident response procedures with triage, instant rollback, and postmortem. Vercel Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Vercel incident response procedures with triage, instant rollback, and postmortem.

When should I use Vercel Incident Runbook?

Vercel Incident Runbook fits situations like: responding to Vercel-related outages; investigating production errors; running post-incident reviews for deployment failures; with phrases like vercel incident.

How do I install Vercel Incident Runbook in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill vercel-incident-runbook -a claude-code`. Or copy the skill folder (skills/.curated/vercel-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/vercel-incident-runbook in your project. Claude Code loads it when a task matches its description.

How do I install Vercel Incident Runbook in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill vercel-incident-runbook -a codex`. Or copy the skill folder (skills/.curated/vercel-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/vercel-incident-runbook in your project. Codex loads it when a task matches its description.

Can I use Vercel Incident Runbook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill vercel-incident-runbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vercel-incident-runbook, .gemini/skills/vercel-incident-runbook, .github/skills/vercel-incident-runbook and .opencode/skills/vercel-incident-runbook in your project.

What does Vercel Incident Runbook need to run?

Going by SKILL.md and its folder, Vercel Incident Runbook needs the command-line tools its instructions call (vercel, jq, curl and git) and credentials named VERCEL_TOKEN. Our summary lists: A credential in VERCEL_TOKEN. Its frontmatter pre-approves these tools: Read, Grep, Bash(vercel:*), Bash(curl:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Vercel Incident Runbook access the network?

SKILL.md names 3 domains. In commands or code: api.vercel.com and vercel-status.com; the agent is likely to contact these when it follows the instructions. As links in the text: vercel.com. This is read from the text; nothing was executed.

Is Vercel Incident Runbook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vercel Incident Runbook use?

Vercel Incident Runbook is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vercel Incident Runbook use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 443 tokens, read only when the agent opens those files.

What are the alternatives to Vercel Incident Runbook?

Skills that share tags, products or a category with Vercel Incident Runbook: Learnings (kortix-ai/suna, 20k stars), Oncall (pigweed-project/pigweed, 548 stars), Axiom SRE Investigator (openclaw/clawhub, 9.5k stars) and Activation Governance Chaos Rollout (Ali-Marandi/DataSense, 107 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vercel Incident Runbook?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.