Agent skill

Langchain Incident Runbook

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Triage LangChain 1.0 / LangGraph 1.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment.

MITAuto-check passedDevOps & Cloud

Install Langchain Incident Runbook

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-incident-runbook --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langchain-incident-runbook .claude/skills/langchain-incident-runbook && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
langchain-incident-runbook
GitHub stars
2.8k
Token cost
~3.8k tokens
SKILL.md length
1,609 words
Files
6 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Triage LangChain 1.0 / LangGraph 1.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment.

  • Works in 5 steps: Define the LLM-specific SLO set → Triage decision tree: which root path? → Provider outage: detect, circuit-break,… → …
  • With langchain incident
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Langchain Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Triage LangChain 1.0 / LangGraph 1.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment. Use during an on-call page, in a post-mortem, or writing the team's first LLM runbook. Trigger with "langchain incident", "llm on-call", "langchain slo", "langchain outage", "langchain cost spike", "langchain agent loop".

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/cost-overrun-response.md`, `references/latency-triage.md` and `references/llm-slos.md`). Compatibility notes: Designed for Claude Code

It sits in DevOps & Cloud, covering Building AI agents, Runbooks and postmortems and Incident response. It works with LangChain and LangGraph. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • With langchain incident
  • Langchain outage
  • Langchain cost spike
  • Langchain agent loop

Example prompts

  • “s first LLM runbook. Trigger with”
  • “llm on-call”
  • “langchain slo”
  • “/langchain-incident-runbook”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Define the LLM-specific SLO set
  2. Triage decision tree: which root path?
  3. Provider outage: detect, circuit-break, fail over
  4. Agent loop containment: stop the bleed before GraphRecursionError
  5. Post-incident: bundle, write up, communicate

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.smith.langchain.com
    • langchain-ai.github.io
    • cloud.google.com
    • sre.google
    • status.anthropic.com
    • status.openai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Langchain Incident Runbook loads about 3.8k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 107 tokens; SKILL.md has 1,609 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,609 words, ~3,760 tokens.

Download SKILL.mdSave it as .claude/skills/langchain-incident-runbook/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
langchain-incident-runbook
description
Triage LangChain 1.0 / LangGraph 1.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment. Use during an on-call page, in a post-mortem, or writing the team's first LLM runbook. Trigger with "langchain incident", "llm on-call", "langchain slo", "langchain outage", "langchain cost spike", "langchain agent loop".
allowed-tools
Read
compatibility
Designed for Claude Code
version
2.7.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, langchain, langgraph, python, langchain-1.0, sre, incident-response, slo, on-call

LangChain Incident Runbook

Overview

3:07am. PagerDuty: "LangChain p95 latency > 10s for 5 minutes." You open LangSmith, filter by service="triage-agent" over the last 15 minutes, and the first trace is 43 seconds long — an agent is on step 24 of 25 iterations, bouncing between the same two tools on a vague user prompt ("help me with my account"). The cost dashboard shows $400 spent in the last 10 minutes, up from a $6/hour baseline. This is P10: create_react_agent defaults to recursion_limit=25 with no cost cap; vague prompts never converge; the spend hits before GraphRecursionError surfaces. First move is not to push a code fix — it is to flip recursion_limit=5 via config reload and add a middleware token-budget cap per session, then deal with the stuck sessions.

Or: same alert, different signature. p95 is healthy at 1.8s, but p99 is 12s and spiky. The spikes correlate with instance starts in Cloud Run. P36: Python + LangChain + embedding preloads = 5–15s cold start; Cloud Run scales to zero by default, so first-request p99 is 10x p95. First move is --min-instances=1 (or a keepalive pinger), not more CPU.

The shape of the page decides the first move. This runbook gives you:

  1. The LLM-specific SLO set most teams do not have: p95 TTFT <1s, p99 total latency <10s, error-rate <0.5%, cost-per-req <$0.05 — with Prometheus burn-rate recording rules that page on user-visible regression.
  2. A triage decision tree with three root paths (latency / cost / error-rate), each with a 3-step diagnostic and first-response action.
  3. Provider outage runbook wired to .with_fallbacks(backup) so failover is a config flip, not a code change.
  4. Agent-loop containment via recursion_limit tuning and middleware token-budget caps so runaway agents stop burning cost before the GraphRecursionError.
  5. Post-incident debug bundle (cross-ref langchain-debug-bundle) and write-up template.

Pinned: langchain-core 1.0.x, langgraph 1.0.x, langsmith 0.3+. Primary pain anchors: P10 (agent runaway), P36 (cold start). Adjacent: P29 (per-process rate limiter), P30 (max_retries=6 means 7 attempts), P31 (Anthropic cache RPM).

Prerequisites

  • LangSmith workspace with tracing enabled (free tier is fine for runbook work)
  • Prometheus + Alertmanager or equivalent (Datadog, Grafana Cloud) — burn-rate rules assume PromQL
  • langchain-observability skill applied — metrics are grounded in what LangSmith callbacks emit
  • A backup model factory from langchain-rate-limits — the failover playbook assumes .with_fallbacks() is already wired
  • On-call rotation + PagerDuty (or equivalent) integrated with the alert pipeline

Instructions

Step 1 — Define the LLM-specific SLO set

HTTP-style SLOs miss what users actually feel. Define four, publish them, wire burn-rate alerts to the symptom.

SLOThresholdAlert conditionFirst-response action
p95 TTFT (time to first streamed token)<1sburn-rate > 2% over 5minCheck streaming is enabled; check provider status; check cold start (P36)
p99 total latency<10sburn-rate > 5% over 5minCheck agent loop depth (P10); cold start (P36); provider latency
Error rate (5xx + uncaught exceptions)<0.5%burn-rate > 1% over 5minCheck provider 429/500; auth token; schema drift on structured output
Cost per request<$0.05 (tier-dependent)p95 spend/req > $0.20 over 15minCheck agent recursion (P10); retry rate (P30); token-use per req

Prometheus recording rule pattern for p99 latency burn-rate (replicate for TTFT, error-rate, and cost):

yaml
groups:
  - name: langchain_slo
    interval: 30s
    rules:
      - record: langchain:p99_latency_5m
        expr: histogram_quantile(0.99, sum(rate(langchain_request_duration_seconds_bucket[5m])) by (le, service))
      - alert: LangChainP99LatencyBurn
        expr: langchain:p99_latency_5m > 10
        for: 5m
        labels: { severity: page, team: llm }
        annotations:
          summary: "LangChain p99 > 10s for {{ $labels.service }}"
          runbook: "https://runbooks/langchain-incident-runbook#latency"

See LLM SLOs for the canonical set (free / paid / enterprise tiers), burn-rate recipes (fast + slow), and a TTFT-specific rule that requires streaming to be instrumented.

Step 2 — Triage decision tree: which root path?

The alert name tells you the root path. Do not mix diagnostics across paths — the first-response action differs.

Alert fired
  ├── Latency (p95/p99 breach, TTFT breach)
  │     ├── 1. Provider status page (Anthropic, OpenAI) green? → if red, Step 3
  │     ├── 2. Cold start pattern? (p99 >> p95, correlates with instance starts) → P36
  │     └── 3. Streaming configured? (TTFT only makes sense with .stream/.astream)
  │
  ├── Cost (spend/req or absolute spend/hour breach)
  │     ├── 1. Agent recursion depth? (LangSmith: max steps per trace) → P10
  │     ├── 2. Retry rate elevated? (callback log: attempt count / logical call) → P30
  │     └── 3. Token-use per req regression? (input + output tokens from callbacks)
  │
  └── Error rate (5xx + uncaught exceptions)
        ├── 1. Provider 429/500 spike? (distinguish client 4xx from provider 5xx)
        ├── 2. Auth? (API key rotation, expired token, org quota exhausted)
        └── 3. Schema drift on structured output? (Pydantic ValidationError in traces)

For each leaf, Latency Triage and Cost Overrun Response give the LangSmith filter query, the exact metric to inspect, and the remediation.

Step 3 — Provider outage: detect, circuit-break, fail over

Detection precedes failover. Do not flip fallbacks on an application bug.

  1. Detect via three signals — all three should agree before declaring a provider outage:
    • Vendor status page watcher (status.anthropic.com, status.openai.com) — poll every 30s, surface into Slack
    • In-app canary probe — a 1-req/min call to each configured provider with a trivial prompt, tracked as a separate SLO
    • Error-rate spike on the primary provider in your own metrics (distinguishes a real outage from your app's bug)
  2. Circuit-break the primary: a CircuitBreaker middleware (see langchain-middleware-patterns if available, or a simple aiobreaker-backed runnable) opens after N consecutive APIError / APITimeoutError within a window. Once open, calls skip the primary and go straight to the backup. This bounds the latency cost of a down provider.
  3. Fail over via .with_fallbacks(backup) — the fallback chain is already wired (see langchain-rate-limits). During an outage, either flip a feature flag that swaps the default factory, or temporarily set the primary's max_retries=0 so the chain reaches the fallback immediately.
  4. Comms: post a user-facing status page entry ("Degraded performance on feature X — monitoring upstream provider") and an internal Slack update with the canary graph attached. Provider Outage Playbook has the full comms template and the circuit-breaker middleware snippet.
Step 4 — Agent loop containment: stop the bleed before GraphRecursionError

P10 is the most common cost-spike cause. create_react_agent defaults to recursion_limit=25, meaning 25 model calls per user turn — with Claude Sonnet at ~$3/MTok input, a 10k-token tool-call loop burns real money per minute.

Three containment layers, applied in order:

  1. Set recursion_limit per agent depth — interactive chat agents rarely need more than 5–8 steps; background research agents can justify 15; never leave the default 25 in production.

    python
    from langgraph.prebuilt import create_react_agent
    
    agent = create_react_agent(
        llm, tools,
        recursion_limit=8,  # P10 — was default 25
    )
  2. Middleware token-budget cap per session — a callback that tracks cumulative input + output tokens for a session id and raises a custom BudgetExceeded exception once the cap is hit. The agent terminates cleanly; the user sees a polite "I could not finish this task in budget, try rephrasing" instead of a spinning UI until GraphRecursionError.

  3. Circuit on repeated tool calls — a LangGraph edge that routes to END when the same tool has been called with the same args twice in a row. This is a cheap heuristic for "the agent is stuck in a loop."

Cost Overrun Response has the middleware implementation, the repeat-tool edge pattern, and a per-tenant budget enforcement example.

Show full SKILL.md (634 more words)Show less
Step 5 — Post-incident: bundle, write up, communicate

Within 30 minutes of all-clear:

  1. Debug bundle — capture the failing LangSmith trace URL(s), the Prometheus dashboard screenshot at the breach window, the agent's config (recursion_limit, model id, max_retries), and the provider's status page state at incident time. Cross-reference langchain-debug-bundle if present.
  2. Write-up template — timeline (detection → triage → mitigation → all-clear), root cause in one sentence with pain-catalog anchor (e.g. "P10: recursion_limit=25 default, vague prompt, no token cap"), permanent fix ticket link, follow-up SLO tuning.
  3. Comms — close the user-facing status page entry, post a short Slack summary (one-paragraph timeline + link to write-up), schedule the post-mortem review in the next weekly SRE sync.

Output

  • LLM SLO set (TTFT, p99 latency, error-rate, cost-per-req) published and alerted via Prometheus burn-rate rules
  • Triage decision tree posted in the runbook with three root paths (latency / cost / error-rate) and first-response actions per leaf
  • Provider outage playbook with circuit breaker + .with_fallbacks(backup) wired to a feature flag for one-flip failover
  • recursion_limit set per agent depth (never default 25 in prod); middleware token-budget cap per session
  • Post-incident debug-bundle template + write-up template wired to the on-call workflow

Error Handling

SymptomLikely causeFirst-response action
p95 latency breach, TTFT degradedStreaming disabled, or provider-side latencyVerify .stream()/.astream() used; check provider status page
p99 >> p95, correlates with instance startsCloud Run cold start (P36)--min-instances=1, CPU-always-allocated billing, preload imports
Cost-per-req spike, agent traces show 20+ stepsrecursion_limit=25 default + vague prompt (P10)recursion_limit=5–8, add middleware token-budget cap
Cost spike, callback log shows 7 attempts per logical callmax_retries=6 inflates cost 7x (P30)max_retries=2 + circuit breaker; log retries via callbacks
429 storm despite requests_per_second=10 on each of N workersInMemoryRateLimiter is per-process (P29)Switch to RedisRateLimiter or provider-side quota
Anthropic 429 while token budget has headroomCache RPM throttled separately (P31)Client-side semaphore on RPM, not token count; monitor cached-read vs uncached separately
Error-rate spike, all on primary providerProvider outageCanary probe confirms; flip failover to .with_fallbacks(backup) via flag
ValidationError surge on structured outputSchema drift — model added fieldsConfigDict(extra="ignore") on the Pydantic schema (see langchain-sdk-patterns)
Agent never terminates, no GraphRecursionError yetStuck in tool-call loopAdd "repeated tool call" edge routing to END; raise BudgetExceeded from middleware

Examples

On-call page: cost spike from agent runaway

PagerDuty: "cost-per-req > $0.20 for 15 minutes." LangSmith filtered to the last 15 minutes shows average trace depth = 22 steps (baseline 4). Single tenant, single conversation pattern — a user who asked an open-ended question the agent cannot resolve. First-response action: flip recursion_limit=5 via config reload (no deploy), add session to the blocklist in middleware, post internal Slack with the trace URL.

See Cost Overrun Response for the middleware token-budget implementation and the per-tenant budget pattern.

On-call page: p99 latency spike during traffic ramp

p95 healthy at 1.8s, p99 at 12s, spikes correlate with Cloud Run instance starts — classic P36. First-response action: gcloud run services update <svc> --min-instances=1, verify heavy imports are at module top level, schedule follow-up ticket to move embedding preload to a warm-up hook.

See Latency Triage for the cold-start detection recipe and the p95-vs-p99 attribution decision tree.

On-call page: provider outage mid-day

Anthropic status page goes red. Canary probe error-rate jumps from 0% to 100% on Anthropic, stays at 0% on OpenAI. Flip the failover flag — the .with_fallbacks(backup=ChatOpenAI(...)) chain (already wired via langchain-rate-limits) takes over. Post user-facing status entry, monitor cost (OpenAI pricing differs — watch cost-per-req SLO), revert when upstream recovers.

See Provider Outage Playbook for the circuit-breaker middleware, the canary probe snippet, and the user-comms template.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/.curated/langchain-incident-runbook of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/cost-overrun-response.md
  • references/latency-triage.md
  • references/llm-slos.md
  • references/one-pager.md
  • references/provider-outage-playbook.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Langchain Incident Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Langchain Incident Runbook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Langchain Incident Runbook this skilljeremylongshore/tons-of-skills-marketplace2.8k—~3.8kAutomated safety check: PassMIT
SRE EngineerJeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT
Incident ResponderDokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI508—~706Automated safety check: PassCustom licence
Eng Runbooksanqiufong/slides-from-anything1321 repos~380Automated safety check: PassApache-2.0
Incident Commanderborghei/Claude-Skills891—~1.8kAutomated safety check: PassMIT
Incident Slo Runbookmajiayu000/spellbook287—~460Automated safety check: PassMIT

Similar skills

  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 7 days ago
    DevOps & CloudAuto-check passed
  • Incident Responder

    Dokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI

    Expert SRE incident responder specializing in rapid problem resolution.

    508 GitHub stars~706 tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed
  • Eng Runbook

    sanqiufong/slides-from-anything

    An engineering runbook — service overview, alerts table, dashboards links, common procedures with copy-pasteable commands, on-call rotation, and an incident-response checklist.

    132 GitHub starsUsed in 1 repo~380 tokens
    DevOps & CloudAuto-check passed
  • Incident Commander

    borghei/Claude-Skills

    Production incident response. An agent skill from borghei/Claude-Skills.

    891 GitHub stars~1.8k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed
  • Incident Slo Runbook

    majiayu000/spellbook

    Create or audit SLOs, SLIs, alert rules, incident response steps, escalation paths, postmortems, operational runbooks, and customer-impact communication.

    287 GitHub stars~460 tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure.

    2k GitHub stars~931 tokensUpdated 23 days ago
    DevOps & CloudAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Langchain Incident Runbook

What does Langchain Incident Runbook do?

Triage LangChain 1.0 / LangGraph 1.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment. Langchain Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment.

When should I use Langchain Incident Runbook?

Langchain Incident Runbook fits situations like: with langchain incident; langchain outage; langchain cost spike; langchain agent loop.

How do I install Langchain Incident Runbook in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a claude-code`. Or copy the skill folder (skills/.curated/langchain-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-incident-runbook in your project. Claude Code loads it when a task matches its description.

How do I install Langchain Incident Runbook in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a codex`. Or copy the skill folder (skills/.curated/langchain-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-incident-runbook in your project. Codex loads it when a task matches its description.

Can I use Langchain Incident Runbook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-incident-runbook, .gemini/skills/langchain-incident-runbook, .github/skills/langchain-incident-runbook and .opencode/skills/langchain-incident-runbook in your project.

What does Langchain Incident Runbook need to run?

SKILL.md names no scripts, command-line tools or credentials: Langchain Incident Runbook is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read. Compatibility (from SKILL.md): Designed for Claude Code.

Does Langchain Incident Runbook access the network?

SKILL.md names 6 domains. As links in the text: docs.smith.langchain.com, langchain-ai.github.io, cloud.google.com, sre.google, status.anthropic.com and status.openai.com. This is read from the text; nothing was executed.

Is Langchain Incident Runbook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Langchain Incident Runbook use?

Langchain Incident Runbook is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Langchain Incident Runbook use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.3k tokens, read only when the agent opens those files.

What are the alternatives to Langchain Incident Runbook?

Skills that share tags, products or a category with Langchain Incident Runbook: SRE Engineer (Jeffallan/claude-skills, 12k stars), Incident Responder (Dokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI, 508 stars), Eng Runbook (sanqiufong/slides-from-anything, 132 stars) and Incident Commander (borghei/Claude-Skills, 891 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Langchain Incident Runbook?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.