SRE Engineer
Jeffallan/claude-skills
Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.
Agent skill
by jeremylongshore in jeremylongshore/tons-of-skills-marketplace
Triage LangChain 1.0 / LangGraph 1.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-incident-runbook --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/langchain-incident-runbook .claude/skills/langchain-incident-runbook && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "langchain-incident-runbook" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-incident-runbook into .claude/skills/langchain-incident-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-incident-runbook", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-incident-runbookType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-incident-runbook --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.curated/langchain-incident-runbook .agents/skills/langchain-incident-runbook && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "langchain-incident-runbook" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-incident-runbook into .agents/skills/langchain-incident-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-incident-runbook", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-incident-runbook --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.curated/langchain-incident-runbook .cursor/skills/langchain-incident-runbook && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "langchain-incident-runbook" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-incident-runbook into .cursor/skills/langchain-incident-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-incident-runbook", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jeremylongshore/tons-of-skills-marketplace.git --path skills/.curated/langchain-incident-runbook--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-incident-runbook --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.curated/langchain-incident-runbook .gemini/skills/langchain-incident-runbook && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "langchain-incident-runbook" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-incident-runbook into .gemini/skills/langchain-incident-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-incident-runbook", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-incident-runbookInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.curated/langchain-incident-runbook .github/skills/langchain-incident-runbook && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "langchain-incident-runbook" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-incident-runbook into .github/skills/langchain-incident-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-incident-runbook", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace langchain-incident-runbook --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.curated/langchain-incident-runbook .opencode/skills/langchain-incident-runbook && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "langchain-incident-runbook" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/langchain-incident-runbook into .opencode/skills/langchain-incident-runbook/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "langchain-incident-runbook", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
langchain-incident-runbookTriage LangChain 1.0 / LangGraph 1.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment.
Langchain Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace. Triage LangChain 1.0 / LangGraph 1.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment. Use during an on-call page, in a post-mortem, or writing the team's first LLM runbook. Trigger with "langchain incident", "llm on-call", "langchain slo", "langchain outage", "langchain cost spike", "langchain agent loop".
Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/cost-overrun-response.md`, `references/latency-triage.md` and `references/llm-slos.md`). Compatibility notes: Designed for Claude Code
It sits in DevOps & Cloud, covering Building AI agents, Runbooks and postmortems and Incident response. It works with LangChain and LangGraph. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.smith.langchain.comlangchain-ai.github.iocloud.google.comsre.googlestatus.anthropic.comstatus.openai.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Langchain Incident Runbook loads about 3.8k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 107 tokens; SKILL.md has 1,609 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,609 words, ~3,760 tokens.
.claude/skills/langchain-incident-runbook/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.3:07am. PagerDuty: "LangChain p95 latency > 10s for 5 minutes." You open LangSmith,
filter by service="triage-agent" over the last 15 minutes, and the first trace
is 43 seconds long — an agent is on step 24 of 25 iterations, bouncing between
the same two tools on a vague user prompt ("help me with my account"). The cost
dashboard shows $400 spent in the last 10 minutes, up from a $6/hour baseline.
This is P10: create_react_agent defaults to recursion_limit=25 with no cost
cap; vague prompts never converge; the spend hits before GraphRecursionError
surfaces. First move is not to push a code fix — it is to flip
recursion_limit=5 via config reload and add a middleware token-budget cap per
session, then deal with the stuck sessions.
Or: same alert, different signature. p95 is healthy at 1.8s, but p99 is 12s and
spiky. The spikes correlate with instance starts in Cloud Run. P36: Python +
LangChain + embedding preloads = 5–15s cold start; Cloud Run scales to zero by
default, so first-request p99 is 10x p95. First move is --min-instances=1
(or a keepalive pinger), not more CPU.
The shape of the page decides the first move. This runbook gives you:
.with_fallbacks(backup) so failover is a
config flip, not a code change.recursion_limit tuning and middleware
token-budget caps so runaway agents stop burning cost before the
GraphRecursionError.langchain-debug-bundle) and write-up
template.Pinned: langchain-core 1.0.x, langgraph 1.0.x, langsmith 0.3+. Primary pain
anchors: P10 (agent runaway), P36 (cold start). Adjacent: P29 (per-process
rate limiter), P30 (max_retries=6 means 7 attempts), P31 (Anthropic cache RPM).
langchain-observability skill applied — metrics are grounded in what LangSmith callbacks emitbackup model factory from langchain-rate-limits — the failover playbook assumes .with_fallbacks() is already wiredHTTP-style SLOs miss what users actually feel. Define four, publish them, wire burn-rate alerts to the symptom.
| SLO | Threshold | Alert condition | First-response action |
|---|---|---|---|
| p95 TTFT (time to first streamed token) | <1s | burn-rate > 2% over 5min | Check streaming is enabled; check provider status; check cold start (P36) |
| p99 total latency | <10s | burn-rate > 5% over 5min | Check agent loop depth (P10); cold start (P36); provider latency |
| Error rate (5xx + uncaught exceptions) | <0.5% | burn-rate > 1% over 5min | Check provider 429/500; auth token; schema drift on structured output |
| Cost per request | <$0.05 (tier-dependent) | p95 spend/req > $0.20 over 15min | Check agent recursion (P10); retry rate (P30); token-use per req |
Prometheus recording rule pattern for p99 latency burn-rate (replicate for TTFT, error-rate, and cost):
groups:
- name: langchain_slo
interval: 30s
rules:
- record: langchain:p99_latency_5m
expr: histogram_quantile(0.99, sum(rate(langchain_request_duration_seconds_bucket[5m])) by (le, service))
- alert: LangChainP99LatencyBurn
expr: langchain:p99_latency_5m > 10
for: 5m
labels: { severity: page, team: llm }
annotations:
summary: "LangChain p99 > 10s for {{ $labels.service }}"
runbook: "https://runbooks/langchain-incident-runbook#latency"See LLM SLOs for the canonical set (free / paid / enterprise tiers), burn-rate recipes (fast + slow), and a TTFT-specific rule that requires streaming to be instrumented.
The alert name tells you the root path. Do not mix diagnostics across paths — the first-response action differs.
Alert fired
├── Latency (p95/p99 breach, TTFT breach)
│ ├── 1. Provider status page (Anthropic, OpenAI) green? → if red, Step 3
│ ├── 2. Cold start pattern? (p99 >> p95, correlates with instance starts) → P36
│ └── 3. Streaming configured? (TTFT only makes sense with .stream/.astream)
│
├── Cost (spend/req or absolute spend/hour breach)
│ ├── 1. Agent recursion depth? (LangSmith: max steps per trace) → P10
│ ├── 2. Retry rate elevated? (callback log: attempt count / logical call) → P30
│ └── 3. Token-use per req regression? (input + output tokens from callbacks)
│
└── Error rate (5xx + uncaught exceptions)
├── 1. Provider 429/500 spike? (distinguish client 4xx from provider 5xx)
├── 2. Auth? (API key rotation, expired token, org quota exhausted)
└── 3. Schema drift on structured output? (Pydantic ValidationError in traces)For each leaf, Latency Triage and Cost Overrun Response give the LangSmith filter query, the exact metric to inspect, and the remediation.
Detection precedes failover. Do not flip fallbacks on an application bug.
status.anthropic.com, status.openai.com) —
poll every 30s, surface into SlackCircuitBreaker middleware (see
langchain-middleware-patterns if available, or a simple
aiobreaker-backed runnable) opens after N consecutive APIError /
APITimeoutError within a window. Once open, calls skip the primary and go
straight to the backup. This bounds the latency cost of a down provider..with_fallbacks(backup) — the fallback chain is already
wired (see langchain-rate-limits). During an outage, either flip a feature
flag that swaps the default factory, or temporarily set the primary's
max_retries=0 so the chain reaches the fallback immediately.P10 is the most common cost-spike cause. create_react_agent defaults to
recursion_limit=25, meaning 25 model calls per user turn — with Claude Sonnet
at ~$3/MTok input, a 10k-token tool-call loop burns real money per minute.
Three containment layers, applied in order:
Set recursion_limit per agent depth — interactive chat agents rarely
need more than 5–8 steps; background research agents can justify 15; never
leave the default 25 in production.
from langgraph.prebuilt import create_react_agent
agent = create_react_agent(
llm, tools,
recursion_limit=8, # P10 — was default 25
)Middleware token-budget cap per session — a callback that tracks
cumulative input + output tokens for a session id and raises a custom
BudgetExceeded exception once the cap is hit. The agent terminates
cleanly; the user sees a polite "I could not finish this task in budget,
try rephrasing" instead of a spinning UI until GraphRecursionError.
Circuit on repeated tool calls — a LangGraph edge that routes to END
when the same tool has been called with the same args twice in a row. This
is a cheap heuristic for "the agent is stuck in a loop."
Cost Overrun Response has the middleware implementation, the repeat-tool edge pattern, and a per-tenant budget enforcement example.
Within 30 minutes of all-clear:
langchain-debug-bundle if present..with_fallbacks(backup)
wired to a feature flag for one-flip failoverrecursion_limit set per agent depth (never default 25 in prod); middleware
token-budget cap per session| Symptom | Likely cause | First-response action |
|---|---|---|
| p95 latency breach, TTFT degraded | Streaming disabled, or provider-side latency | Verify .stream()/.astream() used; check provider status page |
| p99 >> p95, correlates with instance starts | Cloud Run cold start (P36) | --min-instances=1, CPU-always-allocated billing, preload imports |
| Cost-per-req spike, agent traces show 20+ steps | recursion_limit=25 default + vague prompt (P10) | recursion_limit=5–8, add middleware token-budget cap |
| Cost spike, callback log shows 7 attempts per logical call | max_retries=6 inflates cost 7x (P30) | max_retries=2 + circuit breaker; log retries via callbacks |
429 storm despite requests_per_second=10 on each of N workers | InMemoryRateLimiter is per-process (P29) | Switch to RedisRateLimiter or provider-side quota |
| Anthropic 429 while token budget has headroom | Cache RPM throttled separately (P31) | Client-side semaphore on RPM, not token count; monitor cached-read vs uncached separately |
| Error-rate spike, all on primary provider | Provider outage | Canary probe confirms; flip failover to .with_fallbacks(backup) via flag |
ValidationError surge on structured output | Schema drift — model added fields | ConfigDict(extra="ignore") on the Pydantic schema (see langchain-sdk-patterns) |
Agent never terminates, no GraphRecursionError yet | Stuck in tool-call loop | Add "repeated tool call" edge routing to END; raise BudgetExceeded from middleware |
PagerDuty: "cost-per-req > $0.20 for 15 minutes." LangSmith filtered to the
last 15 minutes shows average trace depth = 22 steps (baseline 4). Single
tenant, single conversation pattern — a user who asked an open-ended question
the agent cannot resolve. First-response action: flip recursion_limit=5 via
config reload (no deploy), add session to the blocklist in middleware, post
internal Slack with the trace URL.
See Cost Overrun Response for the middleware token-budget implementation and the per-tenant budget pattern.
p95 healthy at 1.8s, p99 at 12s, spikes correlate with Cloud Run instance
starts — classic P36. First-response action: gcloud run services update <svc> --min-instances=1, verify heavy imports are at module top level,
schedule follow-up ticket to move embedding preload to a warm-up hook.
See Latency Triage for the cold-start detection recipe and the p95-vs-p99 attribution decision tree.
Anthropic status page goes red. Canary probe error-rate jumps from 0% to 100%
on Anthropic, stays at 0% on OpenAI. Flip the failover flag — the
.with_fallbacks(backup=ChatOpenAI(...)) chain (already wired via
langchain-rate-limits) takes over. Post user-facing status entry, monitor
cost (OpenAI pricing differs — watch cost-per-req SLO), revert when upstream
recovers.
See Provider Outage Playbook for the circuit-breaker middleware, the canary probe snippet, and the user-comms template.
create_react_agent and recursion_limitdocs/pain-catalog.md (primary: P10, P36; adjacent: P29, P30, P31)langchain-debug-bundle (post-incident capture), langchain-observability (SLO metrics source), langchain-rate-limits (.with_fallbacks chain), langchain-cost-tuning (token budget caps), langchain-deploy-integration (cold-start fix)© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/.curated/langchain-incident-runbook of jeremylongshore/tons-of-skills-marketplace.
Open the folder on GitHubat commit cfae287
Langchain Incident Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Langchain Incident Runbook this skilljeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~3.8k | Automated safety check: Pass | MIT | |
| SRE EngineerJeffallan/claude-skills | 12k | — | ~1.7k | Automated safety check: Pass | MIT | |
| Incident ResponderDokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI | 508 | — | ~706 | Automated safety check: Pass | Custom licence | |
| Eng Runbooksanqiufong/slides-from-anything | 132 | 1 repos | ~380 | Automated safety check: Pass | Apache-2.0 | |
| Incident Commanderborghei/Claude-Skills | 891 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Incident Slo Runbookmajiayu000/spellbook | 287 | — | ~460 | Automated safety check: Pass | MIT |
Jeffallan/claude-skills
Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.
Dokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI
Expert SRE incident responder specializing in rapid problem resolution.
sanqiufong/slides-from-anything
An engineering runbook — service overview, alerts table, dashboards links, common procedures with copy-pasteable commands, on-call rotation, and an incident-response checklist.
borghei/Claude-Skills
Production incident response. An agent skill from borghei/Claude-Skills.
majiayu000/spellbook
Create or audit SLOs, SLIs, alert rules, incident response steps, escalation paths, postmortems, operational runbooks, and customer-impact communication.
cbrock84/headcount
Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
jeremylongshore/tons-of-skills-marketplace
Execute proactive auto-loading: automatically detects and loads agents.md files.
jeremylongshore/tons-of-skills-marketplace
Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.
jeremylongshore/tons-of-skills-marketplace
Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.
Categories
Triage LangChain 1.0 / LangGraph 1.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment. Langchain Incident Runbook is an agent skill from jeremylongshore/tons-of-skills-marketplace.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment.
Langchain Incident Runbook fits situations like: with langchain incident; langchain outage; langchain cost spike; langchain agent loop.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a claude-code`. Or copy the skill folder (skills/.curated/langchain-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/langchain-incident-runbook in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a codex`. Or copy the skill folder (skills/.curated/langchain-incident-runbook in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/langchain-incident-runbook in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill langchain-incident-runbook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/langchain-incident-runbook, .gemini/skills/langchain-incident-runbook, .github/skills/langchain-incident-runbook and .opencode/skills/langchain-incident-runbook in your project.
SKILL.md names no scripts, command-line tools or credentials: Langchain Incident Runbook is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read. Compatibility (from SKILL.md): Designed for Claude Code.
SKILL.md names 6 domains. As links in the text: docs.smith.langchain.com, langchain-ai.github.io, cloud.google.com, sre.google, status.anthropic.com and status.openai.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Langchain Incident Runbook is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Langchain Incident Runbook: SRE Engineer (Jeffallan/claude-skills, 12k stars), Incident Responder (Dokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI, 508 stars), Eng Runbook (sanqiufong/slides-from-anything, 132 stars) and Incident Commander (borghei/Claude-Skills, 891 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.
Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.