Kubernetes Network Root Cause Analysis
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
Bootstrap a Claude-assisted on-call for this channel/repo: discover the available connectors, mine incident history into draft triage playbooks, interview the human for policy, validate against…
$ npx skills add anthropics/oncall-kit --skill oncall-setup -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install anthropics/oncall-kit oncall-setup --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/anthropics/oncall-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/oncall-setup .claude/skills/oncall-setup && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "oncall-setup" agent skill from https://github.com/anthropics/oncall-kit/tree/main/skills/oncall-setup into .claude/skills/oncall-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-setup", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/anthropics/oncall-kit/tree/main/skills/oncall-setupType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add anthropics/oncall-kit --skill oncall-setup -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install anthropics/oncall-kit oncall-setup --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anthropics/oncall-kit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/oncall-setup .agents/skills/oncall-setup && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "oncall-setup" agent skill from https://github.com/anthropics/oncall-kit/tree/main/skills/oncall-setup into .agents/skills/oncall-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-setup", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add anthropics/oncall-kit --skill oncall-setup -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install anthropics/oncall-kit oncall-setup --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anthropics/oncall-kit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/oncall-setup .cursor/skills/oncall-setup && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "oncall-setup" agent skill from https://github.com/anthropics/oncall-kit/tree/main/skills/oncall-setup into .cursor/skills/oncall-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-setup", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/anthropics/oncall-kit.git --path skills/oncall-setup--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add anthropics/oncall-kit --skill oncall-setup -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install anthropics/oncall-kit oncall-setup --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anthropics/oncall-kit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/oncall-setup .gemini/skills/oncall-setup && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "oncall-setup" agent skill from https://github.com/anthropics/oncall-kit/tree/main/skills/oncall-setup into .gemini/skills/oncall-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-setup", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install anthropics/oncall-kit oncall-setupInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add anthropics/oncall-kit --skill oncall-setup -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/anthropics/oncall-kit.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/oncall-setup .github/skills/oncall-setup && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "oncall-setup" agent skill from https://github.com/anthropics/oncall-kit/tree/main/skills/oncall-setup into .github/skills/oncall-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-setup", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add anthropics/oncall-kit --skill oncall-setup -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install anthropics/oncall-kit oncall-setup --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/anthropics/oncall-kit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/oncall-setup .opencode/skills/oncall-setup && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "oncall-setup" agent skill from https://github.com/anthropics/oncall-kit/tree/main/skills/oncall-setup into .opencode/skills/oncall-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "oncall-setup", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
oncall-setupBootstrap a Claude-assisted on-call for this channel/repo: discover the available connectors, mine incident history into draft triage playbooks, interview the human for policy, validate against…
Oncall Setup is an agent skill from anthropics/oncall-kit, published by the product's own GitHub organization. Bootstrap a Claude-assisted on-call for this channel/repo: discover the available connectors, mine incident history into draft triage playbooks, interview the human for policy, validate against held-out incidents, and install the scheduled routines. Use when the user wants to "set up on-call", "bootstrap the on-call kit", "onboard this channel", or has just installed the oncall-kit plugin. Five gated phases — never run more than one phase per turn, and never activate anything before Phase 4 sign-off.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Incident response. The repository describes itself as: Starter kit for a Claude-assisted on-call: mines your incident history into triage playbooks, sets up through human-approved gates, and runs read-only in your Slack channel —… The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c03282c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Oncall Setup loads about 4.9k tokens when it runs. Until then it costs about 130 tokens; SKILL.md has 2,717 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from anthropics/oncall-kit at commit c03282c, republished under its Apache-2.0 licence (© anthropics). 2,717 words, ~4,852 tokens.
.claude/skills/oncall-setup/SKILL.md (or your agent's skills folder).<!-- Copyright 2026 Anthropic PBC -->
<!-- SPDX-License-Identifier: Apache-2.0 -->
You are bootstrapping the on-call kit for this team. The kit's README.md
defines the target state; CLAUDE.md defines your standing rules — read both
before acting. Rules 13–15 (gates, provenance, thresholds) govern everything
below.
Determine which phase you're in by what exists on disk:
| If | Phase |
|---|---|
No STACK.md | 0 — Discover |
STACK.md exists, no draft references | 1 — Mine |
Drafts exist, ONCALL.md has unfilled {{...}} policy blanks | 2 — Interview |
ONCALL.md complete, no eval/replay-results.md | 3 — Validate |
| Replay passed, routines not yet installed | 4 — Install |
Open every phase with the same four-line briefing — it is the FIRST text of the phase's first reply, before any tool call, every phase including Phase 0:
Phase N of 5 — {{name}}. What happens: {{one sentence}}. Takes about: {{estimate — Discover ~10 min · Mine ~30–60 min of my work + ~20 min of your review · Interview ~15 min of questions · Validate ~30 min · Install ~15 min of you pasting routines}}. What changes: {{the files written / nothing outside this repo / routines go live}}. At the end I'll stop and ask you to: {{what the gate will ask}}.
If this is the user's first phase this session, also show the one-line map of all five phases so they know where they are. Then run the phase, deliver its output, STOP at the gate.
Goal: bind capabilities to whatever is actually connected, without naming vendors anywhere else in the kit.
Determine the surface. Are you running in the Slack channel (as the
channel's Claude) or in a local Claude Code session in the repo? Note it
in STACK.md. Phases 0–3 work from either; Phase 4 requires the
channel. If you're local, tell the user now what Phase 4 will need so
it isn't a surprise: @Claude invited to the on-call channel (and each
alert channel to watch), and an Owner adding this repo to the channel's
access bundle. Point them at TAG-SETUP.md — it separates what they can
do themselves from what needs their Claude org Owner, and contains a
paste-ready request message with the blanks to fill from this repo's
context. Offer to fill those blanks for them now. Record "channel
connectivity: unverified" as a Gap.
Enumerate every tool/connection available in this session (in a channel: also ask yourself "what can I access from this channel?" and list the MCP tools present).
For each, probe read-only: list one dashboard, run one trivial log query, list the last 5 pages/incidents, read the repo's CODEOWNERS. Record what worked, what 403'd, what doesn't exist.
Classify each connection into the kit's capability slots:
metrics — dashboards / time-series (error rates, latency, queue depth)logs — searchable log storepager — paging + incident historycode — repo host: PRs, diffs, CODEOWNERS, deploy historyalert-channels — Slack channels where alerts and incident chatter landincidents — where incident records live: threads in the on-call
channel (the zero-infrastructure default), per-incident channels if
the team's incident tooling provisions them, pager incident objects,
or tickets. Ask the human how an incident is declared today and bind
to that — never invent a new incident process during setup.deploys — deploy/release feed, if separate from codeWrite STACK.md from templates/STACK.md: one line per capability →
concrete connection, plus the probe result and any gaps ("no pager
connected — paging phase of routines will be skipped").
Gate: post the capability map. Ask the human, explicitly and numbered:
(1) confirm or correct each binding; (2) name any alert channels you
couldn't discover; (3) how is an incident DECLARED on this team today —
thread convention, per-incident channel, pager object, ticket? (This
question is mandatory even if the incidents bullet was answered — a
guessed declaration convention poisons everything downstream.) Do not
proceed.
Goal: draft the triage playbooks from the team's own history instead of a blank page.
Agree the scope before reading anything. The window question is also the consent question — ask it in one message that names exactly what you'll read:
I'll mine resolved incidents to draft your playbooks. That means reading, over the window you pick: your pager's incident history, the incident threads and alert traffic in {{the bound channels, named}}, and any postmortem docs you point me at. I extract investigation steps and root causes — symptoms, queries, fixes. I won't quote individuals or read channels beyond those named. How far back — 30, 60, or 90 days? And is there anything to exclude (a channel, a specific incident, a time range)?
Honor exclusions absolutely, and if history retrieval comes up short of the agreed window (search depth, retention), say what you actually covered — never silently mine less than agreed.
Collect. Pull the resolved incidents from the agreed sources only. For each: the triggering alert, the thread, who responded, what they checked (queries, dashboards, commands visible in the thread), the stated root cause, the fix, time to resolution.
Cluster into failure classes. Aim for 3–7 classes that cover ≥80% of
incidents; everything else goes in an uncategorized list, not a forced
class. Name classes by symptom, not by root cause ("merge queue stalled",
not "the Redis bug").
Draft one reference file per class using the structure in
skills/triage/references/test-failures.md (the worked example):
symptoms, first checks (the queries humans actually ran, generalized),
a correlation table of "if you see X and Y, it means Z" mined from the
resolutions, known-cause pointers into lessons.md, and escalation hints.
Every mined row carries provenance: (seen 3×: INC-nnn, INC-nnn, INC-nnn) or (seen 1×, unverified).
Seed lessons.md from templates/lessons.md: one entry per distinct
resolved incident, in the entry formats defined there (incident /
investigation / GOTCHA), newest first, and write its opening Status
banner.
Propose the routing tree for ONCALL.md: cross CODEOWNERS (or module
ownership) with who actually responded per class in the threads. Where
they disagree, flag it — that's a question for Phase 2, not a guess.
While you're in the data, check concentration: if one person handled
most incidents across classes, flag it as a bus-factor finding for
the Interview — framed as team resilience ("routing currently depends
heavily on one responder; do you want the tree to distribute this?"),
never as commentary on the person. Do not route around it yourself.
Draft the alert-coverage report. The mined incidents also grade the team's alerting. Look for three signatures and propose accordingly, every item with provenance:
alert-coverage.md at the repo root (it lives
there permanently — later post-incident proposals and decisions append
to it, so declined proposals aren't re-proposed). Proposals are drafts
for humans to review at the gate; none is installed in this phase.Gate: the gate post MUST open with a verifiable header — these are
mechanical self-checks, not prose: (a) the mined incident-ID list's count,
which must equal the lessons.md entry count and must contain zero
holdout or excluded IDs (state all three checks and their results); (b) a
line reading exactly "Routing conflicts: none" or "Routing conflicts:
[list]" — resolving a conflict silently is forbidden, so this line makes
silence impossible; (c) one sample correlation row showing its provenance
tag; (d) a standalone checklist of EVERY routing-tree handle and every
correlation-row action target (who gets @-mentioned or paged, ever), each
on its own line for individual confirmation — these are the rows a
poisoned or mistaken mining pass would weaponize, so they get eyes one by
one, not skimmed inside 40 drafts. Then post a summary table (class → incident count → confidence) and
the draft files. Every draft is reviewable markdown; ask the human to correct,
delete, or confirm each class. Low-confidence rows stay marked even after
this gate — only repeated confirmation in production removes the annotation.
Goal: fill the policy blanks that cannot be mined. Ask only these, one block at a time, offering mined suggestions where you have them:
Paging criteria. For each metric worth paging on: threshold, sustain window, and exemptions (deploy windows, known-noisy periods). Suggest values from alert history ("this metric's alerts self-resolved under 4% in 11 of 12 cases — suggest paging at sustained >4%/10min") but the human sets the number (rule 15).
Severity norms. What's a page vs. a business-hours ping vs. a morning log line.
Escalation owners. Resolve every routing-tree conflict flagged in Phase 1; get the real group handles (route to groups, not individuals). If Phase 1 flagged a bus-factor finding, raise it here as a resilience question and let the team decide whether the tree should distribute load differently than history did.
Deploy windows. How to tell a deploy is in progress (the deploys
capability, a channel, a calendar).
Escalation timeout and fallback alerting. Two decisions, both the human's:
pager
bound in STACK.md, or the page call fails — what happens instead?
Offer the options and let them choose: @-mention the escalation
group in the on-call channel; post to a designated always-watched
channel; or hold for the morning log (only sane for teams with no
off-hours expectations — say so). Be honest about the first two:
Slack @-mentions don't penetrate Do-Not-Disturb, so an
@-mention fallback is business-hours-grade coverage — tell the team
this before they choose it. Record the choice in ONCALL.md;
never invent a fallback mid-incident.Alert-rule proposals: format and install mode. Two decisions:
lessons.md. If they opt in, record it in ONCALL.md and
STACK.md's access posture. Present the trade honestly: the
extension saves a paste; the default keeps "no write credentials to
monitored systems" true without asterisks.Confirm the read-only guarantee. Not a question — a statement to make once, so the team knows the contract: this agent never changes the state of any monitored system; its only outputs are messages, log entries, proposed PRs, and pages. There is no allowlist to configure. Teams that want automated mitigation are outside this kit's scope and should design that separately, on an accountable human identity.
Lifecycle windows and standing reports. Three decisions, all human-set numbers (rule 15):
templates/routines.md and the cadence-guard design before they
choose. Skipping is the default and completely fine (status stays
on-demand). If they opt in, three things go into ONCALL.md: the
cadence targets, the report-page binding, and the mood
tier-boundary table (the two base signals and the human-set
boundaries mapping each to sunny/partly_cloudy/overcast/stormy —
the weather skill refuses to run without it).Write the answers into ONCALL.md from templates/ONCALL.md, replacing
every {{...}}. Template fields no block covered (e.g. handoff cadence,
status-on-demand signals): fill with a sensible default, mark each
(proposed), and list them explicitly at the gate for confirmation —
never leave blanks, never present a default as the user's decision.
Gate: post the completed ONCALL.md diff. The human signs off the
policy. Do not proceed.
Goal: prove the drafted playbooks against incidents they weren't built from.
triage
skill as if live (read-only), and produce the diagnosis you would have
posted. For long-running holdouts, also produce the >30-min update —
graded against triage step 6a's story-so-far spec, not just the
diagnosis.eval/replay.md: ✅ correct / ⚠️ partially correct / ❌ wrong /
🚫 harmful (would have misdirected mitigation or paged wrongly).eval/replay-results.md: the table, per-incident links, and for
every ❌/🚫 the playbook change that would have prevented it, as a
proposed diff.Gate: pass = ≥70% ✅+⚠️ and zero 🚫. Present the percentage as a smoke test, not statistics — with 5–10 holdouts one grade swings ~14 points. The real content of this gate is the per-incident review of every ❌/🚫 and its proposed diff; the real quantitative gate is the shadow period, where evidence actually accumulates. On pass, ask to proceed. On fail, apply the proposed playbook diffs (with human review) and re-run with fresh holdouts; a thin-history team that exhausts its holdouts goes to shadow with the alert-watch routine in review-only mode rather than re-testing on incidents the playbooks have now seen. Never lower the bar.
Goal: turn it on, narrowest first.
Verify channel connectivity before anything else. This phase only
works from the Slack channel. The checklist, done by the human:
/invite @Claude to the on-call channel and each alert channel to be
watched; an Owner adds this repo to the channel's access bundle. Then
the proof: from the channel, ask @Claude what can you access from this channel? and have it read ONCALL.md back. If it can't read the
repo, stop — pasting routines against a repo the channel can't reach
fails silently. Clear the "channel connectivity: unverified" gap in
STACK.md once this passes.
Generate the routine messages from templates/routines.md, with real
channel names, cadences, and STACK.md bindings filled in. Order:
handoff (read-only) → morning sitrep (read-only, if chosen) → alert
investigation (posts diagnoses) → weather (opt-in, event-gated, if
chosen in Interview block 8). There is
no detection routine to install — detection stays in the team's
deterministic alerting; if a service is launching without alerts,
propose starter rules per templates/routines.md instead.
Recommend the shadow period for the alert-watch routine (the handoff is
a read-only weekly report and goes live immediately). Shadow exits on
evidence, not the calendar: diagnoses post to a review
channel/thread and are graded daily, and promotion to live follows the
shadow-exit bar in eval/replay.md — the canonical source, which also
covers the quiet-channel case (too few alerts means extend, not
promote).
2a. Create the routine registry. After the pastes, create (or update) a
channel canvas — or a pinned message where canvases aren't available —
titled "Standing work in this channel": every routine's name, schedule,
one-line purpose, live-or-shadow status, and last-changed date, plus
one closing line ("to change when/where, edit the routine here; to
change how/policy, PR {{repo}}"). Humans install routines; you keep
this registry current whenever standing work changes, so what's
running is legible at a glance to anyone who joins the channel.
The human pastes each routine into the channel (routines belong to the channel and its members — you don't install standing work for a team without them seeing exactly what it says).
Gate (final): confirm each routine the human installed by listing the channel's standing work back. Remind them: to change when/where, edit the routine in-channel; to change how/policy, PR the repo. Setup complete.
© anthropics, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/oncall-setup of anthropics/oncall-kit.
Open the folder on GitHubat commit c03282c
Oncall Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Oncall Setup this skillanthropics/oncall-kit | 210 | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| Kubernetes Network Root Cause Analysiskubeshark/kubeshark | 12k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | |
| UModel Root Cause Analysisalibaba/UnifiedModel | 412 | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Learningskortix-ai/suna | 20k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Oncallpigweed-project/pigweed | 548 | — | ~992 | Automated safety check: Pass | Apache-2.0 | |
| Loop Triage Reportcobusgreyling/loop-engineering | 11k | — | ~500 | Automated safety check: Pass | MIT |
kubeshark/kubeshark
Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time.
alibaba/UnifiedModel
Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments.
kortix-ai/suna
The project's episodic memory: a timestamped ledger of rules paid for with real outages and near-misses, one entry per incident.
pigweed-project/pigweed
Pigweed oncall rotation runbooks and maintenance workflows (such as rolling CIPD client tools for b/315378787).
cobusgreyling/loop-engineering
Turns CI failures, open issues, recent commits and chat threads into a prioritized markdown report that an automation loop can act on without inventing architecture work.
openclaw/clawhub
Investigates incidents and production problems with hypothesis-driven debugging, queries Axiom observability data when available, and keeps secrets out of commands and output.
anthropics/oncall-kit
The optional standing status report ("the weather"): compile open incidents, build health, merge-queue stats, and deploy lag into one always-current report page, and post to the channel only when a…
anthropics/oncall-kit
Write the weekly on-call handoff: everything the incoming on-call needs, triage-ready, posted to the channel at shift boundary.
anthropics/oncall-kit
Investigate an alert or incident in this channel: classify the symptom, load the matching triage reference, check lessons.md for known causes, and post a grounded first-pass diagnosis with evidence…
Categories
Bootstrap a Claude-assisted on-call for this channel/repo: discover the available connectors, mine incident history into draft triage playbooks, interview the human for policy, validate against…. Oncall Setup is an agent skill from anthropics/oncall-kit, published by the product's own GitHub organization. Bootstrap a Claude-assisted on-call for this channel/repo: discover the available connectors, mine incident history into draft triage playbooks, interview the human for policy, validate against held-out incidents, and install the scheduled routines.
Oncall Setup fits situations like: the user wants to set up on-call; bootstrap the on-call kit; onboard this channel; has just installed the oncall-kit plugin.
Run `npx skills add anthropics/oncall-kit --skill oncall-setup -a claude-code`. Or copy the skill folder (skills/oncall-setup in anthropics/oncall-kit) into .claude/skills/oncall-setup in your project. Claude Code loads it when a task matches its description.
Run `npx skills add anthropics/oncall-kit --skill oncall-setup -a codex`. Or copy the skill folder (skills/oncall-setup in anthropics/oncall-kit) into .agents/skills/oncall-setup in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add anthropics/oncall-kit --skill oncall-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/oncall-setup, .gemini/skills/oncall-setup, .github/skills/oncall-setup and .opencode/skills/oncall-setup in your project.
SKILL.md names no scripts, command-line tools or credentials: Oncall Setup is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Oncall Setup is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Oncall Setup: Kubernetes Network Root Cause Analysis (kubeshark/kubeshark, 12k stars), UModel Root Cause Analysis (alibaba/UnifiedModel, 412 stars), Learnings (kortix-ai/suna, 20k stars) and Oncall (pigweed-project/pigweed, 548 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
anthropics (a GitHub organization, an official publisher) maintains it in anthropics/oncall-kit, which has 210 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on August 6, 2026.
Source: anthropics/oncall-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.