Agent skill

CI Debug

by yonatangross in yonatangross/orchestkit

Diagnose a failing CI run against an 11-pattern playbook. An agent skill from yonatangross/orchestkit.

MITAuto-check: notesDevOps & Cloud

Install CI Debug

skills CLI
$ npx skills add yonatangross/orchestkit --skill ci-debug -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yonatangross/orchestkit ci-debug --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/ci-debug .claude/skills/ci-debug && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ci-debug
GitHub stars
288
Token cost
~2.7k tokens
SKILL.md length
1,006 words
Files
1
Skills in repo
107
Repo updated
First seen
Licence
MIT

At a glance

Diagnose a failing CI run against an 11-pattern playbook. An agent skill from yonatangross/orchestkit.

  • Works in 4 steps: Resolve the failing job → Fetch the failing log → Classify against the playbook → …
  • A specific PR check
  • SKILL.md covers Input, Execution, Headless invocation… and CRITICAL guardrails, plus 4 more sections
  • Calls gh, git and python; reaches github.com

What it does

CI Debug is an agent skill from yonatangross/orchestkit. Diagnose a failing CI run against an 11-pattern playbook. Classifies the failure, cites the relevant memory entry, proposes the exact fix command — but NEVER applies without explicit user approval. Use when a specific PR check or GitHub Actions run failed and you want a diagnosis instead of speculation. Don't use for org-wide CI sweeps (that's /status) or for app-level test failures (the playbook is CI-infra-specific).

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Claude Code 2.1.277+ — uses gh CLI for GitHub Actions log inspection. Works in interactive sessions and headless claude -p --bare invocations (e.g…

It sits in DevOps & Cloud, covering Failing and flaky tests and CI/CD. It works with GitHub Actions and pnpm. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.

When your agent uses it

  • A specific PR check
  • GitHub Actions run failed and you want a diagnosis instead of speculation
  • Org-wide CI sweeps (thats /status)
  • For app-level test failures (the playbook is CI-infra-specific)

Example prompts

  • “t use for org-wide CI sweeps (that”
  • “/ci-debug”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Claude Code 2.1.277+ — uses `gh` CLI for GitHub Actions log inspection. Works in interactive sessions and headless `claude -p --bare` invocations (e.g. /ork:ci-sentinel).
  • Pre-approved tools (allowed-tools): Bash, Read, Grep, Glob

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Resolve the failing job
  2. Fetch the failing log
  3. Classify against the playbook
  4. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 1f8d8f3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • git
    • python
    • claude
    • pnpm
    • uv
    • node
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code 2.1.277+ — uses `gh` CLI for GitHub Actions log inspection. Works in interactive sessions and headless `claude -p --bare` invocations (e.g. /ork:ci-sentinel).

    From compatibility in the SKILL.md frontmatter.

Context cost

CI Debug loads about 2.7k tokens when it runs. Until then it costs about 108 tokens; SKILL.md has 1,006 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Grep, Glob

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from yonatangross/orchestkit at commit 1f8d8f3, republished under its MIT licence (© yonatangross). 1,006 words, ~2,718 tokens.

Download SKILL.mdSave it as .claude/skills/ci-debug/SKILL.md (or your agent's skills folder).
name
ci-debug
description
Diagnose a failing CI run against an 11-pattern playbook. Classifies the failure, cites the relevant memory entry, proposes the exact fix command — but NEVER applies without explicit user approval. Use when a specific PR check or GitHub Actions run failed and you want a diagnosis instead of speculation. Don't use for org-wide CI sweeps (that's /status) or for app-level test failures (the playbook is CI-infra-specific).
allowed-tools
Bash, Read, Grep, Glob
compatibility
Claude Code 2.1.277+ — uses `gh` CLI for GitHub Actions log inspection. Works in interactive sessions and headless `claude -p --bare` invocations (e.g. /ork:ci-sentinel).
license
MIT
argument-hint
<PR-number | run-URL | job-URL>
context
fork
background
false
disable-model-invocation
false
user-invocable
true
skills
github-operations, memory
model
sonnet
metadata.category
workflow-automation
metadata.origin
/insights audit 2026-05-11 — recurring CI-debug pattern across 12 sessions in 3 weeks

/ci-debug — classify a failing CI run

Direct response to the recurring CI-debug pattern surfaced by /insights: ~12 sessions in 3 weeks doing the same classification dance. This skill encodes the 11 patterns so the dance becomes a lookup.

Input

User invokes with one of:

  • PR number: /ci-debug 822 (default repo from context; ask if ambiguous)
  • Run URL: /ci-debug https://github.com/owner/repo/actions/runs/12345
  • Job URL: /ci-debug https://github.com/owner/repo/actions/runs/X/job/Y

Execution

1. Resolve the failing job
bash
# From PR number:
gh pr checks <n> --repo <owner>/<repo> --json bucket,link,name \
  --jq '.[] | select(.bucket=="fail") | "\(.name)|\(.link)"'

# From run URL:
gh api repos/<owner>/<repo>/actions/runs/<run-id>/jobs \
  --jq '.jobs[] | select(.conclusion=="failure")
                | {id, name, runner_name, started_at, completed_at,
                   steps: [.steps[] | select(.conclusion=="failure") | {name, number}]}'

If multiple jobs failed, pick the one with the shortest duration — root cause is usually the first failure; later jobs cascade.

No job in the fail bucket but a check won't settle? If gh pr checks shows zero fail-bucket entries yet a status sits in pending that never resolves (and gh pr view --json mergeStateStatus returns UNSTABLE while mergeable=MERGEABLE), this is a stuck external status, not a failure — jump straight to Pattern #11. There is no failing log to fetch; classify on the commit-status metadata (gh api repos/<o>/<r>/commits/<sha>/status).

2. Fetch the failing log
bash
gh api repos/<owner>/<repo>/actions/jobs/<job_id>/logs 2>&1 \
  | grep -iE '(error|fail|ERR_|CONFLICT|Process completed with exit code)' \
  | head -30

Capture the FIRST distinct error message (later lines often echo).

3. Classify against the playbook

Walk the patterns in order. First match wins.

#PatternSignature in logsMemory refProposed fix
1Billing blockrunner_name empty + steps[] empty + ~3s duration + annotation: "recent account payments have failed or your spending limit needs to be increased"billing-surface-hosted-vs-self-hosted.mdOrg admin → Settings → Billing & plans → raise limit / update card. No code change.
2Root-lockfile driftERR_PNPM_OUTDATED_LOCKFILE mentioning <ROOT>/typescript/<pkg>/package.jsonpnpm-lock-root-vs-workspace-duality.mdpnpm install --lockfile-only && git add pnpm-lock.yaml && git commit && git push.
3uv.lock drifterror: The lockfile at uv.lock needs to be updatedchangeset-release-uv-lock-drift.mdcd python && uv lock then commit.
4ci-shared.yml missing permissionsstartup_failure pattern (empty runner_name + steps[]=[] + ~3s) BUT billing is resolvedci-shared-permissions-block-required.mdAdd permissions: { contents: read, packages: read } to the caller workflow.
5YAML python embedYAML parse error pointing at a multi-line block scalar with python -cyaml-python-embed.mdRewrite python -c as a separate shell script invocation; never inline multi-line python in YAML.
6actionlint shellcheck false-positiveaudit/actionlint job failing with SC2086/SC2046 on workflow YAMLs you didn't touchaudit-actionlint-triggers-on-workflow-edit.mdNot required check; safe to merge past if the warnings predate your change. Optional: add shellcheck disable comments.
7macOS BSD date %3N%3N printed literally in CI output / arithmetic failsmacos-bsd-date-no-percent-3N.mdReplace date +%s%3N with node -e 'console.log(Date.now())' or python3 -c 'import time; print(int(time.time()*1000))'.
8Runner pnpm Rosetta arch driftpnpm install fails with "wrong-arch native bin" / dlopen error on a self-hosted runnerrunner-pnpm-rosetta-arch-drift.mdRestart the affected runner pool; root cause is node x64↔arm64 flips storing wrong-arch native bins in shared cache.
9Shallow clone false divergencegit status reports diverged but PR was actually mergedshallow-clone-false-divergence.mdgit fetch origin <branch> --unshallow then gh pr view --merge-commit to verify.
10Publish run cancelledPublish-tag workflow run shows conclusion=cancelled; artifact never landspublish-runs-cancelled-need-redrive.mdRe-fire via gh workflow run publish-python.yml -f tag=<tag> (adjust for your publish workflow).
11Vercel status orphaned (path-skip)No job in the fail bucket, but Vercel appears as a commit status (not a check-run) stuck state=pending with created_at == updated_at and no terminal update; all GitHub Actions checks green; mergeStateStatus=UNSTABLE + mergeable=MERGEABLE on an unprotected base branchvercel-pending-orphaned-on-path-skip.mdNot a failure — cosmetic. Vercel posted a pending status then skipped the build (project-root path filter, e.g. a docs-only change that never touches apps/web), orphaning the status. Safe to merge: gh pr merge <n> --repo <owner>/<repo> --squash. Permanent fix: the Vercel project's Ignored Build Step must exit 0 AND report success for skipped paths so the status flips instead of dangling.

The memory references point at user-curated memory files (~/.claude/projects/<project>/memory/*.md). If your memory doesn't have them yet, the signature column is enough to classify — the memory citation is a nice-to-have, not required.

Show full SKILL.md (397 more words)Show less
4. Report

For a matched pattern:

markdown
## CI Debug: <repo> · <pr-or-run-ref>

**Failing job:** `<job name>` (<duration>s) on runner `<runner_name>`
**Failing step:** <step name> (#<step number>)
**Error excerpt:**
\`\`\`
<first 3 lines of grep'd error>
\`\`\`

**Classification:** Pattern #<n> — <pattern name>
**Reference:** memory `<memory-file.md>`

**Proposed fix:**
<exact commands, one per line>

**Will I apply this?** No — awaiting your approval. Reply "go" to ship.

For an UNMATCHED failure:

markdown
## CI Debug: <repo> · <pr-or-run-ref> · NOVEL

**Failing step:** <step name>
**Unique log lines:**
\`\`\`
<top 10 distinct error lines>
\`\`\`

This doesn't match any of the 11 playbook patterns. Surfacing the raw
evidence for your read. Once you identify the root cause, consider
adding it to the playbook (in this SKILL.md) so the next run catches
it automatically.

Headless invocation (permission mode)

/ci-debug REQUIRES Bash tool access — gh pr checks, gh run view, and gh api .../jobs/.../logs are how it fetches evidence. In headless claude -p runs (ci-sentinel, cron):

  • Use --permission-mode acceptEdits — the headless "use tools without prompting" mode. The skill is read-only by design, so this grants nothing risky.
  • NEVER dontAsk — it silently REFUSES permission-requiring tools (including Bash), so every analysis returns empty output with no error (#1862 Bug C).
  • claude -p reports auth/permission failures as JSON on stdout, not stderr — capture and log both streams on failure (see shared/rules/cc-bare-auth-gotcha.md for the auth side).

CRITICAL guardrails

  • NEVER auto-apply a fix. This skill proposes; the user approves. Wasted cycles from misclassification are exactly the failure mode we're hardening against.
  • NEVER skip the fetch-log step. Without the actual error text, classification is guessing. If logs are gone (>90 days old, run deleted), say so and stop.
  • Cite the exact memory entry (filename + section heading where relevant) so the user can verify the analogy.
  • Each classification must have falsifying evidence — the signature must match. Don't pattern-match on hope.

When to invoke

  • A specific CI run / job / PR check has failed and the user wants the diagnosis.
  • The user pastes a run URL or "PR #N is failing — why?".

When NOT to invoke

  • For broad "what's failing across the org?" — use /status instead.
  • For application-level test failures (the playbook is CI-infra-specific, not e.g. pytest failures).
  • Before reading the actual log — don't speculate.

Adding patterns

When a novel CI failure surfaces (the UNMATCHED report fires), add a row to the table above:

  1. Signature: the unique-enough log line / metadata combination. Must be falsifiable.
  2. Memory ref: name the memory file you wrote with the full incident write-up.
  3. Proposed fix: the EXACT command line. No prose, no "you might want to". The skill produces commands the user runs; ambiguity defeats the purpose.

Commit the SKILL.md change as docs(ci-debug): add pattern #N (<short-name>). The ci-sentinel workflow picks up new patterns automatically on its next sweep — no separate plumbing needed.

  • Composes with — ci-sentinel (the autonomous hourly trigger for this skill against open red PRs).
  • Anti-pattern — manually re-running gh pr checks and eyeballing the logs across N repos. That's exactly the toil this skill kills.
  • Upstream — /status for org-wide sweeps that surface WHICH PRs are red; this skill answers WHY.

© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/skills/ci-debug of yonatangross/orchestkit.

Open the folder on GitHubat commit 1f8d8f3

Compare with similar skills

CI Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CI Debug compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CI Debug this skillyonatangross/orchestkit288—~2.7kAutomated safety check: NotesMIT
Babysit PRZenUml/web-sequence150—~871Automated safety check: PassMIT
CI Checkadam-s/intercept189—~612Automated safety check: PassMIT
CI Workflow Guidesgl-project/sglang37k2 repos~5.5kAutomated safety check: PassApache-2.0
ReleaseSma1lboy/rove146—~3.1kAutomated safety check: WarnMIT
Manor CI Triagemanor-os/manor-ai161—~533Automated safety check: PassCustom licence

Similar skills

  • Babysit PR

    ZenUml/web-sequence

    Monitor and diagnose GitHub Actions checks on ZenUML web-sequence PRs, fixing code-caused CI failures when appropriate.

    150 GitHub stars~871 tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • CI Check

    adam-s/intercept

    Run local CI checks and verify GitHub Actions status. An agent skill from adam-s/intercept.

    189 GitHub stars~612 tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • CI Workflow Guide

    sgl-project/sglang

    Guide to SGLang CI workflow orchestration — stage ordering, fail-fast, gating, partitioning, execution modes, and debugging CI failures.

    37k GitHub starsUsed in 2 repos~5.5k tokens
    DevOps & CloudAuto-check passed
  • Release

    Sma1lboy/rove

    Autonomously cut a Rove (@sma1lboy/rove) release end-to-end — detect the semver bump from pending changesets (flagging an upstream minor you didn't intend), run the release gates, dispatch the…

    146 GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check: warnings
  • Manor CI Triage

    manor-os/manor-ai

    A skill your agent uses when Manor GitHub Actions, .github/workflows/ci.yml, OSS smoke/regression jobs, web source smoke, frontend build, lint, or public CI failure logs need diagnosis or repair.

    161 GitHub stars~533 tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • CI Failure Triage and Repair

    Chachamaru127/claude-code-harness

    Diagnoses failing CI pipelines and tests, deciding first whether the test or the implementation is at fault, and hands hard cases to a dedicated fixer subagent.

    3.2k GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check: notes

More from yonatangross/orchestkit

All 107 skills in this repo
  • API Design

    yonatangross/orchestkit

    API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.

    288 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Architecture Decision Record

    yonatangross/orchestkit

    ADR templates in the Nygard format with context, decision, consequences, and alternatives.

    288 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Audit Full

    yonatangross/orchestkit

    Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.

    288 GitHub stars~3.5k tokensUpdated yesterday
    Auto-check: notes
  • Code Review Playbook

    yonatangross/orchestkit

    Structured review processes, conventional comments, language-specific checklists, and feedback templates.

    288 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Create PR

    yonatangross/orchestkit

    Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.

    288 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check: notes
  • Explore

    yonatangross/orchestkit

    Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.

    288 GitHub stars~3.9k tokensUpdated yesterday
    Auto-check: notes

Questions about CI Debug

What does CI Debug do?

Diagnose a failing CI run against an 11-pattern playbook. An agent skill from yonatangross/orchestkit. CI Debug is an agent skill from yonatangross/orchestkit. Diagnose a failing CI run against an 11-pattern playbook.

When should I use CI Debug?

CI Debug fits situations like: A specific PR check; GitHub Actions run failed and you want a diagnosis instead of speculation; org-wide CI sweeps (thats /status); for app-level test failures (the playbook is CI-infra-specific).

How do I install CI Debug in Claude Code?

Run `npx skills add yonatangross/orchestkit --skill ci-debug -a claude-code`. Or copy the skill folder (src/skills/ci-debug in yonatangross/orchestkit) into .claude/skills/ci-debug in your project. Claude Code loads it when a task matches its description.

How do I install CI Debug in Codex?

Run `npx skills add yonatangross/orchestkit --skill ci-debug -a codex`. Or copy the skill folder (src/skills/ci-debug in yonatangross/orchestkit) into .agents/skills/ci-debug in your project. Codex loads it when a task matches its description.

Can I use CI Debug in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill ci-debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ci-debug, .gemini/skills/ci-debug, .github/skills/ci-debug and .opencode/skills/ci-debug in your project.

What does CI Debug need to run?

Going by SKILL.md and its folder, CI Debug needs the command-line tools its instructions call (gh, git, python, claude, pnpm and uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, Grep, Glob. Compatibility (from SKILL.md): Claude Code 2.1.277+ — uses `gh` CLI for GitHub Actions log inspection. Works in interactive sessions and headless `claude -p --bare` invocations (e.g. /ork:ci-sentinel)..

Does CI Debug access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is CI Debug safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does CI Debug use?

CI Debug is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CI Debug use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to CI Debug?

Skills that share tags, products or a category with CI Debug: Babysit PR (ZenUml/web-sequence, 150 stars), CI Check (adam-s/intercept, 189 stars), CI Workflow Guide (sgl-project/sglang, 37k stars) and Release (Sma1lboy/rove, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CI Debug?

yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 288 GitHub stars. The repository holds 107 skills in this directory. The repository was last updated on October 6, 2026.

Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.