Agent skill

Diagnose a Red Run

by different-ai in different-ai/openwork

Classifies a failing test, typecheck or CI job before any code changes, by recording the failure and running a clean control to show whether it was already broken.

Custom licenceAuto-check passedTesting & QA

Install Diagnose a Red Run

skills CLI
$ npx skills add different-ai/openwork --skill diagnose-a-red-run -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install different-ai/openwork diagnose-a-red-run --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/different-ai/openwork.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.opencode/skills/diagnose-a-red-run .claude/skills/diagnose-a-red-run && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
diagnose-a-red-run
GitHub stars
24k
Token cost
~779 tokens
SKILL.md length
395 words
Files
1
Skills in repo
33
Repo updated
First seen
Licence
Custom licence

At a glance

Classifies a failing test, typecheck or CI job before any code changes, by recording the failure and running a clean control to show whether it was already broken.

  • A test, typecheck, lint or CI job has just gone red
  • SKILL.md covers Capture the branch failure, Run a clean control, Classify by signature and Red core journey (pr-proof.yml…, plus 1 more section
  • Calls git and tsx
  • Deciding whether a failure was already present on the base branch

What it does

The skill starts by recording the exact command, commit SHA, exit code and the passed, failed and skipped counts, and quoting the first actionable failure rather than summarizing it. It classifies the check as a testkit spec, typecheck, build, lint or CI job. To call a failure pre-existing it requires a clean control: a detached git worktree of origin/dev with the same prerequisites, tool versions, environment and flags, where the same command is run again.

A failure counts as pre-existing only when the control shows the same failure. Otherwise it is labeled introduced, environment-specific or unresolved, and a control that does not reproduce it is not proof of flakiness until both sides are repeated. A signature list maps symptoms to causes, for example an INVALID_ORIGIN 403 pointing at trusted origins, a fresh_auth_required 403 meaning the session aged, and a timeout with an on-screen dump that names the state.

It also has steps for a red core journey job, starting with the job summary table, and for environment forensics: kill by port rather than process name, treat EADDRINUSE alongside a healthy response as a surviving zombie process, and kill the process group for Electron respawns.

When your agent uses it

  • A test, typecheck, lint or CI job has just gone red
  • Deciding whether a failure was already present on the base branch
  • Investigating a timeout, a flaky run or a leftover process holding a port

Example prompts

  • “The typecheck job failed on my branch. Was this already broken on dev?”
  • “Diagnose why the core journey job timed out and classify the failure.”
  • “The dev server reports EADDRINUSE but the health check returns 200; find what is still running.”

Requirements

  • A git repository with an origin/dev branch for the clean control

What it can do on your machine

Read from SKILL.md and the folder at commit 6325c02. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • tsx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Diagnose a Red Run loads about 779 tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 395 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~779

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 395 words (~779 tokens).

“counts. Quote the first actionable failure; do not summarize it away.”

— opening of SKILL.md by different-ai, Custom licence
name
diagnose-a-red-run

Read the full SKILL.md on GitHub

Files

Just SKILL.md in .opencode/skills/diagnose-a-red-run of different-ai/openwork.

Open the folder on GitHubat commit 6325c02

Compare with similar skills

Diagnose a Red Run next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Diagnose a Red Run compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Diagnose a Red Run this skilldifferent-ai/openwork24k—~779Automated safety check: PassCustom licence
cmux Package Test Bisectmanaflow-ai/cmux28k—~1.5kAutomated safety check: PassCustom licence
CI Failure Triage and RepairChachamaru127/claude-code-harness3.2k1 repos~1.1kAutomated safety check: NotesMIT
AI Bug Triagepetrkindlmann/qa-skills170—~5.2kAutomated safety check: PassMIT
Regression Root Cause AnalyzerArabelaTso/Skills-4-SE253—~3.4kAutomated safety check: PassApache-2.0
TiDB Test Diff Triagepingcap/tidb41k—~498Automated safety check: PassApache-2.0

Similar skills

  • cmux Package Test Bisect

    manaflow-ai/cmux

    Finds which commit broke a failing Swift package suite in the cmux repo by bisecting on CI, then judges per test whether it went stale or the code regressed.

    28k GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check passed
  • CI Failure Triage and Repair

    Chachamaru127/claude-code-harness

    Diagnoses failing CI pipelines and tests, deciding first whether the test or the implementation is at fault, and hands hard cases to a dedicated fixer subagent.

    3.2k GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check: notes
  • AI Bug Triage

    petrkindlmann/qa-skills

    Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation.

    170 GitHub stars~5.2k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Regression Root Cause Analyzer

    ArabelaTso/Skills-4-SE

    Locate root causes of failing regression tests by analyzing code changes, error messages, and test dependencies.

    253 GitHub stars~3.4k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Investigates TiDB plan or test-result diffs that the change does not explain, ruling out failpoint setup and merge effects before expected outputs are updated.

    41k GitHub stars~498 tokensUpdated today
    Testing & QAAuto-check passed
  • Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.

    6.7k GitHub stars~4.4k tokensUpdated today
    Testing & QAAuto-check passed

More from different-ai/openwork

All 33 skills in this repo
  • OpenWork Desktop CDP Driver

    different-ai/openwork

    Drives a running OpenWork desktop window over CDP from the shell to evaluate JS, take screenshots, start sessions and send prompts for hand checks.

    24k GitHub stars~465 tokensUpdated today
    Auto-check passed
  • Fake Model Provider Faults

    different-ai/openwork

    Makes the desktop app's model provider fail on demand, with refused connections, resets, stalls and HTTP 4xx and 5xx errors, so error and retry states can be reproduced.

    24k GitHub stars~642 tokensUpdated today
    Auto-check passed
  • OpenWork Model Alias Manager

    different-ai/openwork

    Manages OpenWork's inference model aliases, discounts and overlays over the upstream OpenRouter catalog, and triggers the GitHub workflow that refreshes base models.

    24k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Reproduce Chat States

    different-ai/openwork

    Fires known chat states in the running OpenWork desktop app, such as provider errors, retries and tool steps, so you can check how each renders.

    24k GitHub stars~673 tokensUpdated today
    Auto-check passed
  • Attaches OpenCode browser tools to the OpenWork Electron dev app through CDP to explore its UI, send a composer task and debug, not to give test verdicts.

    24k GitHub stars~780 tokensUpdated today
    Auto-check passed
  • Daytona Sandbox Operations

    different-ai/openwork

    Covers Daytona CLI setup, sandbox debugging, keeping a sandbox alive and which credentials the CLI uses, for when Daytona itself is the problem rather than the tests.

    24k GitHub stars~917 tokensUpdated today
    Auto-check passed

Works with

Questions about Diagnose a Red Run

What does Diagnose a Red Run do?

Classifies a failing test, typecheck or CI job before any code changes, by recording the failure and running a clean control to show whether it was already broken. The skill starts by recording the exact command, commit SHA, exit code and the passed, failed and skipped counts, and quoting the first actionable failure rather than summarizing it. It classifies the check as a testkit spec, typecheck, build, lint or CI job.

When should I use Diagnose a Red Run?

Diagnose a Red Run fits situations like: A test, typecheck, lint or CI job has just gone red; deciding whether a failure was already present on the base branch; investigating a timeout, a flaky run or a leftover process holding a port.

How do I install Diagnose a Red Run in Claude Code?

Run `npx skills add different-ai/openwork --skill diagnose-a-red-run -a claude-code`. Or copy the skill folder (.opencode/skills/diagnose-a-red-run in different-ai/openwork) into .claude/skills/diagnose-a-red-run in your project. Claude Code loads it when a task matches its description.

How do I install Diagnose a Red Run in Codex?

Run `npx skills add different-ai/openwork --skill diagnose-a-red-run -a codex`. Or copy the skill folder (.opencode/skills/diagnose-a-red-run in different-ai/openwork) into .agents/skills/diagnose-a-red-run in your project. Codex loads it when a task matches its description.

Can I use Diagnose a Red Run in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add different-ai/openwork --skill diagnose-a-red-run -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/diagnose-a-red-run, .gemini/skills/diagnose-a-red-run, .github/skills/diagnose-a-red-run and .opencode/skills/diagnose-a-red-run in your project.

What does Diagnose a Red Run need to run?

Going by SKILL.md and its folder, Diagnose a Red Run needs the command-line tools its instructions call (git and tsx). Our summary lists: A git repository with an origin/dev branch for the clean control.

Does Diagnose a Red Run access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Diagnose a Red Run safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Diagnose a Red Run use?

Diagnose a Red Run has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Diagnose a Red Run use?

About 779 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Diagnose a Red Run?

Skills that share tags, products or a category with Diagnose a Red Run: cmux Package Test Bisect (manaflow-ai/cmux, 28k stars), CI Failure Triage and Repair (Chachamaru127/claude-code-harness, 3.2k stars), AI Bug Triage (petrkindlmann/qa-skills, 170 stars) and Regression Root Cause Analyzer (ArabelaTso/Skills-4-SE, 253 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Diagnose a Red Run?

different-ai (a GitHub organization) maintains it in different-ai/openwork, which has 23,987 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 10, 2026.

Source: different-ai/openwork on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.