Turns a suspected bug into a proven one. An agent skill from hashgraph-online/awesome-codex-plugins.

Apache-2.0Auto-check passedDevelopment

Install Red Green Proof

skills CLI
$ npx skills add hashgraph-online/awesome-codex-plugins --skill red-green-proof -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hashgraph-online/awesome-codex-plugins red-green-proof --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hashgraph-online/awesome-codex-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/RooAGI/red-green-proof/skills/red-green-proof .claude/skills/red-green-proof && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
red-green-proof
GitHub stars
1.3k
Token cost
~1.9k tokens
SKILL.md length
1,133 words
Files
1
Skills in repo
716
Repo updated
First seen
Licence
Apache-2.0

At a glance

Turns a suspected bug into a proven one. An agent skill from hashgraph-online/awesome-codex-plugins.

  • Works in 5 steps: Verify the cause. Do not infer it. → Write the test. Watch it fail. → Fix it. → …
  • Investigating an incident
  • SKILL.md covers Picking the target, The loop, When you cannot get a true red and Reporting, plus 2 more sections
  • Calls git

What it does

Red Green Proof is an agent skill from hashgraph-online/awesome-codex-plugins. Turns a suspected bug into a proven one. Verify the cause against reality before claiming it, write a test that FAILS on the current code, apply the fix, watch it pass — then revert the fix and confirm the test goes red again, because a test that passes both ways proves nothing. Invoked bare after a debugging conversation, it takes the target from context rather than asking. Use when fixing a bug, investigating an incident, hardening a flaky area, or when asked to "add tests that reveal the bug". Triggers on…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development. The repository describes itself as: A curated list of awesome OpenAI Codex / ChatGPT plugins, skills, and resources. The 1 Codex Marketplace. See live plugins at: https://hol.org/plugins/best-codex-plugins. The licence is Apache-2.0.

When your agent uses it

  • Investigating an incident
  • Hardening a flaky area
  • Asked to add tests that reveal the bug
  • /red-green-proof

Example prompts

  • “add tests that reveal the bug”
  • “/red-green-proof”
  • “prove the bug”
  • “/red-green-proof”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Verify the cause. Do not infer it.
  2. Write the test. Watch it fail.
  3. Fix it.
  4. Revert the fix. Confirm red. ← the whole point
  5. Run everything.

What it can do on your machine

Read from SKILL.md and the folder at commit 3e1456a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Red Green Proof loads about 1.9k tokens when it runs. Until then it costs about 167 tokens; SKILL.md has 1,133 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~167
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hashgraph-online/awesome-codex-plugins at commit 3e1456a, republished under its Apache-2.0 licence (© hashgraph-online). 1,133 words, ~1,921 tokens.

Download SKILL.mdSave it as .claude/skills/red-green-proof/SKILL.md (or your agent's skills folder).
name
red-green-proof
description
Turns a suspected bug into a proven one. Verify the cause against reality before claiming it, write a test that FAILS on the current code, apply the fix, watch it pass — then revert the fix and confirm the test goes red again, because a test that passes both ways proves nothing. Invoked bare after a debugging conversation, it takes the target from context rather than asking. Use when fixing a bug, investigating an incident, hardening a flaky area, or when asked to "add tests that reveal the bug". Triggers on "/red-green-proof", "prove the bug", "red-green", "make the test fail first", "is that test load-bearing", "reveal the bug with a test".

Red-Green-Proof

A bug is not fixed because the tests pass. A bug is fixed when a test fails without your fix and passes with it, and you have watched it do both.

Most "regression tests" written after a fix never had the chance to fail. They are decoration. This skill is the discipline that separates a real test from a decorative one.

Picking the target

With an argument, that is the target: /red-green-proof the status endpoint reports finished runs as running.

With no argument, take the target from the conversation. The usual case is that you have just spent a long stretch investigating something — the bug is already on screen. Do not ask what to work on. Scan back through the session and pick, in this order:

  1. A defect just identified but not yet fixed.
  2. A fix applied without a failing test to back it — the most valuable target, because the test still has to be proven load-bearing.
  3. A test written but never verified red.
  4. Something described as "still open", "not fixed yet", or "characterization only".

State your pick in one line and start:

Target: the buffer discards events when the flush write fails (checkpointBuffer.ts:271). Verifying before writing the test.

Only ask the user if two or more candidates are genuinely equal in priority — then list them as a short numbered choice and stop. If several related defects came up, handle them one at a time through the full loop rather than batching; each needs its own red.

If nothing in the conversation qualifies, say so and ask for a target rather than inventing one.

The loop

Run these in order. Do not skip 1, and never skip 4.

1. Verify the cause. Do not infer it.

Before you write a line of test code, establish what actually happened using evidence you can point at: the real record from the database or API, the real log line, the actual source of the function you are blaming.

Say which of these you have:

  • Proven — I read the value / ran the code / pulled the record.
  • Inferred — consistent with the evidence, but I have not confirmed it.
  • Unknown — I cannot determine this from what is available.

State the label out loud. If the honest answer is Unknown, say so and name the one artifact that would settle it. A confident wrong cause costs more than an admitted gap, because it sends the fix to the wrong place.

Two traps worth naming:

  • A plausible mechanism is not the mechanism. Rank candidates, then go rule them out one at a time by reading code, not by reasoning about it.
  • Check whether the "bug" was deliberate. git log -S '<the exact line>' and read the commit. If a test already asserts the current behaviour, someone may have wanted it. Understand why before you invert it.
2. Write the test. Watch it fail.

The test must fail against the code as it is right now, before any fix exists.

Name it after the defect, not the function: reveals bug: stale status shadows a terminal event, not test inferStatus.

Assert the observable consequence a user or caller would see — the wrong status, the lost record, the 500 — not an internal call count. Call counts pass for the wrong reasons.

If it does not fail, you have not reproduced the bug. Go back to step 1.

3. Fix it.

Smallest change that makes the failing test pass. If the fix needs to be large, the test was probably too broad — narrow it and split.

4. Revert the fix. Confirm red. ← the whole point

Put the buggy code back, run the test, and watch it fail. Then restore the fix.

bash
cp path/to/fixed.ts /tmp/fix.bak
# revert the fix (edit, patch, or git stash the single file)
<run the specific test>          # MUST fail here
cp /tmp/fix.bak path/to/fixed.ts
<run the specific test>          # green again

If it stayed green, the test is fake. Not "probably fine" — fake. Rewrite it and repeat.

This is not paranoia. It catches a specific, common failure: a test that exercises a path adjacent to the bug rather than the bug. Two real examples:

  • A concurrency test modelled two writers but had no await between each writer's read and its write. Both ran to completion atomically, never interleaved, and the test passed with the lock removed. Fix: add latency between read and write, mirroring the real network gap.
  • A fault-isolation test asserted "no 500 when the store fails", but the injected failure was a missing record, which the loader turns into null rather than throwing. The error path never ran. Fix: inject a real 500.

Both looked correct on the page. Only the revert exposed them.

Show full SKILL.md (384 more words)Show less
5. Run everything.

The full suite, plus type checks. A fix that corrects one behaviour often contradicts a test that pinned the old one — which is a finding, not an annoyance. Go read that test and decide deliberately (see step 1).

When you cannot get a true red

Sometimes the buggy code cannot be invoked: it is a closure inside a 3,000-line factory, a route handler, a private method. Do not fake a behavioural test. Pick one of these and label it in the file header:

TypeUse whenHonesty requirement
Extract, then testThe logic can be moved to a module with injected dependenciesPreferred. Do this if the extraction is small.
Structural testThe call site cannot be reached, but the invariant is visible in the sourceScan the source and assert the property (e.g. "every write cycle sits inside the lock"). It will name offenders itself. Say in the header that it is structural and should be replaced when the code becomes reachable.
Model testTesting an interaction between components you cannot both instantiateMirror the real code field-for-field, cite the file:line it models, and note in the header that it is a model. Extra important to run step 4 — models are where fake tests hide.
Characterization testDocumenting current behaviour you are not fixing yetName it as such. Do not call it a regression test. Flag that it will invert when the bug is fixed.

Reporting

When you present the work, include:

  • What you proved vs inferred vs still do not know.
  • The revert output — the actual failure line, not a claim that it failed.
  • Which tests are behavioural, structural, model, or characterization.
  • Anything you got wrong earlier in the investigation, corrected explicitly.

Do not describe a test as "revealing the bug" unless you have seen it red.

Anti-patterns

  • Writing the test after the fix and never reverting → decoration.
  • Asserting a mock was called instead of the user-visible outcome → passes for the wrong reason.
  • Widening a test until it passes → deleting the signal.
  • Inverting a pre-existing test because it blocks you → check the history first; it may encode a decision.
  • Reporting "all green" as success when nothing was ever red.

One-line version

Red first, green second, red again on purpose, then green — and say which parts you actually proved.

© hashgraph-online, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/RooAGI/red-green-proof/skills/red-green-proof of hashgraph-online/awesome-codex-plugins.

Open the folder on GitHubat commit 3e1456a

Compare with similar skills

Red Green Proof next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Red Green Proof compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Red Green Proof this skillhashgraph-online/awesome-codex-plugins1.3k—~1.9kAutomated safety check: PassApache-2.0
Vercel Composition Patternssupabase/supabase111k58 repos~726Automated safety check: PassMIT
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
Typescript Advanced Typesrolling-scopes/rsschool-app10k25 repos~4.2kAutomated safety check: PassMPL-2.0
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k4 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Official

    React composition patterns that scale. An agent skill from supabase/supabase.

    111k GitHub starsUsed in 58 repos~726 tokens
    DevelopmentAuto-check passed
  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Typescript Advanced Types

    rolling-scopes/rsschool-app

    Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.

    10k GitHub starsUsed in 25 repos~4.2k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed

More from hashgraph-online/awesome-codex-plugins

All 716 skills in this repo
  • Anime Reaction Gif

    hashgraph-online/awesome-codex-plugins

    Create original anime-style reaction stickers as looping GIFs and MP4 previews, using generated character pose sheets and timed key poses.

    1.3k GitHub stars~922 tokensUpdated yesterday
    Auto-check passed
  • Calibredb

    hashgraph-online/awesome-codex-plugins

    Manage and query Calibre libraries with the calibredb CLI (local paths or Calibre Content server URLs).

    1.3k GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Rust API Test Harness

    hashgraph-online/awesome-codex-plugins

    A skill your agent uses when adding, changing, testing, or debugging Rust HTTP APIs and services, especially when Codex needs black-box integration tests, random-port app startup, real database test…

    1.3k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Art

    hashgraph-online/awesome-codex-plugins

    Make a studio's game look like something at build time — a cover from a real frame of the game (free), painted covers, backdrops, textures and character plates from image models through the…

    1.3k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Calle

    hashgraph-online/awesome-codex-plugins

    Use CALL-E from Codex through the calle CLI. An agent skill from hashgraph-online/awesome-codex-plugins.

    1.3k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Game Balance Economy

    hashgraph-online/awesome-codex-plugins

    Balance game difficulty, resources, rewards, probability, progression, economies, and dominant strategies.

    1.3k GitHub stars~618 tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Red Green Proof

What does Red Green Proof do?

Turns a suspected bug into a proven one. An agent skill from hashgraph-online/awesome-codex-plugins. Red Green Proof is an agent skill from hashgraph-online/awesome-codex-plugins. Turns a suspected bug into a proven one.

When should I use Red Green Proof?

Red Green Proof fits situations like: investigating an incident; hardening a flaky area; asked to add tests that reveal the bug; /red-green-proof.

How do I install Red Green Proof in Claude Code?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill red-green-proof -a claude-code`. Or copy the skill folder (plugins/RooAGI/red-green-proof/skills/red-green-proof in hashgraph-online/awesome-codex-plugins) into .claude/skills/red-green-proof in your project. Claude Code loads it when a task matches its description.

How do I install Red Green Proof in Codex?

Run `npx skills add hashgraph-online/awesome-codex-plugins --skill red-green-proof -a codex`. Or copy the skill folder (plugins/RooAGI/red-green-proof/skills/red-green-proof in hashgraph-online/awesome-codex-plugins) into .agents/skills/red-green-proof in your project. Codex loads it when a task matches its description.

Can I use Red Green Proof in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hashgraph-online/awesome-codex-plugins --skill red-green-proof -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/red-green-proof, .gemini/skills/red-green-proof, .github/skills/red-green-proof and .opencode/skills/red-green-proof in your project.

What does Red Green Proof need to run?

Going by SKILL.md and its folder, Red Green Proof needs the command-line tools its instructions call (git).

Does Red Green Proof access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Red Green Proof safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Red Green Proof use?

Red Green Proof is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Red Green Proof use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Red Green Proof?

Skills that share tags, products or a category with Red Green Proof: Vercel Composition Patterns (supabase/supabase, 111k stars), Finishing a Development Branch (obra/superpowers, 297k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars) and PR Babysitter (openinterpreter/openinterpreter, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Red Green Proof?

hashgraph-online (a GitHub organization) maintains it in hashgraph-online/awesome-codex-plugins, which has 1,267 GitHub stars. The repository holds 716 skills in this directory. The repository was last updated on October 10, 2026.

Source: hashgraph-online/awesome-codex-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.