Agent skill

Benchmark Fp Fn Audit

by millionco in millionco/react-doctor

Audit React Doctor against ReactBench or similar diagnostic benchmark corpora for confirmed false positives, false negatives, taxonomy gaps, and verifier artifacts.

Custom licenceAuto-check passedDevelopment

Install Benchmark Fp Fn Audit

skills CLI
$ npx skills add millionco/react-doctor --skill benchmark-fp-fn-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install millionco/react-doctor benchmark-fp-fn-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/millionco/react-doctor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchmark-fp-fn-audit .claude/skills/benchmark-fp-fn-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-fp-fn-audit
GitHub stars
15k
Token cost
~2k tokens
SKILL.md length
748 words
Files
3 (incl. references)
Skills in repo
13
Repo updated
First seen
Licence
Custom licence

At a glance

Audit React Doctor against ReactBench or similar diagnostic benchmark corpora for confirmed false positives, false negatives, taxonomy gaps, and verifier artifacts.

  • Works in 5 steps: Inventory the corpus → Build diagnostic distributions → Inspect high-impact clusters → …
  • Analyzing rd.log
  • SKILL.md covers Corpus and required resources, Evidence rules and Workflow
  • Calls rg; reaches react.doctor

What it does

Benchmark Fp Fn Audit is an agent skill from millionco/react-doctor. Audit React Doctor against ReactBench or similar diagnostic benchmark corpora for confirmed false positives, false negatives, taxonomy gaps, and verifier artifacts. Use when analyzing rd.log, rd-before.json, rd-after.json, model.patch, result.json, reward/test logs, rule distributions, or when asked to perform a second adversarial pass over React Doctor benchmark findings.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/benchmark-artifacts.md`).

It sits in Development. It works with React. The repository describes itself as: Your agent writes bad React. This catches it.

When your agent uses it

  • Analyzing rd.log
  • Reward/test logs
  • Rule distributions
  • Asked to perform a second adversarial pass over React Doctor benchmark findings

Example prompts

  • “/benchmark-fp-fn-audit”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Inventory the corpus
  2. Build diagnostic distributions
  3. Inspect high-impact clusters
  4. Perform the independent second pass
  5. Produce audit artifacts

What it can do on your machine

Read from SKILL.md and the folder at commit 5b301c9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • rg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • react.doctor

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Fp Fn Audit loads about 2k tokens when it runs, and up to ~2.6k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 748 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 748 words (~1,974 tokens).

“Perform an evidence-backed audit of React Doctor diagnostics across a benchmark corpus. Read the complete rule documentation, quantify the distribution of failures, inspect every relevant trial artifact, and independently perform a second pass for additional false positives and false negatives.”

— opening of SKILL.md by millionco, Custom licence
name
benchmark-fp-fn-audit

Read the full SKILL.md on GitHub

Files

SKILL.md and 2 other files (references) in .agents/skills/benchmark-fp-fn-audit of millionco/react-doctor.

  • SKILL.md
  • agents/openai.yaml
  • references/benchmark-artifacts.md

Open the folder on GitHubat commit 5b301c9

Compare with similar skills

Benchmark Fp Fn Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Fp Fn Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Fp Fn Audit this skillmillionco/react-doctor15k—~2kAutomated safety check: PassCustom licence
Vercel Composition Patternssupabase/supabase111k58 repos~726Automated safety check: PassMIT
Electron Multi-Process ArchitectureiOfficeAI/AionUi33k1 repos~1.8kAutomated safety check: PassApache-2.0
Nx Generatenomcopter/react-mosaic4.8k7 repos~1.9kAutomated safety check: PassCustom licence
@pierre/diffs Code Renderingpierrecomputer/pierre6.3k2 repos~803Automated safety check: PassApache-2.0
OpenTUI Terminal Interfacescline/cline70k—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    React composition patterns that scale. An agent skill from supabase/supabase.

    111k GitHub starsUsed in 58 repos~726 tokens
    DevelopmentAuto-check passed
  • Tells the agent where new code belongs in an Electron multi-process project and which APIs each process may use, with rules for new bridges, services, agents and workers.

    33k GitHub starsUsed in 1 repo~1.8k tokens
    DevelopmentAuto-check passed
  • Nx Generate

    nomcopter/react-mosaic

    Generate code using nx generators. An agent skill from nomcopter/react-mosaic.

    4.8k GitHub starsUsed in 7 repos~1.9k tokens
    DevelopmentAuto-check passed
  • @pierre/diffs Code Rendering

    pierrecomputer/pierre

    Guides an agent through using @pierre/diffs to render syntax-highlighted files and diffs, and to build editing and review surfaces in React or plain JavaScript.

    6.3k GitHub starsUsed in 2 repos~803 tokens
    DevelopmentAuto-check passed
  • Helps build terminal user interfaces with OpenTUI using its core imperative API or its React and Solid reconcilers, with references for layout, keyboard, animation and testing.

    70k GitHub stars~1.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Claude Code Skill

    codeaashu/claude-code

    Development conventions and architecture guide for the Claude Code CLI repository.

    3.4k GitHub starsUsed in 1 repo~2.7k tokens
    DevelopmentAuto-check passed

More from millionco/react-doctor

All 13 skills in this repo
  • Run Parity

    millionco/react-doctor

    Compare React Doctor diagnostics for a GitHub pull request (PR) with Vercel Sandbox.

    15k GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Deslop

    millionco/react-doctor

    Simplify and refine recently modified code while preserving functionality.

    15k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Find Similar Functions

    millionco/react-doctor

    Use truffler to find similar or pre-existing JavaScript/TypeScript symbols before implementing new code, especially helpers, utilities, parsers, formatters, scanners, fuzzy matchers, and other…

    15k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Improve Threejs

    millionco/react-doctor

    Audit and fix Three.js and React Three Fiber apps for frame-loop performance, GPU memory leaks, scene-graph correctness, and visual defects like z-fighting, shadow acne, wrong color space, and…

    15k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Performance

    millionco/react-doctor

    Diagnose React runtime performance with React Doctor traces, live render outlines, Long Animation Frames, interaction timing, and component render evidence.

    15k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Product Thinking

    millionco/react-doctor

    Think like a product manager before changing React Doctor's public surface — CLI commands/flags, the 0–100 score, config (doctor.config.), the JSON report schema, package APIs…

    15k GitHub stars~3.1k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Benchmark Fp Fn Audit

What does Benchmark Fp Fn Audit do?

Audit React Doctor against ReactBench or similar diagnostic benchmark corpora for confirmed false positives, false negatives, taxonomy gaps, and verifier artifacts. Benchmark Fp Fn Audit is an agent skill from millionco/react-doctor. Audit React Doctor against ReactBench or similar diagnostic benchmark corpora for confirmed false positives, false negatives, taxonomy gaps, and verifier artifacts.

When should I use Benchmark Fp Fn Audit?

Benchmark Fp Fn Audit fits situations like: analyzing rd.log; reward/test logs; rule distributions; asked to perform a second adversarial pass over React Doctor benchmark findings.

How do I install Benchmark Fp Fn Audit in Claude Code?

Run `npx skills add millionco/react-doctor --skill benchmark-fp-fn-audit -a claude-code`. Or copy the skill folder (.agents/skills/benchmark-fp-fn-audit in millionco/react-doctor) into .claude/skills/benchmark-fp-fn-audit in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Fp Fn Audit in Codex?

Run `npx skills add millionco/react-doctor --skill benchmark-fp-fn-audit -a codex`. Or copy the skill folder (.agents/skills/benchmark-fp-fn-audit in millionco/react-doctor) into .agents/skills/benchmark-fp-fn-audit in your project. Codex loads it when a task matches its description.

Can I use Benchmark Fp Fn Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add millionco/react-doctor --skill benchmark-fp-fn-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-fp-fn-audit, .gemini/skills/benchmark-fp-fn-audit, .github/skills/benchmark-fp-fn-audit and .opencode/skills/benchmark-fp-fn-audit in your project.

What does Benchmark Fp Fn Audit need to run?

Going by SKILL.md and its folder, Benchmark Fp Fn Audit needs the command-line tools its instructions call (rg).

Does Benchmark Fp Fn Audit access the network?

SKILL.md names 1 domain. In commands or code: react.doctor; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Benchmark Fp Fn Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Fp Fn Audit use?

Benchmark Fp Fn Audit has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Benchmark Fp Fn Audit use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 621 tokens, read only when the agent opens those files.

What are the alternatives to Benchmark Fp Fn Audit?

Skills that share tags, products or a category with Benchmark Fp Fn Audit: Vercel Composition Patterns (supabase/supabase, 111k stars), Electron Multi-Process Architecture (iOfficeAI/AionUi, 33k stars), Nx Generate (nomcopter/react-mosaic, 4.8k stars) and @pierre/diffs Code Rendering (pierrecomputer/pierre, 6.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Fp Fn Audit?

millionco (a GitHub organization) maintains it in millionco/react-doctor, which has 14,992 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 10, 2026.

Source: millionco/react-doctor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.