Agent skill

Evidence Audit

by WrongStack in WrongStack/WrongStack

Evidence-led audit/fix rounds for any codebase in any language.

MITAuto-check passedDevelopment

Install Evidence Audit

skills CLI
$ npx skills add WrongStack/WrongStack --skill evidence-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install WrongStack/WrongStack evidence-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/WrongStack/WrongStack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/core/skills/evidence-audit .claude/skills/evidence-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evidence-audit
GitHub stars
370
Token cost
~3.5k tokens
SKILL.md length
1,769 words
Files
2
Skills in repo
38
Repo updated
First seen
Licence
MIT

At a glance

Evidence-led audit/fix rounds for any codebase in any language.

  • Works in 6 steps: Candidate → proof → Patch → Verify → …
  • The user asks to audit
  • SKILL.md covers Setup (every round, before any…, Procedure, Environment boundaries… and Cross-language defect shapes, plus 2 more sections
  • Calls git

What it does

Evidence Audit is an agent skill from WrongStack/WrongStack. Evidence-led audit/fix rounds for any codebase in any language. Finds only defects that a runnable proof reproduces on current code, applies the narrowest in-scope patch, verifies with a second proof, promotes high-risk proofs to permanent regression tests, and reports exact validation results. Use whenever the user asks to audit, bug-hunt, "find real bugs in", "prove and fix", "continue the audit round on", or "what is actually broken in" a package, module, directory, or service, even without the word "audit"…

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Development, covering Refactoring. The repository describes itself as: An AI coding agent that reads your code, edits files, runs commands, and reasons through bugs — across a terminal REPL, a full-screen TUI, and a browser UI, while you keep your… The licence is MIT.

When your agent uses it

  • The user asks to audit
  • Find real bugs in
  • Continue the audit round on
  • What is actually broken in a package

Example prompts

  • “find real bugs in”
  • “prove and fix”
  • “continue the audit round on”
  • “/evidence-audit”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Candidate → proof
  2. Patch
  3. Verify
  4. Promote
  5. Validate (in this order, report each exactly)
  6. Hand-off

What it can do on your machine

Read from SKILL.md and the folder at commit 57f6018. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evidence Audit loads about 3.5k tokens when it runs. Until then it costs about 174 tokens; SKILL.md has 1,769 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~174
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from WrongStack/WrongStack at commit 57f6018, republished under its MIT licence (© WrongStack). 1,769 words, ~3,496 tokens.

Download SKILL.mdSave it as .claude/skills/evidence-audit/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
evidence-audit
description
Evidence-led audit/fix rounds for any codebase in any language. Finds only defects that a runnable proof reproduces on current code, applies the narrowest in-scope patch, verifies with a second proof, promotes high-risk proofs to permanent regression tests, and reports exact validation results. Use whenever the user asks to audit, bug-hunt, "find real bugs in", "prove and fix", "continue the audit round on", or "what is actually broken in" a package, module, directory, or service, even without the word "audit", and for turning .temp_files proof scripts into regression tests. Do not use for a single already-reported bug, an unscoped security claim, or general refactoring.
required-capabilities
filesystem.read, filesystem.write, execution.shell, verification.run, version-control.manage

Evidence-led audit

Invoke as $evidence-audit <scope>. The text after the skill name is the scope. Every round produces the same thing: a ledger of findings where each finding has a proof that fails on current code, a patch that touches only the scope, and a verifier that passes after the patch. Nothing without a proof becomes a patch. Nothing without a verifier is called fixed. This holds in every language and every repo: an unproven "improvement" is a regression risk with no upside.

Setup (every round, before any analysis)

  1. Read the repo's agent instructions if present (AGENTS.md, CONTRIBUTING.md, or equivalent at the root and in the scope). Run git status --short and record the output: these are pre-existing changes you must preserve and never touch.
  2. Detect the toolchain from the repo, not from assumptions. Record, with the file that proves each: language(s) and version, package/build tool, test runner and how the project invokes it, type checker (if any), formatter/linter, and whether a race detector, sanitizer, or coverage tool is already configured. Use only what is already there. Do not install a test framework, mutation tool, or linter.
  3. Resolve the scope. A scope is concrete when it maps to specific files. If the prompt gives none, or it is vague ("the backend", "everything"):
    • Do a shallow structural scan only: directory layout, entry points, modules with thin or no tests, recent churn from git log --stat -20. Do not read implementations in depth yet.
    • List 5–8 numbered candidate scopes, each with path(s), one line on what it does, and one line on why it is risky (concurrency, I/O, error handling, no tests, recent churn, many callers).
    • Ask the user to pick a number or give a path, print Awaiting scope., and stop. Do not choose for them.
  4. Print Scope confirmed: <paths>. Everything you patch must be inside it. Files outside may be read for callers, types, and configs only.
  5. If .temp_files/ledger_<scope-slug>.md already exists, this is a continued round: read it, keep its IDs, start new findings from the next free ID. Never renumber.
  6. Read the scope's source and its nearest tests before forming any candidate. Write down the invariants the code is supposed to keep: what must be true after each await/yield/callback, who owns each resource, what state a caller may observe mid-operation, which inputs are trusted. Real defects are violations of invariants the author assumed but never checked.

Procedure

1. Candidate → proof

Keep the ledger at .temp_files/ledger_<scope-slug>.md. All proofs live in .temp_files/ at the repo root (create it if needed) and are written in the project's own language using its own runner or a plain executable in that language. For each candidate defect:

  • Write .temp_files/prove_F<ID>_<slug>.<ext>. It must run against current, unmodified code and contain:
    • the failing case, printing EXPECTED: and ACTUAL: side by side and ending with PROBLEM CONFIRMED + non-zero exit;
    • an unaffected control: a nearby input that behaves correctly, so the proof shows the defect is specific and not an environment artifact.
    • For a performance claim: a fixed input size, enough repetitions to be stable, printed numbers, and a stated baseline for what "acceptable" means. A timing without a baseline is not a finding.
  • Run it. Paste the output into the ledger. If it does not reproduce, the candidate is dropped: record it under "Not reproduced" in one line and move on. Do not keep it as a hunch.
  • Classify each confirmed finding: bug / perf / resource-leak / race / ownership / input-handling. Record location (file:line, symbol), impact, and severity with a one-line reason.

Candidate discovery is bounded: aim for 1–3 confirmed findings per round. More than that usually means the round should split into two scopes.

2. Patch
  • Make the smallest scope-only change that addresses the root cause named in the proof. Write the root cause in the ledger before editing.
  • No public API, config format, wire/serialization format, CLI flag, schema, or persisted-data changes unless the bug is in that surface; say so explicitly if it is.
  • No drive-by refactors, renames, or formatting outside the lines the fix needs. The diff must be traceable line-for-line to a finding ID.
  • No new dependencies.
  • Do not patch anything that exists only as an audit observation. If you notice something unproven while patching, add it as a candidate for step 1, not to the diff.
3. Verify
  • Write .temp_files/verify_F<ID>_<slug>.<ext>: the same reproduction plus at least two nearby edge cases (empty, boundary, repeated call, out-of-order completion, failure injected at the dependency). It ends with FIX VERIFIED + zero exit.
  • Re-run the original prove_F<ID>_* and confirm it now prints PROBLEM NOT REPRODUCED. Paste both outputs into the ledger.
  • For race / ownership / lifecycle findings, the verifier must use gated completion order: explicit barriers, channels, futures, latches, or deferreds that force the stale task to complete after the stop/restart. Sleep-based timing is not accepted as verification in any language. Run under the project's race detector or sanitizer if one is configured.
  • Run the verifier at least 3 times in a row. A verifier that is not deterministic is not accepted.
4. Promote

For findings rated High or Critical, or any race / ownership finding, promote the strongest proof into the project's permanent test suite next to the code it covers, following the project's test naming, placement, and assertion conventions. .temp_files artifacts are evidence for this round; the regression test is what protects the branch next month. Record the test path in the ledger. Medium/Low findings may stay as .temp_files verifiers unless the user asks otherwise.

5. Validate (in this order, report each exactly)
  1. The focused proof/verify/regression tests for every finding.
  2. The scope's own test target (package, module, crate, directory, as the project defines it).
  3. Compile / typecheck / build, as the project defines it.
  4. Formatter and linter, scoped to the changed files, using the project's configured tools.
  5. git diff --check.

Run broader suites serially when tests share writers (temp dirs, daemons, ports, databases); parallel runs produce false failures in such repos. Paste the command and the result line for each step. If a step was not run, say NOT RUN and why. Never imply it passed.

If the sandbox blocks a command (network, write outside the workspace, process spawn), request approval for that exact command rather than working around it. A workaround changes what was validated.

Show full SKILL.md (726 more words)Show less
6. Hand-off
  • Leave all prove_*, verify_*, and ledger files in place.
  • No git commit, push, reset, stash, checkout --, or .temp_files cleanup unless the user explicitly authorizes it in this session.
  • Show git status --short again and confirm the pre-existing changes from Setup step 1 are unchanged.
  • End with the report (below), then Awaiting instructions for the next turn. If the user asked for continued rounds, propose the next scope as one line with a reason, not a roadmap.

Environment boundaries (record, do not infer)

  • Permission errors on temp files, daemon/service startup failures, port conflicts, missing system libraries, and flaky network are environment contamination until a focused reproduction ties them to the patch. Compare: does the focused proof pass while the broad suite fails on an environment error? Then rely on focused + scoped test + build + lint + diff evidence and state the broad-suite boundary explicitly. Do not label the product "fixed" or "regressed" from a contaminated run.
  • Security claims that need a real boundary (sandbox escape, privilege bypass, injection into a live service) require that boundary to exist in the environment. If it does not, stop that line of inquiry, write the boundary in the ledger, and do not infer a result from static reading.
  • A proof that depends on network, a live external service, wall-clock timing, or the host's locale/timezone is not a proof here. Rewrite it with a fake at the project's own seam, or drop the finding.

Cross-language defect shapes

Recurring defect shapes that survive regardless of language. Check for them explicitly where the scope has async, concurrency, resources, or external input.

SymptomCauseFix shapeVerifier shape
Stale async work publishes into a stopped/restarted/replaced ownerState accepted after an await/callback without re-checking identity, generation, or livenessCapture identity+generation before the suspension point, compare after, drop on mismatchGated completion order: start → stop/replace → release the stale task → assert nothing published
Cleanup removes another instance's file, lock, socket, or recordOwnership assumed after a failed exclusive create or an uncertain existence probeClean or reclaim only on proven ownership (token/PID/generation) or proven absence; fail closed on unknown errorsTwo instances contend; assert the loser never deletes the winner's resource; injected unknown error aborts cleanup
Handler, callback, or listener runs on a disposed/closed ownerRegistration outlives the owner; close didn't unregister or cancel in-flight workClose unregisters and flips a closed flag every entry point checksClose, then fire the event; assert no-op and no panic/throw
Resource count grows across repeated operations (handles, goroutines/threads, listeners, connections, memory)Release missing on an error or early-return pathMove release into finally/defer/RAII/context manager; make release idempotentN iterations with injected failure; assert count unchanged before vs after
Partial write, torn record, or replay divergenceWriter committed an incomplete or out-of-order record on error; reader trusts arrival orderWrite complete records atomically; sequence by generation, not arrivalReplay into fresh state after one injected failure; deep-equal against live state
Off-by-one or boundary mishandlingInclusive/exclusive mismatch at 0, 1, length, maxFix the comparison; assert the boundary in a comment only if non-obviousProof at exactly the boundary, control one step inside
Error swallowed or converted to a defaultCatch-all returns nil/zero/empty instead of propagatingPropagate, or handle the one specific case and propagate the restInject the error at the seam; assert it reaches the caller
Unvalidated input reaches a sensitive operationTrust boundary assumed, not enforcedValidate at the boundary; reject, don't sanitize-and-continueMalformed/hostile input; assert rejection and no side effect

Report structure

ALWAYS end the round with this exact layout, taken from the ledger:

## Round <N> — <scope>
### Summary
<findings confirmed / dropped / promoted; residual risk in one sentence>

### Toolchain
<language/version, test runner cmd, build cmd, lint cmd, race/sanitizer if any — with the file each was read from>

### Findings
#### F<ID> — <type> — <severity> — <title>
- Location:
- Root cause:
- Proof (before): .temp_files/prove_F<ID>_<slug>.<ext> → <key lines>
- Patch: <files:lines changed>
- Proof (after): .temp_files/verify_F<ID>_<slug>.<ext> → <key lines>; original proof → PROBLEM NOT REPRODUCED
- Regression test: <path> | not promoted (<reason>)

### Not reproduced (dropped)
<one line each, or "none">

### Validation
- Focused tests: <cmd> → <result>
- Scope tests: <cmd> → <result>
- Build/typecheck: <cmd> → <result>
- Format/lint (scoped): <cmd> → <result>
- git diff --check: <result>
- NOT RUN: <full suite / live / security / other> — <reason>

### Boundaries
<environment contamination observed, security lines stopped, shared-branch notes>

### Working tree
<git status --short; confirmation that pre-existing changes are untouched>

### Next scope (if rounds continue)
<one line + reason>

Verification checklist (self-check before reporting)

  • Toolchain was read from the repo and cited; nothing was installed.
  • Every finding has a prove_ that printed PROBLEM CONFIRMED on unmodified code, with an unaffected control (and a stated baseline for perf).
  • Every finding has a verify_ that printed FIX VERIFIED three times in a row, and its prove_ now prints PROBLEM NOT REPRODUCED.
  • Race/ownership verifiers use gated completion order, not sleeps; race detector/sanitizer used if configured.
  • Diff contains only lines traceable to a finding ID; no files outside scope; no new dependencies; pre-existing git status changes untouched.
  • High/Critical and race/ownership findings have a permanent regression test path.
  • Validation steps are reported with exact commands and results; unrun steps say NOT RUN.
  • No commit/push/reset/cleanup performed without explicit authorization.
  • Ledger on disk matches the report; IDs are stable across rounds.

© WrongStack, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in packages/core/skills/evidence-audit of WrongStack/WrongStack.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 57f6018

Compare with similar skills

Evidence Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evidence Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evidence Audit this skillWrongStack/WrongStack370—~3.5kAutomated safety check: PassMIT
Dep Validateblokadaorg/blokada3.3k—~5.7kAutomated safety check: NotesMPL-2.0
Effect TSmattiacerutti/supernova187—~2.8kAutomated safety check: PassMIT
Skill Writingmillionco/expect3.6k—~1.8kAutomated safety check: PassCustom licence
Go Rigmudrii/openclaw-dashboard458—~2.8kAutomated safety check: PassMIT
AI Development Guideshinpr/claude-code-workflows691—~3.9kAutomated safety check: PassMIT

Similar skills

  • Dep Validate

    blokadaorg/blokada

    A skill your agent uses to validate risky dependency bumps end to end as a local or cloud-launched agent.

    3.3k GitHub stars~5.7k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Effect TS

    mattiacerutti/supernova

    Write idiomatic Effect v4 TypeScript following official best practices from effect-solutions and the Effect source.

    187 GitHub stars~2.8k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Skill Writing

    millionco/expect

    Write and improve agent skills (SKILL.md files). An agent skill from millionco/expect.

    3.6k GitHub stars~1.8k tokensUpdated 5 mo ago
    DevelopmentAuto-check passed
  • Go Rig

    mudrii/openclaw-dashboard

    A skill your agent uses when building, reviewing, or refactoring Go code that must follow strict design discipline — ATDD/TDD workflow, explicit dependency injection, package-boundary discipline…

    458 GitHub stars~2.8k tokensUpdated 15 days ago
    DevelopmentAuto-check passed
  • AI Development Guide

    shinpr/claude-code-workflows

    Applies language-agnostic and backend technical decision criteria, anti-pattern detection, debugging, and quality gates.

    691 GitHub stars~3.9k tokensUpdated 6 days ago
    DevelopmentAuto-check passed
  • Go Rig

    mudrii/openclaw-dashboard

    A skill your agent uses when building, reviewing, or refactoring Go code in this repository.

    458 GitHub stars~682 tokensUpdated 15 days ago
    DevelopmentAuto-check passed

More from WrongStack/WrongStack

All 38 skills in this repo
  • Design Craft

    WrongStack/WrongStack

    Design or substantially improve user-facing interfaces with a product-specific visual direction, content hierarchy, and rendered critique.

    370 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Design Critique

    WrongStack/WrongStack

    A skill your agent uses to audit an interface that already exists and say precisely why it looks generated, templated, or unfinished — a scored rubric across composition, typography, color, states…

    370 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Mailbox Bridge

    WrongStack/WrongStack

    A skill your agent uses when external coding agents (Claude Code, Aider, custom scripts) need to participate in the project's shared WrongStack mailbox, or when a user asks to "expose the mailbox"…

    370 GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Multi Agent

    WrongStack/WrongStack

    A skill your agent uses whenever work can be split across multiple AI agents running in parallel, or when orchestrating leader/worker patterns in WrongStack.

    370 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Web Platform Baseline

    WrongStack/WrongStack

    Use this skill before asserting that a CSS, HTML or accessibility capability is available, unavailable, or the right tool — it carries dated, refreshable platform facts and refuses to let stale…

    370 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Wrongstack Mailbox

    WrongStack/WrongStack

    A skill your agent uses when the user wants to communicate with WrongStack's shared project mailbox from outside WrongStack — read messages sent by WrongStack agents, send replies, broadcast to all…

    370 GitHub stars~3.5k tokensUpdated today
    Auto-check passed

Questions about Evidence Audit

What does Evidence Audit do?

Evidence-led audit/fix rounds for any codebase in any language. Evidence Audit is an agent skill from WrongStack/WrongStack. Evidence-led audit/fix rounds for any codebase in any language.

When should I use Evidence Audit?

Evidence Audit fits situations like: the user asks to audit; find real bugs in; continue the audit round on; what is actually broken in a package.

How do I install Evidence Audit in Claude Code?

Run `npx skills add WrongStack/WrongStack --skill evidence-audit -a claude-code`. Or copy the skill folder (packages/core/skills/evidence-audit in WrongStack/WrongStack) into .claude/skills/evidence-audit in your project. Claude Code loads it when a task matches its description.

How do I install Evidence Audit in Codex?

Run `npx skills add WrongStack/WrongStack --skill evidence-audit -a codex`. Or copy the skill folder (packages/core/skills/evidence-audit in WrongStack/WrongStack) into .agents/skills/evidence-audit in your project. Codex loads it when a task matches its description.

Can I use Evidence Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add WrongStack/WrongStack --skill evidence-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evidence-audit, .gemini/skills/evidence-audit, .github/skills/evidence-audit and .opencode/skills/evidence-audit in your project.

What does Evidence Audit need to run?

Going by SKILL.md and its folder, Evidence Audit needs the command-line tools its instructions call (git).

Does Evidence Audit access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evidence Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evidence Audit use?

Evidence Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evidence Audit use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evidence Audit?

Skills that share tags, products or a category with Evidence Audit: Dep Validate (blokadaorg/blokada, 3.3k stars), Effect TS (mattiacerutti/supernova, 187 stars), Skill Writing (millionco/expect, 3.6k stars) and Go Rig (mudrii/openclaw-dashboard, 458 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evidence Audit?

WrongStack (a GitHub organization) maintains it in WrongStack/WrongStack, which has 370 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on October 7, 2026.

Source: WrongStack/WrongStack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.