Agent skill

Deep Audit

by pedrohcgs in pedrohcgs/claude-code-my-workflow

Comprehensive adversarial audit of a theory, proof, math/econ paper, codebase, or set of claims — decompose into components, fan out independent skeptics that must return CONCRETE defects…

MITAuto-check: notesAgent Workflows

Install Deep Audit

skills CLI
$ npx skills add pedrohcgs/claude-code-my-workflow --skill deep-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pedrohcgs/claude-code-my-workflow deep-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/deep-audit .claude/skills/deep-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deep-audit
GitHub stars
1.7k
Token cost
~2.9k tokens
SKILL.md length
1,488 words
Files
2 (incl. references)
Skills in repo
59
Repo updated
First seen
Licence
MIT

At a glance

Comprehensive adversarial audit of a theory, proof, math/econ paper, codebase, or set of claims — decompose into components, fan out independent skeptics that must return CONCRETE defects…

  • Correctness must be bulletproof and single-pass
  • SKILL.md covers When to reach for this, The method, Failure-mode lenses (adapt to… and The fresh-eyes pass (final gate), plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Round-by-round review is too slow and too shallow

What it does

Deep Audit is an agent skill from pedrohcgs/claude-code-my-workflow. Comprehensive adversarial audit of a theory, proof, math/econ paper, codebase, or set of claims — decompose into components, fan out independent skeptics that must return CONCRETE defects, adjudicate every finding with a separate judge, fix all confirmed defects, then re-verify. Use when correctness must be bulletproof and single-pass or round-by-round review is too slow and too shallow. Invoke for "audit this rigorously", "find ALL the bugs/gaps", "make this rock solid", "converge faster on correctness".

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/repo-infrastructure-audit.md`).

It sits in Agent Workflows. The repository describes itself as: A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols. The licence is MIT.

When your agent uses it

  • Correctness must be bulletproof and single-pass
  • Round-by-round review is too slow and too shallow

Example prompts

  • “audit this rigorously”
  • “find ALL the bugs/gaps”
  • “make this rock solid”
  • “/deep-audit”

Requirements

  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Write, Edit, Agent, Task

What it can do on your machine

Read from SKILL.md and the folder at commit ae72617. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Write
    • Edit
    • Agent
    • Task

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deep Audit loads about 2.9k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 130 tokens; SKILL.md has 1,488 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Write, Edit, Agent, Task

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pedrohcgs/claude-code-my-workflow at commit ae72617, republished under its MIT licence (© pedrohcgs). 1,488 words, ~2,890 tokens.

Download SKILL.mdSave it as .claude/skills/deep-audit/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
deep-audit
description
Comprehensive adversarial audit of a theory, proof, math/econ paper, codebase, or set of claims — decompose into components, fan out independent skeptics that must return CONCRETE defects, adjudicate every finding with a separate judge, fix all confirmed defects, then re-verify. Use when correctness must be bulletproof and single-pass or round-by-round review is too slow and too shallow. Invoke for "audit this rigorously", "find ALL the bugs/gaps", "make this rock solid", "converge faster on correctness".
allowed-tools
Read, Grep, Glob, Bash, Write, Edit, Agent, Task
disable-model-invocation
true
metadata.protocol
threat-prioritization

Deep adversarial audit

A convergent alternative to slow round-by-round review. Instead of one reviewer finding one or two issues per pass, fan out many independent skeptics over the whole artifact at once, adjudicate what they find, fix everything confirmed, and re-verify. Modeled on the multi-agent methodology behind hard formal-proof efforts (diverse independent portfolio, adversarial throughout, concrete evidence only, synthesize-challenge-repeat).

When to reach for this

  • The artifact is dense enough that a single review keeps surfacing new issues each pass (the tell that round-by-round is the wrong tool).
  • Correctness is the priority and the cost of a missed defect is high (a paper going to a top venue, a proof, a security-sensitive change, a migration).
  • The user asked to "fix ALL of it", "be deeper", "converge faster", "100% rock solid".

Requires the user to have opted into multi-agent orchestration (they asked for a workflow / deep audit / to fan out agents, or ultracode is on). If they haven't, propose it and its rough cost first.

The method

1. Decompose (diverse portfolio). Break the artifact into components by idea, not by section: each independent claim, lemma, estimator, subsystem, invariant. Add cross-cutting failure-mode lenses (see below). Aim for coverage such that every load-bearing claim is attacked by at least one agent that is looking straight at it. Don't tell the agents your favored reading — preserve independence so they don't all converge on the same attractive-but-wrong conclusion.

2. Fan out adversarial finders (one per component). Each finder is prompted to refute, defaulting to "there is a bug," and must ground every claim in the actual text/code (read it, don't paraphrase from memory). Hard rules, borrowed from what works:

  • Concrete findings only. Every finding = exact location (file:line / label + quoted text) + one-sentence defect + a failing case (specific inputs/configuration → wrong output, or the exact missing hypothesis).
  • Reject status reports, "looks fine", "this is standard/routine", vague optimism, and "the global step is straightforward."
  • A fix that re-imposes the same difficulty elsewhere, or assumes its own conclusion, is not a fix — flag it.
  • If, after genuinely attacking, nothing is found, the agent must state the specific attacks it ran and why each closed — not just "clean."

3. Adjudicate every finding (independent judge). A separate judge re-opens each cited location and decides CONFIRMED / REFUTED / DOWNGRADED, skeptical of both the artifact and the finding. This kills false positives (misreads, hypotheses that are actually present elsewhere, failing cases that don't arise under the stated conditions) — the step that keeps the fix list honest.

4. Synthesize. Dedup by location, rank fatal > major > minor, and hand back one clean defect list. Nothing is accepted as an issue until it survives this.

5. Fix all confirmed, then re-verify. Apply every confirmed fix (you, in the main loop — fixing needs care and judgment). Then re-audit the touched spots and check that no fix created a new defect. Repeat waves until two consecutive audit passes come back empty (fallback cap: 5 waves; a finding that survives waves N and N+2 goes to the user rather than a third patch). Don't stop after the first wave.

Failure-mode lenses (adapt to domain)

Beyond per-component attacks, sweep these cross-cutting modes explicitly — they are where real defects hide:

  • Overclaim: the headline/abstract claims more than the theorems/tests actually deliver.
  • Scope creep in a proof: a pointwise result used where a uniform one is needed; a both-correct property stated unqualified; a special-case argument invoked generally.
  • Silent hypotheses: a differentiability/density/continuity/boundedness/positivity condition used but never stated (in math), or an un-checked precondition/invariant (in code).
  • Edge cases: atoms/ties, endpoints/unbounded support, empty/degenerate inputs, boundary of the parameter space.
  • Circularity: an assumption that assumes its own conclusion; a result that cites itself; a citation that gives less than claimed (read the cited source).
  • Internal contradiction: a definition/notation used two ways; a table cell contradicting a proposition; main text vs appendix disagreement; a dangling/wrong cross-reference.

The fresh-eyes pass (final gate)

Every targeted wave inherits the blind spots of whoever wrote its prompts: focus hints, fix history, and expected failure modes all prime the auditors toward known territory. After all targeted waves and fixes are done, run one cold audit with little to no context: independent auditors given ONLY the artifact and a minimal instruction ("find concrete defects: location + failing case"), with no cluster assignments, no history, no special-focus lists. Diversify only the entry point (main-text-first as a journal referee would; appendix-first; tables/claims-first; a single deep dive of the auditor's own choosing). Adjudicate as usual. Clean fresh-eyes pass + clean targeted coverage + green mechanical battery is the closure standard; a fresh-eyes finding that targeted waves missed is also a diagnosis of the prompt set — add the missed failure mode to the lenses.

Full inventory — never sample

For a paper/proof artifact: enumerate every formal statement first (grep \begin{theorem|proposition|lemma|corollary} + labels) and assign each proof to a verifier — coverage must be 100% of load-bearing statements, not "a few proofs of the reviewer's choice." Sampling converges linearly and stochastically; inventories converge in one wave. Group tightly-coupled small lemmas into clusters; big proofs get their own verifier. Each verifier returns, besides findings, a steps-verified list and a hypotheses ledger (used-vs-stated; used-but-unstated is a finding).

Show full SKILL.md (641 more words)Show less

The mechanical battery (the highest-yield check)

Written arguments can read soundly while the object they define is wrong. For every estimating equation, influence-function identity, identification claim, and population moment, write an executable check that computes the population object on adversarial toy designs — truncation (censoring endpoint below the outcome endpoint), interior atoms, misspecified nuisances, boundary/overlap failure — and asserts the claimed centering/identity numerically (analytic or fine-grid/large-N with fixed seed). Keep the scripts as a permanent test directory in the repo with a README; rerun after any change to the corresponding formula. A 5-line population computation catches classes of defects (tail-renormalized roots, sign flips, mass-deficit weighting) that neither careful reading nor model consensus reliably finds.

Fix hygiene

  • New math introduced by fixes is un-audited math: every fix wave is followed by a verification wave over exactly the fixed spots before anything is declared closed.
  • In workflow synthesis, match findings to verdicts by INDEX (require the judge to return verdicts in the findings' order), never by location string — judges paraphrase locations and silent drops follow. Verify the synthesized summary against the journal before acting on it.

Blocked routes are outcomes

If a component cannot be fixed under the stated assumptions, that is a finding, not a failure of the audit: report the exact remaining gap (the precise missing hypothesis or broken step) and the honest options (weaken the claim, add the hypothesis, restrict scope). Do not search for a favorable reading, and do not let an agent paper over a theorem-strength gap as "routine."

Orchestration

  • Fan out with parallel Agent calls in one message (one finder per component), then one judge per component over that component's findings, then synthesize; see orchestrator-protocol.md. Where the Workflow tool is available (e.g. an ultracode session), pipeline(components, finder, judge) is an optional accelerator that judges each component's findings the moment its finder returns (no barrier); it is never a requirement. Return the confirmed list; do the fixing yourself afterward.
  • Model division: finders = the strong execution/analysis model (find and attack); judges = the strong adjudication model. Match to the local convention in model-routing.md (here: both roles run on the Opus tier, and a judge gets more effort before any change of tier; the Fable tier is a per-session choice, never the fleet default). Keep judges to one-per-component (adjudicating all that component's findings at once) to conserve the judging budget.
  • Set finder effort high; give each the exact labels/locations to read and its specific attack list.
  • Persist. Don't return "best effort" or a list of why it's hard. Return the confirmed defects (and, once fixed, a clean re-audit) — or the single strongest remaining gap stated exactly.

Prompt skeletons

Finder: "You are a HOSTILE referee auditing ONE component. Read the ACTUAL text at {locations}. Attack: {failure modes}. Return CONCRETE findings only (location + defect + failing case); no 'looks fine'/'routine'/vague. A fix that re-imposes the difficulty isn't a fix. If clean, list the specific attacks you ran and why each closed."

Judge: "An adversarial referee returned these findings on component X. For each, open the cited location, verify against what the text ACTUALLY says and its proof, mark CONFIRMED/REFUTED/DOWNGRADED. Skeptical of both the artifact and the finding."

Shared doctrine lives in one place

Three things this audit depends on are not restated here, because they are the same rules every other verification surface uses and a third copy would drift:

Auditing this repository itself

For the repo-infrastructure application — surface-sync, skill/agent/rule integrity, hook and script review, doc-vs-reality drift — see references/repo-infrastructure-audit.md. Start with ./scripts/backtest.sh: the mechanical battery is already written, and an agent should never hand-check what a script decides.

© pedrohcgs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .claude/skills/deep-audit of pedrohcgs/claude-code-my-workflow.

  • SKILL.md
  • references/repo-infrastructure-audit.md

Open the folder on GitHubat commit ae72617

Compare with similar skills

Deep Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deep Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deep Audit this skillpedrohcgs/claude-code-my-workflow1.7k—~2.9kAutomated safety check: NotesMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
Hook Development for Claude Code Pluginsanthropics/claude-plugins-official38k10 repos~4.1kAutomated safety check: NotesApache-2.0
Using Superpowersfarm-fe/farm5.6k35 repos~1.4kAutomated safety check: PassMIT
Executing Plans Inlineobra/superpowers297k2 repos~5.1kAutomated safety check: PassMIT
Skill CreatorAzure/azqr79589 repos~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Hook Development for Claude Code Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.

    38k GitHub starsUsed in 10 repos~4.1k tokens
    Agent WorkflowsAuto-check: notes
  • Using Superpowers

    farm-fe/farm

    A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions

    5.6k GitHub starsUsed in 35 repos~1.4k tokens
    Agent WorkflowsAuto-check passed
  • Executing Plans Inline

    obra/superpowers

    Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.

    297k GitHub starsUsed in 2 repos~5.1k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    795 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed

More from pedrohcgs/claude-code-my-workflow

All 59 skills in this repo
  • Devils Advocate

    pedrohcgs/claude-code-my-workflow

    Adversarial 5-7 question challenge to a deck's pedagogical choices — ordering, prerequisites, cognitive load, motivation.

    1.7k GitHub starsUsed in 2 repos~641 tokens
    Auto-check passed
  • Vaccinate

    pedrohcgs/claude-code-my-workflow

    Qualify a check before it is allowed to clear anything — prove it can detect the failure it is meant to catch.

    1.7k GitHub stars~2.1k tokensUpdated 12 days ago
    Auto-check: notes
  • Compile Latex

    pedrohcgs/claude-code-my-workflow

    Compile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex).

    1.7k GitHub starsUsed in 1 repo~492 tokens
    Auto-check: notes
  • Context Status

    pedrohcgs/claude-code-my-workflow

    Show current context status and session health. An agent skill from pedrohcgs/claude-code-my-workflow.

    1.7k GitHub starsUsed in 1 repo~613 tokens
    Auto-check: notes
  • Capture Environment

    pedrohcgs/claude-code-my-workflow

    Snapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt /…

    1.7k GitHub stars~2.8k tokensUpdated 12 days ago
    Auto-check: notes
  • Checkpoint

    pedrohcgs/claude-code-my-workflow

    Save a structured state snapshot before stopping or handing off.

    1.7k GitHub stars~2.8k tokensUpdated 12 days ago
    Auto-check: notes

Categories

Questions about Deep Audit

What does Deep Audit do?

Comprehensive adversarial audit of a theory, proof, math/econ paper, codebase, or set of claims — decompose into components, fan out independent skeptics that must return CONCRETE defects…. Deep Audit is an agent skill from pedrohcgs/claude-code-my-workflow. Comprehensive adversarial audit of a theory, proof, math/econ paper, codebase, or set of claims — decompose into components, fan out independent skeptics that must return CONCRETE defects, adjudicate every finding with a separate judge, fix all confirmed defects, then re-verify.

When should I use Deep Audit?

Deep Audit fits situations like: correctness must be bulletproof and single-pass; round-by-round review is too slow and too shallow.

How do I install Deep Audit in Claude Code?

Run `npx skills add pedrohcgs/claude-code-my-workflow --skill deep-audit -a claude-code`. Or copy the skill folder (.claude/skills/deep-audit in pedrohcgs/claude-code-my-workflow) into .claude/skills/deep-audit in your project. Claude Code loads it when a task matches its description.

How do I install Deep Audit in Codex?

Run `npx skills add pedrohcgs/claude-code-my-workflow --skill deep-audit -a codex`. Or copy the skill folder (.claude/skills/deep-audit in pedrohcgs/claude-code-my-workflow) into .agents/skills/deep-audit in your project. Codex loads it when a task matches its description.

Can I use Deep Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pedrohcgs/claude-code-my-workflow --skill deep-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deep-audit, .gemini/skills/deep-audit, .github/skills/deep-audit and .opencode/skills/deep-audit in your project.

What does Deep Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Deep Audit is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Write, Edit, Agent, Task.

Does Deep Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Deep Audit safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Deep Audit use?

Deep Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deep Audit use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.

What are the alternatives to Deep Audit?

Skills that share tags, products or a category with Deep Audit: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 297k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deep Audit?

pedrohcgs (a GitHub user) maintains it in pedrohcgs/claude-code-my-workflow, which has 1,653 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on September 27, 2026.

Source: pedrohcgs/claude-code-my-workflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.