Agent skill

Audit Analysis

by claesbackman in claesbackman/AI-research-feedback

Adversarially audit changed analysis code against a base ref, hunting for correctness errors in sample construction, merges, variable construction, silent failures, and clustering or fixed effects.

MITAuto-check: notesDevelopment

Install Audit Analysis

skills CLI
$ npx skills add claesbackman/AI-research-feedback --skill audit-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install claesbackman/AI-research-feedback audit-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/claesbackman/AI-research-feedback.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Skills/audit-analysis .claude/skills/audit-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit-analysis
GitHub stars
491
Token cost
~1.1k tokens
SKILL.md length
621 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Adversarially audit changed analysis code against a base ref, hunting for correctness errors in sample construction, merges, variable construction, silent failures, and clustering or fixed effects.

  • Works in 3 steps: Establish scope → Launch the auditor → Relay without softening
  • Tasks that involve Reproducible research
  • SKILL.md covers Phase 1: Establish scope, Phase 2: Launch the auditor and Phase 3: Relay without softening
  • Calls git

What it does

Audit Analysis is an agent skill from claesbackman/AI-research-feedback. Adversarially audit changed analysis code against a base ref, hunting for correctness errors in sample construction, merges, variable construction, silent failures, and clustering or fixed effects. Runs in an isolated subagent. Use before circulating results or submitting. This is not a reproducibility or paper-to-code review — use review-paper-code for that.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Reproducible research and Subagents. It works with Git. The repository describes itself as: A collection of Claude Code skills for academic research review. These tools were developed by Claes Bäckman. The licence is MIT.

When your agent uses it

  • Tasks that involve Reproducible research
  • Tasks that involve Subagents

Example prompts

  • “/audit-analysis”

Requirements

  • Pre-approved tools (allowed-tools): Bash, Read, Grep, Glob, Agent

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Establish scope
  2. Launch the auditor
  3. Relay without softening

What it can do on your machine

Read from SKILL.md and the folder at commit d129756. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Grep
    • Glob
    • Agent

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit Analysis loads about 1.1k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 621 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Grep, Glob, Agent

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from claesbackman/AI-research-feedback at commit d129756, republished under its MIT licence (© claesbackman). 621 words, ~1,116 tokens.

Download SKILL.mdSave it as .claude/skills/audit-analysis/SKILL.md (or your agent's skills folder).
name
audit-analysis
description
Adversarially audit changed analysis code against a base ref, hunting for correctness errors in sample construction, merges, variable construction, silent failures, and clustering or fixed effects. Runs in an isolated subagent. Use before circulating results or submitting. This is not a reproducibility or paper-to-code review — use review-paper-code for that.
allowed-tools
Bash, Read, Grep, Glob, Agent
argument-hint
[optional: base ref, default main]
disable-model-invocation
true

Audit Analysis Code

Find errors in changed empirical code before a referee does.

The audit runs in a subagent with a clean context. That isolation is the point: whoever wrote the code — including this session, if it helped — must not be able to steer the findings. Do not read the changed files yourself before launching, do not form a view, and do not answer the auditor's questions mid-run.

Phase 1: Establish scope

Set BASE from $ARGUMENTS if given, otherwise main.

Run, and stop with a short explanation if any of the first three fail:

  • git rev-parse --git-dir — must be a repository
  • git rev-parse --verify BASE — the base ref must exist
  • git diff --stat BASE — if empty, there is nothing to audit
  • git log BASE..HEAD --oneline — may legitimately be empty when the work is uncommitted, or when HEAD is BASE and only the working tree has changed. Note it and drop the commit-message check from the audit.

Report to the user in two or three lines: base ref, number of changed files, number of changed lines, and whether commit messages are available. Then launch immediately.

Phase 2: Launch the auditor

One Agent call, subagent_type: "general-purpose". Substitute BASE and pass this verbatim:

Review empirical research code adversarially. The author wants it broken now rather than by a referee. Read git log BASE..HEAD and git diff BASE, then the changed files in full. Follow variables built outside the diff.

Check, and report on each of:

  • Claims vs. code: do comments and commit messages match what runs? Quote both sides of any disagreement.
  • Sample: N before and after every filter, merge, and collapse. Take N from logs; write "N unverified" where there is no log. Flag undocumented drops.
  • Merges: key, uniqueness on the side that needs it, fate of unmatched observations, whether _merge is inspected, duplicate id-period pairs after.
  • Variables: trace every regressor and outcome. Units, logs vs. levels, deflation, lag alignment. Does construction match the name?
  • Silent failures: missings coerced to zero, if x > 0 true on missing, destring ... force, replace that changes nothing, loops that skip. In Python, fillna(0), silent dtype coercion, chained assignment.
  • Estimation: clustering level and cluster count, what the fixed effects absorb, weights, whether estimation N matches the sample traced above.

Each finding: file, line, quoted excerpt, what is wrong, consequence for the results. Tag CONFIRMED (visible in the code) or SUSPECTED (needs the data). Style and naming are not findings. Order by consequence, worst first, ten max. Then one line per category: what you found, or that you found nothing. Close with the one thing you could not check without the data. Change nothing.

Show full SKILL.md (185 more words)Show less

If the diff exceeds roughly 1,500 changed lines, run two auditors in parallel instead — one taking claims, sample, and merges, the other taking variables, silent failures, and estimation — and concatenate their findings. Do not split a smaller diff; the categories inform each other.

Phase 3: Relay without softening

Pass the findings through in the order returned, worst first. Do not reclassify a SUSPECTED finding as fine, do not add reassurance, and do not open with what the code gets right. The user asked for errors.

Drop any finding that lacks a file, a line, and a quoted excerpt, and tell the user how many you dropped. Unanchored findings are the failure mode this design exists to catch — an auditor told to find errors will manufacture them if nothing forces it to point at code.

Reproduce the per-category coverage lines verbatim, including the categories that came back clean, and the closing line about what could not be checked without the data. A clean category is a claim the auditor is on the record for.

Fix nothing. If the user wants repairs, that is a separate request.

© claesbackman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in Skills/audit-analysis of claesbackman/AI-research-feedback.

Open the folder on GitHubat commit d129756

Compare with similar skills

Audit Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit Analysis this skillclaesbackman/AI-research-feedback491—~1.1kAutomated safety check: NotesMIT
Cursor Composer Task DelegateChachamaru127/claude-code-harness3.2k—~4.4kAutomated safety check: NotesMIT
Firewood Reviewava-labs/firewood153—~2.1kAutomated safety check: NotesCustom licence
Tutti Architecture Reviewtutti-os/tutti3.8k—~2.3kAutomated safety check: PassApache-2.0
PR Monitoring Loopelastic/terraform-provider-elasticstack210—~4.1kAutomated safety check: PassMIT
Parallel Adversarial Change Reviewben-manes/caffeine18k—~1.9kAutomated safety check: NotesApache-2.0

Similar skills

  • Cursor Composer Task Delegate

    Chachamaru127/claude-code-harness

    Hands one implementation task to Cursor Composer in an isolated git worktree, then reviews its diff and cherry-picks the result into the main branch.

    3.2k GitHub stars~4.4k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes
  • Firewood Review

    ava-labs/firewood

    A skill your agent uses when reviewing ava-labs/firewood code changes — pull request or local workspace.

    153 GitHub stars~2.1k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Review tutti git diffs for project structure, layering, module ownership, and duplicate event-center infrastructure by planning focused architecture review tasks, then having the main agent…

    3.8k GitHub stars~2.3k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • PR Monitoring Loop

    elastic/terraform-provider-elasticstack

    Official

    Monitor GitHub pull requests through a subagent-based loop that watches CI checks, review comments, PR comments, review state, merge conflicts, and branch freshness.

    210 GitHub stars~4.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Runs three parallel reviewers on a diff or branch, one blind, one design-aware and one matching past bug patterns, then triages their findings.

    18k GitHub stars~1.9k tokensUpdated 2 days ago
    DevelopmentAuto-check: notes
  • Veomni Review

    ByteDance-Seed/VeOmni

    Pre-PR code review gate. An agent skill from ByteDance-Seed/VeOmni.

    2.2k GitHub stars~1.7k tokensUpdated 7 days ago
    DevelopmentAuto-check passed

More from claesbackman/AI-research-feedback

All 10 skills in this repo
  • Explorable Deck

    claesbackman/AI-research-feedback

    Build a Quarto reveal.js slide deck in the explorable-explanation style (Nicky Case) — one idea per slide, assertion titles, a concrete running example, run-time SVG stages the presenter drives…

    491 GitHub stars~2.9k tokensUpdated 12 days ago
    Auto-check passed
  • PDF To Markdown

    claesbackman/AI-research-feedback

    Split a PDF into chunks and convert it to readable markdown text.

    491 GitHub stars~1.8k tokensUpdated 12 days ago
    Auto-check passed
  • Review Paper Light

    claesbackman/AI-research-feedback

    Run a fast 2-agent pre-submission check for an economics paper — focuses on contribution, identification, and causal overclaiming.

    491 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check: notes
  • Paper Version

    claesbackman/AI-research-feedback

    Convert a LaTeX research paper into a policy brief, 1-page summary, or 5-page summary for a general audience, with factual review and a standalone HTML page for GitHub Pages.

    491 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check: notes
  • Review Grant

    claesbackman/AI-research-feedback

    Run a 6-agent pre-submission panel review for a grant proposal targeting a specified funder or program

    491 GitHub starsUsed in 1 repo~5.6k tokens
    Auto-check: notes
  • Review Paper Checks

    claesbackman/AI-research-feedback

    Run a fast 3-agent mechanical check of an economics paper — spelling and grammar, internal consistency and cross-references, and unsupported claims.

    491 GitHub stars~4.3k tokensUpdated 12 days ago
    Auto-check: notes

Works with

Categories

Questions about Audit Analysis

What does Audit Analysis do?

Adversarially audit changed analysis code against a base ref, hunting for correctness errors in sample construction, merges, variable construction, silent failures, and clustering or fixed effects. Audit Analysis is an agent skill from claesbackman/AI-research-feedback. Adversarially audit changed analysis code against a base ref, hunting for correctness errors in sample construction, merges, variable construction, silent failures, and clustering or fixed effects.

When should I use Audit Analysis?

Audit Analysis fits situations like: tasks that involve Reproducible research; tasks that involve Subagents.

How do I install Audit Analysis in Claude Code?

Run `npx skills add claesbackman/AI-research-feedback --skill audit-analysis -a claude-code`. Or copy the skill folder (Skills/audit-analysis in claesbackman/AI-research-feedback) into .claude/skills/audit-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Audit Analysis in Codex?

Run `npx skills add claesbackman/AI-research-feedback --skill audit-analysis -a codex`. Or copy the skill folder (Skills/audit-analysis in claesbackman/AI-research-feedback) into .agents/skills/audit-analysis in your project. Codex loads it when a task matches its description.

Can I use Audit Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add claesbackman/AI-research-feedback --skill audit-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-analysis, .gemini/skills/audit-analysis, .github/skills/audit-analysis and .opencode/skills/audit-analysis in your project.

What does Audit Analysis need to run?

Going by SKILL.md and its folder, Audit Analysis needs the command-line tools its instructions call (git). Its frontmatter pre-approves these tools: Bash, Read, Grep, Glob, Agent.

Does Audit Analysis access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Audit Analysis safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Audit Analysis use?

Audit Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit Analysis use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audit Analysis?

Skills that share tags, products or a category with Audit Analysis: Cursor Composer Task Delegate (Chachamaru127/claude-code-harness, 3.2k stars), Firewood Review (ava-labs/firewood, 153 stars), Tutti Architecture Review (tutti-os/tutti, 3.8k stars) and PR Monitoring Loop (elastic/terraform-provider-elasticstack, 210 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit Analysis?

claesbackman (a GitHub user) maintains it in claesbackman/AI-research-feedback, which has 491 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on September 25, 2026.

Source: claesbackman/AI-research-feedback on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.