Audit Replication
brycewang-stanford/Auto-Empirical-Research-Skills
Validate the replication package for the sewage-house-prices project.
Enforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs.
$ npx skills add pedrohcgs/claude-code-my-workflow --skill audit-reproducibility -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install pedrohcgs/claude-code-my-workflow audit-reproducibility --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/audit-reproducibility .claude/skills/audit-reproducibility && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "audit-reproducibility" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/audit-reproducibility into .claude/skills/audit-reproducibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-reproducibility", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/audit-reproducibilityType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add pedrohcgs/claude-code-my-workflow --skill audit-reproducibility -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install pedrohcgs/claude-code-my-workflow audit-reproducibility --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/audit-reproducibility .agents/skills/audit-reproducibility && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "audit-reproducibility" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/audit-reproducibility into .agents/skills/audit-reproducibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-reproducibility", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pedrohcgs/claude-code-my-workflow --skill audit-reproducibility -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install pedrohcgs/claude-code-my-workflow audit-reproducibility --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/audit-reproducibility .cursor/skills/audit-reproducibility && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "audit-reproducibility" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/audit-reproducibility into .cursor/skills/audit-reproducibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-reproducibility", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/pedrohcgs/claude-code-my-workflow.git --path .claude/skills/audit-reproducibility--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add pedrohcgs/claude-code-my-workflow --skill audit-reproducibility -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install pedrohcgs/claude-code-my-workflow audit-reproducibility --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/audit-reproducibility .gemini/skills/audit-reproducibility && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "audit-reproducibility" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/audit-reproducibility into .gemini/skills/audit-reproducibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-reproducibility", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install pedrohcgs/claude-code-my-workflow audit-reproducibilityInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add pedrohcgs/claude-code-my-workflow --skill audit-reproducibility -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/audit-reproducibility .github/skills/audit-reproducibility && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "audit-reproducibility" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/audit-reproducibility into .github/skills/audit-reproducibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-reproducibility", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pedrohcgs/claude-code-my-workflow --skill audit-reproducibility -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install pedrohcgs/claude-code-my-workflow audit-reproducibility --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/audit-reproducibility .opencode/skills/audit-reproducibility && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "audit-reproducibility" agent skill from https://github.com/pedrohcgs/claude-code-my-workflow/tree/main/.claude/skills/audit-reproducibility into .opencode/skills/audit-reproducibility/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-reproducibility", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
audit-reproducibilityEnforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs.
Audit Reproducibility is an agent skill from pedrohcgs/claude-code-my-workflow. Enforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Report PASS/FAIL per claim against tolerance thresholds. Use before submission and before releasing a replication package.
Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `evals/cases/manuscript-not-oracle.md` and `evals/cases/tolerance-before-comparison.md`).
It sits in Research & Science, covering Econometrics and empirical research, Reproducible research and Database administration. It works with Python. The repository describes itself as: A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ae72617. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGrepGlobWriteBashAgentTaskMonitorFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Audit Reproducibility loads about 6.4k tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 2,738 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Grep, Glob, Write, Bash, Agent, Task, MonitorAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from pedrohcgs/claude-code-my-workflow at commit ae72617, republished under its MIT licence (© pedrohcgs). 2,738 words, ~6,429 tokens.
.claude/skills/audit-reproducibility/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Compare numeric claims in a manuscript (point estimates, standard errors, p-values, counts) against the actual outputs produced by the analysis pipeline. Report PASS / FAIL per claim against the tolerance thresholds defined in .claude/rules/replication-protocol.md.
Core principle: If the paper says ATT = -1.632 (0.584) and the code produces -1.628 (0.591), we verify — numerically — that the difference is within the documented tolerance. No more "looks close enough" eyeballing.
Two directions, not one. Vertically, each claim is checked against the output that produced it. Horizontally, it is checked against every other artifact that displays the same number — the supplement table, the slide deck, the poster. The vertical check passes contentedly while a deck quotes last month's value; only the horizontal one catches that. Declared displays live in the passport's appears_in list — see replication-protocol.md → The horizontal check.
/commit. Pair with a pre-commit invocation on manuscript + analysis changes.$0 — path to the manuscript (.tex, .qmd, .md, .pdf). Required.$1 — path to the outputs directory. Defaults to output/, where every language's pipeline writes (R, Stata via /stata-replication, Python). Recognised alternatives: _targets/objects/ (R targets workflows), any directory the user-specified outputs live in. If output/ does not exist but a pre-v2.6 scripts/<lang>/_outputs/ does, use that and say so in the report.replication-protocol.md for the tolerance thresholds currently in effect.Rscript scripts/R/00_run_all.R) before auditing.sessionInfo.txt or equivalent environment capture exists in the outputs dir.Parse the manuscript for numeric claims. Patterns to match:
ATT = -1.632 (0.584), $\beta = 0.342$ (0.091), hat{\tau} = 1.28** with starred significance& -1.632$^{***}$ & 0.584 & in LaTeX table environmentsour sample of 2,847 firms, $N = 2{,}847$mean = 0.423, SD = 0.087p < 0.01, $p = 0.003$Record each claim as a tuple:
{
claim_id: "Table2_col3_ATT",
location: "Table 2, Column 3, row 'Treatment'",
kind: "point_estimate" | "standard_error" | "p_value" | "count" | "percentage",
reported_value: -1.632,
uncertainty: 0.584, # only for point estimates
significance_stars: 3, # 0-3 or None
raw_context: "the ATT estimate of -1.632 (0.584) indicates..."
}Write the extracted claims to quality_reports/reproducibility_claims_[manuscript-name].json so the user can review the extraction before audit.
Scan $1 for corresponding values. Priority order:
.rds files — readRDS(path)$coef[["treatment"]] style lookups. Can use Rscript -e "saveRDS(summary(readRDS(...)), '/tmp/audit.rds')" to extract..tex tables — parse LaTeX table cells directly; match on column headers + row labels..csv summary files — pandas/readr parse, key-value lookup..out / .log files (Stata, regress output) — regex extraction..json — direct key lookup.Record each extracted result:
{
source: "output/results.rds",
lookup_key: "fit_main$coefficients['treated']",
value: -1.628,
uncertainty: 0.591,
p_value: 0.005
}For each passport claim, read its appears_in list and pull the value as displayed at each entry:
{
claim_id: "C3",
displays: [
{ path: "manuscript.tex", locator: "Table 1, Col 2", display_precision: 3, shown: 0.342 },
{ path: "Slides/Lecture04_Results.tex", locator: "frame 'Main result'", display_precision: 2, shown: 0.34 }
]
}Rules for this pass:
location: is one of the displays, not a separate thing — if appears_in repeats it, that is one display, not two.locator — is recorded as shown: NOT_FOUND. It resolves to FAIL in Phase 4c. Never skip it: a locator that has quietly stopped matching is exactly where a stale number hides.appears_in list is horizontally unchecked, not horizontally clean. Report it that way (see Phase 5) rather than silently passing it.Use fuzzy heuristics when exact labels don't match:
"treatment effect" ~ "ATT" ~ "treated")raw_context field (table number, row label, description)For every claim, produce a match candidate with a confidence score. Claims below 0.7 confidence get flagged as "UNMATCHED — manual review needed" rather than silently passing.
For each matched claim, apply the thresholds from replication-protocol.md:
| Kind | Tolerance | Example |
|---|---|---|
| Integers (N, counts) | Exact | 2,847 must equal 2,847 |
| Point estimates | abs(reported - computed) < 0.01 | -1.632 vs -1.628 → diff = 0.004 → PASS |
| Standard errors | abs(reported - computed) < 0.05 | 0.584 vs 0.591 → diff = 0.007 → PASS |
| P-values | Same significance level | p<0.01 and p<0.01 → PASS; p<0.01 and p=0.03 → FAIL |
| Percentages | ±0.1pp | 42.3% vs 42.35% → PASS |
Respect any tolerance overrides the user has written into their replication-protocol.md fork (they may loosen for MC noise or tighten for administrative data).
A tolerance check resolves to one of four dispositions:
A mismatch is not automatically a failure. In applied work the most common out-of-tolerance result is a defensible alternative spec, not a bug — reghdfe vs feols clustering df, a different bandwidth-selection rule, a different MC seed/reps, or display rounding. The skill's job is to stage the disagreement for a human auditor, not to pronounce the code right and the paper wrong. (The df-adjustment note in "Stata-specific notes" below is the canonical example of a named alternative.)
The manuscript is not the oracle. When the computed value disagrees with the manuscript, do not presume the code is correct and the paper stale — nor the reverse. A refactor may have broken a previously-correct table (the on-disk output is the buggy one), or the paper may carry an old number. The computed value is a challenger, not ground truth. Report a mismatch as "one of {paper, code} must change — isolate which," never "revert the code to match the paper." This prevents the trap of reverting a genuine bug-fix just to make the paper 'reproduce.'
A FAIL may be downgraded to EXPLAINED only when a specific named alternative is recorded for that exact claim — in the passport entry's notes: field (passport mode) or the audit report's author-note column (default mode). Example of a valid note:
"reghdfe vs feols clustering-df adjustment; under the reghdfe small-sample correction the published value is −1.19, within rounding of the script's −1.187. CODE-CORRECTED pending."
The author is the auditor: the skill stages the two-sided comparison (reported value and computed value, both shown); the human writes the one-line named alternative; the skill records it and thereafter respects it. Tag the resolution PAPER-CORRECTED, CODE-CORRECTED, or DEFENSIBLE-ALTERNATIVE.
Hard floor — never downgradable to EXPLAINED:
(Citation/existence claims are out of scope here — /verify-claims owns those, and applies the same named-alternative softening on its side.)
Reuse the two-strikes rule from review-paper --adversarial and summary-parity.md: if the same claim is downgraded to EXPLAINED in two consecutive audits without ever being corrected to PASS (the author keeps invoking the alternative but never updates paper or code), stop treating it as quietly resolved. Surface it prominently in Phase 5 — "this contested number has been EXPLAINED twice but never corrected" — so a standing disagreement can't hide behind a recorded note indefinitely. In passport mode, detect this by comparing the current status/notes against the prior audit's.
Phase 4 compared the claim to the code. This phase compares the claim's displays to each other, pairwise over the displays collected in Phase 2b:
display_precision values. 0.34 in a deck against 0.342 in the paper agrees at 2 decimals; 0.29 against 0.342 does not.tolerance: block governs the vertical comparison, where two measurements are compared. Two displays of one number are two copies — after the rounding in step 1 they either match or they do not.NOT_FOUND. Never downgradable to EXPLAINED: a named alternative spec explains why code and paper differ and has nothing to say about why two copies of one number differ. If a display genuinely shows a different specification, it is a different claim and belongs in its own passport entry.appears_in list. Reported, non-blocking; the fix is to declare the displays.replication-protocol.md → Anti-patterns).Write quality_reports/reproducibility_audit_[manuscript-name].md:
# Reproducibility Audit: [Manuscript Title]
**Date:** [YYYY-MM-DD]
**Manuscript:** [path]
**Outputs directory:** [path]
**Tolerance source:** .claude/rules/replication-protocol.md
## Summary
| Status | Count |
|---|---|
| PASS (vertical and horizontal) | N |
| FAIL (diff > tolerance, no named alternative) | M |
| EXPLAINED (out of tolerance, named alternative recorded) | E |
| UNMATCHED (manual review) | K |
| HORIZONTAL DRIFT (declared displays disagree, or a display not found) | H |
| HORIZONTALLY UNCHECKED (claim declares no `appears_in`) | U |
| **Overall verdict** | **PASS / FAIL** (FAIL iff M > 0 or H > 0; EXPLAINED does not fail the audit) |
## PASS (all within tolerance)
| Claim | Reported | Computed | Diff | Tolerance |
|---|---|---|---|---|
| Table2_col3_ATT | -1.632 (0.584) | -1.628 (0.591) | 0.004 / 0.007 | 0.01 / 0.05 |
## FAIL (outside tolerance — BLOCKER)
| Claim | Reported | Computed | Diff | Tolerance | Location in paper | Author note (name a concrete alternative to downgrade → EXPLAINED) |
|---|---|---|---|---|---|---|
## EXPLAINED (out of tolerance; defensible named alternative recorded — non-blocking, carry into response-to-referees)
| Claim | Reported | Computed | Named alternative (why the gap is defensible) | Resolution |
|---|---|---|---|---|
| Table3_col2_ATT | -1.187 | -1.19 | reghdfe vs feols clustering-df adjustment | DEFENSIBLE-ALTERNATIVE |
## UNMATCHED (manual review)
| Claim | Raw context | Candidate sources |
|---|---|---|
## HORIZONTAL DRIFT (one number, two values — BLOCKER)
| Claim | Display A | Display B | Compared at | Values | Which side moved |
|---|---|---|---|---|---|
| C3 | manuscript.tex — Table 1, Col 2 | Slides/Lecture04_Results.tex — frame 'Main result' | 2 decimals | 0.34 vs 0.29 | deck older than the claim's output_file |
## HORIZONTALLY UNCHECKED (no `appears_in` declared)
| Claim | Primary display | Other artifacts to declare |
|---|---|---|
## Environment
[sessionInfo excerpt]
## Next steps
1. Resolve each FAIL row — either correct the manuscript, rerun the analysis, or (if the gap is a defensible alternative spec) record a concrete named alternative to downgrade it to EXPLAINED.
2. Resolve each HORIZONTAL DRIFT row — regenerate the lagging display from the claim's output, never by retyping the other display's value. A `NOT_FOUND` display is fixed by correcting the `locator`, not by dropping the entry.
3. Review UNMATCHED rows — add explicit lookup keys or widen the search scope.
4. Declare `appears_in` for the HORIZONTALLY UNCHECKED rows — every artifact a reader sees the number in.
5. Review EXPLAINED rows before submission — each should map to a sentence in the response-to-referees.
6. After zero FAILs and zero HORIZONTAL DRIFT (EXPLAINED rows allowed), the paper is replication-ready.The report above and the passport are rewritten on every run, so neither keeps a history. The log does. After every run — default and passport mode alike — append one block to quality_reports/replication-log.md. It is committed, so a co-author or data editor can see what was checked, against which commit, and how each number was obtained, without reading the code.
>>; never edit or reorder a past block. A correction is a new run, not an edit. The repo-hygiene gate fails a commit that edits or removes a committed line — at the pre-commit hook against the last commit, and in CI against the branch the work merges into. It proves no entry was edited, not that every run was logged. A disclosure redaction is the one exception: commit it with ALLOW_LOG_REWRITE=1 and the reason in the commit message.-dirty when anything outside quality_reports/ is uncommitted (untracked files included) — so a verdict on uncommitted code says so, and the audit's own report files do not trigger it. Outside git, or before the first commit, the stamp is no-commit.confidential-data.md).LOG=quality_reports/replication-log.md
[ -f "$LOG" ] || printf '%s\n' "# Replication log" "" \
"Append-only record of /audit-reproducibility runs: what was checked, against which commit, and how each number was computed. Never edit a past entry; a correction is a new run." "" > "$LOG"
# -dirty = anything uncommitted outside quality_reports/, untracked files included (the audit writes its own files in quality_reports/)
if REV=$(git rev-parse --short HEAD 2>/dev/null); then
[ -n "$(git status --porcelain -- . ':(exclude)quality_reports' 2>/dev/null)" ] && REV="$REV-dirty"
else
REV="no-commit" # not a git repository, or nothing committed yet
fi
printf '## %s — %s @ %s\n\n' "$(date +%F)" "<manuscript path>" "$REV" >> "$LOG"
# Quoted heredoc: nothing below is expanded, so an accessor's `$` or backtick is written as-is.
cat >> "$LOG" <<'EOF'
Outputs: `<outputs dir>` · Verdict: **<PASS|FAIL>** (<M> FAIL, <H> horizontal drift, <E> EXPLAINED, <K> unmatched)
| Claim | Location | Reported | Computed | How computed | Tolerance | Verdict |
|---|---|---|---|---|---|---|
| Table2_col3_ATT | main.tex: Table 2, col 3 | -1.632 | -1.628 | `readRDS("output/results.rds")$coef[["treatment"]]` | 0.01 | PASS |
EOFOne row per audited claim, in the order of the report.
/commit pre-commit gate — see replication-protocol.md for the enforcement pattern. EXPLAINED rows do NOT count as FAIL and never trigger exit 1 — they are surfaced, not blocking. The gate keeps its full teeth for genuine FAILs (no named alternative) and for fabricated/UNMATCHED claims.The skill compares manuscript claims against outputs in three source-language ecosystems. All three write to the same output/ directory:
| Source | Default outputs dir | Read-output via | Common claim sources |
|---|---|---|---|
| R (default) | output/ | readRDS(), arrow::read_parquet(), vroom::vroom() | .rds / .parquet / .csv / tinytable .tex |
| Stata (v1.9.0) | output/ | haven::read_dta() from R, or pyreadstat.read_dta() from Python | .dta / esttab .tex / .smcl log values |
| Python | output/ (or _targets/) | pandas.read_parquet, pickle.load | .parquet / .pickle / .csv |
Stata-specific notes (v1.9.0):
.dta outputs are read via haven::read_dta() (R), pyreadstat.read_dta() (Python), or by parsing the corresponding esttab .tex if the table-cell value is what the manuscript cites.\input{output/tab_main.tex} is the strongest provenance signal — the cell value comes mechanically from the .do file. Match the location in the .tex to the regression call in 03_analyze.do.reghdfe and base reg, cluster(). If a SE mismatches at the 2nd decimal, the tolerance in replication-protocol.md covers it; if it mismatches at the 1st decimal, investigate the df adjustment.When quality_reports/passports/<paper-slug>.yaml exists, the skill operates in passport mode: instead of emitting a one-shot report, it reads, updates, and rewrites the passport file in place.
claims: entry in the passport, perform the same numeric audit as the default mode (extract reported value from manuscript at location, locate computed value at source_file:source_line / output_file:output_field, compare against tolerance:), then the horizontal sweep of Phases 2b and 4c over the entry's appears_in list.status in place:notes does not name a concrete alternative; or two declared displays disagree; or an appears_in entry could not be located in its file. Record the discrepancy in notes — reported vs computed for a vertical FAIL, the two display values and their paths for a horizontal one. Blocks (exit 1).notes already records a specific named alternative spec (not blank, not "unclear"). The skill reads notes on its next run and resolves the same out-of-tolerance claim to EXPLAINED instead of FAIL — surfaced, non-blocking. The hard floor still applies: an UNMATCHED claim or a note without a named alternative stays FAIL.source_file, output_file, or any appears_in path has a modification time later than last_verified_on, mark STALE and re-run the audit logic (after the rerun, status becomes PASS / FAIL / EXPLAINED — STALE is transient).last_verified_on and last_verified_by: "/audit-reproducibility" per claim.paper.last_audit at the top level.If a claim in the manuscript is detected that has no matching passport entry, emit an UNVERIFIED warning — the author should add it (passport scope is author-curated, not auto-populated, to avoid bad inferences).
Passport mode does NOT delete passport entries. If a claim disappears from the manuscript, the passport entry remains with a STALE status — the author decides whether to delete (claim retracted) or update the entry's location (claim moved).
Nor does it add appears_in entries on its own. A display the sweep happens to notice — the same value in a deck or a supplement that the passport never declared — is reported as a suggestion in the HORIZONTALLY UNCHECKED table; the author declares it. Same reasoning as the no-auto-populate rule above: an inferred display that is actually a different quantity would fail the horizontal check forever.
See .claude/rules/replication-protocol.md "Claims Provenance: passport.yaml" for the full schema and integration points (/commit, /review-paper).
.claude/rules/replication-protocol.md — the tolerance contract + passport schema.templates/passport-template.yaml — starter file to copy for a new paper; the appears_in list is where displays are declared..claude/hooks/claim-reconcile.py — the event-driven nudge: writing a tracked script, output, or declared display names the claims to re-audit. It counts declarations; this skill does the comparing..claude/skills/review-r/SKILL.md — catches code-style issues; this skill catches NUMERICAL reproducibility..claude/skills/diagnose/SKILL.md — when a claim resolves to FAIL and you need to localize which pipeline step produced the out-of-tolerance value, hand off to /diagnose (single-claim root-cause: reproduce → minimise → bisect)..claude/skills/review-paper/SKILL.md — content review; pair with this skill for a full pre-submission audit..claude/skills/replication-package/SKILL.md — gates on this skill before assembling the AEA DCAS deposit..claude/skills/capture-environment/SKILL.md · .claude/skills/disclosure-check/SKILL.md — environment capture + restricted-data screening downstream.-1.632 is reproducible. Whether -1.632 is the RIGHT estimand is a review-paper / domain-reviewer question.sessionInfo.txt capture lets a reviewer see the env; pinning versions is on the user (via renv.lock or a DESCRIPTION file).When /audit-reproducibility is asked to verify all numeric claims in a paper, the safest approach is to re-run the full pipeline (00_run_all.R or equivalent) and compare the regenerated outputs to the manuscript values. For pipelines that take more than a couple of minutes, background-launch the rerun and use Anthropic's Monitor tool (Apr 2026 Week 15) to stream stdout. The audit can react to errors mid-stream rather than waiting for the entire pipeline to finish before noticing a failed step.
© pedrohcgs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in .claude/skills/audit-reproducibility of pedrohcgs/claude-code-my-workflow.
Open the folder on GitHubat commit ae72617
Audit Reproducibility next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Audit Reproducibility this skillpedrohcgs/claude-code-my-workflow | 1.7k | — | ~6.4k | Automated safety check: Notes | MIT | |
| Audit Replicationbrycewang-stanford/Auto-Empirical-Research-Skills | 4.6k | — | ~984 | Automated safety check: Notes | Custom licence | |
| Data Depositbrycewang-stanford/Auto-Empirical-Research-Skills | 4.6k | — | ~1.2k | Automated safety check: Notes | Custom licence | |
| Ectheory Replication And Data Policyfranklee16/academic-research-skills | 223 | 1 repos | ~983 | Automated safety check: Pass | None | |
| Qe Replication And Data Policyfranklee16/academic-research-skills | 223 | 1 repos | ~1.2k | Automated safety check: Pass | None | |
| Stata C Pluginsdylantmoore/stata-skill | 291 | 1 repos | ~5.8k | Automated safety check: Pass | Custom licence |
brycewang-stanford/Auto-Empirical-Research-Skills
Validate the replication package for the sewage-house-prices project.
brycewang-stanford/Auto-Empirical-Research-Skills
Prepare a replication package for the sewage-house-prices project.
franklee16/academic-research-skills
A skill your agent uses to handle reproducibility and the supplementary-material file for an Econometric Theory (ET) paper — ET has no mandatory data/code archive (theory journal); the relevant…
franklee16/academic-research-skills
A skill your agent uses to assemble a Quantitative Economics (QE) replication package that passes the Econometric Society Data Editor's pre-acceptance reproducibility check — raw data, code…
dylantmoore/stata-skill
Develop high-performance C/C++ plugins for Stata using the stplugin.h SDK.
yushui2022/MathModel-Skill
Generates result-evidence contracts, tables and runnable q1 to q3 modeling code scaffolds for a math modeling paper from a model route, a data plan and cleaned data.
pedrohcgs/claude-code-my-workflow
Adversarial 5-7 question challenge to a deck's pedagogical choices — ordering, prerequisites, cognitive load, motivation.
pedrohcgs/claude-code-my-workflow
Qualify a check before it is allowed to clear anything — prove it can detect the failure it is meant to catch.
pedrohcgs/claude-code-my-workflow
Compile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex).
pedrohcgs/claude-code-my-workflow
Show current context status and session health. An agent skill from pedrohcgs/claude-code-my-workflow.
pedrohcgs/claude-code-my-workflow
Snapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt /…
pedrohcgs/claude-code-my-workflow
Save a structured state snapshot before stopping or handing off.
Works with
Categories
Enforce the replication-protocol.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs. Audit Reproducibility is an agent skill from pedrohcgs/claude-code-my-workflow.md rule by cross-checking numeric claims in a manuscript against the actual R / Stata / Python outputs.
Audit Reproducibility fits situations like: tasks that involve Econometrics and empirical research; tasks that involve Reproducible research; tasks that involve Database administration.
Run `npx skills add pedrohcgs/claude-code-my-workflow --skill audit-reproducibility -a claude-code`. Or copy the skill folder (.claude/skills/audit-reproducibility in pedrohcgs/claude-code-my-workflow) into .claude/skills/audit-reproducibility in your project. Claude Code loads it when a task matches its description.
Run `npx skills add pedrohcgs/claude-code-my-workflow --skill audit-reproducibility -a codex`. Or copy the skill folder (.claude/skills/audit-reproducibility in pedrohcgs/claude-code-my-workflow) into .agents/skills/audit-reproducibility in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pedrohcgs/claude-code-my-workflow --skill audit-reproducibility -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-reproducibility, .gemini/skills/audit-reproducibility, .github/skills/audit-reproducibility and .opencode/skills/audit-reproducibility in your project.
Going by SKILL.md and its folder, Audit Reproducibility needs the command-line tools its instructions call (git). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Grep, Glob, Write, Bash, Agent, Task, Monitor.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Audit Reproducibility is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Audit Reproducibility: Audit Replication (brycewang-stanford/Auto-Empirical-Research-Skills, 4.6k stars), Data Deposit (brycewang-stanford/Auto-Empirical-Research-Skills, 4.6k stars), Ectheory Replication And Data Policy (franklee16/academic-research-skills, 223 stars) and Qe Replication And Data Policy (franklee16/academic-research-skills, 223 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
pedrohcgs (a GitHub user) maintains it in pedrohcgs/claude-code-my-workflow, which has 1,655 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on September 27, 2026.
Source: pedrohcgs/claude-code-my-workflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.