Agent skill

Replication Audit

by flonat in flonat/flonat-research

Map claims in a literature to independent replications, robustness checks, failures, and unresolved evidence gaps.

MITAuto-check passedDatabases

Install Replication Audit

skills CLI
$ npx skills add flonat/flonat-research --skill replication-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install flonat/flonat-research replication-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/replication-audit .claude/skills/replication-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
replication-audit
GitHub stars
145
Token cost
~2.3k tokens
SKILL.md length
783 words
Files
1
Skills in repo
83
Repo updated
First seen
Licence
MIT

At a glance

Map claims in a literature to independent replications, robustness checks, failures, and unresolved evidence gaps.

  • Works in 6 steps: Findings Inventory → Replication Search → Classification → …
  • Assessing the empirical reliability of a body of findings rather than reproducing one projects code
  • SKILL.md covers Output Path, When to Use, When NOT to Use and Input, plus 4 more sections
  • Calls bash

What it does

Replication Audit is an agent skill from flonat/flonat-research. Map claims in a literature to independent replications, robustness checks, failures, and unresolved evidence gaps. Use when assessing the empirical reliability of a body of findings rather than reproducing one project's code. For package rerunnability, use $replication-package.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering Database administration and Econometrics and empirical research. The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.

When your agent uses it

  • Assessing the empirical reliability of a body of findings rather than reproducing one projects code
  • Tasks that involve Database administration
  • Tasks that involve Econometrics and empirical research

Example prompts

  • “/replication-audit”

Requirements

  • Pre-approved tools (allowed-tools): Read, Write, Edit, Glob, Grep, Bash(uv*), Bash(uv:*), Task, WebSearch, WebFetch, Bash(paperpile*)

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Findings Inventory
  2. Replication Search
  3. Classification
  4. Dependency Mapping
  5. Risk Assessment
  6. Output

What it can do on your machine

Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Glob
    • Grep
    • Bash(uv*)
    • Bash(uv:*)
    • Task
    • WebSearch
    • WebFetch

    …and 1 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Replication Audit loads about 2.3k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 783 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 783 words, ~2,332 tokens.

Download SKILL.mdSave it as .claude/skills/replication-audit/SKILL.md (or your agent's skills folder).
name
replication-audit
description
Map claims in a literature to independent replications, robustness checks, failures, and unresolved evidence gaps. Use when assessing the empirical reliability of a body of findings rather than reproducing one project's code. For package rerunnability, use $replication-package.
allowed-tools
Read, Write, Edit, Glob, Grep, Bash(uv*), Bash(uv:*), Task, WebSearch, WebFetch, Bash(paperpile*)
argument-hint
[topic, .bib file, or paper directory]
skill-dependencies
method-audit

Replication Audit

Examine which findings in a literature have been replicated, failed to replicate, or never been tested. Flag papers that build on non-replicated foundations.

Don't build your dissertation on a foundation made of sand.

Output Path

Per rules/review-artefact-routing.md (auto-loads in research projects (path-scoped to paper-*/ and paper/)):

  • Source slug: replication-audit
  • Write reports to: reviews/<scope>/replication-audit/<YYYY-MM-DD-HHMM>.md inside the project, where <scope> is the paper slug (e.g., paper-jtp) for paper-level reviews or _project for project-level reviews. Path is relative to the research project root, not the Task-Management repo.
  • Never at project root (./CRITIC-REPORT.md-style filenames are forbidden — pre-rule layout).
  • Idempotency: if the run timestamp already exists, append a same-run descriptor ({timestamp}-revision.md, {timestamp}-r2.md) — never overwrite the same <YYYY-MM-DD-HHMM> path.
  • Index update: if reviews/INDEX.md exists, write a one-line entry under "Latest per source" pointing at the new file. Otherwise review-recap will rebuild the index next time it runs.
  • Infrastructure repos (Task-Management, atlas-workspace, etc.): this section does not apply — the path-scoped rule won't load there.

When to Use

  • Before committing to a theoretical framework — is the evidence solid?
  • Writing a literature review — need to distinguish robust from fragile findings
  • Preparing a replication study — need to know what's already been tested
  • Evaluating a paper for peer review — are its cited foundations sound?

When NOT to Use

  • Your own results — use the referee2-reviewer agent
  • Methodological comparison — use method-audit
  • Finding papers — use an installed scholarly-search workflow first

Input

A .bib file, PDF directory, topic, or list of key findings to audit. Works best with a focused set of influential papers whose findings underpin a line of research.

Workflow

Phase 1: Findings Inventory

From the corpus (assembled as in other corpus skills), extract the key empirical findings — not paper summaries, but specific claims with effect sizes where available:

FindingPaperEffectNMethod
"X causes Y"Author (Year)d = 0.5200RCT
"A predicts B"Author (Year)r = 0.31,500Survey

Focus on findings that other papers depend on — the claims that, if wrong, would undermine subsequent work.

For each key finding, search for replication attempts:

  1. Forward citations — use scholarly scholarly-citations on the original paper's DOI
  2. Keyword search — search for "replication" + key terms from the finding via scholarly scholarly-search
  3. Replication databases — search for the paper on:
    • ReplicationWiki (via web search)
    • Many Labs projects
    • Replication registries in the relevant field
  4. Author self-replication — check if the original authors replicated in a new sample

Dispatch rule. If ≥5 key findings need auditing, batch findings into groups of 4–5 per sub-agent. Each sub-agent runs steps 1–4 for its batch (scholarly-citations, scholarly-search, web searches) and writes replication evidence to /tmp/replication-audit-<n>.json. Main context merges and proceeds to Phase 3 classification. For <5 findings, sequential searches in main context are fine. See _shared/cli-dispatch-policy.md.

For each replication found, record:

  • Result: Successful / Failed / Partial / Conceptual (similar but different design)
  • Sample: How does it compare to the original (size, population)?
  • Method: Exact replication or conceptual?
  • Effect size: Same direction? Same magnitude?
  • Published where? (Replication failures in top journals carry more weight)
Show full SKILL.md (284 more words)Show less
Phase 3: Classification

Classify each finding:

StatusCriteria
ReplicatedSuccessfully replicated by at least one independent team
Multiply replicatedReplicated 3+ times across different samples/contexts
Failed to replicateAt least one serious replication attempt found null or opposite results
ContestedSome replications succeed, others fail — mixed evidence
Never testedNo known replication attempts (most common and most concerning)
UnreplicableData/method too expensive, proprietary, or impractical to replicate
Phase 4: Dependency Mapping

Map which papers in the corpus depend on each finding:

Finding: "X causes Y" (Author, 2015) — STATUS: Failed to replicate
├── Paper A (2017) — builds entire model on this finding
├── Paper B (2019) — uses this as a control variable
└── Paper C (2021) — cites this as motivation but doesn't depend on it

Flag any paper whose core contribution depends on a non-replicated or failed finding.

Phase 5: Risk Assessment

For each "never tested" finding, estimate replication risk:

Risk factorIncreases concern
Small sample (N < 100)High
P-value just below 0.05High
Surprising/counterintuitive resultMedium
Complex interaction effectsMedium
Single study, no robustness checksHigh
Author has other failed replicationsMedium
Published in a journal with low replication standardsMedium
Phase 6: Output

Write to REPLICATION-AUDIT.md in the project directory.

Output Format

markdown
# Replication Audit: [Topic]

**Date:** YYYY-MM-DD
**Corpus:** [N] papers
**Key findings audited:** [N]
**Status breakdown:** Replicated: X | Failed: Y | Never tested: Z | Contested: W

## Summary

[2-3 sentences: overall replication health of this literature]

## Finding-by-Finding Audit

### 1. "[Finding statement]" — Author (Year)

**Original:** N = [X], Effect = [Y], Method = [Z]
**Status:** [Replicated / Failed / Never tested / Contested]

**Replication evidence:**
- [Author (Year)] — [Result] — N = [X], Effect = [Y]
- [Author (Year)] — [Result] — N = [X], Effect = [Y]

**Depends on this:** [Papers in corpus that build on this finding]
**Risk level:** [Low / Medium / High / Critical]

### 2. "[Finding statement]" — Author (Year)
...

## Replication Status Matrix

| Finding | Original (Year) | Replicated? | Times tested | Risk |
|---------|----------------|-------------|-------------|------|

## Dependency Risk Map

Papers building on shaky foundations:

| Paper | Depends on | Status of dependency | Risk to paper's claims |
|-------|-----------|---------------------|----------------------|

## Recommendations

### Safe foundations (build on these)
- [Finding] — multiply replicated, robust across contexts

### Proceed with caution
- [Finding] — replicated once, small samples

### Avoid or re-test
- [Finding] — failed to replicate or never tested despite high risk factors

### Replication opportunities
- [Finding] — never tested, high-impact if confirmed, feasible to replicate with [data/method]

Log to REVIEW-STATE.md (final step)

Write the replication audit to reviews/<scope>/replication-audit/<YYYY-MM-DD-HHMM>.md (mkdir -p reviews/<scope>/replication-audit/ first), where <scope> is the paper slug or _project. Then append a row to the project's REVIEW-STATE.md:

bash
bash <skills-root>/_shared/review-state-log.sh \
  --check replication-audit \
  --paper "<paper-{venue} dir, or — for project-level audits>" \
  --verdict "<PASS|PARTIAL|FAIL>" \
  --score "<replicated-count>/<total-findings-checked>" \
  --open-issues "<failed-or-untested-count>/<total-findings-checked>" \
  --report "reviews/<scope>/replication-audit/<YYYY-MM-DD-HHMM>.md" \
  --notes "<one-line: e.g. '12/15 robust; 2 failed; 1 never tested'>" \
  [--trigger "pre-submission-report|review-cluster"]
  • Verdict: PASS if every findings is robustly replicated; PARTIAL if some replicated, some not; FAIL if foundational findings failed to replicate.
  • Score: robustly-replicated findings / total checked.
  • Open issues: (failed + never-tested) / total at run time.
  • Trigger: pass orchestrator name only if invoked as a sub-agent. Otherwise omit.

Schema: the installed shared resource shared/review-state-schema.md.

Cross-References

SkillWhen to use instead/alongside
method-auditFor broader methodological comparison (not replication-specific)
weakness-scannerFor logical and argumentative weaknesses (not replication status)
Installed scholarly-search workflowTo find the replication studies identified in this audit
split-pdfTo deep-read any replication study found

© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/replication-audit of flonat/flonat-research.

Open the folder on GitHubat commit da27600

Compare with similar skills

Replication Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Replication Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Replication Audit this skillflonat/flonat-research145—~2.3kAutomated safety check: PassMIT
Ajps Replication And Verificationfranklee16/academic-research-skills2231 repos~1.3kAutomated safety check: PassNone
Jie Replication And Data Policyfranklee16/academic-research-skills2231 repos~1.1kAutomated safety check: PassNone
Jole Replication And Data Policyfranklee16/academic-research-skills2231 repos~1.3kAutomated safety check: PassNone
Ectj Replication And Data Policyfranklee16/academic-research-skills2231 repos~394Automated safety check: PassNone
Rof Replication And Data Policybrycewang-stanford/Awesome-Journal-Skills1.2k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Ajps Replication And Verification

    franklee16/academic-research-skills

    A skill your agent uses when building the replication package for an American Journal of Political Science (AJPS) manuscript.

    223 GitHub starsUsed in 1 repo~1.3k tokens
    DatabasesAuto-check passed
  • Jie Replication And Data Policy

    franklee16/academic-research-skills

    A skill your agent uses when preparing the replication package for a Journal of International Economics (JIE) manuscript — JIE requires that all materials needed to replicate published papers…

    223 GitHub starsUsed in 1 repo~1.1k tokens
    DatabasesAuto-check passed
  • Jole Replication And Data Policy

    franklee16/academic-research-skills

    A skill your agent uses when assembling the data and code deposit for an accepted Journal of Labor Economics (JOLE) paper — the JOLE Dataverse Repository, the AER data-availability policy (adopted…

    223 GitHub starsUsed in 1 repo~1.3k tokens
    DatabasesAuto-check passed
  • Ectj Replication And Data Policy

    franklee16/academic-research-skills

    A skill your agent uses when preparing The Econometrics Journal replication files, README, software versions, data documentation, seeds, proprietary-data exemptions, and OUP Supporting Information…

    223 GitHub starsUsed in 1 repo~394 tokens
    DatabasesAuto-check passed
  • Rof Replication And Data Policy

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when preparing the Review of Finance (RoF) replication package — code, data, pseudo-datasets for proprietary sources such as Datastream or Bankscope, log files, the Data…

    1.2k GitHub stars~1.4k tokensUpdated 11 days ago
    DatabasesAuto-check passed
  • Jape Replication And Data Policy

    franklee16/academic-research-skills

    A skill your agent uses when assembling the mandatory JAE Data Archive deposit for an accepted Journal of Applied Econometrics paper — plain-ASCII/CSV data with a readme, the programs that replicate…

    223 GitHub starsUsed in 1 repo~671 tokens
    DatabasesAuto-check passed

More from flonat/flonat-research

All 83 skills in this repo
  • Latex Posters

    flonat/flonat-research

    Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.

    145 GitHub stars~1.5k tokensUpdated 8 days ago
    Auto-check: notes
  • Skill Creator

    flonat/flonat-research

    Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.

    145 GitHub stars~4.4k tokensUpdated 8 days ago
    Auto-check passed
  • DOCX

    flonat/flonat-research

    Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.

    145 GitHub stars~1.2k tokensUpdated 8 days ago
    Auto-check passed
  • PDF

    flonat/flonat-research

    Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.

    145 GitHub stars~488 tokensUpdated 8 days ago
    Auto-check passed
  • Init Project Orchestration

    flonat/flonat-research

    Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.

    145 GitHub stars~1.6k tokensUpdated 8 days ago
    Auto-check passed
  • Pre Commit Audit

    flonat/flonat-research

    Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

    145 GitHub stars~2.8k tokensUpdated 8 days ago
    Auto-check: notes

Categories

Questions about Replication Audit

What does Replication Audit do?

Map claims in a literature to independent replications, robustness checks, failures, and unresolved evidence gaps. Replication Audit is an agent skill from flonat/flonat-research. Map claims in a literature to independent replications, robustness checks, failures, and unresolved evidence gaps.

When should I use Replication Audit?

Replication Audit fits situations like: assessing the empirical reliability of a body of findings rather than reproducing one projects code; tasks that involve Database administration; tasks that involve Econometrics and empirical research.

How do I install Replication Audit in Claude Code?

Run `npx skills add flonat/flonat-research --skill replication-audit -a claude-code`. Or copy the skill folder (skills/replication-audit in flonat/flonat-research) into .claude/skills/replication-audit in your project. Claude Code loads it when a task matches its description.

How do I install Replication Audit in Codex?

Run `npx skills add flonat/flonat-research --skill replication-audit -a codex`. Or copy the skill folder (skills/replication-audit in flonat/flonat-research) into .agents/skills/replication-audit in your project. Codex loads it when a task matches its description.

Can I use Replication Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill replication-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/replication-audit, .gemini/skills/replication-audit, .github/skills/replication-audit and .opencode/skills/replication-audit in your project.

What does Replication Audit need to run?

Going by SKILL.md and its folder, Replication Audit needs the command-line tools its instructions call (bash). Its frontmatter pre-approves these tools: Read, Write, Edit, Glob, Grep, Bash(uv*), Bash(uv:*), Task, WebSearch, WebFetch, Bash(paperpile*).

Does Replication Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Replication Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Replication Audit use?

Replication Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Replication Audit use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Replication Audit?

Skills that share tags, products or a category with Replication Audit: Ajps Replication And Verification (franklee16/academic-research-skills, 223 stars), Jie Replication And Data Policy (franklee16/academic-research-skills, 223 stars), Jole Replication And Data Policy (franklee16/academic-research-skills, 223 stars) and Ectj Replication And Data Policy (franklee16/academic-research-skills, 223 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Replication Audit?

flonat (a GitHub user) maintains it in flonat/flonat-research, which has 145 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.

Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.