Agent skill

Stata Replication

by pedrohcgs in pedrohcgs/claude-code-my-workflow

End-to-end Stata replication pipeline — scaffolds numbered .do files in scripts/stata/, executes them via the stata-mcp MCP server, captures logs and outputs to output/, and produces…

MITAuto-check: notesResearch & Science

Install Stata Replication

skills CLI
$ npx skills add pedrohcgs/claude-code-my-workflow --skill stata-replication -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pedrohcgs/claude-code-my-workflow stata-replication --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/stata-replication .claude/skills/stata-replication && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
stata-replication
GitHub stars
1.6k
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
855 words
Files
1
Skills in repo
59
Repo updated
First seen
Licence
MIT

At a glance

End-to-end Stata replication pipeline — scaffolds numbered .do files in scripts/stata/, executes them via the stata-mcp MCP server, captures logs and outputs to output/, and produces…

  • Works in 5 steps: Pre-flight → Scaffold the pipeline → Execute (unless --no-execute) → …
  • User says stata replication
  • SKILL.md covers When to use, When NOT to use, Prerequisite: stata-mcp… and Workflow, plus 4 more sections
  • Calls claude

What it does

Stata Replication is an agent skill from pedrohcgs/claude-code-my-workflow. End-to-end Stata replication pipeline — scaffolds numbered .do files in scripts/stata/, executes them via the stata-mcp MCP server, captures logs and outputs to output/, and produces publication-ready tables (esttab) and figures (graph export). Mirrors /data-analysis for R-first projects. Use when user says "stata replication", "set up Stata pipeline", "scaffold the .do files", "run Stata analysis", "AEA replication package in Stata", or when a project's analysis language is Stata not R.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Econometrics and empirical research and Database administration. It works with Model Context Protocol. The repository describes itself as: A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols. The licence is MIT.

When your agent uses it

  • User says stata replication
  • Set up Stata pipeline
  • Scaffold the .do files
  • Run Stata analysis

Example prompts

  • “stata replication”
  • “set up Stata pipeline”
  • “scaffold the .do files”
  • “/stata-replication”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Glob, Grep, Bash, Agent, Task

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Pre-flight
  2. Scaffold the pipeline
  3. Execute (unless --no-execute)
  4. Verify
  5. (optional): R cross-check

What it can do on your machine

Read from SKILL.md and the folder at commit ae72617. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Glob
    • Grep
    • Bash
    • Agent
    • Task

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • aeadataeditor.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Stata Replication loads about 2.1k tokens when it runs. Until then it costs about 130 tokens; SKILL.md has 855 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~130
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Glob, Grep, Bash, Agent, Task

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pedrohcgs/claude-code-my-workflow at commit ae72617, republished under its MIT licence (© pedrohcgs). 855 words, ~2,061 tokens.

Download SKILL.mdSave it as .claude/skills/stata-replication/SKILL.md (or your agent's skills folder).
name
stata-replication
description
End-to-end Stata replication pipeline — scaffolds numbered `.do` files in `scripts/stata/`, executes them via the `stata-mcp` MCP server, captures logs and outputs to `output/`, and produces publication-ready tables (esttab) and figures (graph export). Mirrors `/data-analysis` for R-first projects. Use when user says "stata replication", "set up Stata pipeline", "scaffold the .do files", "run Stata analysis", "AEA replication package in Stata", or when a project's analysis language is Stata not R.
allowed-tools
Read, Write, Edit, Glob, Grep, Bash, Agent, Task
argument-hint
[paper-or-data-pointer] [--from-r] [--no-execute]
disable-model-invocation
true
metadata.author
Claude Code Academic Workflow
metadata.version
1.0.0

/stata-replication — Stata pipeline scaffold + execution

Build a complete Stata replication pipeline in scripts/stata/: numbered .do files following .claude/rules/stata-code-conventions.md, executed via the stata-mcp MCP server, with outputs landing in output/.

When to use

  • Your project's analysis language is Stata (not R). Common in econ field experiments, RCT studies, and any AEA submission where the original replication package is Stata.
  • You're porting an R-first project to Stata for an AEA submission.
  • You're adding a Stata robustness check to an R-first paper.
  • You want a one-command reproduction: do scripts/stata/99_run_all.do.

When NOT to use

  • Your project is R-first. Use /data-analysis.
  • Your project is Python-first. Neither this skill nor /data-analysis is the right fit; consider extending the convention rule for Python or porting one of these skills.
  • You're doing quick exploratory work. The numbered-pipeline scaffold is for replication packages, not scratch notebooks.

Prerequisite: stata-mcp installed

This skill requires the stata-mcp MCP server. Install once per user:

bash
claude mcp add stata-mcp --scope user -- uvx stata-mcp

The MCP server provides command-guarded Stata execution (refuses destructive operations like !/shell/erase), RAM monitoring, and Stata Language Server pairing. Maintained by SepineTam.

If stata-mcp is not installed, the skill halts at Phase 0 with installation instructions.

Workflow

Phase 0: Pre-flight
  1. Verify stata-mcp is registered in the user's MCP configuration. If not → halt with install instructions.
  2. Verify Stata is installed locally (the MCP server cannot run without it). Output stata version to confirm.
  3. Confirm scripts/stata/ directory exists or can be created.
  4. Read .claude/rules/stata-code-conventions.md — every emitted .do file follows this convention.
  5. If --from-r flag is set, locate the existing R pipeline at scripts/R/ and use it as a translation source. Apply the Stata → R pitfalls table from replication-protocol.md in reverse.
Phase 1: Scaffold the pipeline

Emit (or update) these files in scripts/stata/, each conforming to the header convention from stata-code-conventions.md:

scripts/stata/
├── 00_install.do        # ssc install, set globals, paths, sessionInfo capture
├── 01_clean.do          # raw → cleaned panel
├── 02_descriptive.do    # summary tables, balance (iebaltab), attrition
├── 03_analyze.do        # main regression specs (reghdfe / ivreg2 as needed)
├── 04_robustness.do     # alt specs, sensitivity
├── 05_tables_figures.do # esttab .tex outputs + graph export PDFs
└── 99_run_all.do        # do "01_clean.do" / do "02_..." / ...

If the paper or data source suggests specific specs (e.g., DiD with reghdfe, IV with ivreg2, RD with rdrobust), tailor 03_analyze.do accordingly.

Phase 2: Execute (unless --no-execute)

For each script in numbered order:

  1. Dispatch to stata-mcp to execute the .do file.
  2. Capture the log (Stata writes to output/NN_log.smcl per the header convention) and the resulting .dta / .tex / .pdf outputs.
  3. If a script fails, first append its specification to the ledger (step 4) with Status failed and the error in Why, then halt — do NOT auto-fix unless the failure is trivial (typo flagged by Stata at parse time). For substantive failures (insufficient observations, singular matrices, missing covariates), surface to the user.
  4. Append every specification each estimation .do file ran — kept, dropped, or failed — to quality_reports/spec-ledger.md, with the same columns, commit stamp and append-only block as /data-analysis Phase 3 ("Specification ledger"). A failed run is a row too, with Status failed and the error in Why.

For long-running scripts (> 2 minutes), use the Monitor tool to stream stdout — same pattern documented in /data-analysis and /audit-reproducibility.

Phase 3: Verify
  1. Confirm every expected output exists in output/.
  2. Check output/sessionInfo_stata.txt was captured (package versions).
  3. Run /audit-reproducibility if a manuscript exists — it reads Stata .dta outputs via haven/pyreadstat.
  4. Report scripts run, outputs produced, any warnings from Stata.
Show full SKILL.md (342 more words)Show less
Phase 4 (optional): R cross-check

If --from-r was set, run the R version of the same analysis (assumed to live at scripts/R/) and compare:

  • Point estimates: should match to ~0.01 (per replication-protocol.md tolerance).
  • Standard errors: should match to ~0.05 (clustering df adjustments can differ slightly between Stata and R).
  • Sample sizes: must match exactly.

Discrepancies are surfaced for the user to investigate — typical culprits: clustering df, default options (logit vs probit for PS), bootstrap seed handling.

Companion skills

  • /data-analysis — R analogue. Same pipeline shape, different language.
  • /audit-reproducibility — reads both .rds and .dta outputs. Cross-checks manuscript claims against the produced values.
  • /review-paper — if the paper exists and cites tables/figures produced by this pipeline, /review-paper auto-invokes /audit-reproducibility (per cross-artifact-review.md).

Anti-patterns

  • Hand-editing .dta files. Never. All transformations happen via the .do files; .dta outputs are derived and reproducible.
  • Skipping the 99_run_all.do. This is the AEA-mandated one-command entry point. Build it even for small projects.
  • Using , robust by default. Use , cluster(id) at the appropriate level — see stata-code-conventions.md §6.
  • Hand-formatting tables in LaTeX. Use esttab and \input{} — see stata-code-conventions.md §4.
  • Pinning Stata version in only one .do file. Every .do file starts with version 18 per the convention.

Cross-references

Long-running fits / batch reruns: use the Monitor tool (Apr 2026)

Long Stata fits (multi-hour bootstrap with cluster bootstrap, large reghdfe with millions of observations, simulation studies) should be background-launched and tailed with the Monitor tool — same pattern as /data-analysis and /audit-reproducibility for R / Python. The .do file logs to SMCL (output/NN_log.smcl). Monitor does not attach to a background job or its stderr: only the stdout of the command you give it becomes events. So run Monitor on a command that tails the log and filters for progress lines and Stata errors, e.g. tail -f output/NN_log.smcl | grep --line-buffered -E '\{err\}|r\([0-9]+\);|<your progress marker>', so Claude can react to errors mid-stream (a multi-hour run needs persistent: true, then TaskStop once the job ends).

© pedrohcgs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/stata-replication of pedrohcgs/claude-code-my-workflow.

Open the folder on GitHubat commit ae72617

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in pedrohcgs/claude-code-my-workflow, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Stata Replication next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Stata Replication compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Stata Replication this skillpedrohcgs/claude-code-my-workflow1.6k1 repos~2.1kAutomated safety check: NotesMIT
Stata AuditSepineTam/mcp-for-stata264—~1.2kAutomated safety check: PassAGPL-3.0
Diagnostic DofileSepineTam/mcp-for-stata264—~1.2kAutomated safety check: PassAGPL-3.0
Referee2scunning1975/MixtapeTools4702 repos~2.4kAutomated safety check: PassNone
Stata DiscoverSepineTam/mcp-for-stata264—~1.7kAutomated safety check: PassAGPL-3.0
Rfc Impl GeneratorSepineTam/mcp-for-stata264—~1.1kAutomated safety check: PassAGPL-3.0

Similar skills

  • Stata Audit

    SepineTam/mcp-for-stata

    Inspect, validate, summarize, and render local Stata-MCP audit evidence under .statamcp.

    264 GitHub stars~1.2k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Diagnostic Dofile

    SepineTam/mcp-for-stata

    A skill your agent uses when the user needs to inspect, audit, or diagnose the safety of a Stata do-file.

    264 GitHub stars~1.2k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Referee2

    scunning1975/MixtapeTools

    Systematic audit and review by Referee 2. An agent skill from scunning1975/MixtapeTools.

    470 GitHub starsUsed in 2 repos~2.4k tokens
    Research & ScienceAuto-check passed
  • Stata Discover

    SepineTam/mcp-for-stata

    A skill your agent uses when you need to find Stata on the user's machine or configure stata-mcp to use it.

    264 GitHub stars~1.7k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Rfc Impl Generator

    SepineTam/mcp-for-stata

    Generate RFC and IMPL documents from a user-provided feature/fix description.

    264 GitHub stars~1.1k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Stata Skill

    SepineTam/mcp-for-stata

    A packaged Stata Runner skill via official MCP-for-Stata server including statado, adopackageinstall, help, readlog and getdatainfo tools.

    264 GitHub stars~2.7k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed

More from pedrohcgs/claude-code-my-workflow

All 59 skills in this repo
  • Devils Advocate

    pedrohcgs/claude-code-my-workflow

    Adversarial 5-7 question challenge to a deck's pedagogical choices — ordering, prerequisites, cognitive load, motivation.

    1.6k GitHub starsUsed in 2 repos~641 tokens
    Auto-check passed
  • Vaccinate

    pedrohcgs/claude-code-my-workflow

    Qualify a check before it is allowed to clear anything — prove it can detect the failure it is meant to catch.

    1.6k GitHub stars~2.1k tokensUpdated 11 days ago
    Auto-check: notes
  • Compile Latex

    pedrohcgs/claude-code-my-workflow

    Compile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex).

    1.6k GitHub starsUsed in 1 repo~492 tokens
    Auto-check: notes
  • Context Status

    pedrohcgs/claude-code-my-workflow

    Show current context status and session health. An agent skill from pedrohcgs/claude-code-my-workflow.

    1.6k GitHub starsUsed in 1 repo~613 tokens
    Auto-check: notes
  • Disclosure Check

    pedrohcgs/claude-code-my-workflow

    Pre-screen analysis outputs (tables, figures, logs) built on restricted or confidential data for statistical-disclosure-limitation problems before any release.

    1.6k GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check: notes
  • Capture Environment

    pedrohcgs/claude-code-my-workflow

    Snapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt /…

    1.6k GitHub stars~2.8k tokensUpdated 11 days ago
    Auto-check: notes

Questions about Stata Replication

What does Stata Replication do?

End-to-end Stata replication pipeline — scaffolds numbered .do files in scripts/stata/, executes them via the stata-mcp MCP server, captures logs and outputs to output/, and produces…. Stata Replication is an agent skill from pedrohcgs/claude-code-my-workflow.do files in scripts/stata/, executes them via the stata-mcp MCP server, captures logs and outputs to output/, and produces publication-ready tables (esttab) and figures (graph export).

When should I use Stata Replication?

Stata Replication fits situations like: user says stata replication; set up Stata pipeline; scaffold the .do files; run Stata analysis.

How do I install Stata Replication in Claude Code?

Run `npx skills add pedrohcgs/claude-code-my-workflow --skill stata-replication -a claude-code`. Or copy the skill folder (.claude/skills/stata-replication in pedrohcgs/claude-code-my-workflow) into .claude/skills/stata-replication in your project. Claude Code loads it when a task matches its description.

How do I install Stata Replication in Codex?

Run `npx skills add pedrohcgs/claude-code-my-workflow --skill stata-replication -a codex`. Or copy the skill folder (.claude/skills/stata-replication in pedrohcgs/claude-code-my-workflow) into .agents/skills/stata-replication in your project. Codex loads it when a task matches its description.

Can I use Stata Replication in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pedrohcgs/claude-code-my-workflow --skill stata-replication -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stata-replication, .gemini/skills/stata-replication, .github/skills/stata-replication and .opencode/skills/stata-replication in your project.

What does Stata Replication need to run?

Going by SKILL.md and its folder, Stata Replication needs the command-line tools its instructions call (claude). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Write, Edit, Glob, Grep, Bash, Agent, Task.

Does Stata Replication access the network?

SKILL.md names 2 domains. As links in the text: github.com and aeadataeditor.github.io. This is read from the text; nothing was executed.

Is Stata Replication safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Stata Replication use?

Stata Replication is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Stata Replication use?

About 2.1k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Stata Replication?

Skills that share tags, products or a category with Stata Replication: Stata Audit (SepineTam/mcp-for-stata, 264 stars), Diagnostic Dofile (SepineTam/mcp-for-stata, 264 stars), Referee2 (scunning1975/MixtapeTools, 470 stars) and Stata Discover (SepineTam/mcp-for-stata, 264 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Stata Replication?

pedrohcgs (a GitHub user) maintains it in pedrohcgs/claude-code-my-workflow, which has 1,645 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on September 27, 2026.

Source: pedrohcgs/claude-code-my-workflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.