Stress-test a finding against the choices you did not make. An agent skill from pedrohcgs/claude-code-my-workflow.

MITAuto-check: notesTesting & QA

Install Challenge

skills CLI
$ npx skills add pedrohcgs/claude-code-my-workflow --skill challenge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pedrohcgs/claude-code-my-workflow challenge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pedrohcgs/claude-code-my-workflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/challenge .claude/skills/challenge && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
challenge
GitHub stars
1.7k
Token cost
~1.9k tokens
SKILL.md length
882 words
Files
4 (incl. references)
Skills in repo
59
Repo updated
First seen
Licence
MIT

At a glance

Stress-test a finding against the choices you did not make. An agent skill from pedrohcgs/claude-code-my-workflow.

  • Works in 6 steps: Enumerate the forks, before running… → Run the grid → Report the distribution, not the winner → …
  • The user says is this robust
  • SKILL.md covers Preconditions, Step 1 — Enumerate the forks,…, Step 2 — Run the grid and Step 3 — Report the…, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Challenge is an agent skill from pedrohcgs/claude-code-my-workflow. Stress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/fork-catalog.md` and `references/sensitivity-statistics.md`).

It sits in Testing & QA, covering Load testing and Statistics. The repository describes itself as: A ready-to-fork Claude Code template for academics using LaTeX/Beamer + R. Multi-agent review, quality gates, adversarial QA, and replication protocols. The licence is MIT.

When your agent uses it

  • The user says is this robust
  • Challenge this result
  • Specification curve
  • How sensitive is this

Example prompts

  • “is this robust”
  • “challenge this result”
  • “specification curve”
  • “/challenge”

Requirements

  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Write, Edit, Agent, Task

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Enumerate the forks, before running anything
  2. Run the grid
  3. Report the distribution, not the winner
  4. Attack the identifying assumption
  5. Placebo and falsification
  6. Write the ledger entry

What it can do on your machine

Read from SKILL.md and the folder at commit ae72617. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Write
    • Edit
    • Agent
    • Task

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Challenge loads about 1.9k tokens when it runs, and up to ~3.2k if it reads all its reference files. Until then it costs about 177 tokens; SKILL.md has 882 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~177
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Write, Edit, Agent, Task

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pedrohcgs/claude-code-my-workflow at commit ae72617, republished under its MIT licence (© pedrohcgs). 882 words, ~1,931 tokens.

Download SKILL.mdSave it as .claude/skills/challenge/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
challenge
description
Stress-test a finding against the choices you did not make. Enumerates the discrete forks a competent analyst could have taken (measure definition, sample filter, control set, clustering level, weighting, functional form), runs the specification grid, and reports the distribution rather than a point estimate — then attacks the identifying assumption with named, computable sensitivity statistics. Use when the user says "is this robust", "challenge this result", "specification curve", "multiverse", "how sensitive is this", "what if I'd used a different measure", "stress-test my estimate", or before a result becomes a headline claim. NOT a reviewer of prose or code — it challenges the CLAIM.
allowed-tools
Read, Grep, Glob, Bash, Write, Edit, Agent, Task
argument-hint
[script or results file] [--forks N] [--dry-run]
disable-model-invocation
true
metadata.protocol
threat-prioritization

Challenge — does the result survive the choices you didn't make?

A single specification is one draw from a distribution you never looked at.

Why this exists, measured rather than asserted. In a controlled study, 150 autonomous agents were given the same data and the same questions. Effect-size interquartile ranges reached ~10.7 %/yr, and the spread concentrated in discrete measure-choice forks — not in estimation noise. Within a measure family, agents agreed to ~0.25 %/yr. Two findings from that study shape this skill:

  • AI peer review left the spread essentially unchanged. Review catches errors; it does not reduce analytical-choice variance. A clean referee report is not robustness.
  • Exposure to exemplar papers collapsed the spread by 80–99 % — convergence by imitation, not by correctness. Herding is not agreement.

So the spread has to be measured, not reviewed away.

Preconditions

  • A working baseline specification that runs and produces the headline estimate.
  • The estimate's estimand stated in words — "the ATT for units treated in 2015, over 2013–2019, on the treated population". If you cannot state it, stop: you cannot challenge a claim you have not defined.
  • A fork budget (--forks, default 64). Grid size is the product of your choices; it grows faster than intuition.

Step 1 — Enumerate the forks, before running anything

List every point where a competent, honest analyst could have chosen differently. Do this before seeing any alternative result, and write it down — the list is the pre-registration of the challenge.

ForkTypical alternatives
Measure definitionlevel vs rate vs share; dollar vs count; stock vs flow
Sample filterbalanced vs unbalanced; trimming rules; inclusion windows
Control setnone / baseline / baseline+trends / interacted
Clustering levelunit / treatment-assignment / two-way
Weightingunweighted / population / inverse-propensity
Functional formlevels / logs / IHS / Poisson
Winsorizationnone / 1% / 5%

Some forks change the estimand, not just the estimate — averaging over them is meaningless. references/fork-catalog.md labels every fork; record estimand forks separately and say so in the report.

Ship --dry-run first. Print the grid size and an estimated runtime before executing anything. A 6-fork grid with 3 options each is 729 fits.

Step 2 — Run the grid

One fit per cell, same seed, same data build. Persist every cell — including failures. A specification that does not converge is information about fragility, not a cell to drop.

Record per cell: the fork coordinates, the point estimate, the standard error, N, and the convergence status.

Step 3 — Report the distribution, not the winner

  • Specification curve: estimates sorted, with the fork coordinates shown underneath so a reader can see which choices move the result.
  • The share of specifications with the same sign, and the share significant at conventional levels. Report both; they answer different questions.
  • Which fork drives the spread. This is the payload. "The result is robust except to the choice between dollar and share volume" is a far more useful sentence than a robustness paragraph.
  • Your baseline's percentile in its own distribution. If the headline sits at the 97th percentile of specifications you yourself called defensible, say so.

The descriptive curve is not a test — read it as a description of fragility. If you need inference over the whole curve, use specification-curve analysis's joint permutation test (Simonsohn, Simmons & Nelson 2020), which supplies the sharp null the picture alone lacks.

Show full SKILL.md (346 more words)Show less

Step 4 — Attack the identifying assumption

The grid varies what you can vary. The identifying assumption is what you cannot test — so bound it instead, with a named, computable statistic. See references/sensitivity-statistics.md.

ConcernStatistic
Unobserved confoundingE-value; Cinelli–Hazlett robustness value
Selection on observables → unobservablesOster δ (with a stated R²max)

Rows for the causal-identification designs are deliberately absent (unvetted-methods veto): populate them from your field's canonical sources after vetting.

Label every statistic executable-here or describe-and-cite. Honesty about what your environment can actually run is itself a verification step; a cited-but-unrun statistic is not evidence.

Step 5 — Placebo and falsification

Where a falsification test exists, run it: a negative-control outcome that should show nothing, a negative-control exposure, a timing placebo. A passed placebo is weak positive evidence; a failed placebo is strong negative evidence. Report both with equal prominence.

Step 6 — Write the ledger entry

Append to the specification-search ledger (verification-ladder.md rung 5): the fork list, the grid size, the distribution summary, which forks moved the result, the sensitivity statistics with their values, and every attempt including the failures.

Pre-commit the interpretation. Before running the grid, write down what result would SUPPORT and what would WEAKEN the claim. The ledger is the arbiter. A robustness exercise interpreted after the fact is not a robustness exercise.

Anti-patterns

  • Running until something interesting appears. That is the pathology this skill exists to prevent. Fixed fork list, fixed budget, stated stopping rule.
  • Reporting only the specifications that agree — a curve showing only supporting cells is a fishing expedition with better graphics.
  • Treating a wide curve as failure. Wide is a finding. Publish it and say what drives it.
  • Adding forks nobody would defend to pad the denominator and dilute the fragile cells.

Reference files

FileRead when
references/sensitivity-statistics.mda challenge rests on an untestable identifying assumption and needs a computable bound
references/fork-catalog.mdenumerating forks for a design you have not challenged before

Cross-references

© pedrohcgs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in .claude/skills/challenge of pedrohcgs/claude-code-my-workflow.

  • SKILL.md
  • evals/marker.txt
  • references/fork-catalog.md
  • references/sensitivity-statistics.md

Open the folder on GitHubat commit ae72617

Compare with similar skills

Challenge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Challenge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Challenge this skillpedrohcgs/claude-code-my-workflow1.7k—~1.9kAutomated safety check: NotesMIT
Icml Experimentsbrycewang-stanford/Awesome-Journal-Skills1.2k—~843Automated safety check: PassMIT
Jbes Identification Strategyfranklee16/academic-research-skills2231 repos~1kAutomated safety check: PassNone
Amr Data Analysisfranklee16/academic-research-skills2231 repos~1.5kAutomated safety check: PassNone
Restat Identificationbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.6kAutomated safety check: PassMIT
Statistical Data Analysislingzhi227/agent-research-skills386—~886Automated safety check: PassNone

Similar skills

  • Icml Experiments

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when stress-testing ICML experimental evidence before submission or rebuttal, including strong tuned baselines, mechanism-isolating ablations, seed variance and confidence…

    1.2k GitHub stars~843 tokensUpdated 12 days ago
    Research & ScienceAuto-check passed
  • Jbes Identification Strategy

    franklee16/academic-research-skills

    A skill your agent uses when the methodological core of a Journal of Business & Economic Statistics (JBES) paper is the bottleneck — assumptions, regularity conditions, asymptotic theory, and Monte…

    223 GitHub starsUsed in 1 repo~1k tokens
    Data & AnalyticsAuto-check passed
  • Amr Data Analysis

    franklee16/academic-research-skills

    A skill your agent uses when stress-testing the LOGIC of an Academy of Management Review (AMR) theory manuscript — checking logical coherence, running thought experiments and counterfactuals…

    223 GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Restat Identification

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when the causal-identification or measurement strategy is the bottleneck for a The Review of Economics and Statistics (REStat) manuscript — a DID / RD / IV / shift-share…

    1.2k GitHub stars~1.6k tokensUpdated 12 days ago
    Research & ScienceAuto-check passed
  • Statistical Data Analysis

    lingzhi227/agent-research-skills

    Writes statistical analysis code for experimental data, runs it through a four-round review, and reports effect sizes, p-values and confidence intervals.

    386 GitHub stars~886 tokensUpdated 7 mo ago
    Data & AnalyticsAuto-check passed
  • Runs exploratory data analysis on tabular data after you confirm each column's measurement level, then writes CSV tables and a narrative summary.

    108 GitHub stars~1.1k tokensUpdated 15 days ago
    Data & AnalyticsAuto-check passed

More from pedrohcgs/claude-code-my-workflow

All 59 skills in this repo
  • Devils Advocate

    pedrohcgs/claude-code-my-workflow

    Adversarial 5-7 question challenge to a deck's pedagogical choices — ordering, prerequisites, cognitive load, motivation.

    1.7k GitHub starsUsed in 2 repos~641 tokens
    Auto-check passed
  • Vaccinate

    pedrohcgs/claude-code-my-workflow

    Qualify a check before it is allowed to clear anything — prove it can detect the failure it is meant to catch.

    1.7k GitHub stars~2.1k tokensUpdated 11 days ago
    Auto-check: notes
  • Compile Latex

    pedrohcgs/claude-code-my-workflow

    Compile a Beamer LaTeX slide deck with XeLaTeX (3 passes + bibtex).

    1.7k GitHub starsUsed in 1 repo~492 tokens
    Auto-check: notes
  • Context Status

    pedrohcgs/claude-code-my-workflow

    Show current context status and session health. An agent skill from pedrohcgs/claude-code-my-workflow.

    1.7k GitHub starsUsed in 1 repo~613 tokens
    Auto-check: notes
  • Capture Environment

    pedrohcgs/claude-code-my-workflow

    Snapshot the computational environment for a replication package — detects the analysis stack (R / Stata / Python) and emits the right lockfiles (renv.lock + sessionInfo.txt, requirements.txt /…

    1.7k GitHub stars~2.8k tokensUpdated 11 days ago
    Auto-check: notes
  • Checkpoint

    pedrohcgs/claude-code-my-workflow

    Save a structured state snapshot before stopping or handing off.

    1.7k GitHub stars~2.8k tokensUpdated 11 days ago
    Auto-check: notes

Questions about Challenge

What does Challenge do?

Stress-test a finding against the choices you did not make. An agent skill from pedrohcgs/claude-code-my-workflow. Challenge is an agent skill from pedrohcgs/claude-code-my-workflow. Stress-test a finding against the choices you did not make.

When should I use Challenge?

Challenge fits situations like: the user says is this robust; challenge this result; specification curve; how sensitive is this.

How do I install Challenge in Claude Code?

Run `npx skills add pedrohcgs/claude-code-my-workflow --skill challenge -a claude-code`. Or copy the skill folder (.claude/skills/challenge in pedrohcgs/claude-code-my-workflow) into .claude/skills/challenge in your project. Claude Code loads it when a task matches its description.

How do I install Challenge in Codex?

Run `npx skills add pedrohcgs/claude-code-my-workflow --skill challenge -a codex`. Or copy the skill folder (.claude/skills/challenge in pedrohcgs/claude-code-my-workflow) into .agents/skills/challenge in your project. Codex loads it when a task matches its description.

Can I use Challenge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pedrohcgs/claude-code-my-workflow --skill challenge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/challenge, .gemini/skills/challenge, .github/skills/challenge and .opencode/skills/challenge in your project.

What does Challenge need to run?

SKILL.md names no scripts, command-line tools or credentials: Challenge is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Write, Edit, Agent, Task.

Does Challenge access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Challenge safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Challenge use?

Challenge is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Challenge use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Challenge?

Skills that share tags, products or a category with Challenge: Icml Experiments (brycewang-stanford/Awesome-Journal-Skills, 1.2k stars), Jbes Identification Strategy (franklee16/academic-research-skills, 223 stars), Amr Data Analysis (franklee16/academic-research-skills, 223 stars) and Restat Identification (brycewang-stanford/Awesome-Journal-Skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Challenge?

pedrohcgs (a GitHub user) maintains it in pedrohcgs/claude-code-my-workflow, which has 1,653 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on September 27, 2026.

Source: pedrohcgs/claude-code-my-workflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.