Agent skill

Reproduce First

by vfarcic in vfarcic/dot-agent-deck

Reproduce first, then fix — turn a reported defect into a failing test, fix it, and confirm that same test goes green.

MITAuto-check passedTesting & QA

Install Reproduce First

skills CLI
$ npx skills add vfarcic/dot-agent-deck --skill reproduce-first -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vfarcic/dot-agent-deck reproduce-first --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vfarcic/dot-agent-deck.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/reproduce-first .claude/skills/reproduce-first && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
reproduce-first
GitHub stars
109
Token cost
~2k tokens
SKILL.md length
1,256 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Reproduce first, then fix — turn a reported defect into a failing test, fix it, and confirm that same test goes green.

  • Works in 8 steps: Restate the symptom as the user sees it.… → Find the test that already covers this… → Assert at the user's altitude. The… → …
  • The user describes the software behaving differently from what they expected
  • SKILL.md covers When to use this, The order is the whole point, Process and The traps that have actually…, plus 2 more sections
  • Calls cargo

What it does

Reproduce First is an agent skill from vfarcic/dot-agent-deck. Reproduce first, then fix — turn a reported defect into a failing test, fix it, and confirm that same test goes green. Use whenever the user describes the software behaving differently from what they expected or intended, however they phrase it: a complaint, a neutral observation, a question about whether something is meant to work that way, or an aside that something works "except for" one detail. It applies to a report that arrives mid-task about unrelated work, and to a symptom mentioned in passing. Trigger on…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Failing and flaky tests. The repository describes itself as: A rich terminal dashboard for monitoring and controlling multiple AI coding agent sessions. The licence is MIT.

When your agent uses it

  • The user describes the software behaving differently from what they expected
  • However they phrase it: a complaint
  • A neutral observation
  • A question about whether something is meant to work that way

Example prompts

  • “except for”
  • “/reproduce-first”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Restate the symptom as the user sees it. At their altitude: a card that stays on screen, a tab that never appears, a name that reads…
  2. Find the test that already covers this surface, and EXTEND it. Bias order: extend an existing test > modify an existing test > write a new…
  3. Assert at the user's altitude. The user-visible outcome, not an adjacent artefact — a file on disk, a log line, or a registry entry can…
  4. Run it and confirm it fails FOR THEIR REASON. A test that fails for a setup error (a stub that broke an unrelated path, a form field that…
  5. Add a control that isolates the cause. The nearest thing that should still work — the same close on a card with no worktree, the same…
  6. Fix it.
  7. Watch the same test go green, then prove each fix is load-bearing: revert one change at a time and confirm the test goes red again. If…
  8. Run the wider tier before reporting done: cargo xtask affected-checks --run (for a code change that is cargo fmt --check, rule 2's clippy…

What it can do on your machine

Read from SKILL.md and the folder at commit 9cc3e60. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Reproduce First loads about 2k tokens when it runs. Until then it costs about 166 tokens; SKILL.md has 1,256 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~166
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vfarcic/dot-agent-deck at commit 9cc3e60, republished under its MIT licence (© vfarcic). 1,256 words, ~2,041 tokens.

Download SKILL.mdSave it as .claude/skills/reproduce-first/SKILL.md (or your agent's skills folder).
name
reproduce-first
description
Reproduce first, then fix — turn a reported defect into a failing test, fix it, and confirm that same test goes green. Use whenever the user describes the software behaving differently from what they expected or intended, however they phrase it: a complaint, a neutral observation, a question about whether something is meant to work that way, or an aside that something works "except for" one detail. It applies to a report that arrives mid-task about unrelated work, and to a symptom mentioned in passing. Trigger on the situation, not on any particular wording. Invoke it BEFORE reading code to diagnose the cause and before proposing any fix.
user-invocable
true

Reproduce a reported bug as a failing test, then fix it

When to use this

The user has reported that something is broken. That includes the quiet forms: "it works except for…", "I have to do X twice", "it stopped doing Y", "is it supposed to…?".

Do NOT use it for: a request for a new feature, a question about how something works, or a bug you found yourself while already mid-task with a test already failing for that reason.

The order is the whole point

The first deliverable is a test that fails for the user's reason. Not a fix. Then fix it, and watch that same test pass.

Do not reverse the order, and do not skip to the fix because the cause looks obvious. A fix whose test was written afterwards tends to assert what the code now does rather than what the user asked for. "The cause is obvious" is the single most reliable predictor that it is not — see the case studies below, where two confident diagnoses were both wrong and only the control runs exposed it.

Process

  1. Restate the symptom as the user sees it. At their altitude: a card that stays on screen, a tab that never appears, a name that reads wrong. Not "the registry entry is stale".

  2. Find the test that already covers this surface, and EXTEND it. Bias order: extend an existing test > modify an existing test > write a new one. Only add a brand-new #[spec] id when no catalog entry covers the surface at all — search tests/CATALOG.md for the area before assuming there isn't one.

    Most follow-up reports are the SAME behaviour under a different configuration (a real agent instead of a stand-in, an orchestration instead of a single agent, a command the deck cannot infer an agent type from). That is a case for widening the existing test's coverage to the configuration that actually broke — not a parallel test duplicating 90% of the setup. A new spec id is justified when the mechanism genuinely differs, not when the inputs do.

    This matters more at L2 than anywhere else: every PTY test spawns real binaries, and a suite of near-duplicate e2e tests is slow for everyone and harder to diagnose than one test that names its cases.

  3. Assert at the user's altitude. The user-visible outcome, not an adjacent artefact — a file on disk, a log line, or a registry entry can all be correct while the screen is wrong. If they said "no orchestration tab appeared", the assertion is a tab with live role cards.

  4. Run it and confirm it fails FOR THEIR REASON. A test that fails for a setup error (a stub that broke an unrelated path, a form field that was pre-filled, an ellipsized card title) is not a reproduction. Read the failure message and check it describes their symptom. Fix the harness problem and re-run until the red is the right red.

  5. Add a control that isolates the cause. The nearest thing that should still work — the same close on a card with no worktree, the same dispatch with a fast git — proves the failure is attributable to what you think it is. Without a control you cannot tell "this path is broken" from "this whole feature is broken".

  6. Fix it.

  7. Watch the same test go green, then prove each fix is load-bearing: revert one change at a time and confirm the test goes red again. If reverting a change leaves the test green, that change is not part of the fix — take it out or find what it is really doing.

  8. Run the wider tier before reporting done: cargo xtask affected-checks --run (for a code change that is cargo fmt --check, rule 2's clippy command and cargo test-fast), plus the tests covering what you touched. There is no full-tier obligation before the PR — per CLAUDE.md rule 5, lane 1 runs in CI on every PR, so read that run rather than reproducing it. If your reproduction is an e2e test, run it and its module by filter (cargo test-e2e <filter>, or cargo test-e2e-live <filter> when the test reaches a real agent); do not run either alias unfiltered. A lane-2 test is worth running deliberately: nothing in CI runs one, so if you do not, nobody does.

Show full SKILL.md (536 more words)Show less

The traps that have actually cost time here

Prove the test can fail. An assertion never observed failing is not evidence. This is the cheapest step and it has caught a vacuous assertion repeatedly: a stream-based wait that could never match redrawn chrome; a wait_until_grid capped at the harness's 10s WAIT_TIMEOUT that silently shortened an intended 60s wait; and an AgentRecord.live check that was Some(Idle) for every role within 1.5s of the spawn, before a byte had reached any of those PTYs.

Prefer their configuration over a convenient stand-in. cat roles and print-mode agents prove the plumbing and hide everything else: they cannot tell an agent from a $SHELL, and they never read an orchestrator-context file, so both of those defects shipped green. Where a stand-in is genuinely necessary for cost, say so in the test, name what it stands in for, and add one real-config case beside it.

A stand-in must be narrow. A git stub that slept on every status also hit the deck's own pane-creation path, which has its own 5s budget — the pane never came up and the test failed before reaching what it was about. Key the stand-in to the exact invocation under test.

Reproduce before diagnosing, and diagnose on the reporter's machine, not yours. Environment-shaped bugs do not travel: a whole diagnosis was once built from this server's process table while the user was reporting from a laptop. Ask for the artefacts — the message the pane printed, ~/.local/state/dot-agent-deck/deck.log, command -v — instead of inferring them locally.

If it genuinely cannot be reproduced, say so explicitly and name what is missing, before proposing a fix. "I could not reproduce this, so the fix is unverified" is a legitimate report. Presenting an unreproduced fix as verified is not.

Case studies from this repo

Two stacked defects, one symptom (dispatch/close/001). Reported: closing a dispatched agent left its card behind; a second close removed it. The reporter's guess was that worktree removal blocked the close. The first diagnosis (mine) was a client-side timeout. A control run with the slowness removed still failed — which exposed a different, primary defect underneath: a daemon-spawned card has no local pane until focused, so close_pane returned "not found", the card was preserved by policy, and the agent kept running. Only after fixing that did the reporter's timeout theory become the remaining cause. Both were real; reverting either fix alone turns the test red. Neither would have been found by reading code.

The assertion that proved nothing (orchestration/dispatch/002). The first version passed in 1.5 seconds — impossibly fast for three agent cold boots. It was asserting a field the daemon populates at spawn time. Rewritten to assert what the user looks at (every role named on its own card), it immediately caught a real defect: dispatched role cards were labelled with session UUIDs instead of the role names in the toml.

  • CONTRIBUTING.md — the team-facing statement of this norm, and the TDD loop commands.
  • CLAUDE.md rule 4 (which test tier a change needs) and rule 5 (fast tier per task, plus the tests covering the change; lane 1 runs in CI, and lane 2 runs nowhere but your machine).
  • tests/CATALOG.md — every test's entry records what it does not assert; add yours there.

© vfarcic, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/reproduce-first of vfarcic/dot-agent-deck.

Open the folder on GitHubat commit 9cc3e60

Compare with similar skills

Reproduce First next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Reproduce First compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Reproduce First this skillvfarcic/dot-agent-deck109—~2kAutomated safety check: PassMIT
Swig Testswig/swig6.3k—~2.3kAutomated safety check: PassCustom licence
Triage CI FailureDataDog/datadog-agent3.8k—~2.3kAutomated safety check: PassApache-2.0
Dynamo Jira TicketDynamoDS/Dynamo2k—~1.1kAutomated safety check: PassApache-2.0
Fix Ready PRsfastrepl/anarlog9.5k—~1.4kAutomated safety check: PassMIT
Trx Analysismicrosoft/vstest969—~1.8kAutomated safety check: PassMIT

Similar skills

  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Triage CI Failure

    DataDog/datadog-agent

    Official

    Classify a failed CI as either caused by an active incident, flakiness, or a true code regression.

    3.8k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Dynamo Jira Ticket

    DynamoDS/Dynamo

    Create structured Jira tickets for Dynamo from bug reports, failing tests, or feature requests.

    2k GitHub stars~1.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Fix Ready PRs

    fastrepl/anarlog

    Inspect every open non-draft PR for CI failures and unresolved Cursor Bugbot findings, then fix them on the existing PR branches.

    9.5k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Trx Analysis

    microsoft/vstest

    Official

    Parse and analyze Visual Studio TRX test result files. An agent skill from microsoft/vstest.

    969 GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Wio

    workersio/skills

    Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.

    204 GitHub stars~5.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed

More from vfarcic/dot-agent-deck

All 24 skills in this repo
  • Dispatch Base

    vfarcic/dot-agent-deck

    Bring the base that every dispatched unit is cut from up to date before the first dot-agent-deck dispatch of a batch in this repo, and check the base dispatch reports afterwards.

    109 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Dispatch Shape

    vfarcic/dot-agent-deck

    Choose the shape of a unit you are about to dispatch in this repo — one agent (--single) or a team (--orchestration '<name') — from divisibility criteria instead of asking, and report the shape you…

    109 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Docs Screenshots Review

    vfarcic/dot-agent-deck

    Check that a change to the user-facing docs covers both clients (the TUI and the desktop app) unless the feature exists in only one, and decide whether it needs a new or updated screenshot, then…

    109 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Dot AI Request Dot AI Feature

    vfarcic/dot-agent-deck

    Generate a feature request prompt for another dot-ai project.

    109 GitHub stars~679 tokensUpdated today
    Auto-check passed
  • PR Create

    vfarcic/dot-agent-deck

    Take committed work from a branch to a verified pull request — push, open the PR, settle CI and the automated review, answer and resolve every finding, and hand off.

    109 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Publish Docs

    vfarcic/dot-agent-deck

    Publish the docs site to GHCR with a main-<sha tag and bump site/helm/values.yaml so Argo CD picks it up — without cutting a SemVer release.

    109 GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Categories

Questions about Reproduce First

What does Reproduce First do?

Reproduce first, then fix — turn a reported defect into a failing test, fix it, and confirm that same test goes green. Reproduce First is an agent skill from vfarcic/dot-agent-deck. Reproduce first, then fix — turn a reported defect into a failing test, fix it, and confirm that same test goes green.

When should I use Reproduce First?

Reproduce First fits situations like: the user describes the software behaving differently from what they expected; however they phrase it: a complaint; A neutral observation; A question about whether something is meant to work that way.

How do I install Reproduce First in Claude Code?

Run `npx skills add vfarcic/dot-agent-deck --skill reproduce-first -a claude-code`. Or copy the skill folder (.claude/skills/reproduce-first in vfarcic/dot-agent-deck) into .claude/skills/reproduce-first in your project. Claude Code loads it when a task matches its description.

How do I install Reproduce First in Codex?

Run `npx skills add vfarcic/dot-agent-deck --skill reproduce-first -a codex`. Or copy the skill folder (.claude/skills/reproduce-first in vfarcic/dot-agent-deck) into .agents/skills/reproduce-first in your project. Codex loads it when a task matches its description.

Can I use Reproduce First in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vfarcic/dot-agent-deck --skill reproduce-first -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/reproduce-first, .gemini/skills/reproduce-first, .github/skills/reproduce-first and .opencode/skills/reproduce-first in your project.

What does Reproduce First need to run?

Going by SKILL.md and its folder, Reproduce First needs the command-line tools its instructions call (cargo).

Does Reproduce First access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Reproduce First safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Reproduce First use?

Reproduce First is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Reproduce First use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Reproduce First?

Skills that share tags, products or a category with Reproduce First: Swig Test (swig/swig, 6.3k stars), Triage CI Failure (DataDog/datadog-agent, 3.8k stars), Dynamo Jira Ticket (DynamoDS/Dynamo, 2k stars) and Fix Ready PRs (fastrepl/anarlog, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Reproduce First?

vfarcic (a GitHub user) maintains it in vfarcic/dot-agent-deck, which has 109 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 10, 2026.

Source: vfarcic/dot-agent-deck on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.