Agent skill

Evidence-First Development Loop

by AmazingAng in AmazingAng/old-coder

Replaces line-by-line code review with an approved executable spec and a gauntlet of tests, types, coverage, and mutation checks the code must survive.

MITAuto-check: notesTesting & QA

Install Evidence-First Development Loop

skills CLI
$ npx skills add AmazingAng/old-coder --skill old-coder -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AmazingAng/old-coder old-coder --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AmazingAng/old-coder.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/old-coder .claude/skills/old-coder && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
old-coder
GitHub stars
749
Token cost
~5.4k tokens
SKILL.md length
3,205 words
Files
5 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Replaces line-by-line code review with an approved executable spec and a gauntlet of tests, types, coverage, and mutation checks the code must survive.

  • Works in 6 steps: SPEC — the only thing the human reads… → RED — prove each test can fail → GREEN — minimal implementation → …
  • Doing high-assurance work in money, auth, or concurrency code the user won't read line by line
  • SKILL.md covers The Loop, Anti-Gaming Rules (absolute), Calibration and Independent verification (Tier…, plus 1 more section
  • Calls git

What it does

This skill runs a SPEC to RED to GREEN to REFACTOR to GAUNTLET to EVIDENCE loop, where the human approves only the spec, not the code, and that spec must be written as executable acceptance criteria: concrete inputs, concrete expected outputs, edge cases, and explicit invariants the change must not break, never a vague behavior description.

It is explicit that the gauntlet turns the spec's constraints into executable evidence without proving the spec covers everything that matters, and that evidence reports layered, auditable confidence rather than absolute proof, so skipping a gauntlet step destroys the basis for that trust. When paired with a companion API-contract skill, this skill owns the workflow order and evidence while the other owns the HTTP/JSON contract, run together rather than as two parallel workflows.

When your agent uses it

  • Doing high-assurance work in money, auth, or concurrency code the user won't read line by line
  • Writing an executable spec before implementing a feature
  • Proving a change ran a full gauntlet of tests, types, and mutation checks

Example prompts

  • “Implement this payment refund flow evidence-first, I won't read the code.”
  • “Write an executable spec for this auth change before you touch the implementation.”
  • “Prove this concurrency fix survived the full gauntlet.”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. SPEC — the only thing the human reads before code
  2. RED — prove each test can fail
  3. GREEN — minimal implementation
  4. REFACTOR — clean up under green, assertions frozen
  5. GAUNTLET — the constraint stack
  6. EVIDENCE — the only thing the human reads after code

What it can do on your machine

Read from SKILL.md and the folder at commit a0eb529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evidence-First Development Loop loads about 5.4k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 138 tokens; SKILL.md has 3,205 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~138
When it runs · the whole SKILL.md, loaded when a task matches
~5.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:312
    tree's `.env` or build outputs is not evidence about the landing tree.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AmazingAng/old-coder at commit a0eb529, republished under its MIT licence (© AmazingAng). 3,205 words, ~5,433 tokens.

Download SKILL.mdSave it as .claude/skills/old-coder/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
old-coder
description
Evidence-first development — surround the implementation with an executable spec and a gauntlet of constraints (tests, types, coverage, mutation) so line-by-line review becomes optional. Use when the user explicitly asks for high-assurance or evidence-first work ("reliable", "TDD", "prove it works", "I won't read the code"), or when the change touches high-stakes domains (money, auth, data loss, concurrency, public API). For routine changes where the user just wants normal tests, write good tests directly instead of invoking this loop.

Old Coder: Reliable Coding Under Constraint and Test

The human will NOT read your implementation. Their confidence comes entirely from two artifacts you produce: (1) an executable specification they approve before you write code, and (2) an evidence report proving the code ran the gauntlet. Your job is to make those two artifacts trustworthy enough that line-by-line review becomes optional within the spec's boundaries.

This inverts the normal review model: trust moves from inspection to constraints. Be honest about what that buys: the gauntlet turns the constraints the spec expresses into executable evidence — it cannot show the spec expresses everything that matters, and it is not self-authenticating, because a checker can be unsound and a mapping can claim more than it demonstrates. That is exactly why the human approves the SPEC (the one artifact that breaks the everything-authored-by-the-same-agent correlation), and why EVIDENCE reports layered, auditable confidence, never absolute proof. Every shortcut you take against the gauntlet destroys the only basis of trust.

Composition with old-coder-api: when both skills apply, this skill owns workflow order, SPEC approval, the gauntlet, and EVIDENCE; old-coder-api owns the HTTP/JSON contract. Run its scope check and API gates while drafting SPEC, turn the surviving constraints and risks into acceptance criteria and checks, then map those checks into EVIDENCE. Do not run two parallel workflows.

The Loop

SPEC → (human approves spec, not code) → RED → GREEN → REFACTOR → GAUNTLET → EVIDENCE
                                          ↑_____________________|
                                              repeat per behavior
1. SPEC — the only thing the human reads before code

Turn the request into executable acceptance criteria before touching implementation files:

  • Write behaviors as Gherkin-style scenarios or a named test list — concrete inputs, concrete expected outputs, edge cases, and error cases. "Handles bad input" is not a spec; divide(1, 0) raises ZeroDivisionError with message X is.
  • Include what the change must NOT do (invariants that must survive: existing tests, public API signatures, performance budgets if stated). These negative constraints are contract clauses like any scenario: each must end up mapped in EVIDENCE to a test, a gauntlet layer, or an explicit skipped-with-reason line — never silently absent from the mapping.
  • The spec doubles as the authorization point: include the setup plan — tools to install, git usage (init? checkpoint commit cadence?), files the gauntlet will add by path, and every new dependency with a one-line justification (prefer the standard library and deps already present; an unjustified package is a spec defect) — so approving the spec authorizes the environment changes in one step instead of N interruptions, and the human can veto a risky package before it is ever installed.
  • Show the spec to the human in plain language and get approval before writing implementation. In autonomous mode, state the spec in your response and proceed — but the correlation-breaking review never happened, so EVIDENCE must record spec approval: not obtained (autonomous run) and claim correspondingly lower confidence; the spec becomes the artifact the human reviews after the fact.
  • An answer to a question is not an approval. If you asked the human to decide something, they answered that question and nothing else. Their answer is an INPUT to the spec, and it CHANGES the spec — so any approval you held before the question is approval of a document that no longer exists. Questions and approval are two exchanges, in that order: fold the answers in, say what changed, show the revised spec, ask again. If you cannot quote the words that approved THIS spec, you do not have approval — an answer to your question, a "go ahead" about some other step, silence, and the request that started the task are none of them approval. The recommended-option shape makes this easy to get wrong: when the human picks the options you recommended, the spec looks unchanged and consent looks implied, and neither is true.
  • The spec is append-only during the task. If implementation reveals the spec was wrong, say so explicitly and revise it visibly — never silently drift.
  • Write the spec to a file and name it by absolute path. A relative path is not clickable in a terminal, so the human cannot open the one artifact they are being asked to approve. Same for EVIDENCE when you get there. The SPEC and Gherkin templates are in references/templates.md.
  • Commit the spec at approval where the repo's git conventions allow it — the setup plan is where that was authorized. Once the approved spec is a commit, later drift is literally a git diff. Without a durable spec, a compaction loses the approved contract while the code it authorized remains, and nobody can check whether a scenario was quietly dropped from the EVIDENCE mapping.
2. RED — prove each test can fail

Write the test for one behavior. Run it and watch it fail before writing the implementation. A test you never saw fail proves nothing — it may be testing nothing. Details that matter in practice:

  • If the module under test doesn't exist yet, create a stub that raises (e.g. NotImplementedError) so the test fails on behavior, not on import — a collection error is a weaker RED than an assertion failure.
  • Related behaviors may share one RED run, as long as each new test is individually observed failing.
  • If a new test passes immediately, it is either vacuous (fix it) or the behavior already exists. Don't just assert which — prove it: break the implementation with a one-off throwaway mutant, watch the test fail, restore. Then record it as pre-existing behavior kept as regression armor.
3. GREEN — minimal implementation

Write the least code that makes the failing test pass. Run the full suite, not just the new test.

4. REFACTOR — clean up under green, assertions frozen

Minimal code is often ugly code. While the suite is green, improve names, extract duplication, and simplify structure. What is frozen is behavioral assertions, not test files wholesale:

  • Implementation refactors touch no test files at all.
  • Test-structure refactors (extracting helpers and fixtures, deduplicating setup) are allowed as a separate step: assertions unchanged, suite green before and after, then rerun mutation to confirm the restructured tests still kill — a refactor that blunts the tests is a silent hole in the gauntlet.
  • Anything that requires editing an assertion isn't refactoring, it's a behavior change and belongs back in SPEC.

Run the suite after each refactor. Repeat RED→GREEN→REFACTOR per behavior.

5. GAUNTLET — the constraint stack

After all spec behaviors are green, run every applicable layer. Scale to the task (see "Calibration"), but never skip a layer silently — if a layer doesn't apply or a tool is unavailable, record that in the evidence report with the reason.

LayerWhat it catchesHow
Full test suiteregressionsproject's test command, zero NEW failures (baseline note below)
Static typeswhole classes of bugstsc / mypy / etc., zero new errors
Lint + formatlatent bugs, driftproject's linter, zero new warnings
Coverage on changed linesuntested code pathsevery changed/added line executed by a test; branch coverage where the tool supports it. Global % is vanity — changed-line coverage is the constraint. This layer must exit nonzero when its threshold is missed (--cov-fail-under, diff-cover --fail-under, equivalent): a layer that prints a percentage and exits 0 is a report, not a gauntlet layer, and it will sit there green while coverage falls
Mutation testingtests that assert nothingprefer the project's mutation tool (mutmut, cosmic-ray, Stryker, PIT…), which generates mutants from the syntax tree and cannot silently skip one. No tool available? Manual mutation, per references/gauntlet.md — introduce 3–5 plausible bugs one at a time; the suite must kill every one; restore after. A hand-rolled runner must prove it executed each mutant: a runner that can report a kill it never ran inflates the score and no red gauntlet will ever surface it
Property-based testsedge cases you didn't imaginefor parsing, math, serialization, anything with invariants (round-trip, idempotence, ordering) — add hypothesis/fast-check properties
Complexity budgetunmaintainable outputnew functions small and single-purpose; if a function needs a paragraph to explain, split it
Real execution"passes tests, doesn't run"actually run the app/CLI/endpoint once on a realistic input, not only the test harness
Supply chain & secretsvulnerable/unnecessary deps, leaked credentialswhen the dependency set changed: audit it (pip-audit / npm audit / govulncheck / cargo-audit) and check licenses; scan the diff for secrets; every new dependency must trace back to its SPEC justification. Also eyeball the capability diff: did the change start using network / subprocess / filesystem / env it didn't before?
Suite healthflaky or order-dependent testsrun the suite in randomized order (pytest-randomly etc.); repeat suspected flakes. Every EVIDENCE number rests on the suite being deterministic — a flaky suite quietly invalidates the report

Baseline note — on a repo with pre-existing failures, record the baseline first (which tests already fail, verbatim) and hold the line at zero NEW failures. Fixing unrelated pre-existing failures is scope creep: surface them, don't silently "improve" them.

Mutation caveat — kills are attributed to whichever test fails first, so a 7/7 kill score validates the suite as a whole, not every layer in it. In Tier 3, rerun the mutants against the property suite alone before claiming the properties verify anything; survivors there mean the invariants have blind spots (a common one: a one-sided invariant like "never exceeds limit" cannot catch fail-closed bugs — pair it with the opposite bound).

Checker note — the gauntlet is only as trustworthy as its checkers, and the dangerous checker failure is fail-open: nothing crashes, the layer prints pass. Off-the-shelf tools (pytest, mypy, tsc…) have earned their failure behavior; home-grown checks — grep gates, custom scripts, the manual mutation runner — have not, so two rules apply to them: (1) fail closed — a crash, an unreadable input, an unexpected exit code, or an item silently skipped inside gate code is a hard failure of the layer, never a pass; no || true, no 2>/dev/null, no bare fallthrough. (2) Prove it can fail before trusting its pass: run it once against a known-bad input (a negative control) and watch it fail — the RED principle applied to checkers, exactly like the throwaway mutant for an immediately-passing test. Record the control in EVIDENCE. Be precise about what that buys: a negative control proves one known-bad case reaches the checker's failure path. It does not prove the checker recognizes every violation of the constraint it claims to enforce. A grep gate can fail closed perfectly and still guard a spelling rather than a behavior. When the gate's coverage is narrower than the rule it serves, say so where the rule is written, rather than letting the rule imply more.

Prove a negative control is itself non-vacuous the same way you prove a test: temporarily remove or break the defence it validates, and watch the control go red. A control that passes with the defence removed is measuring nothing — this is a one-time proof, not a permanent extra layer.

Equivalent-mutant note — with a mutation tool, a survivor is not automatically a failure: some mutants are semantically equivalent to the original and cannot be killed. Classify such survivors as "equivalent, because <reason>" in EVIDENCE rather than adding a meaningless test to kill them — that would violate anti-gaming rule 4. Hand-written mutants (the manual procedure) get no such excuse: you chose them, so choose real bugs.

Show full SKILL.md (1,370 more words)Show less
6. EVIDENCE — the only thing the human reads after code

End with a report the human can trust without opening a single source file (template in references/templates.md):

  • The approved spec, with each behavior mapped to the test that verifies it.
  • Each gauntlet layer: the command run, and its actual result (pasted numbers, not adjectives). "All 47 tests pass, changed-line coverage 100% (31/31 lines), 5/5 manual mutants killed" — never "tests look good".
  • All numbers must come from one final fresh run executed after the last code edit — results from mid-task runs are stale and must not be reported.
  • The report must be reproducible from the repo alone: every command it cites (including the mutation script) must exist as a persisted file in the repo, not in a scratch directory or only in the conversation. Reproducible means: dev-tool versions pinned or recorded, one entry-point command that reruns every layer, and the source state identified (commit SHA, or a source-tree hash when git is absent).
  • Layers skipped, and why.
  • Anything that failed and how it was resolved, honestly. A gauntlet you passed on the first try and a gauntlet you fixed your way through are equally fine; a gauntlet you quietly weakened is the only failure.

Anti-Gaming Rules (absolute)

The gauntlet only creates trust if it cannot be gamed. These are hard rules:

  1. Never weaken a test to make it pass. Don't broaden assertions, add skips, raise tolerances, or delete a failing test. If a test seems wrong, that's a spec conversation — surface it, don't bury it.
  2. Never edit a test and the implementation in the same step to reach green. Change one, run, then the other. Simultaneous edits let you accidentally redefine correctness to match your bug.
  3. Never mock the unit under test or mock so much that the test only exercises the mocks. Mock boundaries (network, clock, filesystem), not logic.
  4. Never chase the coverage number. Coverage is a detector of untested code, not a target. A test added only to touch lines, with no meaningful assertion, is gaming — mutation testing exists precisely to catch this, including yours.
  5. Never report a layer you didn't run. An honest "skipped: no mutation tool in this environment, did manual mutation instead" preserves trust; an invented result destroys the entire scheme.
  6. Failing gauntlet blocks done. You are not finished while any layer fails. If you're genuinely blocked, report the failure verbatim as the outcome.

Calibration

Scale effort to blast radius, and say which tier you chose:

  • Tier 1 — trivial (typo, comment, config value): full suite + lint. No new tests required, but state why the change is untestable or already covered.
  • Tier 2 — normal (bug fix, small feature): full loop. Bug fixes MUST start with a RED test reproducing the bug — the fix is not done until yesterday's bug is tomorrow's regression test.
  • Tier 3 — high stakes (money, auth, data loss, concurrency, public API): start with a short failure model: list the ways this specific change can hurt (race condition, partial write, hostile input, overflow, unbounded growth, failed rollback…), and for each mode add a layer that can actually catch it — race/stress tests for concurrency, fuzzing for parsers, rollback rehearsal for migrations, benchmarks for latency budgets, API-compatibility checks for public libraries, contract tests for service boundaries, logging/metric assertions where silent production failure is a mode (full menu in references/gauntlet.md). Mutation and coverage cannot substitute for these; the generic gauntlet is the floor, not the ceiling. Then: full loop + property-based tests + mutation testing (tool-based if available) + adversarial pass — one explicit step trying to break your own implementation with hostile inputs before declaring done. Failure modes deliberately not covered go in EVIDENCE as known limits. The adversarial pass is you attacking your own work and shares your blind spots; where a spec gap would be expensive, consider independent verification below — a different kind of assurance, not another layer.

Independent verification (Tier 3 option, experimental)

The gauntlet is evidence, not self-authentication: its checkers can be unsound, its mappings can overclaim, and the spec can be incomplete. Human spec approval mitigates only the last, by breaking author correlation, and only before code exists — it does not make a spec complete.

Independent verification answers the rest where the stakes justify it: a fresh-context agent that attacks the finished work before EVIDENCE is signed. It reduces task-context correlation, not model correlation. It is not a gauntlet layer — a layer is an executable check with a machine-evaluable result; this is an agent returning prose a human must judge, spending the one resource this skill otherwise guards. Experimental: the evidence is one case study (references/verifier-case-study.md — for deciding whether to run this, not for the verifier to read), not a benchmark.

The protocol is references/verifier.md. Verification has not been performed until that file has been read in full and executed; missing or unreadable → blocked, never passed. What cannot be traded away:

  • Fresh context, blind first, four inputs only — the task contract, the approved SPEC, an exact source state, the entry point. Never your conversation. The draft EVIDENCE comes after its own results, not before.
  • It fixes nothing. A SPEC gap goes to the human, never to the builder to self-amend.
  • The human grades the findings. Behavioural findings are fixed and re-verified in a new context; description and mapping findings are fixed and disclosed without buying another round. Propose a grade if you like — the human decides any disputed or material one, and approves stopping at the cap. Self-grading is the obvious way to make this rule fail open.
  • Cap at two rounds, more only by explicit approval. The cap does not limit the spending; it makes the spending someone's decision.
  • Verification is source-state-specific. A state no verifier saw is not performed, whatever earlier rounds concluded. Fixing a behavioural finding after the final permitted round therefore ships an unverified state: record that as a declared downgrade and keep the earlier rounds as history.
  • Four states: passed finalizes; failed and blocked do not; not performed finalizes only as a declared downgrade, like an unapproved spec. On Tier 3 it needs no apology — say so and claim less.

Setup

Isolation — do not mutate the user's working tree to do your work. Declare the mechanism in the SPEC, with one line of why: a worktree, a branch, or none — the last only at Tier 1, where the blast radius is a typo. The human vetoes the mechanism at approval rather than discovering it afterwards.

The trap: a fresh worktree contains no gitignored content, so the gauntlet often cannot run there until dependencies are rebuilt. Two outcomes are acceptable — rebuild and run there, or fall back to a branch and record why. Never report green from a tree that never ran the suite.

Where the isolated tree and the tree the change lands in differ by ignored or untracked content, say so in EVIDENCE: a green run in a tree missing the landing tree's .env or build outputs is not evidence about the landing tree.

If the project has no test runner, no linter, or no type checking, set up the minimal standard toolchain for the language first (see references/gauntlet.md). A gauntlet can't run on bare ground. Setup changes the user's environment — packages, config files, lockfiles — so it belongs in the SPEC's setup plan, where spec approval authorizes it in one step; record every environment change actually made in the evidence report. If the user forbids adding tooling, fall back to manual layers (manual mutation, manual execution) and record the reduced confidence honestly.

If the directory is not a git repository, propose git init in the SPEC's setup plan. Version control is itself a gauntlet layer: commit at SPEC and at each GREEN/REFACTOR checkpoint, so mutant restores are verifiable with git diff (not by eyeball), a bad refactor is rolled back instead of debugged, and the final diff shows exactly what changed. Checkpoint commits happen only under that spec-approved authorization (or an explicit user request) — never impose a commit cadence on a repo whose owner hasn't agreed to it. If the user declines or git is unavailable, record that in EVIDENCE — mutant restores then rest on rerunning the suite, a weaker guarantee — and identify the source state with a tree hash instead of a SHA.

© AmazingAng, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/old-coder of AmazingAng/old-coder.

  • SKILL.md
  • references/gauntlet.md
  • references/templates.md
  • references/verifier-case-study.md
  • references/verifier.md

Open the folder on GitHubat commit a0eb529

Compare with similar skills

Evidence-First Development Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evidence-First Development Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evidence-First Development Loop this skillAmazingAng/old-coder749—~5.4kAutomated safety check: NotesMIT
Tapd Iteration RunnerTencentBlueKing/bk-bcs840—~2.5kAutomated safety check: PassCustom licence
Tapd Story ImplementTencentBlueKing/bk-bcs840—~1.2kAutomated safety check: PassCustom licence
Map TDDazalio/map-framework157—~6kAutomated safety check: PassMIT
SPARC Development Methodologyruvnet/ruflo74k2 repos~829Automated safety check: PassMIT
Tbdjlevy/strif131—~3.5kAutomated safety check: PassMIT

Similar skills

  • Tapd Iteration Runner

    TencentBlueKing/bk-bcs

    TAPD 迭代调度器——批量开发一个迭代中的全部需求。自动完成环境初始化、 需求依赖分析、逐需求调度 pipeline 实现、迭代汇总报告四段编排。

    840 GitHub stars~2.5k tokensUpdated 13 days ago
    Testing & QAAuto-check passed
  • Tapd Story Implement

    TencentBlueKing/bk-bcs

    迭代执行流水线代码实现阶段。基于 tasks.md 调用 /speckit.implement 以 TDD 模式完成全部任务。

    840 GitHub stars~1.2k tokensUpdated 13 days ago
    Testing & QAAuto-check passed
  • Map TDD

    azalio/map-framework

    TDD MAP workflow: write tests from the spec FIRST, then implement, so tests validate intent not implementation.

    157 GitHub stars~6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Applies the SPARC method (specification, pseudocode, architecture, refinement, completion) with 17 specialized modes and multi-agent orchestration, from research to deployment.

    74k GitHub starsUsed in 2 repos~829 tokens
    DevelopmentAuto-check passed
  • Tbd

    jlevy/strif

    Git-native issue tracking (beads), coding guidelines, knowledge injection, and spec-driven planning for AI agents.

    131 GitHub stars~3.5k tokensUpdated 4 mo ago
    DevelopmentAuto-check passed
  • Enforces red-green-refactor for Rust work, with idiomatic test patterns, a naming convention and a pre-commit gate of cargo fmt, clippy and test.

    83k GitHub stars~753 tokensUpdated yesterday
    Testing & QAAuto-check: notes

More from AmazingAng/old-coder

  • Old Coder API Design

    AmazingAng/old-coder

    Reviews or designs an HTTP/JSON API's endpoints, auth, pagination, versioning and deprecations, guarding against inventing a bespoke interface or silently breaking consumers.

    749 GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check passed

Questions about Evidence-First Development Loop

What does Evidence-First Development Loop do?

Replaces line-by-line code review with an approved executable spec and a gauntlet of tests, types, coverage, and mutation checks the code must survive. This skill runs a SPEC to RED to GREEN to REFACTOR to GAUNTLET to EVIDENCE loop, where the human approves only the spec, not the code, and that spec must be written as executable acceptance criteria: concrete inputs, concrete expected outputs, edge cases, and explicit invariants the change must not break, never a vague behavior description.

When should I use Evidence-First Development Loop?

Evidence-First Development Loop fits situations like: doing high-assurance work in money, auth, or concurrency code the user won't read line by line; writing an executable spec before implementing a feature; proving a change ran a full gauntlet of tests, types, and mutation checks.

How do I install Evidence-First Development Loop in Claude Code?

Run `npx skills add AmazingAng/old-coder --skill old-coder -a claude-code`. Or copy the skill folder (skills/old-coder in AmazingAng/old-coder) into .claude/skills/old-coder in your project. Claude Code loads it when a task matches its description.

How do I install Evidence-First Development Loop in Codex?

Run `npx skills add AmazingAng/old-coder --skill old-coder -a codex`. Or copy the skill folder (skills/old-coder in AmazingAng/old-coder) into .agents/skills/old-coder in your project. Codex loads it when a task matches its description.

Can I use Evidence-First Development Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AmazingAng/old-coder --skill old-coder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/old-coder, .gemini/skills/old-coder, .github/skills/old-coder and .opencode/skills/old-coder in your project.

What does Evidence-First Development Loop need to run?

Going by SKILL.md and its folder, Evidence-First Development Loop needs the command-line tools its instructions call (git).

Does Evidence-First Development Loop access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evidence-First Development Loop safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Evidence-First Development Loop use?

Evidence-First Development Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evidence-First Development Loop use?

About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.8k tokens, read only when the agent opens those files.

What are the alternatives to Evidence-First Development Loop?

Skills that share tags, products or a category with Evidence-First Development Loop: Tapd Iteration Runner (TencentBlueKing/bk-bcs, 840 stars), Tapd Story Implement (TencentBlueKing/bk-bcs, 840 stars), Map TDD (azalio/map-framework, 157 stars) and SPARC Development Methodology (ruvnet/ruflo, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evidence-First Development Loop?

AmazingAng (a GitHub user) maintains it in AmazingAng/old-coder, which has 749 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on August 18, 2026.

Source: AmazingAng/old-coder on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.