Agent skill

Test First

by NoobyGains in NoobyGains/godmode

A skill your agent uses when implementing any feature or bugfix, before writing implementation code

MITAuto-check passedTesting & QA

Install Test First

skills CLI
$ npx skills add NoobyGains/godmode --skill test-first -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NoobyGains/godmode test-first --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NoobyGains/godmode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-first .claude/skills/test-first && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-first
GitHub stars
107
Token cost
~2.6k tokens
SKILL.md length
1,108 words
Files
2
Skills in repo
34
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when implementing any feature or bugfix, before writing implementation code

  • Implementing any feature
  • SKILL.md covers Overview, When to Use, The Prime Directive and Red-Green-Refactor, plus 10 more sections
  • Calls npm
  • Before writing implementation code

What it does

Test First is an agent skill from NoobyGains/godmode. Use when implementing any feature or bugfix, before writing implementation code

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `testing-anti-patterns.md`).

It sits in Testing & QA, covering Test-driven development. The repository describes itself as: The AI development framework that thinks before it builds. 36 composable skills for Claude Code, Cursor, Codex, and OpenCode. The licence is MIT.

When your agent uses it

  • Implementing any feature
  • Before writing implementation code

Example prompts

  • “/test-first”

What it can do on your machine

Read from SKILL.md and the folder at commit 441103a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test First loads about 2.6k tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 1,108 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NoobyGains/godmode at commit 441103a, republished under its MIT licence (© NoobyGains). 1,108 words, ~2,633 tokens.

Download SKILL.mdSave it as .claude/skills/test-first/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
test-first
description
Use when implementing any feature or bugfix, before writing implementation code

Test-First Development

Overview

Write the test first. Observe it fail. Write the minimum code to pass.

Core principle: If you did not watch the test fail, you have no evidence it validates the correct behavior.

No exceptions. No workarounds. No shortcuts.

When to Use

Required:

  • New features
  • Bug fixes
  • Refactoring
  • Behavior modifications

Exceptions (confirm with your human partner):

  • Disposable prototypes
  • Generated code
  • Configuration files

Tempted to think "skip test-first just this once"? Stop. That is rationalization.

The Prime Directive

NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST

Wrote code before the test? Delete it. Start over.

Zero tolerance:

  • Do not keep it as "reference material"
  • Do not "adjust" it while writing tests
  • Do not glance at it
  • Delete means delete

Build fresh from tests. Full stop.

Red-Green-Refactor

dot
digraph tdd_loop {
    rankdir=LR;
    red [label="RED\nWrite failing test", shape=box, style=filled, fillcolor="#ffcccc"];
    check_red [label="Fails\ncorrectly?", shape=diamond];
    green [label="GREEN\nMinimal code", shape=box, style=filled, fillcolor="#ccffcc"];
    check_green [label="All tests\npass?", shape=diamond];
    refactor [label="REFACTOR\nClean up", shape=box, style=filled, fillcolor="#ccccff"];
    next [label="Next", shape=ellipse];

    red -> check_red;
    check_red -> green [label="yes"];
    check_red -> red [label="wrong\nfailure"];
    green -> check_green;
    check_green -> refactor [label="yes"];
    check_green -> green [label="no"];
    refactor -> check_green [label="stay\ngreen"];
    check_green -> next;
    next -> red;
}
RED - Write the Failing Test

Write one minimal test that describes the expected behavior.

<Good>
```typescript
test('retries failed operations up to 3 times', async () => {
  let callCount = 0;
  const unstableOp = () => {
    callCount++;
    if (callCount < 3) throw new Error('transient failure');
    return 'done';
  };

const outcome = await withRetry(unstableOp);

expect(outcome).toBe('done'); expect(callCount).toBe(3); });

Descriptive name, tests actual behavior, single concern
</Good>

<Bad>
```typescript
test('retry works', async () => {
  const spy = jest.fn()
    .mockRejectedValueOnce(new Error())
    .mockRejectedValueOnce(new Error())
    .mockResolvedValueOnce('done');
  await withRetry(spy);
  expect(spy).toHaveBeenCalledTimes(3);
});

Vague name, validates mock interactions not real behavior </Bad>

Criteria:

  • One behavior per test
  • Descriptive name
  • Uses real code (mocks only when unavoidable)
Verify RED - Watch It Fail

MANDATORY. Never skip.

bash
npm test path/to/test.test.ts

Confirm:

  • Test fails (not errors out)
  • Failure message matches expectations
  • Fails because the feature is missing (not due to typos)

Test passes? You are testing existing behavior. Rewrite the test.

Test errors? Fix the error, re-run until it fails correctly.

GREEN - Minimal Code

Write the simplest possible code that makes the test pass.

<Good>
```typescript
async function withRetry<T>(fn: () => Promise<T>): Promise<T> {
  for (let attempt = 0; attempt < 3; attempt++) {
    try {
      return await fn();
    } catch (e) {
      if (attempt === 2) throw e;
    }
  }
  throw new Error('unreachable');
}
```
Just enough to pass
</Good>
<Bad>
```typescript
async function withRetry<T>(
  fn: () => Promise<T>,
  opts?: {
    maxAttempts?: number;
    backoffStrategy?: 'linear' | 'exponential';
    onAttempt?: (n: number) => void;
  }
): Promise<T> {
  // YAGNI
}
```
Over-engineered
</Bad>

Do not add features, refactor unrelated code, or "improve" beyond what the test requires.

Verify GREEN - Watch It Pass

MANDATORY.

bash
npm test path/to/test.test.ts

Confirm:

  • Test passes
  • Other tests still pass
  • Output is clean (no errors, no warnings)

Test fails? Fix the code, not the test.

Other tests fail? Fix them now.

REFACTOR - Clean Up

Only after green:

  • Eliminate duplication
  • Improve naming
  • Extract helpers

Keep tests green. Do not add new behavior.

Repeat

Next failing test for the next behavior.

Quality Standards

QualityGoodBad
FocusedOne thing. "and" in the name? Split it.test('validates email and domain and whitespace')
DescriptiveName explains the behaviortest('test1')
IntentionalDemonstrates the desired APIObscures what the code should do

Why Sequence Matters

"I'll write tests afterward to confirm it works"

Tests written after code pass immediately. Passing immediately proves nothing:

  • May validate the wrong thing
  • May verify implementation details instead of behavior
  • May miss edge cases you overlooked
  • You never watched it catch the failure

Test-first forces you to observe the failure, proving the test actually verifies something.

"I already manually tested all the edge cases"

Manual testing is ad hoc. You believe you covered everything but:

  • No record of what you tested
  • Cannot rerun when code changes
  • Easy to skip cases under pressure
  • "It worked when I tried it" is not comprehensive

Automated tests are systematic. They execute identically every time.

"Deleting X hours of work is wasteful"

Sunk cost fallacy. The time is already spent. Your choice now:

  • Delete and rebuild with test-first (X more hours, high confidence)
  • Keep it and bolt on tests (30 min, low confidence, probable defects)

The real waste is keeping code you cannot trust. Working code without legitimate tests is technical debt.

"Test-first is dogmatic, being pragmatic means adapting"

Test-first IS pragmatic:

  • Catches defects before commit (faster than debugging later)
  • Prevents regressions (tests catch breakage immediately)
  • Documents behavior (tests demonstrate how code works)
  • Enables refactoring (change freely, tests catch problems)

"Pragmatic" shortcuts = debugging in production = slower.

"Tests after achieve the same goals - it's about the spirit not the ritual"

No. Tests-after answer "What does this do?" Tests-first answer "What should this do?"

Tests-after are biased by your implementation. You verify what you built, not what is required. You check the edge cases you remember, not ones you discover.

Tests-first force edge case discovery before implementing. Tests-after verify you remembered everything (you did not).

30 minutes of tests after is not test-first. You get coverage, but lose proof that tests work.

Show full SKILL.md (461 more words)Show less

Cognitive Traps

RationalizationWhat Is Actually True
"Too simple to test"Simple code breaks. A test takes 30 seconds.
"I'll test afterward"Tests passing immediately prove nothing.
"Tests after achieve the same thing"Tests-after = "what does this do?" Tests-first = "what should this do?"
"I already tested manually"Ad hoc is not systematic. No record, cannot rerun.
"Deleting X hours is wasteful"Sunk cost fallacy. Keeping unverified code is technical debt.
"Keep as reference, write tests first"You will adapt it. That is testing after. Delete means delete.
"Need to explore first"Fine. Discard the exploration, then start with test-first.
"Hard to test = I should skip it"Listen to the test. Hard to test = hard to use. Simplify the design.
"Test-first will slow me down"Test-first is faster than debugging. Pragmatic = test-first.
"Manual testing is faster"Manual testing does not prove edge cases. You retest every change.
"Existing code has no tests"You are improving it. Add tests for what exists.

Guardrails - STOP and Start Over

  • Code before test
  • Test after implementation
  • Test passes immediately
  • Cannot explain why the test failed
  • Tests added "later"
  • Rationalizing "just this once"
  • "I already tested it manually"
  • "Tests after serve the same purpose"
  • "It's about the spirit not the ritual"
  • "Keep as reference" or "adapt existing code"
  • "Already invested X hours, deleting is wasteful"
  • "Test-first is dogmatic, I'm being pragmatic"
  • "This is different because..."

All of these mean: Delete the code. Start over with test-first.

Example: Bug Fix

Bug: Empty email accepted by form

RED

typescript
test('rejects empty email', async () => {
  const result = await processForm({ email: '' });
  expect(result.error).toBe('Email required');
});

Verify RED

bash
$ npm test
FAIL: expected 'Email required', got undefined

GREEN

typescript
function processForm(data: FormData) {
  if (!data.email?.trim()) {
    return { error: 'Email required' };
  }
  // ...
}

Verify GREEN

bash
$ npm test
PASS

REFACTOR Extract a validation pipeline for multiple fields if needed.

Completion Checklist

Before declaring work finished:

  • Every new function/method has a test
  • Watched each test fail before implementing
  • Each test failed for the expected reason (feature missing, not typo)
  • Wrote minimal code to pass each test
  • All tests pass
  • Output is clean (no errors, no warnings)
  • Tests use real code (mocks only when unavoidable)
  • Edge cases and error paths are covered

Cannot check every box? You skipped test-first. Start over.

When Stuck

ProblemSolution
Do not know how to write the testWrite the API you wish existed. Write the assertion first. Ask your human partner.
Test too complicatedDesign too complicated. Simplify the interface.
Must mock everythingCode too coupled. Use dependency injection.
Test setup is enormousExtract helpers. Still heavy? Simplify the design.

Debugging Integration

Bug found? Write a failing test that reproduces it. Follow the test-first cycle. The test both proves the fix and prevents regression.

Never fix bugs without a test.

Testing Anti-Patterns

When adding mocks or test utilities, read @testing-anti-patterns.md to avoid common pitfalls:

  • Validating mock behavior instead of real behavior
  • Adding test-only methods to production classes
  • Mocking without understanding dependency chains

Final Rule

Production code -> test exists and failed first
Otherwise -> not test-first

No exceptions without your human partner's permission.

© NoobyGains, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/test-first of NoobyGains/godmode.

  • SKILL.md
  • testing-anti-patterns.md

Open the folder on GitHubat commit 441103a

Compare with similar skills

Test First next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test First compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test First this skillNoobyGains/godmode107—~2.6kAutomated safety check: PassMIT
TDDfossasia/eventyay-interpretation1.6k28 repos~1.1kAutomated safety check: PassApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
TDDsanity-io/sanity6.4k20 repos~1kAutomated safety check: PassMIT
Test Driven Developmentfarm-fe/farm5.6k49 repos~2.5kAutomated safety check: PassMIT
Tapd Story PipelineTencentBlueKing/bk-bcs840—~2.6kAutomated safety check: PassCustom licence

Similar skills

  • TDD

    fossasia/eventyay-interpretation

    Test-driven development. An agent skill from fossasia/eventyay-interpretation.

    1.6k GitHub starsUsed in 28 repos~1.1k tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • TDD

    sanity-io/sanity

    Official

    Test-driven development with red-green-refactor loop. An agent skill from sanity-io/sanity.

    6.4k GitHub starsUsed in 20 repos~1k tokens
    Testing & QAAuto-check passed
  • A skill your agent uses when implementing any feature or bugfix, before writing implementation code

    5.6k GitHub starsUsed in 49 repos~2.5k tokens
    Testing & QAAuto-check passed
  • Tapd Story Pipeline

    TencentBlueKing/bk-bcs

    单需求实现流水线——把一个 TAPD 需求从零推进到代码提交。自动串联技术澄清、 开发计划、任务拆分、TDD 实现、架构/安全校验、代码提交六个阶段。

    840 GitHub stars~2.6k tokensUpdated 13 days ago
    Testing & QAAuto-check passed
  • Absolute Init

    maddhruv/absolute

    One-time setup for absolute: interview how you want it to behave (output style, autonomy, TDD strictness, spec dir, families) + detect the stack once, then write .absolute.config.json (project…

    218 GitHub starsUsed in 1 repo~3k tokens
    Testing & QAAuto-check passed

More from NoobyGains/godmode

All 34 skills in this repo
  • Activation

    NoobyGains/godmode

    A skill your agent uses when starting any conversation - establishes how to locate and invoke skills, mandating Skill tool usage before ANY response including clarifying questions

    107 GitHub stars~2.4k tokensUpdated 7 mo ago
    Auto-check passed
  • Agent Messaging

    NoobyGains/godmode

    A skill your agent uses when dispatching subagents, composing prompts for teammates, structuring handoff reports, or managing context boundaries between agents.

    107 GitHub stars~3k tokensUpdated 7 mo ago
    Auto-check passed
  • Codebase Research

    NoobyGains/godmode

    A skill your agent uses when building ANY feature within an existing project - search the current codebase for existing patterns, conventions, similar implementations, and established approaches…

    107 GitHub stars~3.2k tokensUpdated 7 mo ago
    Auto-check: notes
  • Completion Gate

    NoobyGains/godmode

    A skill your agent uses when about to declare work done, fixed, or passing, before committing or opening PRs - demands executing verification commands and reading their output before making any…

    107 GitHub stars~1.6k tokensUpdated 7 mo ago
    Auto-check passed
  • Comprehension Check

    NoobyGains/godmode

    A skill your agent uses when implementing any substantial feature, multi-file modification, or architectural change - produces a plain-language walkthrough of every alteration so the developer can…

    107 GitHub stars~1.5k tokensUpdated 7 mo ago
    Auto-check passed
  • Delegated Execution

    NoobyGains/godmode

    A skill your agent uses when executing implementation plans with independent tasks in the current session

    107 GitHub stars~2.4k tokensUpdated 7 mo ago
    Auto-check passed

Categories

Questions about Test First

What does Test First do?

A skill your agent uses when implementing any feature or bugfix, before writing implementation code. Test First is an agent skill from NoobyGains/godmode.

When should I use Test First?

Test First fits situations like: implementing any feature; before writing implementation code.

How do I install Test First in Claude Code?

Run `npx skills add NoobyGains/godmode --skill test-first -a claude-code`. Or copy the skill folder (skills/test-first in NoobyGains/godmode) into .claude/skills/test-first in your project. Claude Code loads it when a task matches its description.

How do I install Test First in Codex?

Run `npx skills add NoobyGains/godmode --skill test-first -a codex`. Or copy the skill folder (skills/test-first in NoobyGains/godmode) into .agents/skills/test-first in your project. Codex loads it when a task matches its description.

Can I use Test First in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NoobyGains/godmode --skill test-first -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-first, .gemini/skills/test-first, .github/skills/test-first and .opencode/skills/test-first in your project.

What does Test First need to run?

Going by SKILL.md and its folder, Test First needs the command-line tools its instructions call (npm).

Does Test First access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test First safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test First use?

Test First is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test First use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test First?

Skills that share tags, products or a category with Test First: TDD (fossasia/eventyay-interpretation, 1.6k stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), TDD (sanity-io/sanity, 6.4k stars) and Test Driven Development (farm-fe/farm, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test First?

NoobyGains (a GitHub user) maintains it in NoobyGains/godmode, which has 107 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on March 9, 2026.

Source: NoobyGains/godmode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.