Agent skill

Building

by romiluz13 in romiluz13/cc10x

A skill your agent uses when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation…

MITAuto-check: notesTesting & QA

Install Building

skills CLI
$ npx skills add romiluz13/cc10x --skill building -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install romiluz13/cc10x building --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/romiluz13/cc10x.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/cc10x/skills/building .claude/skills/building && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
building
GitHub stars
164
Token cost
~2.7k tokens
SKILL.md length
1,449 words
Files
4 (incl. references)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation…

  • Works in 4 steps: Behavioral tests (does the function do… → Edge case tests (empty input, null,… → Integration tests (does it work with… → …
  • Writing production code test-first: the RED-GREEN-REFACTOR cycle
  • SKILL.md covers Reference Files, Test Process Discipline, RED → GREEN → REFACTOR and Study Project Patterns First, plus 10 more sections
  • Calls npx and npm

What it does

Building is an agent skill from romiluz13/cc10x. Use when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation patterns.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/integration-and-live-proof.md`, `references/test-data-and-mocks.md` and `references/testing-patterns.md`).

It sits in Testing & QA, covering Test-driven development. It works with Vitest. The repository describes itself as: The Loop Engine for Claude Code — engineer the loop, not the prompt. 1 router · 9 agents · 16 skills · 4 workflows. Fail-closed gates, test honesty, anti-anchored review. The licence is MIT.

When your agent uses it

  • Writing production code test-first: the RED-GREEN-REFACTOR cycle
  • False-RED detection
  • Vertical slicing
  • Scope escalation

Example prompts

  • “/building”

Requirements

  • Node.js
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash, Grep, Glob, LSP

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Behavioral tests (does the function do what it should?)
  2. Edge case tests (empty input, null, boundary values)
  3. Integration tests (does it work with real dependencies?)
  4. Performance tests (only if performance is a stated requirement)

What it can do on your machine

Read from SKILL.md and the folder at commit f346ebe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash
    • Grep
    • Glob
    • LSP

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Building loads about 2.7k tokens when it runs, and up to ~4.3k if it reads all its reference files. Until then it costs about 49 tokens; SKILL.md has 1,449 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash, Grep, Glob, LSP

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from romiluz13/cc10x at commit f346ebe, republished under its MIT licence (© romiluz13). 1,449 words, ~2,664 tokens.

Download SKILL.mdSave it as .claude/skills/building/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
building
description
Use when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation patterns.
allowed-tools
Read, Write, Edit, Bash, Grep, Glob, LSP
user-invocable
false

Building (Code Generation + TDD)

Iron Law: NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST.

Reference Files

Read only what's needed:

  • references/testing-patterns.md — test structure, isolation, naming; load when writing the first test of a cycle or a test feels awkward to structure
  • references/test-data-and-mocks.md — mock discipline, test data factories; load when a test needs fixtures/factories or you are about to mock anything
  • references/integration-and-live-proof.md — integration test guidance, live verification; load when the slice crosses a service/DB/API boundary or the plan names live proof

Test Process Discipline

  • Always use run mode: CI=true npm test, npx vitest run (NOT npx vitest), CI=true npx jest — watch mode never exits, so the agent hangs waiting for a prompt that never returns
  • Timeout guard: timeout 60s npx vitest run if uncertain about CI=true
  • After TDD cycle: pgrep -f "vitest|jest" || echo "Clean". Kill if found — orphaned watchers hold ports and re-run stale code, producing false greens in later cycles.
  • IDE vs CLI truth: If CLI tests pass with exit 0, trust CLI over IDE/LSP errors (stale cache)

RED → GREEN → REFACTOR

RED — Failing Test First

Write one failing test for the current slice. Run it. RED = a behavioral failure ("X is not a function", "expected 3, received undefined") — never a bare exit code.

False-RED guard (CRITICAL): Exit 1 from an import/syntax/collection ERROR is a broken harness, not a RED — fix the harness and re-run. Record the observed failure reason verbatim.

GREEN — Minimal Code

Write the minimum code to pass the test. No extra features, no abstractions for hypothetical futures. No unrelated test breakage. If existing tests break, fix the code not the tests.

REFACTOR — Clean Up

Improve code quality while keeping tests green. If tests fail during refactor, revert — a red test proves it wasn't a refactor, and debugging forward mixes two changes. Re-run after every refactor step.

Safety-Check Guard (MANDATORY): Never simplify away a safety check during refactoring. Safety checks include:

  • Input validation at trust boundaries (API entry points, user input, external data)
  • Error handling that prevents data loss or corruption
  • Security checks (auth, authorization, sanitization)
  • Accessibility checks (ARIA, keyboard navigation, semantic HTML)

If a safety check seems unnecessary, verify with a test that proves it's dead code before removing. "Looks redundant" is not sufficient evidence.

Vertical Slicing (CRITICAL)

Build in thin vertical slices that cross all layers: UI → API → logic → data → test. A horizontal slice (all UI, then all API, then all logic — or all tests first, then all implementation) defers integration risk to the end and produces untestable layers. Each slice should be independently verifiable and shippable.

Seam Discipline

One seam, one test, one minimal implementation per cycle. Each test is a tracer bullet that responds to what the last cycle taught you — work one vertical slice at a time.

Test only at pre-agreed seams. A seam is the place where a module's interface lives: where you can observe or alter behavior without editing in that place, which is how a test observes behavior without reaching inside (cc10x:codebase-design defines the term). Before writing any test, know which seam you're testing at. Prefer existing seams to new ones; use the highest seam possible; the fewer seams across the codebase, the better (ideal is one). If the plan provides a ### Test Seams subsection or an Interfaces block, draw your seams from there.

Implementation-coupled anti-pattern. A test is implementation-coupled if it mocks internal collaborators, tests private methods, or verifies through a side channel (querying the database instead of using the interface). The tell: the test breaks when you refactor but behavior hasn't changed. Test through the public interface, not internals.

Record your seams (enforced contract fields). Your Router Contract carries two seam fields:

  • TEST_SEAMS: [seam names you actually tested at]
  • SEAM_GATE_STATUS: "confirmed" | "proposed" | "disagreed" | "not_applicable"

Set SEAM_GATE_STATUS as follows:

  • confirmed — the plan provided test_seams and you used them (TEST_SEAMS non-empty, matching the plan).
  • proposed — no plan (direct/no-plan path) OR a legacy plan whose phase omits test_seams; you proposed seams at BUILD_PREFLIGHT (TEST_SEAMS non-empty).
  • disagreed — the plan's proposed seam cannot exercise the phase's real risk. Record the disagreement in DECISIONS and either propose a better seam (TEST_SEAMS non-empty with the better seam) or block on genuine ambiguity (STATUS: FAIL, REMEDIATION_REASON: "Ambiguous test surface — no seam exercises the real risk").
  • not_applicable — build_scope=trivial; no seam expectation.

The router validates these per build_scope (see the contract-override table). This is the enforced gate — not advisory.

Study Project Patterns First

Before writing code: read 2-3 existing similar components in the repo. Match naming, file structure, export style, test patterns. Follow the project's conventions — don't introduce a new pattern when an existing one works.

LSP before writing: Use LSP to find definitions, references, and type information before writing code that interfaces with existing modules.

Scope Escalation (SCOPE_INCREASES)

If the build scope grows beyond the approved phase — new files not in the plan, new dependencies, API contract changes — emit SCOPE_INCREASES: ["new scope item"] in the contract. The router decides whether to escalate to a full BUILD (with planner + reviewer) or approve the expansion.

Decision Checkpoints (return FAIL when triggered):

TriggerAction
Changing >3 files not in planFAIL with extra files named
Choosing between 2+ valid patternsFAIL with competing options
Breaking existing API contractFAIL with impacted callers
Adding dependency not in planFAIL with dependency name
Touching a later planned phase earlyFAIL with skipped phase
Show full SKILL.md (569 more words)Show less

Minimal Diffs

Write minimal diffs. A bug fix doesn't need surrounding cleanup. A one-shot operation doesn't need a helper. Don't add error handling, fallbacks, or validation for scenarios that cannot happen. Trust internal code and framework guarantees in production code — no runtime re-validation for scenarios the types or the framework already exclude; only validate at system boundaries. When your feature depends on a framework behavior, pin it with a test instead: guarantees have edge cases, and the test costs less than the defensive code.

Rationalization Table

ExcuseReality
"Too simple to test"Simple code breaks. Test takes 30 seconds.
"I'll write tests after""After" never comes. Write the test first — it IS the spec.
"It's just a refactor"Refactors break things. Run the tests before AND after.
"The existing tests cover this"Then your new test will pass immediately — that's a false RED.
"I manually tested it"Manual testing doesn't survive the next refactor or CI run.
"Adding tests would slow down delivery"Debugging untested code takes longer than writing the test.
"The framework handles this"Pin the depended-on behavior with a test — don't re-validate at runtime.

Red Flags — STOP and Reconsider

  • You're about to write production code without a failing test
  • You're skipping the RED step because "the test will obviously fail"
  • You're adding error handling for a scenario that can't happen
  • You're introducing an abstraction with only one caller
  • You're changing code unrelated to the current phase
  • You're about to commit without running the full test suite
  • You're considering deleting a test to make the build pass
  • You're adding a dependency not in the plan

Tautological Test Anti-Pattern

A tautological test recomputes the expected value the same way the code does — it passes by construction and can never disagree.

typescript
// BAD — tautological: recomputes expected value using same logic
const expected = items.reduce((sum, x) => sum + x.value, 0);
expect(calculateTotal(items)).toBe(expected);

// GOOD — expected value comes from an independent source of truth
expect(calculateTotal([{value: 10}, {value: 20}, {value: 30}])).toBe(60);

Rule: Expected values must come from a known-good literal, a worked example, or the spec — never from re-running the same algorithm the code uses.

Loop Caps

  • TDD Failure Cap: GREEN fails 3 consecutive times on same test → FAIL with error — three failures means the approach is wrong, not unlucky
  • Build/Lint Loop Cap: Same error recurs after 3 fix attempts → FAIL with error_code + file

Coverage Threshold

If coverage-thresholds.json exists, run coverage and compare. Below thresholds → FAIL. If no thresholds file, skip coverage check.

Test Prioritization

Ranked by bugs caught per token — behavioral tests catch the most; performance tests without a stated requirement are speculative work.

  1. Behavioral tests (does the function do what it should?)
  2. Edge case tests (empty input, null, boundary values)
  3. Integration tests (does it work with real dependencies?)
  4. Performance tests (only if performance is a stated requirement)

Design for Testability

If tests are hard to write, the code is hard to test — fix the code, not the test. Pure functions are easy to test. Side effects are hard. Isolate side effects at boundaries; keep core logic pure.

When Stuck

  • RED won't fail: check if the test is actually exercising the code path
  • GREEN won't pass: re-read the test, check if the assertion matches the requirement
  • Existing tests break: your change has a side effect you didn't expect — revert and isolate

No test runner: a scripted check with real exit codes is TDD evidence; manual browser verification is not. Never fabricate TDD_RED_EXIT or TDD_GREEN_EXIT: leave both null. For a Pure HTML/CSS/JS project with no runner and no scripted check, the rule is: require a runner or block (component-builder and bug-investigator define the exact return).

© romiluz13, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in plugins/cc10x/skills/building of romiluz13/cc10x.

  • SKILL.md
  • references/integration-and-live-proof.md
  • references/test-data-and-mocks.md
  • references/testing-patterns.md

Open the folder on GitHubat commit f346ebe

Compare with similar skills

Building next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Building compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Building this skillromiluz13/cc10x164—~2.7kAutomated safety check: NotesMIT
Agent Harness Testing Methodologyhuiliyi37/Tianshu-harness1.1k—~1kAutomated safety check: NotesApache-2.0
MoAI TDD Workflowmodu-ai/moai-adk1.2k—~3.1kAutomated safety check: PassApache-2.0
Test Coveragescott-fryxell/brayness124—~1.6kAutomated safety check: PassMIT
Odc TestingDouglasNeuroInformatics/OpenDataCapture119—~1.5kAutomated safety check: NotesApache-2.0
Test-Driven Development Enforcerzereight/gitlab-mcp2k1 repos~904Automated safety check: PassMIT

Similar skills

  • Agent Harness Testing Methodology

    huiliyi37/Tianshu-harness

    Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type.

    1.1k GitHub stars~1k tokensUpdated today
    Testing & QAAuto-check: notes
  • MoAI TDD Workflow

    modu-ai/moai-adk

    Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.

    1.2k GitHub stars~3.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Coverage

    scott-fryxell/brayness

    Write Vitest specs for Vue 3 JavaScript (Vite Plus, happy-dom, @vue/test-utils) and analyze V8 coverage + Fallow health to prioritize test-first refactors.

    124 GitHub stars~1.6k tokensUpdated 10 days ago
    Testing & QAAuto-check passed
  • Odc Testing

    DouglasNeuroInformatics/OpenDataCapture

    Test a change in Open Data Capture. An agent skill from DouglasNeuroInformatics/OpenDataCapture.

    119 GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check: notes
  • Enforces strict red-green-refactor, with a failing test first, the minimum code to pass it, then cleanup, and a quick reference for common test runners.

    2k GitHub starsUsed in 1 repo~904 tokens
    Testing & QAAuto-check passed
  • TDD Guide

    alirezarezvani/claude-skills

    Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…

    28k GitHub stars~3.4k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from romiluz13/cc10x

All 22 skills in this repo
  • Diff Driven Docs

    romiluz13/cc10x

    A skill your agent uses when a BUILD phase completes, a commit is staged, or a PR is about to be created, and the diff has not yet been reflected in documentation.

    164 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes
  • Planning

    romiluz13/cc10x

    A skill your agent uses when writing an execution plan or a decision RFC: task decomposition, context references, validation levels, risk-based testing, ADR format, plan completeness gate, and…

    164 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Verification

    romiluz13/cc10x

    A skill your agent uses when judging whether a task reached its goal, not just finished: the gate function, self-critique gate, validation levels, evidence array protocol, and goal-backward lens.

    164 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check: notes
  • Agent Common

    romiluz13/cc10x

    A skill your agent uses when a cc10x agent starts a task: the shared preamble for the memory protocol, the contract format, and the output rules.

    164 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Cc10x Guide

    romiluz13/cc10x

    Answers questions about cc10x itself — what it is, how to install and configure it, how the router, workflows, memory, and hooks operate, and how to troubleshoot.

    164 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Cc10x Router

    romiluz13/cc10x

    Routes build, debug, review, plan, QA, and triage requests through the cc10x workflows (task graphs, workflow artifacts, gates); it is the single entry point for cc10x code work.

    164 GitHub stars~20k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Building

What does Building do?

A skill your agent uses when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation…. Building is an agent skill from romiluz13/cc10x. Use when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation patterns.

When should I use Building?

Building fits situations like: writing production code test-first: the RED-GREEN-REFACTOR cycle; false-RED detection; vertical slicing; scope escalation.

How do I install Building in Claude Code?

Run `npx skills add romiluz13/cc10x --skill building -a claude-code`. Or copy the skill folder (plugins/cc10x/skills/building in romiluz13/cc10x) into .claude/skills/building in your project. Claude Code loads it when a task matches its description.

How do I install Building in Codex?

Run `npx skills add romiluz13/cc10x --skill building -a codex`. Or copy the skill folder (plugins/cc10x/skills/building in romiluz13/cc10x) into .agents/skills/building in your project. Codex loads it when a task matches its description.

Can I use Building in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add romiluz13/cc10x --skill building -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/building, .gemini/skills/building, .github/skills/building and .opencode/skills/building in your project.

What does Building need to run?

Going by SKILL.md and its folder, Building needs the command-line tools its instructions call (npx and npm). Our summary lists: Node.js. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Grep, Glob, LSP.

Does Building access the network?

SKILL.md contains no URLs. Its commands use npx and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Building safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Building use?

Building is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Building use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Building?

Skills that share tags, products or a category with Building: Agent Harness Testing Methodology (huiliyi37/Tianshu-harness, 1.1k stars), MoAI TDD Workflow (modu-ai/moai-adk, 1.2k stars), Test Coverage (scott-fryxell/brayness, 124 stars) and Odc Testing (DouglasNeuroInformatics/OpenDataCapture, 119 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Building?

romiluz13 (a GitHub user) maintains it in romiluz13/cc10x, which has 164 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 7, 2026.

Source: romiluz13/cc10x on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.