Agent Harness Testing Methodology
huiliyi37/Tianshu-harness
Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type.
A skill your agent uses when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation…
$ npx skills add romiluz13/cc10x --skill building -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install romiluz13/cc10x building --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/romiluz13/cc10x.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/cc10x/skills/building .claude/skills/building && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "building" agent skill from https://github.com/romiluz13/cc10x/tree/main/plugins/cc10x/skills/building into .claude/skills/building/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/romiluz13/cc10x/tree/main/plugins/cc10x/skills/buildingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add romiluz13/cc10x --skill building -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install romiluz13/cc10x building --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/romiluz13/cc10x.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/cc10x/skills/building .agents/skills/building && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "building" agent skill from https://github.com/romiluz13/cc10x/tree/main/plugins/cc10x/skills/building into .agents/skills/building/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add romiluz13/cc10x --skill building -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install romiluz13/cc10x building --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/romiluz13/cc10x.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/cc10x/skills/building .cursor/skills/building && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "building" agent skill from https://github.com/romiluz13/cc10x/tree/main/plugins/cc10x/skills/building into .cursor/skills/building/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/romiluz13/cc10x.git --path plugins/cc10x/skills/building--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add romiluz13/cc10x --skill building -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install romiluz13/cc10x building --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/romiluz13/cc10x.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/cc10x/skills/building .gemini/skills/building && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "building" agent skill from https://github.com/romiluz13/cc10x/tree/main/plugins/cc10x/skills/building into .gemini/skills/building/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install romiluz13/cc10x buildingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add romiluz13/cc10x --skill building -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/romiluz13/cc10x.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/cc10x/skills/building .github/skills/building && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "building" agent skill from https://github.com/romiluz13/cc10x/tree/main/plugins/cc10x/skills/building into .github/skills/building/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add romiluz13/cc10x --skill building -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install romiluz13/cc10x building --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/romiluz13/cc10x.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/cc10x/skills/building .opencode/skills/building && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "building" agent skill from https://github.com/romiluz13/cc10x/tree/main/plugins/cc10x/skills/building into .opencode/skills/building/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "building", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
buildingA skill your agent uses when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation…
Building is an agent skill from romiluz13/cc10x. Use when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation patterns.
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/integration-and-live-proof.md`, `references/test-data-and-mocks.md` and `references/testing-patterns.md`).
It sits in Testing & QA, covering Test-driven development. It works with Vitest. The repository describes itself as: The Loop Engine for Claude Code — engineer the loop, not the prompt. 1 router · 9 agents · 16 skills · 4 workflows. Fail-closed gates, test honesty, anti-anchored review. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f346ebe. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditBashGrepGlobLSPFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxnpmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx and npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Building loads about 2.7k tokens when it runs, and up to ~4.3k if it reads all its reference files. Until then it costs about 49 tokens; SKILL.md has 1,449 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, Write, Edit, Bash, Grep, Glob, LSPAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from romiluz13/cc10x at commit f346ebe, republished under its MIT licence (© romiluz13). 1,449 words, ~2,664 tokens.
.claude/skills/building/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Iron Law: NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST.
Read only what's needed:
references/testing-patterns.md — test structure, isolation, naming; load when writing the first test of a cycle or a test feels awkward to structurereferences/test-data-and-mocks.md — mock discipline, test data factories; load when a test needs fixtures/factories or you are about to mock anythingreferences/integration-and-live-proof.md — integration test guidance, live verification; load when the slice crosses a service/DB/API boundary or the plan names live proofCI=true npm test, npx vitest run (NOT npx vitest), CI=true npx jest — watch mode never exits, so the agent hangs waiting for a prompt that never returnstimeout 60s npx vitest run if uncertain about CI=truepgrep -f "vitest|jest" || echo "Clean". Kill if found — orphaned watchers hold ports and re-run stale code, producing false greens in later cycles.Write one failing test for the current slice. Run it. RED = a behavioral failure ("X is not a function", "expected 3, received undefined") — never a bare exit code.
False-RED guard (CRITICAL): Exit 1 from an import/syntax/collection ERROR is a broken harness, not a RED — fix the harness and re-run. Record the observed failure reason verbatim.
Write the minimum code to pass the test. No extra features, no abstractions for hypothetical futures. No unrelated test breakage. If existing tests break, fix the code not the tests.
Improve code quality while keeping tests green. If tests fail during refactor, revert — a red test proves it wasn't a refactor, and debugging forward mixes two changes. Re-run after every refactor step.
Safety-Check Guard (MANDATORY): Never simplify away a safety check during refactoring. Safety checks include:
If a safety check seems unnecessary, verify with a test that proves it's dead code before removing. "Looks redundant" is not sufficient evidence.
Build in thin vertical slices that cross all layers: UI → API → logic → data → test. A horizontal slice (all UI, then all API, then all logic — or all tests first, then all implementation) defers integration risk to the end and produces untestable layers. Each slice should be independently verifiable and shippable.
One seam, one test, one minimal implementation per cycle. Each test is a tracer bullet that responds to what the last cycle taught you — work one vertical slice at a time.
Test only at pre-agreed seams. A seam is the place where a module's interface lives: where you can observe or alter behavior without editing in that place, which is how a test observes behavior without reaching inside (cc10x:codebase-design defines the term). Before writing any test, know which seam you're testing at. Prefer existing seams to new ones; use the highest seam possible; the fewer seams across the codebase, the better (ideal is one). If the plan provides a ### Test Seams subsection or an Interfaces block, draw your seams from there.
Implementation-coupled anti-pattern. A test is implementation-coupled if it mocks internal collaborators, tests private methods, or verifies through a side channel (querying the database instead of using the interface). The tell: the test breaks when you refactor but behavior hasn't changed. Test through the public interface, not internals.
Record your seams (enforced contract fields). Your Router Contract carries two seam fields:
TEST_SEAMS: [seam names you actually tested at]SEAM_GATE_STATUS: "confirmed" | "proposed" | "disagreed" | "not_applicable"Set SEAM_GATE_STATUS as follows:
confirmed — the plan provided test_seams and you used them (TEST_SEAMS non-empty, matching the plan).proposed — no plan (direct/no-plan path) OR a legacy plan whose phase omits test_seams; you proposed seams at BUILD_PREFLIGHT (TEST_SEAMS non-empty).disagreed — the plan's proposed seam cannot exercise the phase's real risk. Record the disagreement in DECISIONS and either propose a better seam (TEST_SEAMS non-empty with the better seam) or block on genuine ambiguity (STATUS: FAIL, REMEDIATION_REASON: "Ambiguous test surface — no seam exercises the real risk").not_applicable — build_scope=trivial; no seam expectation.The router validates these per build_scope (see the contract-override table). This is the enforced gate — not advisory.
Before writing code: read 2-3 existing similar components in the repo. Match naming, file structure, export style, test patterns. Follow the project's conventions — don't introduce a new pattern when an existing one works.
LSP before writing: Use LSP to find definitions, references, and type information before writing code that interfaces with existing modules.
If the build scope grows beyond the approved phase — new files not in the plan, new dependencies, API contract changes — emit SCOPE_INCREASES: ["new scope item"] in the contract. The router decides whether to escalate to a full BUILD (with planner + reviewer) or approve the expansion.
Decision Checkpoints (return FAIL when triggered):
| Trigger | Action |
|---|---|
| Changing >3 files not in plan | FAIL with extra files named |
| Choosing between 2+ valid patterns | FAIL with competing options |
| Breaking existing API contract | FAIL with impacted callers |
| Adding dependency not in plan | FAIL with dependency name |
| Touching a later planned phase early | FAIL with skipped phase |
Write minimal diffs. A bug fix doesn't need surrounding cleanup. A one-shot operation doesn't need a helper. Don't add error handling, fallbacks, or validation for scenarios that cannot happen. Trust internal code and framework guarantees in production code — no runtime re-validation for scenarios the types or the framework already exclude; only validate at system boundaries. When your feature depends on a framework behavior, pin it with a test instead: guarantees have edge cases, and the test costs less than the defensive code.
| Excuse | Reality |
|---|---|
| "Too simple to test" | Simple code breaks. Test takes 30 seconds. |
| "I'll write tests after" | "After" never comes. Write the test first — it IS the spec. |
| "It's just a refactor" | Refactors break things. Run the tests before AND after. |
| "The existing tests cover this" | Then your new test will pass immediately — that's a false RED. |
| "I manually tested it" | Manual testing doesn't survive the next refactor or CI run. |
| "Adding tests would slow down delivery" | Debugging untested code takes longer than writing the test. |
| "The framework handles this" | Pin the depended-on behavior with a test — don't re-validate at runtime. |
A tautological test recomputes the expected value the same way the code does — it passes by construction and can never disagree.
// BAD — tautological: recomputes expected value using same logic
const expected = items.reduce((sum, x) => sum + x.value, 0);
expect(calculateTotal(items)).toBe(expected);
// GOOD — expected value comes from an independent source of truth
expect(calculateTotal([{value: 10}, {value: 20}, {value: 30}])).toBe(60);Rule: Expected values must come from a known-good literal, a worked example, or the spec — never from re-running the same algorithm the code uses.
If coverage-thresholds.json exists, run coverage and compare. Below thresholds → FAIL. If no thresholds file, skip coverage check.
Ranked by bugs caught per token — behavioral tests catch the most; performance tests without a stated requirement are speculative work.
If tests are hard to write, the code is hard to test — fix the code, not the test. Pure functions are easy to test. Side effects are hard. Isolate side effects at boundaries; keep core logic pure.
No test runner: a scripted check with real exit codes is TDD evidence; manual browser verification is not. Never fabricate TDD_RED_EXIT or TDD_GREEN_EXIT: leave both null. For a Pure HTML/CSS/JS project with no runner and no scripted check, the rule is: require a runner or block (component-builder and bug-investigator define the exact return).
© romiluz13, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in plugins/cc10x/skills/building of romiluz13/cc10x.
Open the folder on GitHubat commit f346ebe
Building next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Building this skillromiluz13/cc10x | 164 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Agent Harness Testing Methodologyhuiliyi37/Tianshu-harness | 1.1k | — | ~1k | Automated safety check: Notes | Apache-2.0 | |
| MoAI TDD Workflowmodu-ai/moai-adk | 1.2k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Test Coveragescott-fryxell/brayness | 124 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Odc TestingDouglasNeuroInformatics/OpenDataCapture | 119 | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | |
| Test-Driven Development Enforcerzereight/gitlab-mcp | 2k | 1 repos | ~904 | Automated safety check: Pass | MIT |
huiliyi37/Tianshu-harness
Guides an agent through probing an unfamiliar project's test setup, then choosing a red-light-first testing strategy matched to the task type.
modu-ai/moai-adk
Drives test-first development through the RED, GREEN, REFACTOR cycle, with a config switch that selects between TDD and a DDD workflow for existing code.
scott-fryxell/brayness
Write Vitest specs for Vue 3 JavaScript (Vite Plus, happy-dom, @vue/test-utils) and analyze V8 coverage + Fallow health to prioritize test-first refactors.
DouglasNeuroInformatics/OpenDataCapture
Test a change in Open Data Capture. An agent skill from DouglasNeuroInformatics/OpenDataCapture.
zereight/gitlab-mcp
Enforces strict red-green-refactor, with a failing test first, the minimum code to pass it, then cleanup, and a quick reference for common test runners.
alirezarezvani/claude-skills
Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…
romiluz13/cc10x
A skill your agent uses when a BUILD phase completes, a commit is staged, or a PR is about to be created, and the diff has not yet been reflected in documentation.
romiluz13/cc10x
A skill your agent uses when writing an execution plan or a decision RFC: task decomposition, context references, validation levels, risk-based testing, ADR format, plan completeness gate, and…
romiluz13/cc10x
A skill your agent uses when judging whether a task reached its goal, not just finished: the gate function, self-critique gate, validation levels, evidence array protocol, and goal-backward lens.
romiluz13/cc10x
A skill your agent uses when a cc10x agent starts a task: the shared preamble for the memory protocol, the contract format, and the output rules.
romiluz13/cc10x
Answers questions about cc10x itself — what it is, how to install and configure it, how the router, workflows, memory, and hooks operate, and how to troubleshoot.
romiluz13/cc10x
Routes build, debug, review, plan, QA, and triage requests through the cc10x workflows (task graphs, workflow artifacts, gates); it is the single entry point for cc10x code work.
Works with
Categories
A skill your agent uses when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation…. Building is an agent skill from romiluz13/cc10x. Use when writing production code test-first: the RED-GREEN-REFACTOR cycle, false-RED detection, vertical slicing, scope escalation, test process discipline, and code generation patterns.
Building fits situations like: writing production code test-first: the RED-GREEN-REFACTOR cycle; false-RED detection; vertical slicing; scope escalation.
Run `npx skills add romiluz13/cc10x --skill building -a claude-code`. Or copy the skill folder (plugins/cc10x/skills/building in romiluz13/cc10x) into .claude/skills/building in your project. Claude Code loads it when a task matches its description.
Run `npx skills add romiluz13/cc10x --skill building -a codex`. Or copy the skill folder (plugins/cc10x/skills/building in romiluz13/cc10x) into .agents/skills/building in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add romiluz13/cc10x --skill building -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/building, .gemini/skills/building, .github/skills/building and .opencode/skills/building in your project.
Going by SKILL.md and its folder, Building needs the command-line tools its instructions call (npx and npm). Our summary lists: Node.js. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Grep, Glob, LSP.
SKILL.md contains no URLs. Its commands use npx and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Building is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Building: Agent Harness Testing Methodology (huiliyi37/Tianshu-harness, 1.1k stars), MoAI TDD Workflow (modu-ai/moai-adk, 1.2k stars), Test Coverage (scott-fryxell/brayness, 124 stars) and Odc Testing (DouglasNeuroInformatics/OpenDataCapture, 119 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
romiluz13 (a GitHub user) maintains it in romiluz13/cc10x, which has 164 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 7, 2026.
Source: romiluz13/cc10x on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.