Agent skill

Test Health

by sd0xdev in sd0xdev/sd0x-harness

Holistic test coverage measurement. An agent skill from sd0xdev/sd0x-harness.

MITAuto-check passedTesting & QA

Install Test Health

skills CLI
$ npx skills add sd0xdev/sd0x-harness --skill test-health -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sd0xdev/sd0x-harness test-health --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sd0xdev/sd0x-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-health .claude/skills/test-health && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-health
GitHub stars
192
Token cost
~2.5k tokens
SKILL.md length
759 words
Files
7 (incl. scripts, references)
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Holistic test coverage measurement. An agent skill from sd0xdev/sd0x-harness.

  • Works in 4 steps: Test Inventory: Count test files by… → Coverage Artifacts: Scan for existing… → Trend Delta: Read previous snapshot,… → …
  • : assessing test health
  • SKILL.md covers Trigger, When NOT to Use, Workflow and Modes, plus 11 more sections
  • Runs JavaScript scripts from its folder; calls bash, git and go

What it does

Test Health is an agent skill from sd0xdev/sd0x-harness. Holistic test coverage measurement. Use when: assessing test health, measuring coverage trends, quantitative + qualitative test audit. Not for: running tests (use verify), reviewing test sufficiency only (use codex-test-review), generating tests (use codex-test-gen). Output: multi-dimensional dashboard with coverage metrics + test inventory + trend.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/artifact-formats.md`, `references/test-count-parsers.md` and `references/trend-schema.md`).

It sits in Testing & QA, covering Test generation and Test coverage. The repository describes itself as: The harness layer for Claude Code — a reference implementation of harness engineering with hook-enforced dual review, state-machine gates that survive context compaction, and… The licence is MIT.

When your agent uses it

  • : assessing test health
  • Measuring coverage trends
  • Quantitative + qualitative test audit

Example prompts

  • “/test-health”

Requirements

  • Python 3
  • Node.js
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash(bash:*), Bash(git:*), Bash(node:*), Bash(npm:*), Bash(pnpm:*), Bash(yarn:*), Bash(npx:*), Bash(stat:*), Bash(find:*), Bash(python*:*), Bash(pytest:*), Bash(cargo:*), Bash(go:*), Skill, Agent

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Test Inventory: Count test files by layer using Glob (see references/test-count-parsers.md for layer classification). If --scope…
  2. Coverage Artifacts: Scan for existing coverage artifacts (see references/artifact-formats.md). If --scope specified, scan within scope…
  3. Trend Delta: Read previous snapshot, compute delta (see references/trend-schema.md). Skip if --no-trend flag is set.
  4. Output: Quick Dashboard.

What it can do on your machine

Read from SKILL.md and the folder at commit a4d4bc1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash(bash:*)
    • Bash(git:*)
    • Bash(node:*)
    • Bash(npm:*)
    • Bash(pnpm:*)
    • Bash(yarn:*)
    • Bash(npx:*)

    …and 8 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • git
    • go

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Health loads about 2.5k tokens when it runs, and up to ~5k if it reads all its reference files. Until then it costs about 91 tokens; SKILL.md has 759 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from sd0xdev/sd0x-harness at commit a4d4bc1, republished under its MIT licence (© sd0xdev). 759 words, ~2,531 tokens.

Download SKILL.mdSave it as .claude/skills/test-health/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
test-health
description
Holistic test coverage measurement. Use when: assessing test health, measuring coverage trends, quantitative + qualitative test audit. Not for: running tests (use verify), reviewing test sufficiency only (use codex-test-review), generating tests (use codex-test-gen). Output: multi-dimensional dashboard with coverage metrics + test inventory + trend.
allowed-tools
Read, Grep, Glob, Bash(bash:*), Bash(git:*), Bash(node:*), Bash(npm:*), Bash(pnpm:*), Bash(yarn:*), Bash(npx:*), Bash(stat:*), Bash(find:*), Bash(python*:*), Bash(pytest:*), Bash(cargo:*), Bash(go:*), Skill, Agent

Test Health — Holistic Coverage Measurement

Trigger

  • Keywords: test health, coverage measurement, test metrics, coverage trend, test inventory, holistic test audit

When NOT to Use

ScenarioAlternative
Run tests/verify
Review test sufficiency only/codex-test-review
Generate unit tests/codex-test-gen
Feature-doc coverage only/check-coverage
Context-aware test execution + triage/test-deep

Workflow

mermaid
flowchart TD
    U[User: /test-health] --> M{Mode?}
    M --> |quick| Q[Quick Mode]
    M --> |--full| F[Full Mode]

    Q --> Q1[Test Inventory]
    Q1 --> Q2[Consume Coverage Artifacts]
    Q2 --> Q3[Trend Delta]
    Q3 --> QR[Quick Dashboard]

    F --> A[Phase A: /check-coverage]
    A --> B[Phase B: Coverage Collection]
    B --> C[Phase C: /codex-test-review]
    C --> D[Phase D: Aggregate Dashboard]
    D --> T[Trend Snapshot]
    T --> FR[Full Dashboard]

Modes

ModeTriggerContentDuration
quick (default)/test-healthTest inventory + consume artifacts + trend delta<15s
full/test-health --fullPhase A→B→C→D (feature coverage + instrumentation + qualitative + aggregation)2-5min

Quick Mode Workflow

  1. Test Inventory: Count test files by layer using Glob (see references/test-count-parsers.md for layer classification). If --scope <path> specified, limit Glob to that directory. If verify-runner cache exists (.claude/cache/verify/), read historical logs for test counts.
  2. Coverage Artifacts: Scan for existing coverage artifacts (see references/artifact-formats.md). If --scope specified, scan within scope only. Never execute project commands in quick mode.
  3. Trend Delta: Read previous snapshot, compute delta (see references/trend-schema.md). Skip if --no-trend flag is set.
  4. Output: Quick Dashboard.

Full Mode Workflow

Phase A: Feature Coverage

Resolve docs path using bash scripts/resolve-feature.sh (same cascade as other skills) — the shim over the wrapper, which emits the full shape with scan_error: true rather than a bare {} however the CLI fails: nonzero exit, signal, partial write, or a payload that is not the agreed shape. It cannot cover node itself being unavailable — the shim would exit 127 with no JSON — so treat an empty or non-JSON reply as a failure too. Gate on scan_error !== false before reading anything else — never on === true, because an empty or non-JSON reply carries no such field at all and the stricter test is false for it. Only once the flag is exactly false does any other field mean what it says: the failure payload sets has_tech_spec false along with everything else, so branching on that field first reports an unreadable corpus as a feature with no documents, and the coverage of a real feature disappears behind a reassuring advisory.

PayloadPhase A
scan_error !== false (including an empty or non-JSON reply)Skip, advisory "Phase A skipped: feature docs could not be read (scan_error) — coverage is unknown, not absent"
scan_error: false, has_tech_spec: trueDispatch /check-coverage <docs_path> via Skill tool
scan_error: false, feature unresolved or no tech specSkip, advisory "Phase A skipped: no feature docs detected"
Phase B: Test Inventory + Coverage Collection
  1. Count test files by layer (same as quick mode)
  2. If --collect flag: execute project coverage command (test:coverage or coverage from package.json)
  3. Otherwise: consume existing coverage artifacts (same as quick mode)
  4. Parse test runner stdout for test counts (see references/test-count-parsers.md)
Phase C: Qualitative Review

Dispatch /codex-test-review via Skill tool for 5-dimension quality assessment.

Phase D: Aggregate + Trend
  1. Aggregate all dimensions into full dashboard
  2. Write trend snapshot (see references/trend-schema.md)
  3. Output Full Dashboard
Show full SKILL.md (302 more words)Show less

Coverage Collection Strategy (Consume-First)

PriorityMethodTriggerOutput
1Consume existing artifactDefault (quick + full)source_type: instrumented_artifact
2Run project coverage command--collect flag only (opt-in)source_type: collected_now
3Heuristic proxy (test/source file ratio)No artifact and no --collectsource_type: heuristic

Prohibited: Never auto-install coverage tools (c8, nyc, istanbul, pytest-cov, tarpaulin, jacoco).

Output: Quick Dashboard

markdown
## Test Health (Quick)

### Test Inventory
| Layer | Files | Tests | Source |
|-------|-------|-------|--------|
| Unit  | 25    | 47    | cached_stdout |
| Integration | 1 | 12  | cached_stdout |
| E2E   | 0     | —     | file_count |

### Code Coverage
| Metric | Value | Tool | Freshness |
|--------|-------|------|-----------|
| Lines  | 82.3% | c8   | current   |
| Branches | 76.0% | c8 | current   |

### Trend (vs previous)
| Metric | Previous | Current | Delta |
|--------|----------|---------|-------|
| Line coverage | 80.2% | 82.3% | +2.1% |
| Test count | 57 | 59 | +2 |

### Quick Verdicts
| Dimension | Status |
|-----------|--------|
| Has tests for changed files | OK |
| Coverage artifact exists | OK |
| Trend direction | Improving |

Output: Full Dashboard

markdown
## Test Health Report (Full)

### Phase A: Feature Coverage
(from /check-coverage): 12/15 documented features have tests (80%)

### Phase B: Code Coverage + Inventory
| Layer | Files | Tests | Passed | Failed | Duration |
|-------|-------|-------|--------|--------|----------|
| Unit  | 25    | 47    | 45     | 2      | 12s      |
| Integration | 1 | 12  | 12     | 0      | 45s      |
| E2E   | 0     | 0     | —      | —      | —        |

| Metric | Value | Source | Tool | Freshness |
|--------|-------|--------|------|-----------|
| Lines  | 82.3% | instrumented_artifact | c8 | current HEAD |
| Branches | 76.0% | instrumented_artifact | c8 | current HEAD |

### Phase C: Quality Findings
(from /codex-test-review):
| Dimension | Rating |
|-----------|--------|
| Happy path | 4/5 |
| Error handling | 3/5 |
| Edge cases | 3/5 |
| Mock quality | 4/5 |

### Phase D: Aggregate Dashboard

#### Trend (vs last 5 runs)
| Run | Date | Line Cov | Tests | Delta |
|-----|------|----------|-------|-------|
| a1b2c3d | 04-01 | 82.3% | 59 | +2.1% / +2 |
| f4e5d6c | 03-31 | 80.2% | 57 | -0.5% / +0 |

#### Verdicts
| Dimension | Status | Detail |
|-----------|--------|--------|
| Test inventory | WARN | No E2E tests |
| Code coverage | OK | 82.3% lines (instrumented) |
| Feature coverage | OK | 80% features covered |
| Quality | WARN | 1 P2 finding |
| Trend | OK | Improving over last 3 runs |
| Changed-file coverage | OK | All changed files have tests |

Anti-Coverage-Theater Guardrails

RuleDescription
No composite score in v1Multi-dimensional dashboard, no single blended number
Changed-file focusPrioritize git diff files for coverage check
Source transparencyEvery metric tagged: instrumented / heuristic / missing
Qualitative couplingFull mode always runs Phase C even if quantitative metrics are green
Tool change detectiontool_id change resets trend line
Stale detectionArtifact older than HEAD marked stale

Gate Policy

PolicyBehavior
Advisory (default)Output dashboard + verdicts, do not block
Strict (v2, opt-in)Changed files with zero tests block

v1 implements advisory mode only.

Orchestrator Integration

SkillInteractionRelationship
/check-coveragePhase A: feature-doc coverageSub-step
/codex-test-reviewPhase C: qualitative reviewSub-step
/verifyPhase B: reference output or trigger test:coverageOptional sub-step
/test-deepIndependent (execution + triage)Peer
/pre-pr-auditQuick mode as non-blocking signalConsumer

Cross-Ecosystem Support

EcosystemDetectionCoverage ArtifactTest Count Parser
Node.jspackage.jsoncoverage/ dir (LCOV/Istanbul/Jest)node:test / jest / vitest
Pythonpyproject.toml / setup.pycoverage.xmlpytest
Gogo.modcover.outgo test -json
RustCargo.tomltarpaulin-report.json / cobertura.xmlcargo test
Javabuild.gradle / pom.xmlbuild/reports/jacoco/gradle/maven
Unknown—Scan for lcov.info / cobertura.xmlFile count fallback

Graceful degradation: no artifact + no coverage command = heuristic proxy + source_type: heuristic.

Verification

  • Quick mode completes in <15s without executing project commands
  • Full mode orchestrates Phase A→B→C→D in sequence
  • Coverage artifact consumed correctly (or graceful fallback)
  • Trend snapshot written to .claude/cache/test-health/
  • Dashboard output includes all dimensions with source transparency

References

FilePurpose
references/artifact-formats.mdCoverage artifact formats + scan + freshness
references/trend-schema.mdTrend storage schema + lock + comparison rules
references/test-count-parsers.mdFramework output parsers + layer classification

© sd0xdev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in skills/test-health of sd0xdev/sd0x-harness.

  • SKILL.md
  • references/artifact-formats.md
  • references/test-count-parsers.md
  • references/trend-schema.md
  • scripts/artifact-parser.js
  • scripts/count-parser.js
  • scripts/trend.js

Open the folder on GitHubat commit a4d4bc1

Compare with similar skills

Test Health next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Health compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Health this skillsd0xdev/sd0x-harness192—~2.5kAutomated safety check: PassMIT
E2E Test ThinkerUniClipboard/UniClipboard1.9k—~1.7kAutomated safety check: PassAGPL-3.0
Ralph Coveragejvm-skills/jvm-skills140—~683Automated safety check: PassApache-2.0
Designing TestsCloudAI-X/opencode-workflow275—~2.9kAutomated safety check: PassMIT
Mutation Testingproffesor-for-testing/agentic-qe494—~1.7kAutomated safety check: PassMIT
Write Testsgnomeria/usbtree690—~622Automated safety check: PassMIT

Similar skills

  • E2E Test Thinker

    UniClipboard/UniClipboard

    Analyze the current branch's diff against main and determine which changes are testable via CLI-based end-to-end tests.

    1.9k GitHub stars~1.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Ralph Coverage

    jvm-skills/jvm-skills

    Run Ralph in coverage mode — iteratively write tests for untested classes until coverage targets are met.

    140 GitHub stars~683 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/opencode-workflow

    Guides test strategy, TDD/BDD approaches, test coverage planning, and testing best practices.

    275 GitHub stars~2.9k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Mutation Testing

    proffesor-for-testing/agentic-qe

    Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

    494 GitHub stars~1.7k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Write Tests

    gnomeria/usbtree

    Author tests that match the repo's stack and existing test style, at the cheapest level that catches the regression.

    690 GitHub stars~622 tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Mutation Test

    jmagly/aiwg

    Run mutation testing to validate test quality beyond code coverage.

    220 GitHub stars~3.2k tokensUpdated 2 days ago
    Testing & QAAuto-check passed

More from sd0xdev/sd0x-harness

All 89 skills in this repo
  • Adr

    sd0xdev/sd0x-harness

    Write an Architecture Decision Record (ADR) for a feature — Context / Decision / Status / Consequences / Alternatives, filed as docs/features/<feature/adr-<NNN-<title.md with a 3-digit zero-padded…

    192 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Load PR Review

    sd0xdev/sd0x-harness

    Load GitHub PR review comments into AI session — analyze, triage, plan.

    192 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Next Step

    sd0xdev/sd0x-harness

    Change-aware next step advisor. An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Obsidian CLI

    sd0xdev/sd0x-harness

    Obsidian vault integration via official CLI. An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Orchestrate

    sd0xdev/sd0x-harness

    Agent-driven workflow orchestration (v1 report-only). An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • PR Comment

    sd0xdev/sd0x-harness

    Post friendly review comments to a GitHub PR — prepare locally, preview, then submit as atomic review.

    192 GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Categories

Questions about Test Health

What does Test Health do?

Holistic test coverage measurement. An agent skill from sd0xdev/sd0x-harness. Test Health is an agent skill from sd0xdev/sd0x-harness. Holistic test coverage measurement.

When should I use Test Health?

Test Health fits situations like: : assessing test health; measuring coverage trends; quantitative + qualitative test audit.

How do I install Test Health in Claude Code?

Run `npx skills add sd0xdev/sd0x-harness --skill test-health -a claude-code`. Or copy the skill folder (skills/test-health in sd0xdev/sd0x-harness) into .claude/skills/test-health in your project. Claude Code loads it when a task matches its description.

How do I install Test Health in Codex?

Run `npx skills add sd0xdev/sd0x-harness --skill test-health -a codex`. Or copy the skill folder (skills/test-health in sd0xdev/sd0x-harness) into .agents/skills/test-health in your project. Codex loads it when a task matches its description.

Can I use Test Health in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sd0xdev/sd0x-harness --skill test-health -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-health, .gemini/skills/test-health, .github/skills/test-health and .opencode/skills/test-health in your project.

What does Test Health need to run?

Going by SKILL.md and its folder, Test Health needs JavaScript for the scripts in its folder and the command-line tools its instructions call (bash, git and go). Our summary lists: Python 3; Node.js. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash(bash:*), Bash(git:*), Bash(node:*), Bash(npm:*), Bash(pnpm:*), Bash(yarn:*), Bash(npx:*), Bash(stat:*), Bash(find:*), Bash(python*:*), Bash(pytest:*), Bash(cargo:*), Bash(go:*), Skill, Agent.

Does Test Health access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Health safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Test Health use?

Test Health is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Health use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Test Health?

Skills that share tags, products or a category with Test Health: E2E Test Thinker (UniClipboard/UniClipboard, 1.9k stars), Ralph Coverage (jvm-skills/jvm-skills, 140 stars), Designing Tests (CloudAI-X/opencode-workflow, 275 stars) and Mutation Testing (proffesor-for-testing/agentic-qe, 494 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Health?

sd0xdev (a GitHub user) maintains it in sd0xdev/sd0x-harness, which has 192 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 8, 2026.

Source: sd0xdev/sd0x-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.