Agent skill

Test Deep

by sd0xdev in sd0xdev/sd0x-harness

Context-aware test orchestration. An agent skill from sd0xdev/sd0x-harness.

MITAuto-check: notesTesting & QA

Install Test Deep

skills CLI
$ npx skills add sd0xdev/sd0x-harness --skill test-deep -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sd0xdev/sd0x-harness test-deep --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sd0xdev/sd0x-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-deep .claude/skills/test-deep && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-deep
GitHub stars
192
Token cost
~2.3k tokens
SKILL.md length
784 words
Files
4 (incl. references)
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Context-aware test orchestration. An agent skill from sd0xdev/sd0x-harness.

  • Works in 4 steps: Test Selection → Progressive Ladder → Failure Triage Pipeline → …
  • : smart test selection
  • SKILL.md covers Supplementary Agent, Trigger, When NOT to Use and Prohibited Actions, plus 10 more sections
  • Calls git

What it does

Test Deep is an agent skill from sd0xdev/sd0x-harness. Context-aware test orchestration. Use when: smart test selection, failure triage, progressive test ladder, test failure analysis. Not for: writing tests (use post-dev-test), reviewing tests (use codex-test-review), generating tests (use codex-test-gen), full manual run (use verify). Output: test results + triage report + fixer actions.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/fixer-catalog.md`, `references/test-selection.md` and `references/triage-pipeline.md`).

It sits in Testing & QA, covering Failing and flaky tests, Test generation and Verification before completion. The repository describes itself as: The harness layer for Claude Code — a reference implementation of harness engineering with hook-enforced dual review, state-machine gates that survive context compaction, and… The licence is MIT.

When your agent uses it

  • : smart test selection
  • Progressive test ladder
  • Test failure analysis

Example prompts

  • “/test-deep”

Requirements

  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, Write, Agent, AskUserQuestion, Skill

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Test Selection
  2. Progressive Ladder
  3. Failure Triage Pipeline
  4. Fixer Execution

What it can do on your machine

Read from SKILL.md and the folder at commit c9a2036. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • Write
    • Agent
    • AskUserQuestion
    • Skill

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Deep loads about 2.3k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 87 tokens; SKILL.md has 784 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, Write, Agent, AskUserQuestion, Skill

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sd0xdev/sd0x-harness at commit c9a2036, republished under its MIT licence (© sd0xdev). 784 words, ~2,287 tokens.

Download SKILL.mdSave it as .claude/skills/test-deep/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
test-deep
description
Context-aware test orchestration. Use when: smart test selection, failure triage, progressive test ladder, test failure analysis. Not for: writing tests (use post-dev-test), reviewing tests (use codex-test-review), generating tests (use codex-test-gen), full manual run (use verify). Output: test results + triage report + fixer actions.
allowed-tools
Read, Grep, Glob, Bash, Write, Agent, AskUserQuestion, Skill

Test Deep — Context-Aware Test Orchestration

Supplementary Agent

When test failures occur, dispatch background triage:

Agent({ description: "Analyze test failure root cause and suggest fixes", subagent_type: "verify-app", prompt: Analyze the following test failure: <failure output> Identify root cause and suggest minimal fix. })

Trigger

  • Keywords: context-aware testing, smart test, test orchestration, test triage, test failure analysis, test deep, run relevant tests

When NOT to Use

ScenarioAlternative
Writing new tests/post-dev-test
Reviewing test coverage/codex-test-review
Generating unit tests/codex-test-gen
Full manual test run/verify
Code review/codex-review-fast

Prohibited Actions

❌ git add | git commit | git push — per @rules/git-workflow.md

This skill runs tests and may apply safe fixers, but does not commit. To commit, offer the menu per rules/git-workflow.md § Proactive Offer — a commit option on any real branch, a push option only where review-state.js offer allows one — and invoke /smart-commit --execute on selection; never print the command for the user to copy.

Workflow

mermaid
flowchart TD
    U[User: /test-deep] --> S[Phase 0: Test Selection]
    S --> |git diff mapping| T[Test Targets]
    S --> |no mapping| F1[Framework --changedSince]
    S --> |low confidence| F2[Full Suite]
    T --> L[Phase 1: Progressive Ladder]
    F1 --> L
    F2 --> L
    L --> |unit| R1[Run Unit Tests]
    R1 --> |pass| R2[Run Integration Tests]
    R1 --> |fail| TR[Phase 2: Triage Pipeline]
    R2 --> |pass| R3[Run E2E Tests]
    R2 --> |fail| TR
    R3 --> |pass| DONE[All Pass]
    R3 --> |fail| TR
    TR --> P[Parser: Structured Tags]
    P --> LLM[LLM: Root Cause + Action]
    LLM --> FC[Phase 3: Fixer Catalog Lookup]
    FC --> SG{Safety Gate}
    SG --> |safe| AUTO[Auto-run Fix]
    SG --> |side-effect| ASK[AskUserQuestion]
    SG --> |destructive| BLOCK[Manual Only]
    AUTO --> L
    ASK --> |approved| L

Phase 0: Test Selection

Map code changes to test targets. See references/test-selection.md for full strategy.

Strategy priority:

PriorityMethodWhen
1Git diff filename mappingDefault — map changed files to test files via Glob
2Framework nativeJest --changedSince, Vitest --changed
3Full suite fallbackConfig change, no mapping, --all flag

Steps:

  1. Collect changed files: union of unstaged + staged + untracked (--branch uses merge-base)
  2. Apply filename mapping rules to generate candidate test paths
  3. Glob-confirm each candidate exists
  4. Classify confirmed tests by layer (unit / integration / e2e)

Full suite escalation triggers: config file changed, CI/CD file changed, package dependency changed, no test files mapped, --all flag.

Phase 1: Progressive Ladder

Execute tests layer-by-layer with fail-fast.

LayerDirectory PatternTimeoutFail Behavior
Unittest/unit/**, test/scripts/lib/**, unclassified60sFail-fast → triage
Integrationtest/integration/**300sFail-fast → triage
E2Etest/e2e/**600sEnter triage

Fail-fast rule: Unit fail → skip integration + e2e. Integration fail → skip e2e.

Override: --no-fail-fast disables this behavior (run all layers regardless).

Layer detection: Classify test files by directory prefix. Files not matching any integration/e2e pattern → treat as unit.

Execution: Run tests using project's configured test command (from package.json scripts or CLAUDE.md). Capture stdout/stderr and exit code.

Phase 2: Failure Triage Pipeline

When failures occur, analyze and classify. See references/triage-pipeline.md for full spec.

Step 1: Output Parser

Extract structured tags from test output (no classification — just structure):

TagDescription
exit_codeProcess exit code
error_signatures[]Regex-matched error patterns
failing_tests[]Failed test names
failing_files[]Failed test file paths
env_hints[]Environment clues (testnet, localhost, etc.)
stack_depthStack trace line count
Step 2: LLM Root Cause Analysis

Feed parser tags + compressed output to LLM for classification:

ClassificationDefinitionTypical Action
code_bugLogic error in codeFix code
infraInfrastructure issue (port, dependency)Restart / reinstall
environmentExternal precondition unmetFixer catalog action
flakyNon-deterministic failureRetry + quarantine tag

Mandatory secret redaction before LLM analysis — per @rules/logging.md:

  • API keys, private keys, tokens, passwords, mnemonics, URLs with credentials
  • See references/triage-pipeline.md for redaction patterns
Show full SKILL.md (320 more words)Show less
Step 3: Safety-Gated Action

Route suggested_fixer through safety gate:

TierAuto-run?ConfirmationExamples
safeYesNoneretry, clear_cache
side-effectNoAskUserQuestionreinstall_deps, restart_server
destructiveBlockedManual onlyReset state, drop tables

Default-deny: Unknown fixer → side-effect tier (require confirmation).

Phase 3: Fixer Execution

Lookup fixer from catalog, check tier, execute or prompt. See references/fixer-catalog.md for full catalog.

Core fixers (plugin-shipped):

FixerTierDescription
retrysafeRe-run failing tests
clear_cachesafeClear build/test cache
reinstall_depsside-effectRemove node_modules + reinstall
restart_serverside-effectKill + restart dev server
port_cleanupside-effectKill process on conflicting port

Host extensions: Project-specific fixers in .claude/test-deep/fixers.md. Schema validation at load — missing required fields or invalid tier → rejected with warning.

Fixer loop: After safe/approved fixer runs, re-enter progressive ladder for failed tests only. Max 1 fixer retry per failure to prevent infinite loops.

Session Artifacts

Write per-run artifacts for comparison and debugging.

Location: .claude/cache/test-deep/<runId>/

Run ID format: <timestamp>-<shortSHA>-<pid>

FileContent
metadata.jsonRun config: selected tests, ladder config, changed files
results.jsonPer-layer results: pass/fail, duration, exit codes
triage.jsonParser tags, LLM classification, fixer chosen, outcome
output.logCompressed test output (secret-redacted)

TTL: Keep last 5 runs. Prune older on new run start.

latest symlink: Points to most recent run directory.

Arguments

FlagDefaultDescription
--allfalseForce full test suite
--layer <unit|integration|e2e>allRun only specified layer
--no-fail-fastfalseRun all layers regardless of failures
--no-fixfalseTriage only, skip fixer execution
--focus <path>—Limit test selection to path
--branchfalseUse merge-base diff instead of working tree

Output

markdown
## Test Deep Report

### Test Selection
- Changed files: N
- Mapped test files: N (N unit, N integration, N e2e)
- Selection method: git diff mapping | framework native | full suite

### Results

| Layer | Tests | Passed | Failed | Skipped | Duration |
|-------|-------|--------|--------|---------|----------|

### Failure Triage

| # | Test | Classification | Root Cause | Fixer | Tier |
|---|------|---------------|------------|-------|------|

### Actions Taken
- [N] fixer_id: outcome

### Gate
✅ All Pass | ⛔ N failures pending resolution

Verification Checklist

  • Test selection maps changed files to test targets
  • Progressive ladder respects fail-fast
  • Triage pipeline produces structured classification
  • Safety gate enforces tier-based execution
  • Secret redaction applied before LLM analysis
  • Session artifacts written to cache
  • No git add / git commit / git push executed

References

  • references/test-selection.md — Git diff mapping strategy + full suite escalation
  • references/triage-pipeline.md — Parser tags + LLM prompt + safety gate
  • references/fixer-catalog.md — Core fixers + host extensions + safety tiers
  • @rules/logging.md — Secret redaction policy
  • @rules/docs-writing.md — Output format conventions

© sd0xdev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/test-deep of sd0xdev/sd0x-harness.

  • SKILL.md
  • references/fixer-catalog.md
  • references/test-selection.md
  • references/triage-pipeline.md

Open the folder on GitHubat commit c9a2036

Compare with similar skills

Test Deep next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Deep compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Deep this skillsd0xdev/sd0x-harness192—~2.3kAutomated safety check: NotesMIT
Swig Testswig/swig6.3k—~2.3kAutomated safety check: PassCustom licence
Wioworkersio/skills180—~5.8kAutomated safety check: PassMIT
Offloadimbue-ai/offload125—~3.1kAutomated safety check: PassMIT
Test Fixingdavila7/claude-code-templates32k7 repos~739Automated safety check: PassMIT
Write Testsdzhng/skills1k—~2.4kAutomated safety check: PassMIT

Similar skills

  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Wio

    workersio/skills

    Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing test-suite health.

    180 GitHub stars~5.8k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Offload

    imbue-ai/offload

    Activate when you see offload.toml in a repo, offload referenced in build targets (justfile, Makefile, scripts), or when you need to run a large test suite in parallel.

    125 GitHub stars~3.1k tokensUpdated 12 days ago
    Testing & QAAuto-check passed
  • Test Fixing

    davila7/claude-code-templates

    Run tests and systematically fix all failing tests using smart error grouping.

    32k GitHub starsUsed in 7 repos~739 tokens
    Testing & QAAuto-check passed
  • Write Tests

    dzhng/skills

    Write tests that pin real behavior instead of implementation details, config values, or lucky samples.

    1k GitHub stars~2.4k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Frappe Testing Unit

    Impertio-Studio/Frappe_Claude_Skill_Package

    A skill your agent uses when writing unit tests, integration tests, creating test fixtures, or running tests with bench run-tests.

    187 GitHub stars~3k tokensUpdated 20 days ago
    Testing & QAAuto-check passed

More from sd0xdev/sd0x-harness

All 91 skills in this repo
  • Adr

    sd0xdev/sd0x-harness

    Write an Architecture Decision Record (ADR) for a feature — Context / Decision / Status / Consequences / Alternatives, filed as docs/features/<feature/adr-<NNN-<title.md with a 3-digit zero-padded…

    192 GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Load PR Review

    sd0xdev/sd0x-harness

    Load GitHub PR review comments into AI session — analyze, triage, plan.

    192 GitHub stars~4.4k tokensUpdated yesterday
    Auto-check passed
  • Next Step

    sd0xdev/sd0x-harness

    Change-aware next step advisor. An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Obsidian CLI

    sd0xdev/sd0x-harness

    Obsidian vault integration via official CLI. An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Orchestrate

    sd0xdev/sd0x-harness

    Agent-driven workflow orchestration (v1 report-only). An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • PR Comment

    sd0xdev/sd0x-harness

    Post friendly review comments to a GitHub PR — prepare locally, preview, then submit as atomic review.

    192 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Test Deep

What does Test Deep do?

Context-aware test orchestration. An agent skill from sd0xdev/sd0x-harness. Test Deep is an agent skill from sd0xdev/sd0x-harness. Context-aware test orchestration.

When should I use Test Deep?

Test Deep fits situations like: : smart test selection; progressive test ladder; test failure analysis.

How do I install Test Deep in Claude Code?

Run `npx skills add sd0xdev/sd0x-harness --skill test-deep -a claude-code`. Or copy the skill folder (skills/test-deep in sd0xdev/sd0x-harness) into .claude/skills/test-deep in your project. Claude Code loads it when a task matches its description.

How do I install Test Deep in Codex?

Run `npx skills add sd0xdev/sd0x-harness --skill test-deep -a codex`. Or copy the skill folder (skills/test-deep in sd0xdev/sd0x-harness) into .agents/skills/test-deep in your project. Codex loads it when a task matches its description.

Can I use Test Deep in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sd0xdev/sd0x-harness --skill test-deep -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-deep, .gemini/skills/test-deep, .github/skills/test-deep and .opencode/skills/test-deep in your project.

What does Test Deep need to run?

Going by SKILL.md and its folder, Test Deep needs the command-line tools its instructions call (git). Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, Write, Agent, AskUserQuestion, Skill.

Does Test Deep access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Deep safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Test Deep use?

Test Deep is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Deep use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to Test Deep?

Skills that share tags, products or a category with Test Deep: Swig Test (swig/swig, 6.3k stars), Wio (workersio/skills, 180 stars), Offload (imbue-ai/offload, 125 stars) and Test Fixing (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Deep?

sd0xdev (a GitHub user) maintains it in sd0xdev/sd0x-harness, which has 192 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on October 6, 2026.

Source: sd0xdev/sd0x-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.