Generate tests that do not exist yet. An agent skill from yonatangross/orchestkit.

MITAuto-check: notesTesting & QA

Install Cover

skills CLI
$ npx skills add yonatangross/orchestkit --skill cover -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yonatangross/orchestkit cover --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/cover .claude/skills/cover && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cover
GitHub stars
290
Token cost
~6.3k tokens
SKILL.md length
1,795 words
Files
9 (incl. scripts, references)
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Generate tests that do not exist yet. An agent skill from yonatangross/orchestkit.

  • Code has no tests
  • SKILL.md covers Quick Start, Argument Resolution, Step -0.5: Effort-Aware… and Step -1: MCP Probe + Resume…, plus 8 more sections
  • Runs JavaScript scripts from its folder; calls npx and node
  • Raising coverage after implementation

What it does

Cover is an agent skill from yonatangross/orchestkit. Generate tests that do not exist yet. Analyzes coverage gaps, then writes and runs new test files across three tiers (unit, integration via testcontainers, Playwright E2E), one test-generator agent per tier, healing failures for up to 3 iterations. Use when code has no tests or when raising coverage after implementation. Do NOT use to grade tests that already exist (use /ork:verify) or to run a suite without writing anything new.

Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `references/behaviour-gate.md`, `references/claude-code.md` and `references/coverage-report-template.md`). Compatibility notes: Claude Code 2.1.277+. Requires network access.

It sits in Testing & QA, covering Test generation, Integration testing and Test coverage. It works with Playwright. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.

When your agent uses it

  • Code has no tests
  • Raising coverage after implementation
  • Grade tests that already exist (use /ork:verify)
  • Run a suite without writing anything new

Example prompts

  • “/cover”

Requirements

  • Python 3
  • Node.js
  • Docker
  • Compatibility (from SKILL.md): Claude Code 2.1.277+. Requires network access.
  • Pre-approved tools (allowed-tools): SendMessage, AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, TaskStop, ToolSearch, Workflow, CronCreate, CronDelete, Monitor, PushNotification, mcp__memory__search_nodes, mcp__context7__resolve-library-id, mcp__context7__query-docs

What it can do on your machine

Read from SKILL.md and the folder at commit 02bbf9a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • SendMessage
    • AskUserQuestion
    • Bash
    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Agent
    • TaskCreate

    …and 12 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code 2.1.277+. Requires network access.

    From compatibility in the SKILL.md frontmatter.

Context cost

Cover loads about 6.3k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 110 tokens; SKILL.md has 1,795 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~110
When it runs · the whole SKILL.md, loaded when a task matches
~6.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: SendMessage, AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, Ta

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from yonatangross/orchestkit at commit 02bbf9a, republished under its MIT licence (© yonatangross). 1,795 words, ~6,277 tokens.

Download SKILL.mdSave it as .claude/skills/cover/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
cover
description
Generate tests that do not exist yet. Analyzes coverage gaps, then writes and runs new test files across three tiers (unit, integration via testcontainers, Playwright E2E), one test-generator agent per tier, healing failures for up to 3 iterations. Use when code has no tests or when raising coverage after implementation. Do NOT use to grade tests that already exist (use /ork:verify) or to run a suite without writing anything new.
allowed-tools
SendMessage, AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, TaskStop, ToolSearch, Workflow, CronCreate, CronDelete, Monitor, PushNotification, mcp__memory__search_nodes, mcp__context7__resolve-library-id, mcp__context7__query-docs
compatibility
Claude Code 2.1.277+. Requires network access.
license
MIT
argument-hint
[scope-or-feature]
context
fork
background
false
user-invocable
true
skills
testing-unit, testing-integration, testing-e2e, testing-perf, testing-llm, chain-patterns, memory, quality-gates
effort
high
model
sonnet
metadata.category
workflow-automation
metadata.mcp-server
memory, context7

Cover: Test Suite Generator

Host-neutral workflow. Invoke by skill name (cover). Claude Code slash routing, YAML hook loaders, and .claude/chain live in references/claude-code.md.

Generate comprehensive test suites for existing code with real-service integration testing and automated failure healing.

Note: If disableSkillShellExecution is enabled (CC 2.1.91), the precondition check for vitest/jest won't run. Verify a test runner is installed before proceeding: npx vitest --version or npx jest --version.

Quick Start

bash
cover authentication flow
cover --model=opus payment processing
cover --tier=unit,integration user service
cover --real-services checkout pipeline

Argument Resolution

python
SCOPE = "$ARGUMENTS"  # e.g., "authentication flow"

# Flag parsing
MODEL_OVERRIDE = None
TIERS = ["unit", "integration", "e2e"]  # default: all three
REAL_SERVICES = False

for token in "$ARGUMENTS".split():
    if token.startswith("--model="):
        MODEL_OVERRIDE = token.split("=", 1)[1]
        SCOPE = SCOPE.replace(token, "").strip()
    elif token.startswith("--tier="):
        TIERS = token.split("=", 1)[1].split(",")
        SCOPE = SCOPE.replace(token, "").strip()
    elif token == "--real-services":
        REAL_SERVICES = True
        SCOPE = SCOPE.replace(token, "").strip()

Step -0.5: Effort-Aware Coverage Scaling (CC 2.1.76, env var since 2.1.120)

Read $CLAUDE_EFFORT (CC 2.1.120+) first; explicit --effort= token wins as override. Default high when CC < 2.1.120 and no flag. Pattern matches assess + explore (#1540). Scale test generation depth:

Effort LevelTiers GeneratedAgentsmaxIterations to pass Phase 5
lowUnit only1 agent2 (1 repair + 1 verify)
mediumUnit + Integration2 agents2 (1 repair + 1 verify)
high (default)Unit + Integration + E2E3 agents3 (2 repairs + 1 verify)
xhigh (CC 2.1.111+)Unit + Integration + E2E3 agents3 (the ceiling; xhigh adds agents, not heal passes)

Values are what you pass as maxIterations in Phase 5. The script clamps to [2, 3] (heal-loop.js), so anything outside that range is coerced, and the final iteration always verifies rather than repairing. There is no 4-iteration mode.

Override: Explicit --tier= flag or user selection overrides /effort downscaling.

Step -1: MCP Probe + Resume Check

python
# Probe MCPs (parallel):
# memory is alwaysLoad in .mcp.json (CC 2.1.121+, #1541). Probe below kept as fallback for older CC:
ToolSearch(query="select:mcp__memory__search_nodes")
ToolSearch(query="select:mcp__context7__resolve-library-id")

Write(".claude/chain/capabilities.json", {
  "memory": <true if found>,
  "context7": <true if found>,
  "skill": "cover",
  "timestamp": now()
})

# Resume check:
Read(".claude/chain/state.json")
# If exists and skill == "cover": resume from current_phase
# Otherwise: initialize state

Step 0: Scope & Tier Selection

python
AskUserQuestion(
  questions=[
    {
      "question": "What test tiers should I generate?",
      "header": "Test Tiers",
      "options": [
        {"label": "Full coverage (Recommended)", "description": "Unit + Integration (real services) + E2E"},
        {"label": "Unit + Integration", "description": "Skip E2E, focus on logic and service boundaries"},
        {"label": "Unit only", "description": "Fast isolated tests for business logic"},
        {"label": "E2E only", "description": "Playwright browser tests"}
      ],
      "multiSelect": false
    },
    {
      "question": "Healing strategy for failing tests?",
      "header": "Failure Handling",
      "options": [
        {"label": "Auto-heal (Recommended)", "description": "Fix failing tests up to 3 iterations"},
        {"label": "Generate only", "description": "Write tests, report failures, don't fix"},
        {"label": "Strict", "description": "All tests must pass or abort"}
      ],
      "multiSelect": false
    }
  ]
)

Override TIERS based on selection. Skip this step if --tier= flag was provided.


Finish line. Done means: every generated test runs, the suite is green or each remaining failure is reported with its heal-loop classification, and the coverage report is written. Follow Read("../../shared/rules/long-run-protocol.md"): keep going when a step needs no input from the user, stop and ask only when you can't continue without them or before anything destructive, check each subagent's evidence before accepting it, and mark anything you couldn't confirm with where you looked.

Task Management (MANDATORY)

python
# 1. Create main task IMMEDIATELY
TaskCreate(subject=f"Cover: {SCOPE}", description="Generate comprehensive test suite with real-service testing", activeForm=f"Generating tests for {SCOPE}")

# 2. Create subtasks for each phase
TaskCreate(subject="Discover scope and detect frameworks", activeForm="Discovering test scope")    # id=2
TaskCreate(subject="Analyze coverage gaps", activeForm="Analyzing coverage gaps")                  # id=3
TaskCreate(subject="Generate tests (parallel per tier)", activeForm="Generating tests")            # id=4
TaskCreate(subject="Execute generated tests", activeForm="Running tests")                          # id=5
TaskCreate(subject="Heal failing tests", activeForm="Healing test failures")                       # id=6
TaskCreate(subject="Generate coverage report", activeForm="Generating report")                     # id=7

# 3. Set dependencies for sequential phases
TaskUpdate(taskId="3", addBlockedBy=["2"])  # Analysis needs discovery first
TaskUpdate(taskId="4", addBlockedBy=["3"])  # Generation needs gap map
TaskUpdate(taskId="5", addBlockedBy=["4"])  # Execution needs generated tests
TaskUpdate(taskId="6", addBlockedBy=["5"])  # Healing needs test results
TaskUpdate(taskId="7", addBlockedBy=["6"])  # Report needs healed suite

# 4. Update status as you progress
TaskUpdate(taskId="2", status="in_progress")  # When starting
TaskUpdate(taskId="2", status="completed")    # When done, repeat for each subtask

6-Phase Workflow

PhaseActivitiesOutput
1. DiscoveryDetect frameworks, scan scope, find untested codeFramework map, file list
2. Coverage AnalysisRun existing tests, rank risk targets per tierCoverage baseline, risk targets with why
3. GenerationParallel test-generator agents per tier, then behaviour gateTest files created, gate verdicts
4. ExecutionRun all generated testsPass/fail results
5. HealFix failures, re-run (max 3 iterations)Green test suite
6. ReportCoverage delta, test count, summaryCoverage report
Phase Handoffs
After PhaseHandoff FileKey Outputs
1. Discovery01-cover-discovery.jsonFrameworks, scope files, tier plan
2. Analysis02-cover-analysis.jsonBaseline coverage, risk targets with why
3. Generation03-cover-generation.jsonFiles created, test count per tier, behaviour gate verdicts
5. Heal05-cover-healed.jsonFinal pass/fail, iterations used

Phase 1: Discovery

Detect the project's test infrastructure and scope the work.

python
# PARALLEL, all in ONE message:
# 1. Framework detection (hook handles this, but also scan manually)
Grep(pattern="vitest|jest|mocha|playwright|cypress", glob="package.json", output_mode="content")
Grep(pattern="pytest|unittest|hypothesis", glob="pyproject.toml", output_mode="content")
Grep(pattern="pytest|unittest|hypothesis", glob="requirements*.txt", output_mode="content")

# 2. Real-service infrastructure
Glob(pattern="**/docker-compose*.yml")
Glob(pattern="**/testcontainers*")
Grep(pattern="testcontainers", glob="**/package.json", output_mode="content")
Grep(pattern="testcontainers", glob="**/requirements*.txt", output_mode="content")

# 3. Existing test structure
Glob(pattern="**/tests/**/*.test.*")
Glob(pattern="**/tests/**/*.spec.*")
Glob(pattern="**/__tests__/**/*")
Glob(pattern="**/test_*.py")

# 4. Scope files (what to test)
# If SCOPE specified, find matching source files
Grep(pattern=SCOPE, output_mode="files_with_matches")

Real-service decision:

  • docker-compose*.yml found → integration tests use real services
  • testcontainers in deps → use testcontainers for isolated service instances
  • Neither found + --real-services flag → error: "No docker-compose or testcontainers found. Install testcontainers or remove --real-services flag."
  • Neither found, no flag → integration tests use mocks (MSW/VCR)

Load real-service detection details: Read("references/real-service-detection.md")

Phase 2: Coverage Analysis and Risk Targets

Run existing tests for a baseline, then pick targets by RISK. Coverage is information for the report, never a target.

python
# Baseline: npx vitest run --coverage --reporter=json | pytest --cov=<scope> --cov-report=json | go test -coverprofile=coverage.out ./...
# Rank uncovered code by risk and record WHY each target was chosen:
# changed code, complex branches, error paths, security and money paths, past bugs
# gap_map[tier] = [{"target": "src/billing/refund.ts:roundRefund", "why": "money path, fixed in #812"}]

Output the baseline and the risk-ranked targets, each with its why, immediately (progressive output). Signals and how to find them: Read("references/behaviour-gate.md").

Phase 3: Generation (Parallel Agents)

Spawn test-generator agents per tier. Launch ALL in ONE message with run_in_background=true.

Isolation: spawn each tier agent with Agent(isolation="worktree"), one worktree per tier (unit / integration / e2e) so they don't conflict. The subagent bypass of the worktree-isolation guard was fixed in CC 2.1.154 and completed in 2.1.203; ork's floor is >= 2.1.220, so every supported session gets real isolation. Do not create worktrees by hand before spawning.

The new branch's base comes from the worktree.baseRef setting, never from a hardcoded branch name. ork does not set it: a plugin cannot, and no ork settings file carries it. Unless the operator put "baseRef": "head" in .claude/settings.json or ~/.claude/settings.json, CC's default "fresh" applies: every tier agent branches from origin/<default>, unpushed local commits are invisible to it, and tsc fails with "cannot find module" for code you just wrote. Verify the setting before spawning. Full pattern: Read("../chain-patterns/references/worktree-agent-pattern.md")

python
# Unit tests agent (worktree-isolated)
if "unit" in TIERS:
    Agent(
        subagent_type="ork:test-generator",
        isolation="worktree",
        prompt=f"""Generate unit tests for: {SCOPE}
        Risk targets, each with why: {gap_map["unit"]}
        Framework: {detected_framework}
        Existing tests: {existing_test_files}

        Focus on:
        - AAA pattern (Arrange-Act-Assert)
        - Parametrized tests for multiple inputs
        - MSW/VCR for HTTP mocking (never mock fetch directly)
        - Factory-based test data (FactoryBoy/faker-js)
        - Edge cases: empty input, errors, timeouts, boundary values
        - BEHAVIOUR GATE: write a test ONLY if it asserts an observable result. Never write a
          mock-call-only, assertion-free, tautological, or source-reading test (references/behaviour-gate.md)""",
        run_in_background=True,
        max_turns=50,
        model=MODEL_OVERRIDE
    )

# Integration + E2E agents follow the same pattern:
# - subagent_type="ork:test-generator", isolation="worktree", run_in_background=True, same targets + BEHAVIOUR GATE lines
# - Integration focus: API endpoints (Supertest/httpx), real DB, contract tests (Pact), Zod schema validation
# - E2E focus: Playwright, semantic locators, Page Object Model, axe-core a11y, visual regression
#
# Special case: emulate (Vercel Labs stateful API emulation):
# When integration tests need GitHub/Stripe/Resend/Okta/etc. emulated
# (HMAC webhooks, parallel port isolation, full config from scratch),
# spawn emulate-engineer instead of test-generator for that tier:
# - subagent_type="ork:emulate-engineer" (same isolation="worktree" form)
# - Pairs with emulate-seed skill for seed YAML patterns

Output each agent's results as soon as it returns. Don't wait for all agents.

Phase 3b: Behaviour Gate

Before Phase 4, run node "${CLAUDE_SKILL_DIR}/scripts/check-behaviour-tests.mjs" --json <new test files> over the NEW test files only. Exit 1 lists each rejected test with its rule: (a) only mock-call assertions, (b) no assertion, only tautologies (expect(true).toBe(true)), or only not-to-throw when not throwing is not the stated contract, (c) reads or greps source instead of executing it. Delete each rejected test (the whole file on drop, which means every recognised test was rejected) or rewrite it to assert an observable result, and re-run until exit 0. Never delete an unchecked file or test (no test recognised, or its body is declared elsewhere): review it by hand. Run it again after Phase 5. Record the verdicts in 03-cover-generation.json. Rules and limits: Read("references/behaviour-gate.md").

Focus mode (CC 2.1.101): In focus mode, include the full coverage report (before/after delta, test count per tier, files created) in your final message.

Phase 4: Execution

Run all generated tests and collect results.

python
# Run test commands per tier (PARALLEL if independent):
# Unit: npx vitest run tests/unit/ OR pytest tests/unit/
# Integration: npx vitest run tests/integration/ OR pytest tests/integration/
# E2E: npx playwright test

# Collect: pass count, fail count, error details, coverage delta
Phase 5: Heal Loop

Do NOT hand-roll the loop. Run the real executor:

python
Workflow(
  scriptPath="${CLAUDE_SKILL_DIR}/workflows/heal-loop.js",
  args={"testCommand": "<tier test command>", "tier": "unit", "testGlob": "tests/unit/",
        "maxIterations": 3}   # from the effort table above; omitted defaults to 3
)
# One invocation per tier that has failures.

The iteration bound is enforced by the script, not by instruction. heal-loop.js runs a real counted loop clamped to [2, 3]: each iteration spawns a diagnose agent that actually executes the test command and returns structured pass/fail plus the verbatim failure output, then a repair agent that receives that failure text and edits test files only. It exits early the moment the suite is green.

The final iteration diagnoses but does not repair - there would be no run left to verify that repair. So maxIterations: N means N diagnose runs and N-1 repair passes, and the default 3 means 2 repairs. A ceiling of 1 is coerced to 2, since 1 would mean zero repairs.

Every failure is classified into the taxonomy (assertion, import, setup, timeout, stale-selector, type, flaky, plus source-bug), which selects the fix strategy.

The workflow returns status: "healed" or a structured failure (status: "failed", healed: false) carrying remaining_failures, failure_categories, and the per-iteration ledger. Never report a "failed" result as a success: surface the still-failing tests in the Phase 6 report.

Strategy detail (taxonomy table, fix rules, flaky prevention): Read("references/heal-loop-strategy.md")

Boundary: heal fixes TESTS, not source code. If a test fails because the source code has a bug, report it. Don't silently fix production code.

Show full SKILL.md (711 more words)Show less
Phase 6: Report

Generate coverage report with before/after comparison.

Full report layout (baseline→after table, tests-generated counts, heal iterations, files created, remaining gaps, next-steps commands): Read("references/coverage-report-template.md").

PushNotification on Completion (CC 2.1.110+)

Full cover runs (unit + integration + E2E with heal loop) take 15-45 min. After the Phase 6 report is assembled, call PushNotification(message=f"ork:cover complete, {SCOPE}: {tests_kept} tests kept · {tests_dropped} dropped by behaviour gate · {heal_loops} heal iters", status="proactive"). Full rule: Read("../chain-patterns/rules/push-notification-on-completion.md").

Coverage Drift Monitor (CC 2.1.71)

Optionally schedule a weekly check that the risk targets keep their behaviour tests:

python
# Guard: Skip cron in headless/CI (CLAUDE_CODE_DISABLE_CRON)
# if env CLAUDE_CODE_DISABLE_CRON is set, run a single check instead
CronCreate(
  schedule="0 2 * * 0",
  prompt="Weekly drift check for {SCOPE}: run the tests covering the risk targets in 02-cover-analysis.json.
    Alert if one was deleted or fails, or if newly changed code has no behaviour test. Coverage is reported as information only."
)

Key Principles

  • Output limits (CC 2.1.77+): Opus 4.8 defaults to 64k output tokens (128k upper bound). For large test suites, chunk generation across multiple agent turns if output approaches the limit.
  • Partial reads (CC 2.1.144+): Read returns a [PARTIAL view] first page (not an error) on oversized files. Re-read with explicit offset/limit so coverage analysis sees the whole file.
  • Tests only: never modify production source code, only generate test files
  • Real services when available: prefer testcontainers/docker-compose over mocks for integration tests because mock/prod divergence causes silent failures in production
  • Parallel generation: spawn one test-generator agent per tier in ONE message
  • Heal, don't loop forever: max 3 iterations, then report remaining failures
  • Progressive output: show results as each agent completes
  • Factory over fixtures: use FactoryBoy/faker-js for test data, not hardcoded values
  • Mock at network level: MSW/VCR, never mock fetch/axios directly

Agent Coordination

Context Passing

Each test-generator agent receives: risk targets (each with why) for its tier, the behaviour gate rules, test framework config, real-service infrastructure (testcontainers, docker-compose), and fixture patterns from the project.

Monitor + Partial Results (CC 2.1.98)

Use Monitor for streaming test execution output from background agents:

python
# Stream test suite output in real-time
Bash(command="npm test -- --coverage 2>&1", run_in_background=true)
Monitor(pid=test_task_id)  # Each line → notification

Full pattern reference (until-condition gates, partial-result salvage, TaskOutput vs Monitor decision): Read("../chain-patterns/references/monitor-patterns.md").

Partial results (CC 2.1.98): If a test-generator crashes mid-generation, synthesize what it produced:

python
for agent_result in test_gen_results:
    if "[PARTIAL RESULT]" in agent_result.output:
        # Agent crashed: check if it wrote any test files before dying
        partial_tests = Glob(pattern="**/tests/**/*.test.*", path=agent_result.worktree)
        if partial_tests:
            # 6 passing tests from a crashed agent > 0 tests
            # Copy partial tests to main worktree, run them in Phase 4
            for test_file in partial_tests:
                Bash(command=f"cp {test_file} {main_worktree}/{test_file}")
            # Flag as partial in report

A maxTurns stop is also partial since CC 2.1.246 (summary: "stopped at its N-turn limit (partial result; continue it with SendMessage to the task-id)"); continue that agent with SendMessage instead of re-spawning it.

SendMessage (Test Healing)

Cross-session replies land in the parent (CC 2.1.248): when a subagent sends SendMessage to another session, the reply is delivered to the parent session's conversation, never to the subagent; a subagent sends and moves on, the parent reads the answer. Cross-session SendMessage / ListAgents also work on Bedrock, Vertex and Foundry and with telemetry disabled (CC 2.1.248).

When a generated test fails, the healer agent can request context from the generator:

python
SendMessage(to="test-generator-unit", message="Test user_service_test.py:42 fails: TypeError on mock return. What's the expected shape?")
Skill Chain

Standard chain: implement → cover → verify → commit. Use addBlockedBy between each.

Verification Gate

Before claiming coverage is complete, apply: Read("../../shared/rules/verification-gate.md"). Run the coverage report fresh. "Should pass" is not evidence.

Agent Status Protocol

All test-generator agents report using: Read("../../shared/status-protocol.md"). BLOCKED if tests can't be written due to missing interfaces. NEEDS_CONTEXT if test expectations are unclear.

Quality Bar

Done means all of these hold:

  • coverage report shows the before→after delta from the actual coverage command output, not an estimate
  • every generated test file was executed; final pass/fail counts pasted from the runner
  • the behaviour gate exited 0 on the new test files after Phase 5, and every target states why it was chosen
  • no production source modified (only test files created); a source bug is reported, never silently patched
  • failures healed within the iteration budget (≤3, effort-scaled) or reported explicitly with the remaining-failure count
  • each requested tier (unit/integration/e2e) either produced tests or has a stated reason it was skipped
  • ork:implement: generates tests during implementation (Phase 5); use cover after for deeper coverage
  • ork:verify: grades existing tests 0-10; chain: implement → cover → verify
  • testing-unit / testing-integration / testing-e2e: knowledge skills loaded by test-generator agents
  • ork:commit: commit generated test files

Session recovery (CC 2.1.108+): After idle periods or interruptions, use /recap to restore conversational context alongside checkpoint-resume state. Enabled by default since CC 2.1.110 (even with telemetry disabled).

References

Load on demand with Read("references/<file>"):

FileContent
real-service-detection.mdDocker-compose/testcontainers detection, service startup, teardown
heal-loop-strategy.mdFailure classification, fix patterns, iteration budget
coverage-report-template.mdReport format, delta calculation, gap analysis
behaviour-gate.mdRisk target signals, the three reject rules, checker usage and limits
scripts/check-behaviour-tests.mjsPhase 3b checker: per-test keep or reject verdicts for new test files
workflows/heal-loop.jsPhase 5 executor: script-enforced 3-iteration repair loop (run via the Workflow tool)

Version: 1.3.0 (September 2026): behaviour gate (Phase 3b checker) and risk-based targets replace percentage coverage targets

© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in src/skills/cover of yonatangross/orchestkit.

  • SKILL.md
  • references/behaviour-gate.md
  • references/claude-code.md
  • references/coverage-report-template.md
  • references/heal-loop-strategy.md
  • references/real-service-detection.md
  • scripts/check-behaviour-tests.mjs
  • test-cases.json
  • workflows/heal-loop.js

Open the folder on GitHubat commit 02bbf9a

Compare with similar skills

Cover next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cover compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cover this skillyonatangross/orchestkit290—~6.3kAutomated safety check: NotesMIT
Dotnet Testingnovotnyllc/dotnet-artisan233—~972Automated safety check: PassMIT
Senior QAalirezarezvani/claude-skills28k1 repos~2.1kAutomated safety check: PassMIT
Playwright TestAI-Unified-Process/marketplace142—~3.8kAutomated safety check: WarnApache-2.0
Specialist Integration Test GeneratorHoangNguyen0403/agent-skills-standard571—~548Automated safety check: PassMIT
Write and Verify Playwright Testsappsmithorg/appsmith41k—~2.9kAutomated safety check: NotesApache-2.0

Similar skills

  • Dotnet Testing

    novotnyllc/dotnet-artisan

    Defines .NET test strategy and implementation patterns across xUnit v3 (Facts, Theories, fixtures, IAsyncLifetime), integration testing (WebApplicationFactory, Testcontainers), Aspire testing…

    233 GitHub stars~972 tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Senior QA

    alirezarezvani/claude-skills

    Generates unit tests, integration tests, and E2E tests for React/Next.js applications.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check passed
  • Playwright Test

    AI-Unified-Process/marketplace

    Creates Playwright browser-based tests for Vaadin views using the Drama Finder library for type-safe element wrappers with accessibility-first APIs.

    142 GitHub stars~3.8k tokensUpdated 4 days ago
    Testing & QAAuto-check: warnings
  • Specialist Integration Test Generator

    HoangNguyen0403/agent-skills-standard

    Generates one integration/E2E test from an approved test case spec using existing project patterns.

    571 GitHub stars~548 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.

    41k GitHub stars~2.9k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • E2E Testing

    langflow-ai/langflow

    Write and review Playwright E2E tests for Langflow. An agent skill from langflow-ai/langflow.

    155k GitHub stars~3.3k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from yonatangross/orchestkit

All 108 skills in this repo
  • API Design

    yonatangross/orchestkit

    API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.

    290 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Architecture Decision Record

    yonatangross/orchestkit

    ADR templates in the Nygard format with context, decision, consequences, and alternatives.

    290 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Audit Full

    yonatangross/orchestkit

    Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.

    290 GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Code Review Playbook

    yonatangross/orchestkit

    Structured review processes, conventional comments, language-specific checklists, and feedback templates.

    290 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Create PR

    yonatangross/orchestkit

    Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.

    290 GitHub stars~4.5k tokensUpdated today
    Auto-check: notes
  • Explore

    yonatangross/orchestkit

    Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.

    290 GitHub stars~3.9k tokensUpdated today
    Auto-check: notes

Works with

Categories

Questions about Cover

What does Cover do?

Generate tests that do not exist yet. An agent skill from yonatangross/orchestkit. Cover is an agent skill from yonatangross/orchestkit. Generate tests that do not exist yet.

When should I use Cover?

Cover fits situations like: code has no tests; raising coverage after implementation; grade tests that already exist (use /ork:verify); run a suite without writing anything new.

How do I install Cover in Claude Code?

Run `npx skills add yonatangross/orchestkit --skill cover -a claude-code`. Or copy the skill folder (src/skills/cover in yonatangross/orchestkit) into .claude/skills/cover in your project. Claude Code loads it when a task matches its description.

How do I install Cover in Codex?

Run `npx skills add yonatangross/orchestkit --skill cover -a codex`. Or copy the skill folder (src/skills/cover in yonatangross/orchestkit) into .agents/skills/cover in your project. Codex loads it when a task matches its description.

Can I use Cover in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill cover -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cover, .gemini/skills/cover, .github/skills/cover and .opencode/skills/cover in your project.

What does Cover need to run?

Going by SKILL.md and its folder, Cover needs JavaScript for the scripts in its folder and the command-line tools its instructions call (npx and node). Our summary lists: Python 3; Node.js; Docker. Its frontmatter pre-approves these tools: SendMessage, AskUserQuestion, Bash, Read, Write, Edit, Grep, Glob, Agent, TaskCreate, TaskUpdate, TaskList, TaskStop, ToolSearch, Workflow, CronCreate, CronDelete, Monitor, PushNotification, mcp__memory__search_nodes, mcp__context7__resolve-library-id, mcp__context7__query-docs. Compatibility (from SKILL.md): Claude Code 2.1.277+. Requires network access..

Does Cover access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Cover safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Cover use?

Cover is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cover use?

About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.8k tokens, read only when the agent opens those files.

What are the alternatives to Cover?

Skills that share tags, products or a category with Cover: Dotnet Testing (novotnyllc/dotnet-artisan, 233 stars), Senior QA (alirezarezvani/claude-skills, 28k stars), Playwright Test (AI-Unified-Process/marketplace, 142 stars) and Specialist Integration Test Generator (HoangNguyen0403/agent-skills-standard, 571 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cover?

yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 290 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.

Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.