Agent skill

Build Dod

by danshapiro in danshapiro/kilroy

A skill your agent uses when converting a spec, requirements document, or goal statement into a Definition of Done with acceptance criteria and integration test scenarios

MITAuto-check passedTesting & QA

Install Build Dod

skills CLI
$ npx skills add danshapiro/kilroy --skill build-dod -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install danshapiro/kilroy build-dod --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/danshapiro/kilroy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/build-dod .claude/skills/build-dod && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
build-dod
GitHub stars
221
Token cost
~2.5k tokens
SKILL.md length
1,114 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when converting a spec, requirements document, or goal statement into a Definition of Done with acceptance criteria and integration test scenarios

  • Works in 8 steps: Read the full spec → List deliverables — the artifacts that… → If a digraph/flow diagram is provided,… → …
  • Converting a spec
  • SKILL.md covers Core Principle, Process, If a digraph is provided and Acceptance Criteria, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Build Dod is an agent skill from danshapiro/kilroy. Use when converting a spec, requirements document, or goal statement into a Definition of Done with acceptance criteria and integration test scenarios

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Integration testing, User stories and PRD writing. The licence is MIT.

When your agent uses it

  • Converting a spec
  • Requirements document
  • Goal statement into a Definition of Done with acceptance criteria and integration test scenarios

Example prompts

  • “/build-dod”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Read the full spec
  2. List deliverables — the artifacts that exist when done
  3. If a digraph/flow diagram is provided, extract intended user journeys, decision points, and outcomes from it
  4. Write acceptance criteria — one observable assertion per row
  5. Inventory every user-facing message surface from the spec
  6. Write integration test scenarios that prove the deliverable works end-to-end
  7. Map each AC and each message to the scenario(s) that prove it
  8. Crosscheck — confirm every AC and every message is covered and every scenario is sound

What it can do on your machine

Read from SKILL.md and the folder at commit b55fb0f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Build Dod loads about 2.5k tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 1,114 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from danshapiro/kilroy at commit b55fb0f, republished under its MIT licence (© danshapiro). 1,114 words, ~2,497 tokens.

Download SKILL.mdSave it as .claude/skills/build-dod/SKILL.md (or your agent's skills folder).
name
build-dod
description
Use when converting a spec, requirements document, or goal statement into a Definition of Done with acceptance criteria and integration test scenarios

Build DoD

A DoD converts a spec into pass/fail gates. Its power is in integration tests — scenarios that prove the deliverable works by exercising it the way a user would.

Core Principle

Tests aren't there to be passed. They're there to prove results. Verify the deliverable through integration scenarios that exercise it end-to-end, not through unit tests that verify internals.

Process

  1. Read the full spec
  2. List deliverables — the artifacts that exist when done
  3. If a digraph/flow diagram is provided, extract intended user journeys, decision points, and outcomes from it
  4. Write acceptance criteria — one observable assertion per row
  5. Inventory every user-facing message surface from the spec
  6. Write integration test scenarios that prove the deliverable works end-to-end
  7. Map each AC and each message to the scenario(s) that prove it
  8. Crosscheck — confirm every AC and every message is covered and every scenario is sound

When this skill is used inside an Attractor run, scratch outputs should be written under .ai/runs/$KILROY_RUN_ID/.... Root .ai is not implicitly ingested.

If a digraph is provided

Use the graph in two passes:

  1. Intent pass (primary): Treat the graph as a map of product intent and user flow. Use it to identify the major journeys, decisions, and failure/recovery paths that matter to users.
  2. Flow sanity pass (secondary): Do lightweight topology checks only to catch intent-breaking issues (for example, a required outcome has no reachable path).

Acceptance Criteria

Each AC is a single, testable assertion using observable language: "exists", "returns", "displays", "produces", "exits 0".

Group by concern (e.g. Build, Output, Behavior, Integration). Number hierarchically: AC-1.1, AC-1.2, AC-2.1.

ACs describe what must be true. They are proven by integration test scenarios, not by individual unit tests.

Integration Test Scenarios

Integration tests are the primary verification mechanism. Each scenario exercises the delivered artifact directly, proving multiple acceptance criteria simultaneously.

When a digraph exists, write scenarios around high-level intent coverage, not exhaustive graph traversal. Prefer:

  • Primary happy paths
  • High-risk branches and decision outcomes
  • Critical error/recovery paths
  • Mode transitions that change user experience (for example first run vs returning user)

Test the delivered artifact in its delivery form. At least one scenario must exercise the full delivery path:

  • Browser app → serve it in a browser, confirm it loads and runs
  • CLI tool → invoke the binary, check exit code and output
  • Library → import it and call its public API
  • Data file → consume it with its intended reader

If the deliverable is a browser app and no scenario loads it in a browser, the DoD is incomplete.

Validate every user-facing message. Help text, error messages, status displays, feedback strings, prompts, and warnings are promises to the user. Inventory all of them from the spec, then ensure each one is triggered and validated in at least one scenario:

  • Trigger the condition that should display it
  • Confirm the message appears
  • Confirm what the message says is true (if help says "press ? for help", pressing ? must show help)

This means all messages, not a sample. If the spec describes 20 distinct message surfaces, 20 must be tested.

When one artifact references another, verify both. A source file that references an output is evidence of intent; confirm the output itself is present and valid.

For each primary way the deliverable is used, write a scenario with:

  • Starting state — deterministic inputs (fixed seed, known data, clean environment)
  • Actions — a sequence of operations a real user or consumer would perform
  • Expected outcomes — observable results after each action

Scenarios should cross multiple AC groups. A browser app scenario might cover loading, display, input, and state persistence in one flow.

Each scenario is self-contained — it sets up its own preconditions within the test rather than depending on externally pre-computed inputs or manual preparation.

Each scenario becomes a named automated test in the DoD, with test exits 0 as its verification.

For checks that require judgment, write a concrete semantic verification with:

  • The question to answer
  • The expected answer
  • The evidence to examine (file paths, commands, artifacts)
Show full SKILL.md (452 more words)Show less

Test Evidence Contract (Required)

Every DoD must define deterministic, reviewable test artifacts for each integration scenario.

  • Canonical root: .ai/runs/$KILROY_RUN_ID/test-evidence/latest/
  • Per-scenario folder: .ai/runs/$KILROY_RUN_ID/test-evidence/latest/IT-<id>/
  • Manifest file: .ai/runs/$KILROY_RUN_ID/test-evidence/latest/manifest.json

Manifest entries must map each scenario ID to:

  • scenario status (pass/fail)
  • artifact list (type + path)
  • missing/unreadable artifact notes (if any)

Example manifest shape:

json
{
  "version": 1,
  "scenarios": [
    {
      "id": "IT-1",
      "status": "pass",
      "artifacts": [
        { "type": "log", "path": ".ai/runs/$KILROY_RUN_ID/test-evidence/latest/IT-1/test.log" }
      ],
      "notes": []
    }
  ]
}

Artifact requirements:

  • Every IT-* scenario must produce at least one artifact.
  • Each scenario must declare one surface type in evidence metadata:
    • surface=ui: visually rendered user interface is exercised.
    • surface=non_ui: no visually rendered user interface is exercised.
    • surface=mixed: both visual UI and non-UI interfaces are exercised.
  • Scenarios that exercise a visual UI (surface=ui or surface=mixed) must include screenshot artifacts (.png or .jpg) proving key states.
  • Scenarios that do not exercise visual UI (surface=non_ui) must include text or structured evidence artifacts (for example logs, stdout captures, JSON reports).
  • On test failure, emit best-effort artifacts and a manifest entry anyway; missing artifacts must be explicit in the manifest and treated as findings.

Framework policy:

  • Do not require a specific browser/test framework in the DoD. Require outcomes and evidence, not tool brand.
Scenario sanity checks

Before finalizing each scenario, confirm:

  • Automatable — the test can set up its own state, run, and assert without human intervention or external artifacts
  • Bounded — the scenario has a finite, predictable number of steps (a test that must "play until winning" is unbounded; a test that exercises 5 specific levels via setup commands is bounded)
  • Proportional — effort to implement the test is proportional to the confidence it provides (testing 3 representative cases from a category provides nearly as much confidence as testing all 50)
  • Independent — the scenario produces the same result regardless of execution order or environment state
  • Intent-complete — collectively, scenarios prove the intended user journeys and outcomes described by the spec (and digraph, if present)
  • Evidence-complete — scenario defines deterministic artifact paths and required artifact types under .ai/runs/$KILROY_RUN_ID/test-evidence/latest/IT-<id>/

The Crosscheck

After writing all ACs and integration scenarios, review:

Per scenario:

  1. Confirm the scenario exercises the delivered artifact, not just internal components
  2. Confirm the scenario is automatable, bounded, proportional, and independent
  3. Confirm the scenario crosses multiple AC groups
  4. Confirm the scenario declares required evidence artifacts and deterministic paths

Per AC: 5. Confirm at least one scenario proves this AC 6. If no scenario covers an AC, add coverage or justify the gap

Overall: 7. Confirm at least one scenario tests the deliverable in its delivery form 8. Confirm every user-facing message from the inventory is triggered and validated by at least one scenario 9. Confirm the scenarios collectively cover every AC group 10. Confirm .ai/runs/$KILROY_RUN_ID/test-evidence/latest/manifest.json covers every scenario ID with at least one artifact path 11. If a digraph exists, confirm no required intent path is missing or unreachable

Output Format

markdown
# [Project] — Definition of Done

## Scope

### In Scope
[What the deliverable covers]

### Out of Scope
[Explicit exclusions]

### Assumptions
[Prerequisites and environment]

## Deliverables

| Artifact | Location | Description |
|----------|----------|-------------|
| ... | ... | ... |

## Acceptance Criteria

### [Concern Area]

| ID | Criterion | Covered by |
|----|-----------|------------|
| AC-N.M | [Observable assertion] | IT-X, IT-Y |

## User-Facing Message Inventory

| ID | Message surface | Trigger condition | Covered by |
|----|----------------|-------------------|------------|
| MSG-N | [What the user sees] | [What causes it] | IT-X |

## Test Evidence Contract

| Item | Requirement |
|------|-------------|
| Evidence root | `.ai/runs/$KILROY_RUN_ID/test-evidence/latest/` |
| Scenario folder pattern | `.ai/runs/$KILROY_RUN_ID/test-evidence/latest/IT-<id>/` |
| Manifest | `.ai/runs/$KILROY_RUN_ID/test-evidence/latest/manifest.json` |
| UI scenarios (`surface=ui` or `surface=mixed`) | Include screenshot evidence proving key states |
| Non-UI scenarios (`surface=non_ui`) | Include text/structured evidence (log/stdout/json) |
| Failure behavior | Emit best-effort artifacts and manifest entry; record missing artifacts explicitly |

## Integration Test Scenarios

| ID | Scenario | Steps | Verification | Evidence Artifacts |
|----|----------|-------|--------------|--------------------|
| IT-N | [User journey name] | 1. [action] → [expected] 2. [action] → [expected] ... | `test command` exits 0 | `surface=<ui|non_ui|mixed>`; `[type:path, ...]` |

© danshapiro, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/build-dod of danshapiro/kilroy.

Open the folder on GitHubat commit b55fb0f

Compare with similar skills

Build Dod next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Build Dod compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Build Dod this skilldanshapiro/kilroy221—~2.5kAutomated safety check: PassMIT
AI Test Generationpetrkindlmann/qa-skills163—~4.8kAutomated safety check: PassMIT
Prd V07 Test Planningmattgierhart/PRD-driven-context-engineering179—~3.5kAutomated safety check: NotesMIT
Req To TestArabelaTso/Skills-4-SE253—~2.9kAutomated safety check: PassApache-2.0
Build Scenario Teststamdogood/builder-essential-skills218—~1.7kAutomated safety check: PassMIT
Feature ForgeJeffallan/claude-skills12k—~1.1kAutomated safety check: PassMIT

Similar skills

  • AI Test Generation

    petrkindlmann/qa-skills

    Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.

    163 GitHub stars~4.8k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Prd V07 Test Planning

    mattgierhart/PRD-driven-context-engineering

    Define test cases BEFORE implementation, ensuring every API, business rule, and user journey has verifiable acceptance criteria during PRD v0.7 Build Execution.

    179 GitHub stars~3.5k tokensUpdated 1 mo ago
    Testing & QAAuto-check: notes
  • Req To Test

    ArabelaTso/Skills-4-SE

    Generates comprehensive test scenarios from requirements including BDD/Gherkin scenarios, unit tests, integration tests, and end-to-end test cases.

    253 GitHub stars~2.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Build Scenario Tests

    tamdogood/builder-essential-skills

    Inspect an unfamiliar repository, turn a focused Markdown behavior scenario into a deterministic test in the repository's native test stack, run it, and preserve traceability between intent and code.

    218 GitHub stars~1.7k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Feature Forge

    Jeffallan/claude-skills

    Runs a structured requirements interview to produce a feature specification with EARS requirements, acceptance criteria and an implementation checklist.

    12k GitHub stars~1.1k tokensUpdated 4 days ago
    Product & Project ManagementAuto-check passed
  • 01 Acceptance QA

    ai-driven-dev/framework

    Validate a reviewed candidate's observable behavior against its acceptance criteria and record short named videos as reviewer evidence.

    510 GitHub stars~434 tokensUpdated today
    Product & Project ManagementAuto-check passed

More from danshapiro/kilroy

  • Create Runfile

    danshapiro/kilroy

    A skill your agent uses when authoring or repairing Kilroy run config YAML/JSON files, including DOT-to-provider backend alignment and runtime policy defaults.

    221 GitHub stars~1.4k tokensUpdated 5 mo ago
    Auto-check: notes
  • Release Kilroy

    danshapiro/kilroy

    A skill your agent uses when preparing a Kilroy release — writing release notes, tagging, and publishing via goreleaser on GitHub.

    221 GitHub stars~2k tokensUpdated 5 mo ago
    Auto-check passed
  • Create Dotfile

    danshapiro/kilroy

    A skill your agent uses when authoring or repairing Kilroy Attractor DOT graphs from requirements, with template-first topology, routing guardrails, and validator-clean output.

    221 GitHub stars~6.4k tokensUpdated 5 mo ago
    Auto-check passed
  • Investigating Kilroy Runs

    danshapiro/kilroy

    To diagnose active, stuck, or failed Kilroy Attractor runs, inspect run artifacts (manifest.json, live.json, checkpoint.json, final.json, progress.ndjson), resolve run IDs/log roots, identify…

    221 GitHub stars~3.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Starting A Project

    danshapiro/kilroy

    A skill your agent uses when bootstrapping a new project repository for Kilroy Attractor from a clean directory using existing spec, DoD, graph, and run config artifacts.

    221 GitHub stars~498 tokensUpdated 5 mo ago
    Auto-check passed
  • Using Kilroy

    danshapiro/kilroy

    Operate Kilroy Attractor pipelines end-to-end: ingest English requirements into DOT graphs, validate graph semantics, run and resume pipelines with run config files, configure provider backends…

    221 GitHub stars~4.3k tokensUpdated 5 mo ago
    Auto-check passed

Questions about Build Dod

What does Build Dod do?

A skill your agent uses when converting a spec, requirements document, or goal statement into a Definition of Done with acceptance criteria and integration test scenarios. Build Dod is an agent skill from danshapiro/kilroy.

When should I use Build Dod?

Build Dod fits situations like: converting a spec; requirements document; goal statement into a Definition of Done with acceptance criteria and integration test scenarios.

How do I install Build Dod in Claude Code?

Run `npx skills add danshapiro/kilroy --skill build-dod -a claude-code`. Or copy the skill folder (skills/build-dod in danshapiro/kilroy) into .claude/skills/build-dod in your project. Claude Code loads it when a task matches its description.

How do I install Build Dod in Codex?

Run `npx skills add danshapiro/kilroy --skill build-dod -a codex`. Or copy the skill folder (skills/build-dod in danshapiro/kilroy) into .agents/skills/build-dod in your project. Codex loads it when a task matches its description.

Can I use Build Dod in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add danshapiro/kilroy --skill build-dod -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/build-dod, .gemini/skills/build-dod, .github/skills/build-dod and .opencode/skills/build-dod in your project.

What does Build Dod need to run?

SKILL.md names no scripts, command-line tools or credentials: Build Dod is instructions for the agent only.

Does Build Dod access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Build Dod safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Build Dod use?

Build Dod is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Build Dod use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Build Dod?

Skills that share tags, products or a category with Build Dod: AI Test Generation (petrkindlmann/qa-skills, 163 stars), Prd V07 Test Planning (mattgierhart/PRD-driven-context-engineering, 179 stars), Req To Test (ArabelaTso/Skills-4-SE, 253 stars) and Build Scenario Tests (tamdogood/builder-essential-skills, 218 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Build Dod?

danshapiro (a GitHub user) maintains it in danshapiro/kilroy, which has 221 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on April 27, 2026.

Source: danshapiro/kilroy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.