Agent skill

Validate Delivery

by agentculture in agentculture/culture

Run the confirmed plan's behavioral tests agent-side after assign-to-workforce merges its waves and before summarize-delivery closes the loop, then file what was found — evidence for what passed…

Apache-2.0Auto-check passedTesting & QA

Install Validate Delivery

skills CLI
$ npx skills add agentculture/culture --skill validate-delivery -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentculture/culture validate-delivery --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentculture/culture.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/validate-delivery .claude/skills/validate-delivery && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
validate-delivery
GitHub stars
114
Token cost
~3.3k tokens
SKILL.md length
1,489 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run the confirmed plan's behavioral tests agent-side after assign-to-workforce merges its waves and before summarize-delivery closes the loop, then file what was found — evidence for what passed…

  • Works in 7 steps: Identify the obligations. Read the… → Locate the behavioral tests. Consuming… → Run the tests agent-side. The agent (not… → …
  • The user says validate delivery
  • SKILL.md covers When to invoke, The method, The CLI surface this skill… and Hard rules (do not violate), plus 5 more sections
  • Calls pytest

What it does

Validate Delivery is an agent skill from agentculture/culture. Run the confirmed plan's behavioral tests agent-side after assign-to-workforce merges its waves and before summarize-delivery closes the loop, then file what was found — evidence for what passed, behavioral deltas for what the run added, amended, or removed — as first-class, record-only entries via the devague CLI. Never runs the tests inside the CLI (issue 20); never suppresses a failing or partial outcome. Use when the user says "validate delivery", "run behavioral tests", "check what actually behaves", "file…

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA. It works with pytest. The repository describes itself as: Culture turns isolated stochastic agents into cooperative, inspectable, improvable artificial colleagues. The licence is Apache-2.0.

When your agent uses it

  • The user says validate delivery
  • Run behavioral tests
  • Check what actually behaves
  • Record a behavioral delta

Example prompts

  • “validate delivery”
  • “run behavioral tests”
  • “check what actually behaves”
  • “/validate-delivery”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Identify the obligations. Read the plan's confirmed claims (via the
  2. Locate the behavioral tests. Consuming repos identify behavioral
  3. Run the tests agent-side. The agent (not the devague CLI) executes
  4. File evidence for every obligation checked. For each obligation, file
  5. File behavioral deltas for what changed. When the run's actual
  6. Report faithfully. Summarize what was validated, what passed, what
  7. Hand off to /summarize-delivery. The filed evidence and deltas feed

What it can do on your machine

Read from SKILL.md and the folder at commit 5d12851. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Validate Delivery loads about 3.3k tokens when it runs. Until then it costs about 225 tokens; SKILL.md has 1,489 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~225
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentculture/culture at commit 5d12851, republished under its Apache-2.0 licence (© agentculture). 1,489 words, ~3,338 tokens.

Download SKILL.mdSave it as .claude/skills/validate-delivery/SKILL.md (or your agent's skills folder).
name
validate-delivery
description
Run the confirmed plan's behavioral tests agent-side after assign-to-workforce merges its waves and before summarize-delivery closes the loop, then file what was found — evidence for what passed, behavioral deltas for what the run added, amended, or removed — as first-class, record-only entries via the devague CLI. Never runs the tests inside the CLI (issue #20); never suppresses a failing or partial outcome. Use when the user says "validate delivery", "run behavioral tests", "check what actually behaves", "file evidence", "record a behavioral delta", or after assign-to-workforce merges (or fails to merge) a plan's waves and before summarize-delivery runs. Authored and maintained in agentculture/devague (origin = devague); guildmaster pulls this skill from here and broadcasts it to the AgentCulture mesh — it is NOT vendored from guildmaster like the inbound skills here.
type
command

validate-delivery — run behavioral tests, file evidence and deltas

The skill is named validate-delivery; it is the execution-to-evidence leg of the devague method — the seventh leg in flow order (the eighth origin skill, chronologically), sitting between the two closing execution skills:

text
scope -> think -> challenge -> spec-to-plan -> assign-to-workforce ->
deviate -> validate-delivery -> summarize-delivery

Where /assign-to-workforce fans out a converged plan's waves and /summarize-delivery closes the loop afterward, /validate-delivery runs after waves merge and before the delivery summary is written. It is the gap that used to be filled by memory: the confirmed plan's claims are obligations, and until this skill existed nothing forced the run to check whether the merged code actually behaves as claimed before the summary asserted it did.

When to invoke

Run this skill once a wave (or the whole plan) has merged and there is behavior to check against a claim or an approved deviation — always before /summarize-delivery, never as a substitute for it. It is not gated on a complete run: a partial or failed fan-out is still worth validating for whatever did merge.

The method

  1. Identify the obligations. Read the plan's confirmed claims (via the frame) and any approved /deviate records to see what the run promised — an announcement, an after-state, a success signal, an acceptance criterion. Each one that has a behavioral test backing it is an obligation this leg checks.
  2. Locate the behavioral tests. Consuming repos identify behavioral tests one of two ways (either is valid; pick whichever the repo already uses, and say which one in the filed evidence):
    • a pytest marker, e.g. @pytest.mark.behavioral — run with pytest -m behavioral;
    • a dedicated folder, e.g. tests/behavioral/ or behavioral-tests/ — run that path directly.
  3. Run the tests agent-side. The agent (not the devague CLI) executes the behavioral test suite, or the specific tests relevant to the obligations in scope. This is read-only against the codebase — it does not modify code to make a test pass.
  4. File evidence for every obligation checked. For each obligation, file an evidence record naming the obligation met, the test that asserted the behavior, and the outcome — pass or fail. A failing outcome is filed exactly like a passing one; it is never omitted or reworded into something softer.
  5. File behavioral deltas for what changed. When the run's actual behavior added, amended, or removed a behavior relative to the plan, file a delta record — added / amended / removed — with provenance back to the claim or approved deviation that motivated it, and forward to the evidence record(s) that back it.
  6. Report faithfully. Summarize what was validated, what passed, what failed, and what could not be checked at all (no behavioral test exists for that obligation yet). An unmet obligation is unmet — it is reported as such, not folded into a passing tally or left out of the report.
  7. Hand off to /summarize-delivery. The filed evidence and deltas feed directly into devague summary's Delivery Claims table: evidence strength (coverage / fidelity / execution / sensitivity) is the confidence vocabulary there, and any approved lapse on a claim caps its confidence the same way it always has.

The CLI surface this skill drives

Record-only. The devague CLI never runs a test itself (issue #20) — it only records what the agent already ran and found. The exact verb shapes below are minimal placeholders while the underlying schema lands in a parallel task; treat the verb names as stable and the flags as illustrative, and reconcile against devague explain <move> once that task merges.

MoveWhat it records
devague oblige <cN> --seam "<seam>" --behavior "<behavior>"Files a behavioral obligation against a claim, naming the seam to test and the behavior to assert (snapshots the claim text at filing).
devague evidence --obligation <oN> --test "<ref>" --behavior "<asserted>" --contract "<claim text>" --type <type> --strength <level> --basis "<basis>" --outcome pass|fail [--run-commit <sha> --run-timestamp <ts>]Files an evidence record: obligation met by this test, asserting this behavior, outcome pass or fail (a run reference is required at execution strength and above). llm-origin filings land proposed; the human adjudicates.
devague delta --kind added|amended|removed --behavior "<what changed>" --caused-by <cN|dN> [--evidence <eN> ...]Files a behavioral delta: provenance back to the claim/deviation it diverges from (--caused-by), forward to the evidence that backs it.
devague summary [--pr] [--json]Reads the filed evidence and deltas back into the Delivery Claims table (/summarize-delivery's starting point).

--origin llm on oblige / evidence / delta lands the record proposed — exactly the same anti-fabrication contract as deviate and lapse: an agent's own filing never self-confirms, and only the human's --confirm / --reject moves a proposed record forward. A user-origin filing auto-approves, mirroring deviate and lapse.

Hard rules (do not violate)

  • The CLI never runs tests. devague oblige / evidence / delta are record-only moves — they take the agent's already-obtained result and file it. Running the suite is the agent's job, agent-side, exactly like /summarize-delivery's read-only verification step (issue #20).
  • Unmet is unmet. A failing or unchecked obligation is filed and reported as failing or unchecked — never smoothed into "mostly passing" or silently dropped from the report. This is the direct fix for the motivating failure below: findings discovered only by reading data after the fact, never by a test failing loudly in the record.
  • A partial or failed run is still a valid input. There is no completion precondition — validate whatever merged, report the rest as not yet checkable.
  • llm-origin filings stay proposed until the user confirms. Same anti-fabrication contract as every other origin vocabulary in this method — an agent's own proposal never self-confirms.
  • Provenance both ways. Every evidence record ties back to an obligation (a claim or an approved deviation); every delta ties back to what it diverges from and forward to the evidence that backs it. An untraceable evidence or delta record is not filed.
  • This is not a new gate. Like /deviate, /validate-delivery does not add a fourth standing human gate — it produces the record /summarize- delivery and the final PR review consume; the three gates (spec, implementation split plan, final PR) are unchanged.
  • File the record the moment the thing happens, never at closeout — written late is written flattering (issue 97).
Show full SKILL.md (489 more words)Show less

Worked example

Wave 2 of a plan merged the export --format widget-md verb. The plan's confirmed success_signal claim c9 said "round-tripping a widget through export and back loses no fields." A behavioral test exists for it, marked @pytest.mark.behavioral, plus two more behavioral tests for adjacent claims — one of which fails.

bash
# 1. Identify the obligation (echoes its id, e.g. o1)
devague oblige c9 --seam "widget export round-trip" \
  --behavior "round-tripping a widget loses no fields"

# 2. Locate and run the behavioral tests agent-side (read-only)
pytest -m behavioral -q
# -> tests/behavioral/test_widget_export.py::test_round_trip PASSED
# -> tests/behavioral/test_widget_export.py::test_empty_field_rendering FAILED

# 3. File evidence for each outcome — the failure included, not smoothed over
devague evidence --obligation o1 \
  --test tests/behavioral/test_widget_export.py::test_round_trip \
  --behavior "asserts an exported-then-reimported widget compares equal field by field" \
  --contract "round-tripping a widget loses no fields" \
  --type automated --strength execution \
  --basis "behavioral test ran green at the named commit" \
  --outcome pass --run-commit abc1234 --run-timestamp 2026-08-31T12:00:00
devague evidence --obligation o2 \
  --test tests/behavioral/test_widget_export.py::test_empty_field_rendering \
  --behavior "asserts an absent widget field renders as an empty line" \
  --contract "an absent field renders honestly, never as filler" \
  --type automated --strength execution \
  --basis "behavioral test ran red at the named commit" \
  --outcome fail --run-commit abc1234 --run-timestamp 2026-08-31T12:00:00

# 4. File a delta if the failure reveals a real behavioral divergence
devague delta --kind amended \
  --behavior "empty widget fields render as garbled text, not an empty line" \
  --caused-by c11 --evidence e2

# 5. Report faithfully: c9 is validated; c11's claimed behavior is unmet —
#    say so plainly, hand it to /summarize-delivery as Remaining Work, not
#    as a passing claim.

/summarize-delivery then reads these back — c9's Delivery Claims row cites evidence e1 at high confidence (a passing behavioral test); c11's row is unverified or explicitly failing, never rounded up.

After validating — hand off to /summarize-delivery

Once every obligation in scope has an evidence record (or is reported as not yet checkable) and any behavioral deltas are filed, this leg is done — there is nothing separate to export, the filed records already live in devague state. Continue with the sibling /summarize-delivery skill: its Delivery Claims table reads the evidence and deltas filed here directly (devague summary), so the confidence a claim carries in the final delivery artifact traces back to a test that actually ran, not to memory. Don't stop at "tests ran" — the standing flow is file the evidence, then /summarize-delivery.

The motivating record

The Reasoning Degradation Ledger (devague lapse, issue agentculture/devague#97) exists because of this, cited verbatim: "Four graders failed in that cycle... Every one was found by reading data afterwards; none by a test failing." That gap — a corrections record reconstructed only at the end, from memory, because nothing forced a behavioral check to run and be filed along the way — is exactly what /validate-delivery closes for the execution side, the same way /challenge closes it for the spec side.

The design itself traces to issue agentculture/devague#107, "Suggestion: behavioral validation and a derived current spec," which proposed behavior as the primary contract, four evidence types, a strength ladder, and the current spec as a projection of a behavior ledger rather than a hand-maintained document. This skill is the method-only front door to that idea: it does not implement the full ledger or the derived-spec projection — it establishes where in the flow behavioral checking happens, what gets filed, and how the failure mode #97 documented gets closed instead of rediscovered.

Before and after this leg

text
Previous leg: deviate
Next leg: summarize-delivery

After every successful, non-exempt move, the CLI prints one next: <recommended move> line to stderr — follow it, or run devague status when unsure what comes next. The evidence and deltas filed here also feed devague today's read-only projection of current behavior into the committed docs/current-spec.md.

Provenance

This is a first-party skill — its origin is agentculture/devague, the eighth in the outbound family after /scope, /think, /challenge, /spec-to-plan, /assign-to-workforce, /deviate, and /summarize-delivery, covering the execution-to-evidence leg that runs after a plan's waves merge and before the delivery summary is written. guildmaster pulls it from here and broadcasts it to the AgentCulture mesh; because devague is upstream, it is never re-vendored back from guildmaster's re-broadcast copy. The cite, don't import policy still holds: downstream repos copy it, they don't symlink or depend on it. See docs/skill-sources.md.

© agentculture, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/validate-delivery of agentculture/culture.

Open the folder on GitHubat commit 5d12851

Compare with similar skills

Validate Delivery next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Validate Delivery compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Validate Delivery this skillagentculture/culture114—~3.3kAutomated safety check: PassApache-2.0
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Hermetic Python Unit TestsdimensionalOS/dimos4.6k—~1.4kAutomated safety check: PassCustom licence
Test GuardamElnagdy/guard-skills1.3k2 repos~2.1kAutomated safety check: PassMIT
Pytest Runnersaleor/saleor23k—~251Automated safety check: PassBSD-3-Clause
Port Node Red Nodeoldrev/edgelinkd125—~3kAutomated safety check: PassApache-2.0

Similar skills

  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Guard

    amElnagdy/guard-skills

    Reviews newly written or edited tests against nine rules that cut test bloat, such as mock-heavy checks and near-duplicate cases, before they are committed.

    1.3k GitHub starsUsed in 2 repos~2.1k tokens
    Testing & QAAuto-check passed
  • Pytest Runner

    saleor/saleor

    Run pytest tests with automatic virtual environment activation. Use this skill whenever running tests, executing pytest, or when asked to "run tests", "test…

    23k GitHub stars~251 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Port Node Red Node

    oldrev/edgelinkd

    Port a Node-RED node into EdgeLinkd the way this repo does it: implement the node in Rust under crates/core/src/runtime/nodes, mirror Node-RED's mocha spec as pytest tests under tests/, register the…

    125 GitHub stars~3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • ONNX Runtime Test Runner

    microsoft/onnxruntime

    Official

    Runs and debugs ONNX Runtime tests: Google Test executables for C++ and unittest or pytest for Python, with filters and build-directory guidance.

    22k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed

More from agentculture/culture

All 19 skills in this repo
  • Agent Config

    agentculture/culture

    Show a Culture agent's full configuration in one read-only view: its system-prompt file (CLAUDE.md / AGENTS.md / GEMINI.md), the parallel culture.yaml, and the agent's local .claude/skills index.

    115 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Cicd

    agentculture/culture

    CI/CD lane for culture: branch, commit, push, create PR, wait for automated reviewers, fetch comments, fix or pushback, reply, resolve threads.

    115 GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Communicate

    agentculture/culture

    All agent communication from culture: in-mesh chat (channels, DMs, mentions, knowledge sharing) via culture channel CLI, AND cross-repo hand-off briefs to sibling-repo agents (agentirc, steward…

    115 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Communicate

    agentculture/culture

    Cross-repo + mesh communication: file tracked GitHub issues on sibling repos, comment on existing issues, fetch issues with body + comments to inline current state into briefs, and send live…

    115 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Pypi Maintainer

    agentculture/culture

    Switch a PyPI package install between the production index, TestPyPI pre-release builds, and a local editable checkout.

    115 GitHub stars~711 tokensUpdated today
    Auto-check passed
  • Run Tests

    agentculture/culture

    Run pytest with parallel execution and coverage. An agent skill from agentculture/culture.

    115 GitHub stars~565 tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Validate Delivery

What does Validate Delivery do?

Run the confirmed plan's behavioral tests agent-side after assign-to-workforce merges its waves and before summarize-delivery closes the loop, then file what was found — evidence for what passed…. Validate Delivery is an agent skill from agentculture/culture. Run the confirmed plan's behavioral tests agent-side after assign-to-workforce merges its waves and before summarize-delivery closes the loop, then file what was found — evidence for what passed, behavioral deltas for what the run added, amended, or removed — as first-class, record-only entries via the devague CLI.

When should I use Validate Delivery?

Validate Delivery fits situations like: the user says validate delivery; run behavioral tests; check what actually behaves; record a behavioral delta.

How do I install Validate Delivery in Claude Code?

Run `npx skills add agentculture/culture --skill validate-delivery -a claude-code`. Or copy the skill folder (.claude/skills/validate-delivery in agentculture/culture) into .claude/skills/validate-delivery in your project. Claude Code loads it when a task matches its description.

How do I install Validate Delivery in Codex?

Run `npx skills add agentculture/culture --skill validate-delivery -a codex`. Or copy the skill folder (.claude/skills/validate-delivery in agentculture/culture) into .agents/skills/validate-delivery in your project. Codex loads it when a task matches its description.

Can I use Validate Delivery in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentculture/culture --skill validate-delivery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validate-delivery, .gemini/skills/validate-delivery, .github/skills/validate-delivery and .opencode/skills/validate-delivery in your project.

What does Validate Delivery need to run?

Going by SKILL.md and its folder, Validate Delivery needs the command-line tools its instructions call (pytest).

Does Validate Delivery access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Validate Delivery safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Validate Delivery use?

Validate Delivery is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Validate Delivery use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Validate Delivery?

Skills that share tags, products or a category with Validate Delivery: Adk Verify Snippets (google/adk-python, 22k stars), Hermetic Python Unit Tests (dimensionalOS/dimos, 4.6k stars), Test Guard (amElnagdy/guard-skills, 1.3k stars) and Pytest Runner (saleor/saleor, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Validate Delivery?

agentculture (a GitHub organization) maintains it in agentculture/culture, which has 114 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on October 10, 2026.

Source: agentculture/culture on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.