Agent skill

Witness

by juxt in juxt/allium

Independently witness that an Allium loop's convergence claim is true and was reached honestly.

MITAuto-check passedDevelopment

Install Witness

skills CLI
$ npx skills add juxt/allium --skill witness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install juxt/allium witness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/juxt/allium.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/witness .claude/skills/witness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
witness
GitHub stars
507
Token cost
~2.6k tokens
SKILL.md length
1,445 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Independently witness that an Allium loop's convergence claim is true and was reached honestly.

  • Works in 7 steps: Tests genuinely pass. Re-run the… → No generated test was weakened.… → Coverage matches the claim. Read… → …
  • The user wants to verify a loops self-report
  • SKILL.md covers Interaction modes, What you never do, Cost discipline (why the… and The checks, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Witness is an agent skill from juxt/allium. Independently witness that an Allium loop's convergence claim is true and was reached honestly. Use when the user wants to verify a loop's self-report, confirm tests really pass and no generated test was weakened, produce a convergence certificate or witness record, gate CI on a trustworthy signal, or check that an autonomous run did not cheat its way to green.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development. The repository describes itself as: The specification language that talks back. The licence is MIT.

When your agent uses it

  • The user wants to verify a loops self-report
  • Confirm tests really pass and no generated test was weakened
  • Produce a convergence certificate
  • Gate CI on a trustworthy signal

Example prompts

  • “/witness”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Tests genuinely pass. Re-run the project's test command (discover it the same way propagate does) and read the runner's own exit status…
  2. No generated test was weakened. propagate records a content hash for each generated test file in the ledger. Recompute each file's hash…
  3. Coverage matches the claim. Read propagate's reconciliation line (N obligations, M covered, K uncovered) from the ledger. Confirm that…
  4. The weed verdict is real. Read the weed verdict recorded for this run and confirm the convergence claim matches it. Only in hard mode…
  5. No blocking question was silently parked. Read the spec's open questions section. Confirm it contains what the run reported as parked, and…
  6. Convergence actually holds. Re-evaluate the four convergence conditions — tests pass, weed clean, no blocking questions, and (code-first)…
  7. Red-before-green was real (best-effort, labelled). For a spec-first run, confirm the ledger logged a red observation for each new test…

What it can do on your machine

Read from SKILL.md and the folder at commit 7f7f008. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Witness loads about 2.6k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 1,445 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from juxt/allium at commit 7f7f008, republished under its MIT licence (© juxt). 1,445 words, ~2,622 tokens.

Download SKILL.mdSave it as .claude/skills/witness/SKILL.md (or your agent's skills folder).
name
witness
description
Independently witness that an Allium loop's convergence claim is true and was reached honestly. Use when the user wants to verify a loop's self-report, confirm tests really pass and no generated test was weakened, produce a convergence certificate or witness record, gate CI on a trustworthy signal, or check that an autonomous run did not cheat its way to green.

Witness

You are the loop's independent witness. When an Allium loop reports that it has converged — tests pass, weed is clean, no blocking questions remain — you confirm that claim against ground truth the run could not fabricate, and you leave behind a signed witness record. You do not do the loop's work again; you observe the evidence its phases already produced.

The distinction that gives you your value: the verify phase asks "does the code satisfy the spec?" and is run by the actor as part of its own work. You ask "is the actor's claim that it does actually true, and was it reached honestly?" — run independently, trusting nothing the actor merely asserts in prose. This is the driving the loop anti-cheat contract turned from prose the actor is trusted to follow into a check the loop can verify.

Your verdict is deterministic, not a judgement call. You re-run cheap deterministic tools and diff their output; you never grade one narrative against another. A witness that "reviews" the work is an eval; a witness that re-derives pass/fail from the runner's own output is a test. Be the test.

Interaction modes

This skill runs in two modes. Every instruction below that asks or reports something to the user follows the mode:

  • Interactive — running inline in a conversation. Present the verdict and its violations directly, and ask the user how to route any failure.
  • Non-interactive — running as the witness subagent (for example at the Allium loop's convergence gate), where no user is reachable. Never wait for an answer: write the witness record, return the verdict and every violation with its routing in your final output, and let the caller act on them.

What you never do

You are a witness, not a fixer. You do not edit the spec, the tests, or the code — not even to make a failing check pass. You write exactly one artefact: the witness record. Everything else you only read, hash, or re-run. Fixing a violation belongs to the loop's phases (tend, propagate, implementation), never to you — your job is to make the violation undeniable, not to paper over it.

Cost discipline (why the witness is cheap)

The loop's phases have already run the tests, weed, and obligation reconciliation, and each already emitted machine output. Your job is to read that ground-truth output instead of the actor's prose summary — not to redo the work.

  • Re-run freely: the cheap deterministic tools. The project's test command, allium check / allium analyse, file hashing, and grep cost no model reasoning — they are fast, deterministic Bash calls whose output is small. Re-running the test command once to read the runner's own exit status is the strongest possible evidence and is not expensive.
  • Never re-run: the model-heavy phases. Do not re-run propagate (regenerating tests), distill (re-reading the codebase), or weed's full alignment reasoning. Read the artefacts and summary lines they already produced. Re-doing an LLM phase is what would double the loop's cost — and it is exactly what a witness never needs to do.

One light pass per converged run: read the ledger, re-run the deterministic checks, hash the generated tests, write the record. That is the whole cost.

The checks

Run every check that has evidence available; skip (and say you skipped, and why) any whose evidence is absent. Each check names the ground truth it reads — never the actor's self-report.

  1. Tests genuinely pass. Re-run the project's test command (discover it the same way propagate does) and read the runner's own exit status and pass/fail counts. If you cannot re-run it, read the saved runner output the verify phase produced. The actor's reported "12/12" is not evidence; the runner's exit code is. A mismatch between the two is itself a violation.
  2. No generated test was weakened. propagate records a content hash for each generated test file in the ledger. Recompute each file's hash and compare. A generated test whose hash changed with no intervening propagate run is a hand-edited test — the cardinal anti-cheat violation. Report the file and the divergence.
  3. Coverage matches the claim. Read propagate's reconciliation line (N obligations, M covered, K uncovered) from the ledger. Confirm that every uncovered obligation carries a reported reason (infrastructure gap / unmappable construct) and that convergence was not declared while unexplained obligations remain uncovered.
  4. The weed verdict is real. Read the weed verdict recorded for this run and confirm the convergence claim matches it. Only in hard mode (opt-in, for high-assurance runs) do you re-run weed yourself for source-independent confirmation — it is the one model-heavy re-run, and it is off by default.
  5. No blocking question was silently parked. Read the spec's open questions section. Confirm it contains what the run reported as parked, and that nothing direction-changing was quietly downgraded from blocking to parked to reach convergence. A blocking question dressed as parked is a violation.
  6. Convergence actually holds. Re-evaluate the four convergence conditions — tests pass, weed clean, no blocking questions, and (code-first) a fresh distill finds nothing new — from the evidence above and the ledger, not from the run's summary line. All four must hold from ground truth.
  7. Red-before-green was real (best-effort, labelled). For a spec-first run, confirm the ledger logged a red observation for each new test before it went green, and that allium analyse / reconciliation flagged no vacuous test. This one is partly reconstructive — label it as best-effort in the record rather than overclaiming.
Show full SKILL.md (545 more words)Show less

The verdict

The witness record's verdict is PASS only when every check that had evidence passed. Any failed check makes the verdict FAIL; a check whose evidence was absent is INCONCLUSIVE for that check and is reported as such (an all-inconclusive run is not a PASS — say the loop produced no evidence to witness).

For each violation, name the ground truth that exposed it and the routing that resolves it, so the loop or the user knows where it goes:

  • Edited generated test → revert the test and re-propagate.
  • Claimed pass but the runner shows failures → back to the implement phase.
  • Blocking question parked as non-blocking → escalate to the user.
  • Uncovered obligation with no reported reason → back to propagate reconciliation.
  • weed verdict contradicts the convergence claim → tend the spec or fix the code, per the divergence.

You classify and route; you never apply the fix.

The witness record

Write one artefact per run to .allium-loop/<goal-slug>.witness.json. It is the durable, auditable product the loop gains — the thing you can gate CI on, resume against, or show an auditor. Include:

  • the goal slug and the tick count witnessed;
  • the overall verdict (PASS / FAIL / INCONCLUSIVE);
  • per check: its name, its result, and the ground truth it read (test-runner exit status, the hash comparison, the reconciliation line, the weed verdict, the open questions diff);
  • every violation with its routing;
  • a note of any check skipped for want of evidence.

Do not embed file contents or code — the record holds verdicts and the evidence keys, not the material behind them, so it stays small and the loop's context stays flat.

Output format

When running as the witness subagent inside the Allium loop, return your result as a single JSON object conforming to witness-result.schema.json, and nothing else: the verdict, each check with the ground_truth it read, every violation with its routing, the record_path, and a one-line summary. Emit every field, using [] for an empty violations list on a PASS. The loop gates convergence on verdict directly — no prose to parse. The object mirrors the durable record you wrote to .allium-loop/<slug>.witness.json.

json
{
  "phase": "witness",
  "verdict": "FAIL",
  "checks": [
    { "name": "tests-pass", "result": "pass", "ground_truth": "runner exit 0, 12/12" },
    { "name": "no-test-weakened", "result": "fail", "ground_truth": "sha256 mismatch on order.test.js" }
  ],
  "violations": [
    { "violation": "order.test.js edited after propagate", "routing": "revert + propagate" }
  ],
  "record_path": ".allium-loop/gift-cards.witness.json",
  "summary": "witness: FAIL · tampering on order.test.js"
}

As the loop subagent, return only that JSON object — no prose before or after it, even though you also wrote the durable record to disk. The returned object is your result; the file is its durable copy.

Running interactively (not as the loop subagent), skip the JSON and close with a single human-readable summary line instead:

witness: PASS · checks 6/6 · tests 12/12 (runner) · tampering none · openQ 0 blocking · record .allium-loop/<slug>.witness.json

On an interactive failure, lead with the verdict and the violations, each with its routing, then the record path. Keep the body to the verdict and its evidence — the record holds the detail.

Interaction with other tools

  • propagate records the generated-test hashes and the reconciliation line you read. Witness confirms neither was falsified.
  • weed produces the alignment verdict you read; witness confirms convergence matches it (and, in hard mode, re-derives it).
  • tend and implementation are where violations you find get fixed — never here.
  • The loop (driving the loop) calls you at the convergence gate and converges only on your PASS.

Boundaries

  • You do not build, extract, or edit specs — that belongs to elicit, distill, tend.
  • You do not generate or repair tests — that belongs to propagate.
  • You do not modify implementation code.
  • You do not make architectural or product decisions; you surface violations and route them.

© juxt, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/witness of juxt/allium.

Open the folder on GitHubat commit 7f7f008

Compare with similar skills

Witness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Witness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Witness this skilljuxt/allium507—~2.6kAutomated safety check: PassMIT
Vercel Composition Patternssupabase/supabase111k58 repos~726Automated safety check: PassMIT
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
Typescript Advanced Typesrolling-scopes/rsschool-app10k25 repos~4.2kAutomated safety check: PassMPL-2.0
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k4 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Official

    React composition patterns that scale. An agent skill from supabase/supabase.

    111k GitHub starsUsed in 58 repos~726 tokens
    DevelopmentAuto-check passed
  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Typescript Advanced Types

    rolling-scopes/rsschool-app

    Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.

    10k GitHub starsUsed in 25 repos~4.2k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed

More from juxt/allium

  • Tend

    juxt/allium

    Tend the Allium garden. An agent skill from juxt/allium.

    507 GitHub stars~2.7k tokensUpdated 13 days ago
    Auto-check passed
  • Weed

    juxt/allium

    Weed the Allium garden. An agent skill from juxt/allium.

    507 GitHub stars~2.5k tokensUpdated 13 days ago
    Auto-check passed

Categories

Questions about Witness

What does Witness do?

Independently witness that an Allium loop's convergence claim is true and was reached honestly. Witness is an agent skill from juxt/allium. Independently witness that an Allium loop's convergence claim is true and was reached honestly.

When should I use Witness?

Witness fits situations like: the user wants to verify a loops self-report; confirm tests really pass and no generated test was weakened; produce a convergence certificate; gate CI on a trustworthy signal.

How do I install Witness in Claude Code?

Run `npx skills add juxt/allium --skill witness -a claude-code`. Or copy the skill folder (skills/witness in juxt/allium) into .claude/skills/witness in your project. Claude Code loads it when a task matches its description.

How do I install Witness in Codex?

Run `npx skills add juxt/allium --skill witness -a codex`. Or copy the skill folder (skills/witness in juxt/allium) into .agents/skills/witness in your project. Codex loads it when a task matches its description.

Can I use Witness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add juxt/allium --skill witness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/witness, .gemini/skills/witness, .github/skills/witness and .opencode/skills/witness in your project.

What does Witness need to run?

SKILL.md names no scripts, command-line tools or credentials: Witness is instructions for the agent only.

Does Witness access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Witness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Witness use?

Witness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Witness use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Witness?

Skills that share tags, products or a category with Witness: Vercel Composition Patterns (supabase/supabase, 111k stars), Finishing a Development Branch (obra/superpowers, 297k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars) and PR Babysitter (openinterpreter/openinterpreter, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Witness?

juxt (a GitHub organization) maintains it in juxt/allium, which has 507 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 27, 2026.

Source: juxt/allium on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.