Doc Co-Authoring Workflow
shareAI-lab/Kode-CLI
Guides a three-stage workflow for turning partial context into a clear PRD, RFC or design doc: capture context, draft section by section, then test with a fresh reader.
Run an ASSERT evaluation against a described risk. An agent skill from responsibleai/ASSERT.
$ npx skills add responsibleai/ASSERT --skill run-assert-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install responsibleai/ASSERT run-assert-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/responsibleai/ASSERT.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/run-assert-eval .claude/skills/run-assert-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "run-assert-eval" agent skill from https://github.com/responsibleai/ASSERT/tree/main/.claude/skills/run-assert-eval into .claude/skills/run-assert-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-assert-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/responsibleai/ASSERT/tree/main/.claude/skills/run-assert-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add responsibleai/ASSERT --skill run-assert-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install responsibleai/ASSERT run-assert-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/responsibleai/ASSERT.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/run-assert-eval .agents/skills/run-assert-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "run-assert-eval" agent skill from https://github.com/responsibleai/ASSERT/tree/main/.claude/skills/run-assert-eval into .agents/skills/run-assert-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-assert-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add responsibleai/ASSERT --skill run-assert-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install responsibleai/ASSERT run-assert-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/responsibleai/ASSERT.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/run-assert-eval .cursor/skills/run-assert-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "run-assert-eval" agent skill from https://github.com/responsibleai/ASSERT/tree/main/.claude/skills/run-assert-eval into .cursor/skills/run-assert-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-assert-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/responsibleai/ASSERT.git --path .claude/skills/run-assert-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add responsibleai/ASSERT --skill run-assert-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install responsibleai/ASSERT run-assert-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/responsibleai/ASSERT.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/run-assert-eval .gemini/skills/run-assert-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "run-assert-eval" agent skill from https://github.com/responsibleai/ASSERT/tree/main/.claude/skills/run-assert-eval into .gemini/skills/run-assert-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-assert-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install responsibleai/ASSERT run-assert-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add responsibleai/ASSERT --skill run-assert-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/responsibleai/ASSERT.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/run-assert-eval .github/skills/run-assert-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "run-assert-eval" agent skill from https://github.com/responsibleai/ASSERT/tree/main/.claude/skills/run-assert-eval into .github/skills/run-assert-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-assert-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add responsibleai/ASSERT --skill run-assert-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install responsibleai/ASSERT run-assert-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/responsibleai/ASSERT.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/run-assert-eval .opencode/skills/run-assert-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "run-assert-eval" agent skill from https://github.com/responsibleai/ASSERT/tree/main/.claude/skills/run-assert-eval into .opencode/skills/run-assert-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-assert-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
run-assert-evalRun an ASSERT evaluation against a described risk. An agent skill from responsibleai/ASSERT.
Run Assert Eval is an agent skill from responsibleai/ASSERT. Run an ASSERT evaluation against a described risk. Use when the user wants to evaluate, test, or check an AI agent, LLM app, or model against requirements/policies (e.g. "evaluate my agent for budget violations", "test that the support bot never gives legal advice"). Risks come either from Clarity — recommended, driving the real Clarity MCP tools (runclarity) in-IDE to discover failure modes the user has not considered — or directly from the user as a description, PRD, design doc, threat model, red-team finding…
Its SKILL.md is about 11k tokens, which your agent loads only when the skill is triggered. The skill folder holds 37 other files, including assets (for example `README.md`, `SETUP-CHECKLIST.md` and `assets/dimension-review-template.md`).
It sits in Security, covering Architecture decision records, Threat modeling and PRD writing. The repository describes itself as: Requirement-driven evaluation harness for AI agents and LLM applications. Generate behavior-specific test cases, run them against any target (hosted models, callable wrappers… The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e03aa80. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pythonnpmpipgitFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
raw.githubusercontent.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
AZURE_API_KEYGITHUB_TOKENANTHROPIC_API_KEYOPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Run Assert Eval loads about 11k tokens when it runs. Until then it costs about 210 tokens; SKILL.md has 5,863 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
3. **Provider creds exist** in `.env`. NEVER read or print `.env`. If a run fails- **Don't read, print, or commit** `.env`, credential values, `artifacts/`, traces, `.venv`, or logs.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from responsibleai/ASSERT at commit e03aa80, republished under its MIT licence (© responsibleai). 5,863 words, ~11,079 tokens.
.claude/skills/run-assert-eval/SKILL.md (or your agent's skills folder). This skill also uses 29 other files; get the full folder from GitHub.The user wants evidence of how their agent or model actually behaves. Not for fixing the agent — this skill finds and reports failures.
This skill has two entry modes:
.clarity-protocol/ directory or a fresh
discovery run driven through the Clarity MCP server (run_clarity), in-IDE —
or risks the user supplies directly. Then turn each selected risk into an
atomic config, run the pipeline (Steps 3-5), and report (Step 6).artifacts/results/<suite>/<run>/ and the user asks a question about them
("what are the highlights?", "top 3 examples of the worst failure mode?", "why
did case X fail?"). Skip to Step 6 and answer THAT question from the artifacts —
do not re-run, and do not fall back to the full canned report unless asked.Every eval starts from a risk. There are two supported sources, and the user chooses — never decide for them and never block on Clarity.
Path A — Clarity discovery (recommended — present it first, but never alone). Use an existing
.clarity-protocol/ or a fresh run via the Clarity MCP run_clarity tool.
Clarity's value is finding failure modes the user has not thought of, along
with severity and causal chains. Recommend it whenever the user is unsure what
to measure, is new to the agent, or wants coverage rather than one known bug.
Path B — user-supplied risks. The user names the risk themselves, as prose or by pointing at a PRD, design doc, threat model, red-team finding, incident report, risk assessment, or test plan. This is the right path when they already know what they want measured.
Both paths answer what to test for. Neither answers how. That is the job of the research procedure in Step 3: once a risk is named, it runs a literature review of how that risk has actually been evaluated and turns the findings into the test-set design. The output is not a restatement of the topic — it is how the topic manifests:
scenario over prompt.
max_turns itself is fixed at 6 — the timescale finding chooses the test
mode, not the turn budget.levels when sources support them.This is the difference between a config that names a risk and a config that can actually measure it.
Whenever you need a new risk to measure, and the user has not already named one, offer the choice:
I can find a risk two ways. Clarity interviews you and surfaces failure modes you may not have considered — recommended when you know the agent but aren't sure what to measure. Or you name it directly, in your own words or by pointing me at a PRD, design doc, threat model, red-team finding, or risk assessment — best when you already know what you want measured. Either way I then research how that risk has been evaluated and build the test set from that evidence. Which do you prefer?
An existing .clarity-protocol/ changes the default, never the choice.
Offer it as the recommended option ("I found an existing Clarity protocol with
these risks — measure one of those, or is there a different risk you have in
mind?"), then take the user's answer.
Rules that hold on both paths:
.clarity-protocol/ exists. Never substitute the protocol's risks for one the
user just stated. If you think the protocol covers the same ground, say so and
let them decide; do not decide for them.SETUP-CHECKLIST.md — but if they'd rather not set it up now, take Path B.run_clarity tool,
which returns Clarity's genuine process guide inlined. Path B is not a
degraded impression of Clarity — it is a distinct, structured intake (Step 1b).sample_size. Steps 3-6 are risk-source agnostic; nothing about the
config, run, or report changes.record_failure /
record_suggestion apply only when a protocol exists. On Path B, skip them and
say so once; never treat their absence as an error.Copilot is for answering questions and synthesis — direct answers, failure-mode clustering, cited examples, next actions — with no clicking. The bundled local viewer is for visual exploration — forest plots, baseline compare, facet grouping, and stepping through a transcript with the judge's citations highlighted. Answer in chat when the user asks "what / why / which"; hand off to the viewer (Step 7) when they want to see, read a full transcript, compare runs, or watch a live run.
ASSERT installed: assert-ai --help succeeds. If not, guide install from
PyPI — not an editable install of the user's own repo:
python -m pip install "assert-ai[phoenix]" The target project owns its agent framework dependencies. For repository
examples, install the adjacent requirements.txt; for a customer project,
use that project's existing dependency manifest. target.endpoint needs
aiohttp, which ships directly with ASSERT — no
separate extra to install. Use pip install -e ".[phoenix]" only when the
working directory is a clone of the ASSERT repo itself; inside a customer
repo it installs the wrong package.
Clarity MCP server available (needed only for Path A): the clarity-agent
MCP tools (run_clarity, write_protocol_document, record_failure,
record_suggestion, …) are callable in this session. Clarity is the
risk-discovery engine — the skill drives its real MCP tools, it does not
reimplement it. If the tools are missing, the server is not wired up yet: offer
SETUP-CHECKLIST.md (install clarity-agent with the [mcp] extra, run
clarity embed . to generate .vscode/mcp.json, reload MCP servers) and
confirm the LLM provider is configured (clarity doctor — Clarity supports
GitHub Copilot, Anthropic, OpenAI, Azure AI, and Gemini).
This is not a blocker. If the tools can't be made available, or the user would rather not set them up now, say so plainly and continue on Path B (Step 1b). Never strand the user on MCP setup when they came to measure something.
Provider creds exist in .env. NEVER read or print .env. If a run fails
with an auth error, tell the user which variable NAMES are required
(AZURE_API_KEY, AZURE_API_BASE, OPENAI_API_KEY, GITHUB_TOKEN, ANTHROPIC_API_KEY,
etc.) — never their values.
Ask which path the user wants (see "Choosing a risk source" above), then follow 1a or 1b.
.clarity-protocol/ exists. Do not silently switch to the
protocol's risks..clarity-protocol/ exists → offer it as the
default and say what's in it, but still ask before selecting it: "I found an
existing Clarity protocol covering X and Y — want to measure one of those, or
is there a different risk you have in mind?"Risks come from Clarity's real engine, driven through the Clarity MCP server — never by imitating Clarity's interview from your own head.
.clarity-protocol/ directory already exists in the workspace, and the
user has chosen Path A for this risk, use it directly as the risk source — skip
straight to reading its output below. (Selecting Path A is the user's decision,
made in Step 1; the protocol's presence alone does not make it.)run_clarity. It returns Clarity's real process guide inlined as text.write_protocol_document and
record_failure. Continue until the failure-analysis process has written
.clarity-protocol/failures/failures.md.Read Clarity's output to enumerate risks:
.clarity-protocol/failures/failures.md — the failure modes, causal chains,
and management plans. Each distinct failure mode is one candidate ASSERT behavior..clarity-protocol/summary.md, goal/requirements.md, solution/architecture.md
— target/context for the eval's context field.For the full measurement path — parse → triage → one atomic config per selected
failure → sequential runs → report → close the loop → curate the example — follow
workflows/measure-clarity-failures.md. Use the intake parser
(clarity_intake.py) to convert failures.md into candidate behaviors with
severity→priority mapping and variant-derived stratify dimensions.
Before a fresh discovery run, check the preservation gate.
.clarity-protocol/is gitignored, single-domain scratch;run_clarityoverwrites it, destroying the prior domain'sfailures/,goal/, andsolution/with no git recovery. If a protocol from another domain is present, STOP and let the user export it to a user-owned location or explicitly discard it. Never commit the raw discovery workspace intoexamples/.
Clarity records severity/management-plan signal (the parser maps Critical→P1, High→P2, Medium→P3, ranges→max). Order and annotate by what Clarity actually captured; do not fabricate priorities.
The user already knows what to measure. Your job is to turn their input into the
same candidate-behavior shape clarity_intake.py produces on Path A —
{name, description, severity, priority, source_doc, candidate_dimensions, multi_behavior, suggested_splits} — so Steps 2-6 are identical either way.
contextbehavior.name + behavior.descriptionelicitation_variant stratify dimension — a research seed
for Step 3, not the final setprioritymulti_behavior / suggested_splits
check, applied by hand.Set source_doc to the file you read, or user-described when it came from chat.
Record severity as the user rated it; do not invent a priority they didn't give.
For the full measurement path — triage → one atomic config per selected risk →
sequential runs → report → curate the example — follow
workflows/measure-clarity-failures.md, the same workflow Path A uses. Skip its
Step 1 (Parse): there is no failures.md to parse, so join at Step 2 with the
candidate list you just built. Skip its Step 8 (close the loop in Clarity) too,
unless a .clarity-protocol/ exists.
Then continue to Step 2. Everything downstream is unchanged.
Clarity intentionally over-produces (whole-lifecycle threat modeling). Do NOT auto-generate an eval for every failure mode. Surface the enumerated list (ordered by severity signal) and ask the user which to measure now (e.g. "top-severity only?", or named picks). Carry only the selected risks forward.
On Path B the list is usually short and already chosen — still play it back and confirm scope before generating configs, rather than assuming every risk they mentioned should be measured in this pass.
ASSERT performs best with one atomic behavior per eval. Never bundle multiple risks into one config — bundling makes the reported impermissible behavior violated rate a fuzzy logical-OR and hides per-behavior signal.
Configs are researched, cited, and user-approved — not scaffolded and hoped for.
Follow workflows/research-eval-dimensions.md for
each selected risk. That workflow owns the whole of config generation; do not hand-roll a
config here and do not skip its gates.
Its purpose is narrow and worth stating plainly: the risk already has a name by the
time you arrive here. What the research supplies is how that risk has been evaluated
— the timescale it becomes observable on, whose viewpoint exercises it, and which
conditions change it — expressed as stratify dimensions, behavior_category_count,
judge dimensions, and the scenario vs prompt test mode. It does not re-open
what to measure.
Collect one input before entering it, and never silently default it:
| Input | Rule |
|---|---|
N | Positive integer — how many complete dimension-generation passes to run before deduplication. Ask for it when missing or invalid rather than inferring one. |
What that workflow does, in order:
assert-ai library list / show <name>; prefer
behavior.preset or a copy-in spec from examples/behavior_specs/ over reinventing a
description. This is also what settles the harm's stable slug.generation-isolation-workflow.md)
— once the slug is stable, detects prior generations for it by path only, and asks
before using a new dated directory. It never reads a prior generated YAML.N passes and deduplicate (iterative-dimension-workflow.md)
— N complete passes, then semantic deduplication within each namespace.examples/<domain>/<risk>[_YYYY-MM-DD]/eval_config.yaml, with
inline # sources: citations and a consolidated # References block.Outputs land at examples/<domain>/<risk>[_YYYY-MM-DD]/eval_config.yaml — one directory per
generation, never overwritten. Prefix the eval suite name with a domain slug
(<domain>-<risk>) so artifacts/results/<suite>/ and artifacts/acs/<suite>/ do not
collide across domains.
Two things that workflow will ask you to decide, and that matter downstream:
behavior_category_count is 25 — the standard count, and ASSERT's own default
(DEFAULT_BEHAVIOR_CATEGORY_COUNT). Research shapes which categories are generated,
not how many.sample_size is a question for the user, but it has a hard floor:
≥ behavior_category_count (so ≥25). Below the category count some behavior
categories receive zero cases and are silently unmeasured. Each rate is also
violations / sample_size, so even at the 25 floor one flipped case moves the
number 4 percentage points, and the swing grows as the sample shrinks. The floor
protects coverage, not precision — prefer 50+ when the expected delta is small.
Write the chosen value into the config with an inline review comment
(# min for behavior-category coverage -- user should review; 50+ tightens the signal) so it reads as a floor the user still owns, not a settled default.max_turns is fixed at 6 — ASSERT's default (DEFAULT_TESTER_MAX_TURNS) and
the config template's value. The research does not move it, and it must be
identical in the baseline and governed configs or the "only ACS differs"
comparison breaks. A genuinely single-turn harm is expressed by writing prompt
test cases, not by lowering the turn budget — max_turns is read only for
scenario cases.Judge dimensions are authored from the research, and are added on top of the
built-ins — but never reuse a built-in name. policy_violation and overrefusal are
BUILT_IN_DIMENSIONS (assert_ai/core/judge.py) and are always judged unless explicitly
disabled. Config dimensions merge over the built-ins by name into the same dict, so a
researched dimension called policy_violation silently replaces the built-in rubric — no
warning, no error.
What shadowing does and does not affect. It does not move the headline pair. Both
not_permissible_policy_violation_rate and permissible_policy_violation_rate are
recomputed in assert_ai/results.py from the judge's per-behavior node_judgments — each
relevant node's violated flag, bucketed by that behavior category's permissible value —
so they never read either dimension's verdict. What shadowing does change is the
dimension verdict persisted in the run JSON, and DEFAULT_COMPARE_METRIC
(assert_ai/cli.py), which is still policy_violation. That is a comparability and
stored-artifact problem, not a corrupted headline.
The same applies to judge presets, which expand into that same merged list:
judge.preset: safety-core. Its only two dimensions are named exactly
policy_violation and overrefusal, so it replaces both built-in rubrics — and buys
nothing, because the engine treats that pair as superseded:
_SUPERSEDED_DISPLAY_METRICS (assert_ai/cli.py) hides both from CLI detail surfaces
whenever the permissibility split is available, which for this skill's configs is always.
You would pay judge tokens per dimension per row for two numbers that are never shown and
never reported.safety-extended is safe and recommended for nuanced harms — it defines
harm_actionability and pii_leakage, which collide with nothing and are displayed.The post-write gate rejects both the inline and the preset form.
If live source retrieval is unavailable, say so and stop at the ledger. The evidence
gate cannot be met without it, and a config with remembered or invented citations is worse
than no config. assert-ai init --describe-file … remains available as an explicitly
unvalidated scaffold — never for a measurement you intend to report or govern against.
After generation, show the user the resolved behavior.description, context,
pipeline.judge settings, the systematize / judge models, and the reference list.
Confirm before running.
Help the user set the right target in the config:
target.callable WITH target.trace so the judge can cite tool calls and routing.target.model and target.tools.assert-ai judge-traces --traces <path> --config <path>; do not add a --trace flag to assert-ai run.target.endpoint — the runtime POSTs {"message": ..., "history": [...]} and reads {"response": ...}, so no wrapper code is needed (requires aiohttp). Only write a thin target.callable shim if the service's request/response shape differs. Either way the judge sees only final text, so this is a fallback, not the recommended path.target.callable takes a module.path:function reference. The full signature and
return-type contract lives in docs/targets/callable.md.
Two behaviors that doc does not cover can silently corrupt a run:
history is detected by parameter name, not position. ASSERT introspects the
signature and enables multi-turn only when a parameter is literally named history.
Name it messages, conversation, or chat_history and every scenario silently
degrades to single-turn — the run completes, the viewer renders, and the numbers are
wrong with no warning. Confirm the name before trusting any multi-turn baseline, and
therefore any ACS delta measured against it.sys.path → the config's own
directory → the current working directory → direct file load. An agent.py sitting
beside the YAML config resolves even when the CLI is invoked from the repo root —
but a same-named module earlier on sys.path wins, so prefer a domain-unique module
name over a bare agent.target.trace is not optionalTracing decides how much of the agent the judge can actually see. Per the observability
matrix in docs/targets/callable.md ("What the judge sees, by integration path"):
| Integration path | Signals visible to the judge |
|---|---|
Plain str return | 1 of 8 — final text only |
| LiteLLM-style response | 4 of 8 — adds final tool calls, token usage, model name |
| OTel traces | 8 of 8 — adds intermediate tool calls, routing / sub-agent decisions, intermediate model calls, per-span latency |
So without traces a tool-misuse or wrong-routing failure is largely invisible to scoring —
which is why this skill mandates target.callable with target.trace. You rarely
hand-write spans: ASSERT ships OTel auto-instrumentation for 33 frameworks (LangChain /
LangGraph, CrewAI, OpenAI Agents SDK, DSPy, LlamaIndex, AutoGen, MAF, Pydantic AI, …) as a
single helper call at the top of the callable module — see docs/targets/callable.md
("Recommended: OTel-traced agent (33 frameworks)").
Offer a smoke run first. A suite is 25 prompt + 25 scenario cases, and
plumbing errors (wrong callable, missing credentials, a callable that raises
on its first tool call, tool-schema mismatch, undeployed judge model) surface
only once inference starts. Validate on 3 real cases first:
# 1. artifacts only, no inference cost
assert-ai run --config examples/<domain>/<risk>/eval_config.yaml \
--override inference.enabled=false --override judge.enabled=false
# 2. slice 3 real rows out of the generated test set
python .claude/skills/run-assert-eval/smoke_slice.py \
--config examples/<domain>/<risk>/eval_config.yaml --count 3
# 3. inference + judge on those rows only
assert-ai run --config examples/<domain>/<risk>/eval_config.yaml \
--override run=<run>-smoke \
--override inference.test_set_path=<out path from step 2>If it fails, stop and report — do not start the full run. Three cases is not a
measurement, so never report a rate from a smoke run. Never lower
test_set.sample_size instead: it invalidates the cached test set and does not
produce a subset. Full detail in workflows/measure-clarity-failures.md
Step 5a.
assert-ai run --config examples/<domain>/<risk>/eval_config.yaml --output jsonThis is long-running (systematize -> test_set -> inference -> judge). Stream status
to the user as each stage completes. For N configs, run them sequentially and track
each suite/run. After a smoke run the first two stages report CACHED.
--force-stage <stage>suite and run names from the config for Step 6.Read only structured artifacts. Aggregate from the pre-computed, schema'd files —
never trawl raw Phoenix/OpenTelemetry traces to reconstruct an answer (that bulk,
unguided trace-reading is exactly what the viewer's evidence drawer is for). Reading
the inference_set.jsonl row for a specific case the judge already cited is fine;
bulk trace trawling is not.
Headline rates: run assert-ai results status <suite> <run> for per-dimension
flagged rates (split into prompt and scenario). The headline pair is the
permissibility split: add --json and read not_permissible_policy_violation_rate
(real harm got through) and permissible_policy_violation_rate (the agent broke a
behavior it was allowed to do), each one vote per conversation. Report only this
pair — do not report the built-in policy_violation or overrefusal rates (the raw
policy_violation ORs over ALL violated taxonomy nodes, permissible included, so it
couples with overrefusal and is uninformative here). This same pair is what to
headline in an ACS A/B — harm should drop while permissible stays flat (see
workflows/govern-and-remeasure.md). The viewer exposes the same pair as the
dimension keys policy_violation_not_permissible / policy_violation_permissible,
rendered on screen as Impermissible behavior violated /
Permissible behavior violated.
Researched judge dimensions: configs generated by
workflows/research-eval-dimensions.md carry harm-specific judge dimensions on top of
the built-ins (e.g. harm_actionability, severe_harm_escalation,
longitudinal_harm_pattern). Report each one's flagged rate alongside the
permissibility split, never folded into it — they answer different questions
("did harm get through?" vs "how bad was it when it did?"). Name each dimension as the
config defines it, and quote its rubric when the rate needs interpreting.
Top failing cases: read scores.jsonl from artifacts/results/<suite>/<run>/.
For each dimension with failures, pull 3-5 representative cases with:
verdict.dimensions — which dimensions failedverdict.dimension_justifications — the judge's rationale with cited evidenceverdict.node_judgments — which behavior categories were violated, with reasoningCost and timing: read metrics.json for token usage and elapsed time per stage.
This file contains cost metadata only, not score roll-ups.
For Results Q&A mode, answer the user's specific question from these same artifacts
(e.g. rank dimensions by flagged rate for "top failure mode", then quote
dimension_justifications for the cited examples). Don't emit the full template unless asked.
After reporting, point the user to the bundled viewer for anything visual or self-directed — it went through extensive design iteration and owns the exploration surface Copilot should not replicate:
cd viewer && npm install && npm run dev # then open http://localhost:5174Select the suite and run for forest plots, per-dimension breakdowns, facet grouping,
the permissible vs. not-permissible policy-violation split (also available from
assert-ai results status --json and rendered by results compare),
and a transcript drawer with the judge's [N] citations highlighted on the cited turns.
Suggest it specifically when the user wants to:
assert-ai results compare <suite> <runA> <runB>)manifest.json-driven)See docs/guides/use-local-viewer.md for the full layout.
When a run surfaces impermissible behavior violated failures and the user wants
to fix and prove it, don't stop at prompt-tweaking. Generate a deployable ACS (Agent
Control Specification) policy from the findings and re-run the same eval against
the governed agent to show the failure rate dropped — the ACS delta. This uses
ASSERT's native assert-ai acs generate / validate adapter (no external acs
CLI). It requires a callable target whose high-risk tools can be wrapped
(control.protect_tool); a hosted-model Prompt Agent target has nothing
wrappable. Follow workflows/govern-and-remeasure.md for the full loop
(baseline → acs generate → acs validate → governed run → delta from two
results status --json calls → export each run to standalone HTML → close the
loop in Clarity). results compare --metric policy_violation_not_permissible
and --metric policy_violation_permissible compare either half directly; use
the two status --json rate fields when you need machine-readable counts.
Classify the failure before generating the policy (Step 1a): read the baseline's
verdict.dimension_justifications to decide semantic (output annotator) vs
structural (tool gate), and confirm the harm actually routes through the tool you
plan to gate. Getting that wrong is the main cause of a gate that fires ~0 times.
If the governed run's delta still comes out wrong (no drop, or the permissible-violation / over-gating rate rose),
workflows/diagnose-acs-delta.md is the symptom-indexed diagnostic manual —
match the signature, apply the smallest fix, cap at ~4 attempts.
examples/billing_support_agent/agent.py shows the baseline callable shape; the
governed entrypoint is an output of that workflow, not a checked-in file.
Once the eval is passing locally and the user wants it enforced on PRs, hand off to the wire-assert-ci skill. Do not author CI wiring here; direct the user or agent to the action bootstrap:
read https://raw.githubusercontent.com/responsibleai/assert-ai-action/main/ONBOARD.mdPresent a short summary with this structure:
Headline metrics — the only two rates reported:
not_permissible_policy_violation_rate]permissible_policy_violation_rate]Report the permissibility split (from results status --json). Do not report the
built-in policy_violation or overrefusal rates — the raw policy_violation ORs over
all violated nodes and couples with overrefusal, so neither is reported; the split is
the complete headline.
Researched judge dimensions (when the config declares them), each on its own line with its flagged rate — reported beside the headline pair, never merged into it:
<dimension_name>: X% (N/M cases) — one-line gloss of what its rubric scoresEvidence base: the config's # References list (tag → title → URL), so the
provenance of the dimensions being reported is visible alongside the numbers.
Top failing cases (3-5 per dimension): For each failure:
Suggested next step: one concrete action (e.g. "tighten the system prompt
around X behavior", "add a stratify dimension for Y", or govern the failure with ACS and
re-measure to prove the rate dropped — see Step 8 and
workflows/govern-and-remeasure.md).
Team-maintained docs on main. Prefer linking these over restating their content here —
when they disagree with this skill on product behavior, they win; this skill owns the
methodology (the Clarity → ASSERT → ACS → ASSERT loop) and the traps called out above.
| Doc | Use it for | Step |
|---|---|---|
docs/guides/create-evaluation.md | Authoring an eval config from scratch | 3 |
docs/config/schema.md | Full config field reference | 3 |
docs/targets/callable.md | Callable signature, return types, OTel auto-instrumentation | 4 |
docs/targets/model-and-tools.md | target.model + target.tools shape | 4 |
docs/guides/troubleshooting.md | A run errors, hangs, or produces no scores | 5 |
docs/guides/results.md | Interpreting results and artifacts | 6 |
docs/guides/use-local-viewer.md | Viewer layout and drill-down | 7 |
docs/guides/securing-agents-with-acs.md | The ACS generate → validate → guard → re-run path | 8 |
| Workflow | Owns |
|---|---|
workflows/measure-clarity-failures.md | The full measurement path: parse → triage → config → run → report → close the loop |
workflows/research-eval-dimensions.md | Config generation — evidence-gated dimension research, N passes, approval, cited write |
workflows/iterative-dimension-workflow.md | The N-pass cycle, semantic deduplication, and the approval gate |
workflows/generation-isolation-workflow.md | Path-only prior-generation preflight and isolated output directories |
workflows/evaluation-intent-workflow.md | Optional intake: what decision the eval supports, and for whom |
workflows/govern-and-remeasure.md | The ACS baseline → generate → governed run → delta loop |
workflows/diagnose-acs-delta.md | Symptom-indexed diagnostics when the ACS delta comes out wrong |
Helper scripts at the skill root: clarity_intake.py (parse failures.md),
smoke_slice.py (slice N real rows for a smoke run), plan_generation_path.py
(path-only isolation preflight), validate_dimension_review.py
(render / validate / pre-write / post-write gates).
.clarity-protocol/ or a fresh run_clarity run) and risks they
supply themselves. Recommend Clarity, because it surfaces failure modes they
haven't considered — but never present it as the only route. Any menu, list, or
question you offer that includes a Clarity option must carry the user-supplied
option beside it; a user who doesn't know Path B exists cannot ask for it. Hold
the user-supplied path (Step 1b) to the same bar: atomic behaviors, an explicit
permissible boundary, researched and cited dimensions. Never block a measurement on
Clarity setup.run_clarity returns its genuine process
guide inlined). Step 1b is a distinct structured intake, not a hand-rolled
impression of Clarity.run_clarity / write_protocol_document / record_failure for discovery and record_suggestion to close the loop; never hand the user off to a separate Clarity app and never shell out to a clarity cli process.record_suggestion (or record_decision) back into .clarity-protocol/ noting the failure mode now has a measured baseline and where the eval lives, so Clarity's staleness tracking stays aware of it. With no protocol, skip it silently — and consider offering Clarity as a next step for finding risks this pass didn't cover.assert-ai acs generate), review it (scope the gated tools, tighten conditions), and re-run the same eval against the governed callable to show the delta; needs a wrappable callable target (workflows/govern-and-remeasure.md). Generated policies, guarded targets, and governed configs are local run output by default. Commit them only in the user's own product repo when the user wants a reviewed policy deployed; do not automatically add them to ASSERT's worked examples. Whenever a gate needs a value the model doesn't put in the tool args — a trusted session flag (verification), a trusted comparison value (the caller's own id), a trusted numeric cap, or a running total / prior-call fact — the governed agent must surface that scalar from its session state into the tool-call policy_target so the generated input.policy_target.value.* rule actually fires. ACS evaluates each call in isolation, so multi-call constraints (running totals, ordering, rate limits) are handled by that same injection, not by encoding history in Rego. Free-form content failures (unsafe advice, PII in prose, a verbal-only high-risk promise) and inbound prompt-injection instead use an annotator-based gate at the output/input point, proven by the remeasure delta since offline validate can't run annotators. Never hand-drive an external acs CLI for this loop.<domain>-<risk>, e.g. billing-cross-customer-data-exposure, science-<risk>), so artifacts/results/<suite>/ and artifacts/acs/<suite>/ do not collide. Treat .clarity-protocol/ as uncommitted single-domain scratch; preserve it outside examples/ only when the user asks.examples/<domain>/ containing only what a customer needs to understand and reproduce the ASSERT run:agent.py (+ any real runtime deps it imports, e.g. tools.py / mock_tools.py) — the runnable baseline.README.md — scenario, setup, atomic behaviors, run commands, and result paths.eval_config.yaml — one independently runnable baseline config per behavior, written by
workflows/research-eval-dimensions.md into its own isolated
examples/<domain>/<risk>[_YYYY-MM-DD]/ directory and never overwritten.
Do not commit the dimension-review ledger or its approval stamp — those are working
artifacts under artifacts/dimension-reviews/.
A deliberately curated ACS demonstration may additionally keep the smallest
reviewed policy, guarded target, and governed config needed to reproduce its
claim, but ordinary worked examples must not accumulate generated governance
output. Do not commit generated taxonomies, test sets, result artifacts,
discovery mailboxes, snapshots, protocol archives, or automatic skill output.workflows/research-eval-dimensions.md owns config generation. Every dimension must pass its evidence gate, N passes must complete, and the user must explicitly approve the dimension set before any YAML is written. Silence is not approval. assert-ai init remains available as an explicitly unvalidated scaffold, never for a measurement you intend to report or govern against.uncited — needs review. If live retrieval is unavailable, stop at the ledger and say so.policy_violation and overrefusal are BUILT_IN_DIMENSIONS; config dimensions merge over them by name, so reusing one silently replaces its rubric. This does not move the headline split (results.py recomputes it from node_judgments), but it does change the verdict stored in the run JSON and DEFAULT_COMPARE_METRIC. judge.preset: safety-core does this too, since presets expand into the same merged list — and buys nothing, because the engine treats that pair as superseded and hides it once the permissibility split is available. The pre-write gate rejects a reused name in the review ledger; the post-write gate rejects it in the written config, in both the inline and the preset form.cat, parse, grep, hash, git show, or otherwise inspect a matching prior YAML, and do not infer its contents from size, timestamps, or commit history.results status, scores.jsonl, and metrics.json; hand off to the viewer for visual trace/transcript exploration..env, credential values, artifacts/, traces, .venv, or logs.© responsibleai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 29 other files (assets) in .claude/skills/run-assert-eval of responsibleai/ASSERT.
Open the folder on GitHubat commit e03aa80
Run Assert Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Run Assert Eval this skillresponsibleai/ASSERT | 327 | — | ~11k | Automated safety check: Notes | MIT | |
| Doc Co-Authoring WorkflowshareAI-lab/Kode-CLI | 5.2k | — | ~977 | Automated safety check: Pass | Apache-2.0 | |
| Forensifyalexgreensh/repo-forensics | 187 | — | ~2.5k | Automated safety check: Notes | Custom licence | |
| Schematicblader/schematic | 239 | — | ~2.2k | Automated safety check: Pass | MIT | |
| Osint Methodologyelementalsouls/Claude-OSINT | 2.8k | — | ~8.7k | Automated safety check: Notes | MIT | |
| Shep Workstreamsshep-ai/shep | 264 | — | ~2.5k | Automated safety check: Pass | MIT |
shareAI-lab/Kode-CLI
Guides a three-stage workflow for turning partial context into a clear PRD, RFC or design doc: capture context, draft section by section, then test with a fresh reader.
alexgreensh/repo-forensics
Cross-agent self-inspection of your AI-agent stack. An agent skill from alexgreensh/repo-forensics.
blader/schematic
Reverse engineer a detailed product and technical specification document from a git branch's implementation.
elementalsouls/Claude-OSINT
Comprehensive OSINT methodology for external red-team operations and authorized attack-surface assessments.
shep-ai/shep
A skill your agent uses when a large body of work (a version milestone, an epic, a roadmap, a set of PRDs/design docs) needs to be broken into parallel workstreams and executed with the shep CLI.
VisionForge-OU/foreman
Headless grilling pass that challenges an approved implementation plan against the existing codebase and domain model, then writes an ADR draft and a PRD draft into the Foreman feature directory.
Categories
Run an ASSERT evaluation against a described risk. An agent skill from responsibleai/ASSERT. Run Assert Eval is an agent skill from responsibleai/ASSERT. Run an ASSERT evaluation against a described risk.
Run Assert Eval fits situations like: the user wants to evaluate; check an AI agent; model against requirements/policies (e.g.
Run `npx skills add responsibleai/ASSERT --skill run-assert-eval -a claude-code`. Or copy the skill folder (.claude/skills/run-assert-eval in responsibleai/ASSERT) into .claude/skills/run-assert-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add responsibleai/ASSERT --skill run-assert-eval -a codex`. Or copy the skill folder (.claude/skills/run-assert-eval in responsibleai/ASSERT) into .agents/skills/run-assert-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add responsibleai/ASSERT --skill run-assert-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-assert-eval, .gemini/skills/run-assert-eval, .github/skills/run-assert-eval and .opencode/skills/run-assert-eval in your project.
Going by SKILL.md and its folder, Run Assert Eval needs Python for the scripts in its folder, the command-line tools its instructions call (python, npm, pip and git) and credentials named AZURE_API_KEY, GITHUB_TOKEN, ANTHROPIC_API_KEY and OPENAI_API_KEY. Our summary lists: Python 3; A credential in AZURE_API_KEY; A credential in OPENAI_API_KEY.
SKILL.md names 1 domain. In commands or code: raw.githubusercontent.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Run Assert Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 11k tokens (SKILL.md is roughly 44k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Run Assert Eval: Doc Co-Authoring Workflow (shareAI-lab/Kode-CLI, 5.2k stars), Forensify (alexgreensh/repo-forensics, 187 stars), Schematic (blader/schematic, 239 stars) and Osint Methodology (elementalsouls/Claude-OSINT, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
responsibleai (a GitHub organization) maintains it in responsibleai/ASSERT, which has 327 GitHub stars. The repository was last updated on October 6, 2026.
Source: responsibleai/ASSERT on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.