Agent skill

QA Verdict Judge

by Q00 in Q00/ouroboros

Gives a fast single-pass quality verdict on any artifact, with a score from 0 to 1, PASS, REVISE or FAIL, and suggestions for what to change.

MITAuto-check passedAgent Workflows

Install QA Verdict Judge

skills CLI
$ npx skills add Q00/ouroboros --skill qa -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Q00/ouroboros qa --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Q00/ouroboros.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/qa .claude/skills/qa && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
qa
GitHub stars
6.2k
Token cost
~2.3k tokens
SKILL.md length
945 words
Files
1
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Gives a fast single-pass quality verdict on any artifact, with a score from 0 to 1, PASS, REVISE or FAIL, and suggestions for what to change.

  • Works in 4 steps: Parse the Quality Bar — What EXACTLY… → Assess Dimensions — Correctness,… → Render Verdict — Score (0.0-1.0) with… → …
  • Getting a quick pass or revise verdict on a draft, patch or document
  • SKILL.md covers Usage, How It Works, Instructions and Fallback (No MCP Server), plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The QA judge works on code, documents, API responses, test output or any custom content, and is the lightweight alternative to a three-stage formal evaluation pipeline in the same toolkit. It first pins down the quality bar, taken from acceptance criteria in a seed YAML, from you, or by asking what good means for the artifact, then rates correctness, completeness, quality, intent alignment and domain-specific points.

The verdict is a score with PASS at 0.80 or above, REVISE from 0.40 to 0.79 and FAIL below 0.40, and each maps to a loop action of done, continue or escalate. The artifact can be a file path, inline text or the most recent execution output. When the Ouroboros QA MCP tool is available the agent calls it, and a fallback mode exists for setups without the MCP server.

When your agent uses it

  • Getting a quick pass or revise verdict on a draft, patch or document
  • Checking an artifact against explicit acceptance criteria
  • Deciding whether an agent loop should finish, continue or escalate

Example prompts

  • “ooo qa on docs/api.md with the bar that every endpoint has an example.”
  • “Run a quality check on the output of the last run and tell me whether to revise.”

Requirements

  • The Ouroboros QA MCP tool, or the skill's fallback mode without it

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Parse the Quality Bar — What EXACTLY must be true to pass?
  2. Assess Dimensions — Correctness, Completeness, Quality, Intent Alignment, Domain-Specific
  3. Render Verdict — Score (0.0-1.0) with PASS / REVISE / FAIL
  4. Determine Loop Action — done (pass), continue (revise), escalate (fail)

What it can do on your machine

Read from SKILL.md and the folder at commit 0df5b98. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

QA Verdict Judge loads about 2.3k tokens when it runs. Until then it costs about 13 tokens; SKILL.md has 945 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~13
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Q00/ouroboros at commit 0df5b98, republished under its MIT licence (© Q00). 945 words, ~2,251 tokens.

Download SKILL.mdSave it as .claude/skills/qa/SKILL.md (or your agent's skills folder).
name
qa
description
General-purpose QA verdict for any artifact type

/ouroboros:qa

Standalone quality assessment for any artifact — code, documents, API responses, test output, or custom content. Unlike ooo evaluate (3-stage formal verification pipeline), ooo qa is a fast single-pass verdict with actionable suggestions.

Usage

ooo qa [file_path | artifact_text]
ooo qa                                     # evaluate recent execution output
/ouroboros:qa [file_path | artifact_text]   # plugin mode

Trigger keywords: "ooo qa", "qa check", "quality check"

How It Works

The QA Judge evaluates an artifact against a quality bar and returns a structured verdict:

  1. Parse the Quality Bar — What EXACTLY must be true to pass?
  2. Assess Dimensions — Correctness, Completeness, Quality, Intent Alignment, Domain-Specific
  3. Render Verdict — Score (0.0-1.0) with PASS / REVISE / FAIL
  4. Determine Loop Action — done (pass), continue (revise), escalate (fail)
Verdict Thresholds
Score RangeVerdictLoop Action
>= 0.80PASSdone
0.40 - 0.79REVISEcontinue
< 0.40FAILescalate

Instructions

When the user invokes this skill:

Step 0: Determine execution mode

This skill works in two modes. Determine which one before attempting any tool calls:

  • MCP mode — If the QA MCP tool is available (already exposed, or loadable via discovery), use it:

    tool discovery query: "+ouroboros qa"

    If found (typically named mcp__plugin_ouroboros_ouroboros__ouroboros_qa), proceed with QA Steps below.

  • Fallback mode — Only if the QA MCP tool is genuinely absent (no Ouroboros MCP server) skip to the Fallback section; an empty discovery result for an already-exposed tool is expected — call it directly rather than falling back. This skill is designed to work without MCP setup.

QA Steps (MCP mode)
  1. Determine the artifact to evaluate:

    • If user provides a file path: Read the file with Read tool
    • If user provides inline text: Use that directly
    • If no artifact specified: Look for the most recent execution output in conversation context
    • Ask user if unclear what to evaluate
  2. Determine the quality bar:

    • If a seed YAML is available in context: Extract acceptance criteria from it
    • If user specifies a quality bar: Use that
    • If neither: Ask the user "What does 'good' mean for this artifact?"
  3. Determine artifact type:

    • code — source code files
    • test_output — test results, CI output
    • document — specs, docs, READMEs
    • api_response — API responses, JSON payloads
    • screenshot — visual artifacts
    • custom — anything else

3.5. Acting verification fan-out — probe in parallel, then judge (do not skip for behaviour-bearing artifacts): A text judge can be fooled by a hopeful log line. When the artifact actually does something (code, an app, an API, a UI), fan out empirical probes using the host's native parallel sub-agent primitive — one probe sub-agent per acting modality the runtime actually exposes, all spawned in the same message so they run concurrently:

  • process probe (Bash/shell): run the command / start the app / run the declared smoke commands with bounded timeouts; capture exit codes and real output.
  • browser probe (browser-use tools, when the artifact serves HTTP or is a web UI): load it, click the primary flows, capture what actually renders and any console/network errors.
  • computer-use probe (desktop computer-use tools, when the artifact is a GUI/TUI): drive it like a user, screenshot the observed states.
  • artifact probe (file reads): verify declared files/paths exist with real content, not placeholders. Each probe returns structured evidence only — commands run, observed effects, screenshots/paths, pass/fail per probed behaviour. Every probe also hits the applicable adversarial classes (the QA tool lists them): misleading_output (claimed success vs. real effect), hung_command (bounded timeout?), malformed_input, stale_state, dirty_worktree. Skip a modality only when its tools are absent or the artifact type makes it meaningless — and say which modalities were skipped and why.

Await all probes, then pass the merged evidence into the judge as reference (prefer observed behaviour over source text as the artifact when they disagree). Empirical evidence outranks the judge: if the judge scores PASS but any probe observed the behaviour failing, present the verdict as REVISE/FAIL on that evidence and say so explicitly — a score contradicted by observation is not a pass. If no acting tools are available at all, judge on the text alone but flag that behaviour was not observed.

Show full SKILL.md (345 more words)Show less
  1. Call the ouroboros_qa MCP tool:

    Tool: ouroboros_qa
    Arguments:
      artifact: <the content to evaluate>
      quality_bar: <what 'pass' means>
      artifact_type: "code"  (or other type)
      reference: <observed-behaviour evidence from step 3.5, plus any reference>
      pass_threshold: 0.80  (adjustable)
      seed_content: <seed YAML if available>
  2. Present results clearly:

    • Show the score and verdict prominently
    • List dimension scores
    • Highlight specific differences found
    • Show actionable suggestions
    • End with next step guidance based on verdict:
      • PASS (done): Next: Your artifact meets the quality bar. Proceed with confidence.
      • REVISE (continue): Next: Address the suggestions above, then run ooo qa again to re-check.
      • FAIL (escalate): Next: Fundamental issues detected. Consider ooo interview to re-examine requirements, or ooo unstuck to challenge assumptions.
Iterative QA Loop

For iterative usage, track the qa_session_id and iteration_history from the response meta:

  1. First call returns qa_session_id and iteration_entry in meta
  2. On subsequent calls, pass qa_session_id and accumulated iteration_history
  3. Continue until verdict is pass or fail

In fallback mode, generate a qa-<uuid4_short> session ID on the first run and maintain iteration count in conversation context to preserve the same iterative contract.

Fallback (No MCP Server)

If the MCP server is not available, adopt the ouroboros:qa-judge agent role directly:

  1. Read the canonical agent definition: <project-root>/src/ouroboros/agents/qa-judge.md (This is the same prompt used by the MCP QA tool, ensuring consistent verdicts.)
  2. Run the same acting-verification fan-out as step 3.5 (parallel probe sub-agents per available modality; empirical evidence outranks the judge) before judging behaviour-bearing artifacts.
  3. Follow the QA Judge framework to evaluate the artifact
  4. Output the verdict in the standard format (must match MCP output shape):
QA Verdict [Iteration N]
========================
Session: qa-<id>
Score: X.XX / 1.00 [PASS/REVISE/FAIL]
Verdict: pass/revise/fail
Threshold: 0.80

Dimensions:
  Correctness:      X.XX
  Completeness:     X.XX
  Quality:          X.XX
  Intent Alignment: X.XX
  Domain-Specific:  X.XX

Differences:
  - <specific difference>

Suggestions:
  - <actionable fix>

Reasoning: <1-3 sentence summary>

Loop Action: done/continue/escalate

Example

User: ooo qa src/main.py

QA Verdict [Iteration 1]
============================================================
Session: qa-a1b2c3d4
Score: 0.72 / 1.00 [REVISE]
Verdict: revise
Threshold: 0.80

Dimensions:
  Correctness:           0.85
  Completeness:          0.60
  Quality:               0.75
  Intent Alignment:      0.80
  Domain-Specific:       0.60

Differences:
  - Missing error handling for network timeout in fetch_data()
  - No input validation on user_id parameter
  - Type hints missing on 3 public functions

Suggestions:
  - Add try/except with TimeoutError in fetch_data() (line 42)
  - Add isinstance check for user_id at function entry
  - Add return type annotations to get_user(), fetch_data(), process_result()

Reasoning: Core logic is correct but lacks defensive programming
patterns expected for production code.

Loop Action: continue

Next: Address the suggestions above, then run `ooo qa` again to re-check.

Your final response MUST end with exactly one breadcrumb footer line:

◆ <current state> → next: <recommended action>

Derive <current state> from live session state via ouroboros_session_status when that MCP projection is available; otherwise derive it from this skill's actual outcome. Never use a linear Step N of M footer because Ouroboros is an evolutionary loop. When the next action is genuinely a choice, list 2-3 honest options in the next: clause. The breadcrumb line must be the last line of the response.

© Q00, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/qa of Q00/ouroboros.

Open the folder on GitHubat commit 0df5b98

Compare with similar skills

QA Verdict Judge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

QA Verdict Judge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
QA Verdict Judge this skillQ00/ouroboros6.2k—~2.3kAutomated safety check: PassMIT
Verify and StopJuliusBrussee/caveman110k1 repos~176Automated safety check: PassApache-2.0
Iteration Progress Auditprime-radiant-inc/iterative-development181—~1.1kAutomated safety check: PassApache-2.0
Verification Before CompletionjnMetaCode/superpowers-zh8.3k—~443Automated safety check: PassMIT
Verify Before CompletionYeachan-Heo/oh-my-claudecode40k—~277Automated safety check: PassMIT
Auto Review LoopCurryTang/Amadeus176—~4.3kAutomated safety check: NotesNone

Similar skills

  • Verify and Stop

    JuliusBrussee/caveman

    Prove existing work meets acceptance conditions without expanding scope. Use for validation-only tasks, completion checks, focused gate runs, and last-mile…

    110k GitHub starsUsed in 1 repo~176 tokens
    Agent WorkflowsAuto-check passed
  • Iteration Progress Audit

    prime-radiant-inc/iterative-development

    Checks the quality of behavior evidence after each iteration in three tiers, using two auditor subagents in parallel to review the same work and find gaps.

    181 GitHub stars~1.1k tokensUpdated 4 mo ago
    Agent WorkflowsAuto-check passed
  • Verification Before Completion

    jnMetaCode/superpowers-zh

    Chinese-language rule that bars an agent from claiming work is done, fixed or passing until it has run a verification command and read the output.

    8.3k GitHub stars~443 tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Verify Before Completion

    Yeachan-Heo/oh-my-claudecode

    Has the agent prove that a feature, fix or refactor works, using existing tests first, then narrow commands and manual checks, and report only what was actually verified.

    40k GitHub stars~277 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Auto Review Loop

    CurryTang/Amadeus

    Autonomous multi-round research review loop. An agent skill from CurryTang/Amadeus.

    176 GitHub stars~4.3k tokensUpdated 6 mo ago
    Agent WorkflowsAuto-check: notes
  • Suede MCP Release QA

    JasonColapietro/suede-creator-skills

    Checks a Suede AI MCP server release against a live process: the full JSON-RPC lifecycle, schemas, annotations, malformed input, catalog agreement and install docs.

    127 GitHub stars~2.1k tokensUpdated 3 days ago
    Testing & QAAuto-check passed

More from Q00/ouroboros

All 23 skills in this repo
  • Triages and works through GitHub issues and pull requests in the Q00/ouroboros repo as a maintainer, within a stated review boundary and clear limits on what it may change.

    6.2k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Runs a guided product-manager interview that classifies each question automatically and produces a Product Requirements Document.

    6.2k GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Scans a directory for existing git repositories and worktrees, then registers and manages which ones serve as default context during interviews.

    6.2k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote.

    6.2k GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Starts, monitors or rewinds an evolutionary development loop that refines an ontology and acceptance criteria generation by generation until it converges, using the Ouroboros MCP tools.

    6.2k GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Opens or drives the Ouroboros settings GUI, picking a browser, TUI or chat-based approach depending on whether the user can reach a browser window.

    6.2k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Questions about QA Verdict Judge

What does QA Verdict Judge do?

Gives a fast single-pass quality verdict on any artifact, with a score from 0 to 1, PASS, REVISE or FAIL, and suggestions for what to change. The QA judge works on code, documents, API responses, test output or any custom content, and is the lightweight alternative to a three-stage formal evaluation pipeline in the same toolkit. It first pins down the quality bar, taken from acceptance criteria in a seed YAML, from you, or by asking what good means for the artifact, then rates correctness, completeness, quality, intent alignment and domain-specific points.

When should I use QA Verdict Judge?

QA Verdict Judge fits situations like: getting a quick pass or revise verdict on a draft, patch or document; checking an artifact against explicit acceptance criteria; deciding whether an agent loop should finish, continue or escalate.

How do I install QA Verdict Judge in Claude Code?

Run `npx skills add Q00/ouroboros --skill qa -a claude-code`. Or copy the skill folder (skills/qa in Q00/ouroboros) into .claude/skills/qa in your project. Claude Code loads it when a task matches its description.

How do I install QA Verdict Judge in Codex?

Run `npx skills add Q00/ouroboros --skill qa -a codex`. Or copy the skill folder (skills/qa in Q00/ouroboros) into .agents/skills/qa in your project. Codex loads it when a task matches its description.

Can I use QA Verdict Judge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Q00/ouroboros --skill qa -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/qa, .gemini/skills/qa, .github/skills/qa and .opencode/skills/qa in your project.

What does QA Verdict Judge need to run?

SKILL.md names no scripts, command-line tools or credentials: QA Verdict Judge is instructions for the agent only. Our summary lists: The Ouroboros QA MCP tool, or the skill's fallback mode without it.

Does QA Verdict Judge access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is QA Verdict Judge safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does QA Verdict Judge use?

QA Verdict Judge is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does QA Verdict Judge use?

About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to QA Verdict Judge?

Skills that share tags, products or a category with QA Verdict Judge: Verify and Stop (JuliusBrussee/caveman, 110k stars), Iteration Progress Audit (prime-radiant-inc/iterative-development, 181 stars), Verification Before Completion (jnMetaCode/superpowers-zh, 8.3k stars) and Verify Before Completion (Yeachan-Heo/oh-my-claudecode, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains QA Verdict Judge?

Q00 (a GitHub user) maintains it in Q00/ouroboros, which has 6,189 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 6, 2026.

Source: Q00/ouroboros on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.