Agent skill

Skeptical Triage

by avelikiy in avelikiy/great_cto

Runs a three-round self-challenge plus an arbiter over high-stakes findings, so false positives from reviews, audits and flaky-test verdicts do not become blockers.

MITAuto-check: notesDevelopment

Install Skeptical Triage

skills CLI
$ npx skills add avelikiy/great_cto --skill skeptical-triage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install avelikiy/great_cto skeptical-triage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/avelikiy/great_cto.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skeptical-triage .claude/skills/skeptical-triage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skeptical-triage
GitHub stars
103
Token cost
~2.1k tokens
SKILL.md length
901 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Runs a three-round self-challenge plus an arbiter over high-stakes findings, so false positives from reviews, audits and flaky-test verdicts do not become blockers.

  • Works in 6 steps: Absence of defense → VALID, not… → A constant name is not a verified bound… → Name the line or it does not exist.… → …
  • Checking a P0 or P1 security review finding before it blocks a release
  • SKILL.md covers When to invoke, The 4-step pattern, Hard rules and Confidence scoring, plus 4 more sections
  • Calls jq

What it does

A finding goes through three skeptical review rounds followed by an impartial arbiter that turns the votes into a confidence score. Round one asks whether the premise is true, for example whether an outside attacker can actually reach the flagged code path. Round two checks that every cited defense really exists by finding its implementation line with grep. Round three looks for angles the earlier rounds missed, such as error paths, race windows and test pollution.

Each round returns a small JSON record with a VALID, INVALID or UNCERTAIN verdict, its reasoning and the key fact it rests on, and later rounds see the earlier ones. A table says when to apply the pattern: P0 and P1 review findings, security audit results, flaky-test verdicts and architecture trade-off disputes. It is skipped for hard findings such as a secret committed to source or a confirmed CVE, and for advisory items. The price is a few extra model turns.

When your agent uses it

  • Checking a P0 or P1 security review finding before it blocks a release
  • Deciding whether a failing test is a real regression or a flaky test
  • Settling an architecture decision dispute where both options look reasonable
  • Filtering false positives out of a multi-angle code review

Example prompts

  • “Run skeptical triage on the SQL injection finding in the review report.”
  • “Is this failing checkout test a regression or a flake? Challenge it in three rounds.”
  • “Both options in our ADR look reasonable, so put them through skeptical triage.”

Requirements

  • The Read, Grep, Bash and Glob tools
  • Pre-approved tools (allowed-tools): Read, Grep, Bash, Glob

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Absence of defense → VALID, not UNCERTAIN. If you searched for a defense and did not find one, that is the answer. "Other code probably…
  2. A constant name is not a verified bound — only its resolved value is. Grep for the #define / const declaration.
  3. Name the line or it does not exist. Vague references to "assumptions in this codebase" do not count.
  4. Do not contradict your own conclusion in the same response. If you verified a defense is insufficient, that is the verdict. Stop searching…
  5. Code quality issue ≠ security vulnerability. Data race on diagnostic state, NULL check on internal-only API, UB only in debug builds →…
  6. Trust your own reasoning. If you see the crux on first read, don't manufacture a counter-argument.

What it can do on your machine

Read from SKILL.md and the folder at commit 0e9df12. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Bash
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skeptical Triage loads about 2.1k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 901 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Bash, Glob

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from avelikiy/great_cto at commit 0e9df12, republished under its MIT licence (© avelikiy). 901 words, ~2,101 tokens.

Download SKILL.mdSave it as .claude/skills/skeptical-triage/SKILL.md (or your agent's skills folder).
name
skeptical-triage
description
Reusable 3-round self-challenge + arbiter pattern for filtering false positives from findings/verdicts. Use when the cost of a false-positive gate block exceeds the cost of ~4 extra LLM turns.
allowed-tools
Read, Grep, Bash, Glob
when_to_use
Apply skeptical triage when: - A finding could block a gate (gate:code, gate:ship, gate:qa, gate:arch) and flipping it wrongly wastes CTO time - A verdict is…
effort
medium
paths
docs/**, src/**, lib/**, app/**

Skeptical Triage

Filter false positives from multi-angle review, security audit, QA regression flags, or any high-stakes judgment before it turns into a blocker.

Three rounds of skeptical self-review + an impartial arbiter, with a confidence score from the vote.

When to invoke

CallerFinding typeApply triage?
/reviewAngle 2/4/7/9 P0/P1 (security, SQL, privacy, concurrency)Yes
/review --deepAny angle P0/P1Yes
security-officerCSO audit P0/P1Yes
security-officerSecret in source/git, confirmed CVENo — hard finding
qa-engineerFlaky-test verdict (is this a regression or flake?)Yes
architectADR trade-off dispute (option A vs. B when both look reasonable)Yes
AnyP2/advisoryNo

The 4-step pattern

Run these sequentially. Each round sees prior reasoning. Arbiter sees all rounds.

Round 1 — Reachability / Premise

Question: is the premise true?

  • For security/reliability: can an external attacker reach this code path with untrusted input? Trace input flow backward from the bug site to its origin. If only trusted internal callers → lean INVALID.
  • For regressions: does the failing behavior reproduce from a clean state on the target branch?
  • For ADR trade-offs: is the constraint that forces the choice actually binding? (e.g. "we need <10ms p99" — is that real or aspirational?)

Output: {round: 1, verdict: VALID|INVALID|UNCERTAIN, reasoning: "...", crux: "single key fact"}

Round 2 — Verify cited defenses / counter-evidence

Question: are claimed defenses real and sufficient?

  • Every cited defense → use Grep to find its actual implementation line.
  • Resolve constant names to numeric values. MAX_BUF_SIZE is not a verified bound — #define MAX_BUF_SIZE 64 is.
  • For regressions: is the cited "test covers this" actually asserting the right invariant?
  • For ADR: is the cited benchmark/precedent real (grep for it, read it), or rumored?

If you cannot point to the line that enforces the defense, it does not exist.

Output: same JSON shape, with grep_used: true/false.

Round 3 — Missed angles

Question: what did Rounds 1-2 not consider?

  • Error paths, integer overflow, race windows, different callers, platform differences
  • Do NOT rehash prior rounds — add new evidence or concede
  • For QA: retry logic masking the failure? Test pollution from another test?
  • For ADR: option C that neither reviewer raised?

Output: same JSON shape.

Arbiter

Input: all 3 rounds + original finding/question + source code.

Question: final call — which side has the stronger evidence?

  • Deliver single verdict: VALID|INVALID (no UNCERTAIN — make the call).
  • Deliver one-sentence crux — the key fact the verdict turns on.
  • If 3 prior rounds all said the same thing, only override with overwhelming new evidence and explain why.

Output:

json
{
  "verdict": "VALID",
  "crux": "memcpy at auth.c:142 copies network-controlled len bytes into 64-byte stack buffer with no bound check",
  "reasoning": "Rounds 1 and 3 verified attacker reach; Round 2 found no size check in 50 LOC radius; arbiter confirms no caller clamps len."
}

Hard rules

Burn these into every round's prompt:

  1. Absence of defense → VALID, not UNCERTAIN. If you searched for a defense and did not find one, that is the answer. "Other code probably handles this" is not a valid defense.
  2. A constant name is not a verified bound — only its resolved value is. Grep for the #define / const declaration.
  3. Name the line or it does not exist. Vague references to "assumptions in this codebase" do not count.
  4. Do not contradict your own conclusion in the same response. If you verified a defense is insufficient, that is the verdict. Stop searching for reasons to flip.
  5. Code quality issue ≠ security vulnerability. Data race on diagnostic state, NULL check on internal-only API, UB only in debug builds → INVALID.
  6. Trust your own reasoning. If you see the crux on first read, don't manufacture a counter-argument.
Show full SKILL.md (352 more words)Show less

Confidence scoring

confidence = valid_rounds_before_arbiter / 3
  • 100% (VVV) — 3/3 rounds VALID. Arbiter rubber-stamps unless it finds something brand-new.
  • 67% (VVI or VIV or IVV) — majority VALID. Arbiter breaks tie with new evidence.
  • 33% (IIV or IVI or VII) — majority INVALID. Arbiter usually confirms INVALID.
  • 0% (III) — 3/3 INVALID. Arbiter rarely overrides.

Arbiter overrides the final verdict; confidence reflects the round vote for transparency. Record both in the output so humans can see where the arbiter diverged.

Applying triage results to severity

Once the arbiter returns:

Arbiter verdictConfidenceSeverity action
VALID≥ 50%Keep original severity
VALID< 50%Demote: P0→P1, P1→P2
INVALIDanyRemove from gate tally, record as [FILTERED] in report for audit
UNCERTAIN (only if arbiter could not decide)n/aKeep original severity, flag for manual CTO review

Output schema

Every caller logs triage results to .great_cto/triage-log.jsonl (append-only, one JSON per line):

json
{
  "timestamp": "2026-04-19T12:34:56Z",
  "caller": "review|security-officer|qa-engineer|architect",
  "finding_id": "SEC-042",
  "file": "src/auth.c:142",
  "original_severity": "P0",
  "rounds": [
    {"round": 1, "verdict": "VALID",   "crux": "..."},
    {"round": 2, "verdict": "VALID",   "crux": "...", "grep_used": true},
    {"round": 3, "verdict": "INVALID", "crux": "..."}
  ],
  "arbiter": {"verdict": "VALID", "crux": "..."},
  "confidence": 0.67,
  "final_severity": "P0"
}

This log is how we measure whether triage earns its keep. Review it weekly:

bash
# False-positive rate: how many findings the arbiter flipped to INVALID
jq 'select(.arbiter.verdict=="INVALID")' .great_cto/triage-log.jsonl | wc -l

# Average rounds-to-consensus (did we need all 3 or did R1+R2 agree?)
jq '[.rounds[].verdict] | unique | length' .great_cto/triage-log.jsonl

If FP rate < 10% after 50 triages — triage is filtering noise that wasn't there. Lower threshold or skip triage for that angle. If FP rate > 40% — original review prompt is too trigger-happy; tighten the angle rules.

Token budget

Per triaged finding: ~4 LLM turns (3 rounds + arbiter). At typical review sizes (~5-10 triaged findings per PR), total budget: 20-40 extra turns per /review. Batch when possible — one arbiter can handle multiple findings in a single call if their cruxes are independent.

For cost-sensitive runs (approval-level: auto on a huge PR), consider: triage only P0, leave P1 untriaged. Re-tune based on .great_cto/triage-log.jsonl data.

Anti-patterns

  • Don't triage P2/advisory findings. The whole point is gate decisions. P2 is advisory — let the author see it and move on.
  • Don't let rounds rehash each other. Round 3 prompt must say "add NEW evidence or concede." If 3 rounds produce identical reasoning, you wasted 2 turns.
  • Don't skip the arbiter on UNCERTAIN. If all 3 rounds say UNCERTAIN, the arbiter's job is to decide — not to join the fog.
  • Don't hide arbiter overrides. When the arbiter flips the majority vote, record both confidence (the vote) and final_verdict (the arbiter). Humans deserve to see the disagreement.

© avelikiy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/skeptical-triage of avelikiy/great_cto.

Open the folder on GitHubat commit 0e9df12

Compare with similar skills

Skeptical Triage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skeptical Triage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skeptical Triage this skillavelikiy/great_cto103—~2.1kAutomated safety check: NotesMIT
Codex CLIsundial-org/awesome-openclaw-skills663—~2kAutomated safety check: PassNone
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k4 repos~1.1kAutomated safety check: PassMIT
Code Review Skillawesome-skills/code-review-skill2.1k—~2.8kAutomated safety check: NotesMIT
Requesting Code ReviewHezaoHezao/poirot2495 repos~1.6kAutomated safety check: PassMIT

Similar skills

  • Codex CLI

    sundial-org/awesome-openclaw-skills

    Use OpenAI Codex CLI for coding tasks. An agent skill from sundial-org/awesome-openclaw-skills.

    663 GitHub stars~2k tokensUpdated 7 mo ago
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Code Review Skill

    awesome-skills/code-review-skill

    Provides comprehensive code review guidance for React 19, Vue 3, Angular 17+, Svelte 5, Rust, TypeScript, Java, Java 8, PHP, Ruby, Rails, Python, Django, FastAPI, Go, C/.NET, Kotlin, Swift, Dart…

    2.1k GitHub stars~2.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes
  • Requesting Code Review

    HezaoHezao/poirot

    Pre-commit review: security scan, quality gates, auto-fix. An agent skill from HezaoHezao/poirot.

    249 GitHub starsUsed in 5 repos~1.6k tokens
    DevelopmentAuto-check passed
  • Code Review Specialist

    luongnv89/claude-howto

    Reviews code for security, performance, quality and maintainability, using a checklist, a finding template and two metrics scripts.

    42k GitHub stars~764 tokensUpdated today
    DevelopmentAuto-check passed

More from avelikiy/great_cto

All 27 skills in this repo
  • AnyDesign Design Analyzer

    avelikiy/great_cto

    Analyzes a screenshot, website or Figma file and writes a `design.md` with its token system, component inventory and reconstruction notes, or an `element.md` for one element.

    103 GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Opportunity Solution Tree

    avelikiy/great_cto

    Builds an Opportunity Solution Tree that links one measurable outcome to customer opportunities, candidate solutions and experiments.

    103 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Rewrites a feature-list roadmap into outcome statements that name the customer segment, the result they get and the business impact, grouped into themes.

    103 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Exposed Secret Rotation

    avelikiy/great_cto

    Turns a leaked key, token or password into one tracked rotation task the moment it's spotted, instead of a reminder repeated every session.

    103 GitHub stars~884 tokensUpdated today
    Auto-check: notes
  • Aesthetic Instrument

    avelikiy/great_cto

    greatcto's own committed aesthetic — the instrument panel. An agent skill from avelikiy/great_cto.

    103 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Anti Patterns

    avelikiy/great_cto

    Catalogue of known SDLC anti-patterns that greatcto agents must actively reject when reviewing architecture, plans, code, or post-mortems.

    102 GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check passed

Questions about Skeptical Triage

What does Skeptical Triage do?

Runs a three-round self-challenge plus an arbiter over high-stakes findings, so false positives from reviews, audits and flaky-test verdicts do not become blockers. A finding goes through three skeptical review rounds followed by an impartial arbiter that turns the votes into a confidence score. Round one asks whether the premise is true, for example whether an outside attacker can actually reach the flagged code path.

When should I use Skeptical Triage?

Skeptical Triage fits situations like: checking a P0 or P1 security review finding before it blocks a release; deciding whether a failing test is a real regression or a flaky test; settling an architecture decision dispute where both options look reasonable; filtering false positives out of a multi-angle code review.

How do I install Skeptical Triage in Claude Code?

Run `npx skills add avelikiy/great_cto --skill skeptical-triage -a claude-code`. Or copy the skill folder (skills/skeptical-triage in avelikiy/great_cto) into .claude/skills/skeptical-triage in your project. Claude Code loads it when a task matches its description.

How do I install Skeptical Triage in Codex?

Run `npx skills add avelikiy/great_cto --skill skeptical-triage -a codex`. Or copy the skill folder (skills/skeptical-triage in avelikiy/great_cto) into .agents/skills/skeptical-triage in your project. Codex loads it when a task matches its description.

Can I use Skeptical Triage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add avelikiy/great_cto --skill skeptical-triage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skeptical-triage, .gemini/skills/skeptical-triage, .github/skills/skeptical-triage and .opencode/skills/skeptical-triage in your project.

What does Skeptical Triage need to run?

Going by SKILL.md and its folder, Skeptical Triage needs the command-line tools its instructions call (jq). Our summary lists: The Read, Grep, Bash and Glob tools. Its frontmatter pre-approves these tools: Read, Grep, Bash, Glob.

Does Skeptical Triage access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skeptical Triage safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Skeptical Triage use?

Skeptical Triage is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skeptical Triage use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skeptical Triage?

Skills that share tags, products or a category with Skeptical Triage: Codex CLI (sundial-org/awesome-openclaw-skills, 663 stars), PR Babysitter (openinterpreter/openinterpreter, 69k stars), Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars) and Code Review Skill (awesome-skills/code-review-skill, 2.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skeptical Triage?

avelikiy (a GitHub user) maintains it in avelikiy/great_cto, which has 103 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on October 10, 2026.

Source: avelikiy/great_cto on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.