Evidence-based investigative code review using deductive reasoning to determine what actually happened versus what was claimed.

MITAuto-check passedDevelopment

Install Sherlock Review

skills CLI
$ npx skills add proffesor-for-testing/agentic-qe --skill sherlock-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install proffesor-for-testing/agentic-qe sherlock-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/proffesor-for-testing/agentic-qe.git skills-src && mkdir -p .claude/skills && cp -r skills-src/assets/skills/sherlock-review .claude/skills/sherlock-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sherlock-review
GitHub stars
495
Token cost
~1.8k tokens
SKILL.md length
404 words
Files
3 (incl. scripts)
Skills in repo
93
Repo updated
First seen
Licence
MIT

At a glance

Evidence-based investigative code review using deductive reasoning to determine what actually happened versus what was claimed.

  • Works in 5 steps: OBSERVE: Gather all evidence (code,… → DEDUCE: What does evidence actually show… → ELIMINATE: Rule out what cannot be true → …
  • Verifying implementation claims
  • SKILL.md covers Quick Reference Card, Investigation Template, Minimum Findings Enforcement and Investigation Scenarios, plus 5 more sections
  • Calls git and npm

What it does

Sherlock Review is an agent skill from proffesor-for-testing/agentic-qe. Evidence-based investigative code review using deductive reasoning to determine what actually happened versus what was claimed. Use when verifying implementation claims, investigating bugs, validating fixes, or conducting root cause analysis. Elementary approach to finding truth through systematic observation.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `schemas/output.json` and `scripts/validate-config.json`).

It sits in Development, covering Root cause analysis. It works with Git. The repository describes itself as: Agentic QE Fleet is an open-source AI-powered QA/QE platform designed for use with Coding Agents (works best with Claude Code) featuring specialized agents and skills to support… The licence is MIT.

When your agent uses it

  • Verifying implementation claims
  • Investigating bugs
  • Validating fixes
  • Conducting root cause analysis

Example prompts

  • “/sherlock-review”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. OBSERVE: Gather all evidence (code, tests, history, behavior)
  2. DEDUCE: What does evidence actually show vs. what was claimed?
  3. ELIMINATE: Rule out what cannot be true
  4. CONCLUDE: Does evidence support the claim?
  5. DOCUMENT: Findings with proof, not assumptions

What it can do on your machine

Read from SKILL.md and the folder at commit 1363bc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sherlock Review loads about 1.8k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 404 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from proffesor-for-testing/agentic-qe at commit 1363bc7, republished under its MIT licence (© proffesor-for-testing). 404 words, ~1,772 tokens.

Download SKILL.mdSave it as .claude/skills/sherlock-review/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
sherlock-review
description
Evidence-based investigative code review using deductive reasoning to determine what actually happened versus what was claimed. Use when verifying implementation claims, investigating bugs, validating fixes, or conducting root cause analysis. Elementary approach to finding truth through systematic observation.
category
quality-review
priority
high
tokenEstimate
1100
agents
qe-code-reviewer, qe-security-auditor, qe-performance-validator
implementation_status
optimized
optimization_version
1
last_optimized
2025-12-03
quick_reference_card
true
tags
investigation, evidence-based, code-review, root-cause, deduction
trust_tier
2

Sherlock Review

<default_to_action> When investigating code claims:

  1. OBSERVE: Gather all evidence (code, tests, history, behavior)
  2. DEDUCE: What does evidence actually show vs. what was claimed?
  3. ELIMINATE: Rule out what cannot be true
  4. CONCLUDE: Does evidence support the claim?
  5. DOCUMENT: Findings with proof, not assumptions

The 3-Step Investigation:

bash
# 1. OBSERVE: Gather evidence
git diff <commit>
npm test -- --coverage

# 2. DEDUCE: Compare claim vs reality
# Does code match description?
# Do tests prove the fix/feature?

# 3. CONCLUDE: Verdict with evidence
# SUPPORTED / PARTIALLY SUPPORTED / NOT SUPPORTED

Holmesian Principles:

  • "Data! Data! Data!" - Collect before concluding
  • "Eliminate the impossible" - What cannot be true?
  • "You see, but do not observe" - Run code, don't just read
  • Trust only reproducible evidence </default_to_action>

Quick Reference Card

Evidence Collection Checklist
CategoryWhat to CheckHow
ClaimPR description, commit messagesRead thoroughly
CodeActual file changesgit diff
TestsCoverage, assertionsRun independently
BehaviorRuntime outputExecute locally
TimelineWhen things happenedgit log, git blame
Verdict Levels
VerdictMeaning
✓ TRUEEvidence fully supports claim
⚠ PARTIALLY TRUEClaim accurate but incomplete
✗ FALSEEvidence contradicts claim
? NONSENSICALClaim doesn't apply to context

Investigation Template

markdown
## Sherlock Investigation: [Claim]

### The Claim
"[What PR/commit claims to do]"

### Evidence Examined
- Code changes: [files, lines]
- Tests added: [count, coverage]
- Behavior observed: [what actually happens]

### Deductive Analysis

**Claim**: [specific assertion]
**Evidence**: [what you found]
**Deduction**: [logical conclusion]
**Verdict**: ✓/⚠/✗

### Findings
- What works: [with evidence]
- What doesn't: [with evidence]
- What's missing: [gaps in implementation/testing]

### Recommendations
1. [Action based on findings]

Minimum Findings Enforcement

Every investigation MUST surface at least 3 weighted observations (CRITICAL=3, HIGH=2, MEDIUM=1, LOW=0.5). Elementary observations count at INFORMATIONAL=0.25 weight. A Sherlock investigation that finds nothing is a failed investigation -- Holmes always finds clues.


Investigation Scenarios

Scenario 1: "This Fixed the Bug"

Steps:

  1. Reproduce bug on commit before fix
  2. Verify bug is gone on commit with fix
  3. Check if fix addresses root cause or symptom
  4. Test edge cases not in original report

Red Flags:

  • Fix that just removes error logging
  • Works only for specific test case
  • Workarounds instead of root cause fix
  • No regression test added
Show full SKILL.md (150 more words)Show less
Scenario 2: "Improved Performance by 50%"

Steps:

  1. Run benchmark on baseline commit
  2. Run same benchmark on optimized commit
  3. Compare in identical conditions
  4. Verify measurement methodology

Red Flags:

  • Tested only on toy data
  • Different comparison conditions
  • Trade-offs not mentioned
Scenario 3: "Handles All Edge Cases"

Steps:

  1. List all edge cases in code path
  2. Check each has test coverage
  3. Test boundary conditions
  4. Verify error handling paths

Red Flags:

  • catch {} swallowing errors
  • Generic error messages
  • No logging of critical errors

Example Investigation

markdown
## Case: PR #123 "Fix race condition in async handler"

### Claims Examined:
1. "Eliminates race condition"
2. "Adds mutex locking"
3. "100% thread safe"

### Evidence:
- File: src/handlers/async-handler.js
- Changes: Added `async/await`, removed callbacks
- Tests: 2 new tests for async flow
- Coverage: 85% (was 75%)

### Analysis:

**Claim 1: "Eliminates race condition"**
Evidence: Added `await` to sequential operations. No actual mutex.
Deduction: Race avoided by removing concurrency, not synchronization.
Verdict: ⚠ PARTIALLY TRUE (solved differently than claimed)

**Claim 2: "Adds mutex locking"**
Evidence: No mutex library, no lock variables, no sync primitives.
Verdict: ✗ FALSE

**Claim 3: "100% thread safe"**
Evidence: JavaScript is single-threaded. No worker threads used.
Verdict: ? NONSENSICAL (meaningless in this context)

### Conclusion:
Fix works but not for reasons claimed. Race condition avoided by
making operations sequential, not by adding synchronization.

### Recommendations:
1. Update PR description to accurately reflect solution
2. Add test for concurrent request handling
3. Remove incorrect technical claims

Agent Integration

typescript
// Evidence-based code review
await Task("Sherlock Review", {
  prNumber: 123,
  claims: [
    "Fixes memory leak",
    "Improves performance 30%"
  ],
  verifyReproduction: true,
  testEdgeCases: true
}, "qe-code-reviewer");

// Bug fix verification
await Task("Verify Fix", {
  bugCommit: 'abc123',
  fixCommit: 'def456',
  reproductionSteps: steps,
  testBoundaryConditions: true
}, "qe-code-reviewer");

Agent Coordination Hints

Memory Namespace
aqe/sherlock/
├── investigations/*   - Investigation reports
├── evidence/*         - Collected evidence
├── verdicts/*         - Claim verdicts
└── patterns/*         - Common deception patterns
Fleet Coordination
typescript
const investigationFleet = await FleetManager.coordinate({
  strategy: 'evidence-investigation',
  agents: [
    'qe-code-reviewer',        // Code analysis
    'qe-security-auditor',     // Security claim verification
    'qe-performance-validator' // Performance claim verification
  ],
  topology: 'parallel'
});


Remember

"It is a capital mistake to theorize before one has data." Trust only reproducible evidence. Don't trust commit messages, documentation, or "works on my machine."

The Sherlock Standard: Every claim must be verified empirically. What does the evidence actually show?

© proffesor-for-testing, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in assets/skills/sherlock-review of proffesor-for-testing/agentic-qe.

  • SKILL.md
  • schemas/output.json
  • scripts/validate-config.json

Open the folder on GitHubat commit 1363bc7

Compare with similar skills

Sherlock Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sherlock Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sherlock Review this skillproffesor-for-testing/agentic-qe495—~1.8kAutomated safety check: PassMIT
Code Design Rationale Investigatorcursor/plugins11k9 repos~2.6kAutomated safety check: PassNone
Debugging and Error Recoveryaddyosmani/agent-skills104k1 repos~2.6kAutomated safety check: PassMIT
Root Cause Tracingsandgardenhq/sgai1374 repos~1.4kAutomated safety check: PassCustom licence
Codexqa Rootcause Analyzeropenqa-cn/codexqa152—~2.6kAutomated safety check: PassApache-2.0
Investigate Issueanalogjs/analog3.2k—~2kAutomated safety check: PassMIT

Similar skills

  • Official

    Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs.

    11k GitHub starsUsed in 9 repos~2.6k tokens
    DevelopmentAuto-check passed
  • Debugging and Error Recovery

    addyosmani/agent-skills

    Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.

    104k GitHub starsUsed in 1 repo~2.6k tokens
    DevelopmentAuto-check passed
  • Root Cause Tracing

    sandgardenhq/sgai

    A skill your agent uses when errors occur deep in execution and you need to trace back to find the original trigger - systematically traces bugs backward through call stack, adding instrumentation…

    137 GitHub starsUsed in 4 repos~1.4k tokens
    DevelopmentAuto-check passed
  • Diagnoses exception root causes from stack traces, logs, call-chain dumps, and debug output using the CodexQA CLI for structured repo analysis.

    152 GitHub stars~2.6k tokensUpdated 7 days ago
    DevelopmentAuto-check passed
  • Investigate Issue

    analogjs/analog

    Investigate a GitHub issue end to end — reproduce the reporter's repo or code snippet in an isolated sandbox outside the monorepo, trace the root cause in the source, and draft a reply back to the…

    3.2k GitHub stars~2k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Debug

    gnomeria/usbtree

    Systematic root-cause debugging — reproduce, isolate, fix at the source, prove the fix.

    691 GitHub stars~715 tokensUpdated 1 mo ago
    DevelopmentAuto-check passed

More from proffesor-for-testing/agentic-qe

All 93 skills in this repo
  • Contract Testing

    proffesor-for-testing/agentic-qe

    Consumer-driven contract testing for microservices using Pact, schema validation, API versioning, and backward compatibility testing.

    495 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Mutation Testing

    proffesor-for-testing/agentic-qe

    Test quality validation through mutation testing, assessing test suite effectiveness by introducing code mutations and measuring kill rate.

    495 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Performance Testing

    proffesor-for-testing/agentic-qe

    Profiles application performance under load using k6, Artillery, or JMeter to measure latency, throughput, and error rates.

    495 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Code Review Quality

    proffesor-for-testing/agentic-qe

    Conduct context-driven code reviews focusing on quality, testability, and maintainability.

    495 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Security Testing

    proffesor-for-testing/agentic-qe

    Scans for security vulnerabilities including XSS, SQL injection, CSRF, and auth flaws using OWASP Top 10 methodology.

    495 GitHub stars~2.7k tokensUpdated today
    Auto-check: notes
  • Database Testing

    proffesor-for-testing/agentic-qe

    Database schema validation, data integrity testing, migration testing, transaction isolation, and query performance.

    495 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed

Works with

Categories

Questions about Sherlock Review

What does Sherlock Review do?

Evidence-based investigative code review using deductive reasoning to determine what actually happened versus what was claimed. Sherlock Review is an agent skill from proffesor-for-testing/agentic-qe. Evidence-based investigative code review using deductive reasoning to determine what actually happened versus what was claimed.

When should I use Sherlock Review?

Sherlock Review fits situations like: verifying implementation claims; investigating bugs; validating fixes; conducting root cause analysis.

How do I install Sherlock Review in Claude Code?

Run `npx skills add proffesor-for-testing/agentic-qe --skill sherlock-review -a claude-code`. Or copy the skill folder (assets/skills/sherlock-review in proffesor-for-testing/agentic-qe) into .claude/skills/sherlock-review in your project. Claude Code loads it when a task matches its description.

How do I install Sherlock Review in Codex?

Run `npx skills add proffesor-for-testing/agentic-qe --skill sherlock-review -a codex`. Or copy the skill folder (assets/skills/sherlock-review in proffesor-for-testing/agentic-qe) into .agents/skills/sherlock-review in your project. Codex loads it when a task matches its description.

Can I use Sherlock Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add proffesor-for-testing/agentic-qe --skill sherlock-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sherlock-review, .gemini/skills/sherlock-review, .github/skills/sherlock-review and .opencode/skills/sherlock-review in your project.

What does Sherlock Review need to run?

Going by SKILL.md and its folder, Sherlock Review needs the command-line tools its instructions call (git and npm).

Does Sherlock Review access the network?

SKILL.md contains no URLs. Its commands use git and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Sherlock Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Sherlock Review use?

Sherlock Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sherlock Review use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sherlock Review?

Skills that share tags, products or a category with Sherlock Review: Code Design Rationale Investigator (cursor/plugins, 11k stars), Debugging and Error Recovery (addyosmani/agent-skills, 104k stars), Root Cause Tracing (sandgardenhq/sgai, 137 stars) and Codexqa Rootcause Analyzer (openqa-cn/codexqa, 152 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sherlock Review?

proffesor-for-testing (a GitHub user) maintains it in proffesor-for-testing/agentic-qe, which has 495 GitHub stars. The repository holds 93 skills in this directory. The repository was last updated on October 9, 2026.

Source: proffesor-for-testing/agentic-qe on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.