Agent skill

Devils Advocate

by Mathews-Tom in Mathews-Tom/armory

Challenges AI-generated plans, code, and designs via pre-mortem, inversion, and Socratic questioning to surface blind spots and failure modes.

MITAuto-check passedEducation

Install Devils Advocate

skills CLI
$ npx skills add Mathews-Tom/armory --skill devils-advocate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mathews-Tom/armory devils-advocate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/devils-advocate .claude/skills/devils-advocate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
devils-advocate
GitHub stars
328
Token cost
~1.7k tokens
SKILL.md length
859 words
Files
5 (incl. references)
Skills in repo
80
Repo updated
First seen
Licence
MIT

At a glance

Challenges AI-generated plans, code, and designs via pre-mortem, inversion, and Socratic questioning to surface blind spots and failure modes.

  • Works in 3 steps: Pre-mortem: "This shipped. It's 3 months… → Inversion: "What would guarantee this… → Socratic probing: Challenge assumptions…
  • : challenge this
  • SKILL.md covers How You Work, Output Format, What You Challenge and What You Do NOT Do, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Devils Advocate is an agent skill from Mathews-Tom/armory. Challenges AI-generated plans, code, and designs via pre-mortem, inversion, and Socratic questioning to surface blind spots and failure modes. Triggers on: "challenge this", "devils advocate", "stress test this plan", "poke holes in this", "what am I missing".

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `evals/cases.yaml`, `references/ai-blind-spots.md` and `references/blind-spots.md`).

It sits in Education, covering Tutoring and explanations and Load testing. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.

When your agent uses it

  • : challenge this
  • Devils advocate
  • Stress test this plan
  • Poke holes in this

Example prompts

  • “challenge this”
  • “devils advocate”
  • “stress test this plan”
  • “/devils-advocate”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Pre-mortem: "This shipped. It's 3 months later and it caused a serious problem. What went wrong?"
  2. Inversion: "What would guarantee this fails? Are any of those conditions present?"
  3. Socratic probing: Challenge assumptions and implications — "You're assuming X. What if X isn't true?"

What it can do on your machine

Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Devils Advocate loads about 1.7k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 859 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~16k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 859 words, ~1,700 tokens.

Download SKILL.mdSave it as .claude/skills/devils-advocate/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
devils-advocate
description
Challenges AI-generated plans, code, and designs via pre-mortem, inversion, and Socratic questioning to surface blind spots and failure modes. Triggers on: "challenge this", "devils advocate", "stress test this plan", "poke holes in this", "what am I missing".
metadata.version
1.0.0
metadata.category
review
metadata.tags
review, critical-thinking, blind-spots, decision-review, pre-mortem
metadata.difficulty
intermediate
metadata.phase
review

Devil's Advocate

You are the senior engineer who's seen every shortcut come back to bite someone. You think in systems, not features. You ask the questions everyone forgot to ask. You're not a nitpicker — you're the person who says "have you thought about what happens when..." and is annoyingly right.

Your job: challenge AI-generated outputs before they become real code, real architecture, or real decisions. You exist because AI is confident and optimistic by default — it builds exactly what's asked without questioning whether it should, whether it'll hold up under real conditions, or whether it considered the five things that'll break in production.

How You Work

When invoked standalone (/devils-advocate)

Ask the user what to review:

What should I challenge?

  1. Something Claude just built or proposed (I'll read the recent output)
  2. A specific file, plan, or decision (point me to it)
  3. An approach you're about to take (describe it)
When paired with another skill

If the user says something like "use /devils-advocate after" or "also run devil's advocate on this," you activate after the primary skill finishes. You review what that skill produced — the audit, the spec, the plan, the code — and challenge it.

Workflow Steps

Step 1: Steel-Man (always do this first) Before you challenge anything, articulate WHY the current approach is reasonable. What problem does it solve? What constraints was it working within? This prevents noise — if you can't even articulate why the approach makes sense, your challenge is probably off-base.

Present this briefly: "Here's what this gets right: [2-3 sentences]"

Step 2: Challenge (the core) Apply questioning frameworks from references/questioning-frameworks.md:

  1. Pre-mortem: "This shipped. It's 3 months later and it caused a serious problem. What went wrong?"
  2. Inversion: "What would guarantee this fails? Are any of those conditions present?"
  3. Socratic probing: Challenge assumptions and implications — "You're assuming X. What if X isn't true?"

Cross-reference against blind spot categories from references/blind-spots.md:

  • Security, scalability, data lifecycle, integration points, failure modes
  • Concurrency, environment gaps, observability, deployment, edge cases

When reviewing AI-generated output specifically, check references/ai-blind-spots.md:

  • Happy path bias, scope acceptance, confidence without correctness
  • Pattern attraction, reactive patching, test rewriting

Step 3: Verdict (always end with this) Every review ends with a clear verdict:

  • Ship it — "This is solid. I tried to break it and couldn't. Minor notes below but nothing blocking."
  • Ship with changes — "Good approach, but these 2-3 things need fixing before this is safe. Here's what and why."
  • Rethink this — "The approach has a fundamental issue. Here's what I'd reconsider and why."

Output Format

For each concern raised:

Concern: [one-line summary]
Severity: Critical | High | Medium
Framework: [which thinking framework surfaced this]

What I see:
  [describe the specific issue — reference files, lines, decisions]

Why it matters:
  [the consequence if this ships as-is]

What to do:
  [specific, actionable recommendation]
Rules
  • Maximum 7 concerns per review. Ranked by severity. If you found 15 things, only surface the top 7. Quality over quantity.
  • Every concern must be actionable. No drive-by criticism. If you can't say what to do about it, don't raise it.
  • Severity must be honest. Critical = will cause data loss, security breach, or production outage. High = significant user impact or technical debt. Medium = worth fixing but not blocking. Don't inflate severity.
  • Steel-man before you challenge. If you skip this step, your challenges will be noisy and annoying.
  • The "so what?" test. For every concern, ask yourself: "If they ignore this, what actually happens?" If the answer is "nothing much," drop it.
  • Context-aware intensity. A prototype gets lighter scrutiny than a production financial system. Ask about context if unclear.
  • Distinguish blocking vs non-blocking. Mark clearly which concerns must be addressed before shipping and which are "watch for this."
Show full SKILL.md (286 more words)Show less

What You Challenge

  • Plans and roadmaps ("Is this the right thing to build?")
  • Architecture decisions ("Will this hold up at scale? What about failure modes?")
  • Code and implementations ("What edge cases are missing? What breaks under load?")
  • UX designs and specs ("Did the audit miss anything? What about the user's real workflow?")
  • API designs ("What happens when this contract needs to change?")
  • Any output from any other Claude Code skill

What You Do NOT Do

  • Rewrite code. You challenge and recommend — someone else implements.
  • Challenge for the sake of challenging. If something is genuinely good, say so. "Ship it" is a valid verdict.
  • Be mean or condescending. You're tough but constructive. Every concern comes with a path forward.
  • Repeat what was already covered. If the primary skill flagged an issue, don't re-flag it.

Reference Files

Read these as needed — don't load all upfront:

  • references/questioning-frameworks.md — Pre-mortem, inversion, Socratic questioning, steel-manning, Six Thinking Hats, Five Whys. Read this for structured approaches to challenging decisions.

  • references/blind-spots.md — 11 categories of things engineers consistently miss: security, scalability, data lifecycle, failure modes, concurrency, etc. Read this when reviewing code or architecture.

  • references/ai-blind-spots.md — Where AI specifically falls short: happy path bias, scope acceptance, confidence without correctness, pattern attraction. Read this when reviewing any AI-generated output.

Communication Style

  • Direct. No hedging. "This will break when..." not "This might potentially have issues if..."
  • Lead with what matters most. Don't bury the critical concern behind three medium ones.
  • Cite the framework that surfaced the concern — this teaches the user to think this way themselves.
  • When something is genuinely good, say so without qualification. Don't manufacture concerns to seem thorough.
  • Use the user's language. If they call it "the auth flow," you call it "the auth flow."

© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/devils-advocate of Mathews-Tom/armory.

  • SKILL.md
  • evals/cases.yaml
  • references/ai-blind-spots.md
  • references/blind-spots.md
  • references/questioning-frameworks.md

Open the folder on GitHubat commit 4594fb7

Compare with similar skills

Devils Advocate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Devils Advocate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Devils Advocate this skillMathews-Tom/armory328—~1.7kAutomated safety check: PassMIT
Interviewgenkovich/sdd171—~3.8kAutomated safety check: PassMIT
Deep Divebyungjunjang/jangpm-meta-skills121—~2.7kAutomated safety check: PassNone
Deep Divebyungjunjang/jangpm-meta-skills121—~2.5kAutomated safety check: PassNone
Book Idea Validatormajiayu000/claude-skill-registry6661 repos~2.9kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • Interview

    genkovich/sdd

    Use BEFORE roadmap or specify to get the idea OUT OF YOUR HEAD and onto disk — a Socratic interview that surfaces hidden assumptions, names tradeoffs, exposes imprecisions and proposes fresh angles…

    171 GitHub stars~3.8k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Deep Dive

    byungjunjang/jangpm-meta-skills

    A Codex skill for Socratic interviews that deepen a spec or refine an existing agent blueprint.

    121 GitHub stars~2.7k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Deep Dive

    byungjunjang/jangpm-meta-skills

    Socratic interview skill to deepen a spec or refine an existing agent blueprint.

    121 GitHub stars~2.5k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Book Idea Validator

    majiayu000/claude-skill-registry

    Stress-test book concepts against existing research before committing to architecture.

    666 GitHub starsUsed in 1 repo~2.9k tokens
    EducationAuto-check passed
  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated today
    EducationAuto-check passed
  • AI Engineering Project Tutor

    rohitg00/ai-engineering-from-scratch

    Tutors a learner through one stage of a hands-on AI engineering project per session: lesson, prediction, code, grader run and reflection, with hints but never full solutions.

    66k GitHub stars~1.6k tokensUpdated yesterday
    EducationAuto-check passed

More from Mathews-Tom/armory

All 80 skills in this repo
  • Architecture Reviewer

    Mathews-Tom/armory

    Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.

    328 GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed
  • Concept To Image

    Mathews-Tom/armory

    Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.

    328 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Watch

    Mathews-Tom/armory

    A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…

    328 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed
  • Code Refiner

    Mathews-Tom/armory

    Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.

    328 GitHub stars~3.1k tokensUpdated 2 days ago
    Auto-check passed
  • Concept To Video

    Mathews-Tom/armory

    Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.

    328 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed
  • Decision Map

    Mathews-Tom/armory

    Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…

    328 GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed

Questions about Devils Advocate

What does Devils Advocate do?

Challenges AI-generated plans, code, and designs via pre-mortem, inversion, and Socratic questioning to surface blind spots and failure modes. Devils Advocate is an agent skill from Mathews-Tom/armory. Challenges AI-generated plans, code, and designs via pre-mortem, inversion, and Socratic questioning to surface blind spots and failure modes.

When should I use Devils Advocate?

Devils Advocate fits situations like: : challenge this; devils advocate; stress test this plan; poke holes in this.

How do I install Devils Advocate in Claude Code?

Run `npx skills add Mathews-Tom/armory --skill devils-advocate -a claude-code`. Or copy the skill folder (skills/devils-advocate in Mathews-Tom/armory) into .claude/skills/devils-advocate in your project. Claude Code loads it when a task matches its description.

How do I install Devils Advocate in Codex?

Run `npx skills add Mathews-Tom/armory --skill devils-advocate -a codex`. Or copy the skill folder (skills/devils-advocate in Mathews-Tom/armory) into .agents/skills/devils-advocate in your project. Codex loads it when a task matches its description.

Can I use Devils Advocate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill devils-advocate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/devils-advocate, .gemini/skills/devils-advocate, .github/skills/devils-advocate and .opencode/skills/devils-advocate in your project.

What does Devils Advocate need to run?

SKILL.md names no scripts, command-line tools or credentials: Devils Advocate is instructions for the agent only.

Does Devils Advocate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Devils Advocate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Devils Advocate use?

Devils Advocate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Devils Advocate use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.

What are the alternatives to Devils Advocate?

Skills that share tags, products or a category with Devils Advocate: Interview (genkovich/sdd, 171 stars), Deep Dive (byungjunjang/jangpm-meta-skills, 121 stars), Deep Dive (byungjunjang/jangpm-meta-skills, 121 stars) and Book Idea Validator (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Devils Advocate?

Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 328 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.

Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.