Agent skill

Devils Advocate

by flonat in flonat/flonat-research

Adversarially challenge research assumptions, mechanisms, and arguments in writing.

MITAuto-check passedAgent Workflows

Install Devils Advocate

skills CLI
$ npx skills add flonat/flonat-research --skill devils-advocate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install flonat/flonat-research devils-advocate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/devils-advocate .claude/skills/devils-advocate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
devils-advocate
GitHub stars
145
Token cost
~1.8k tokens
SKILL.md length
738 words
Files
2 (incl. references)
Skills in repo
83
Repo updated
First seen
Licence
MIT

At a glance

Adversarially challenge research assumptions, mechanisms, and arguments in writing.

  • Works in 4 steps: Understand the claim — Read the… → Generate competing hypotheses — If… → Run the debate — Use the multi-turn… → …
  • Stress-testing a claim
  • SKILL.md covers Purpose, When to Use, When NOT to Use and Workflow, plus 6 more sections
  • Calls bash

What it does

Devils Advocate is an agent skill from flonat/flonat-research. Adversarially challenge research assumptions, mechanisms, and arguments in writing. Use when stress-testing a claim or design before committing to it. Not for an interactive oral drill or a full referee report; use $grill-me or a review agent.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/competing-hypotheses.md`).

It sits in Agent Workflows, covering Load testing and Requirements gathering. The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.

When your agent uses it

  • Stress-testing a claim
  • Design before committing to it

Example prompts

  • “/devils-advocate”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Understand the claim — Read the paper/argument being evaluated
  2. Generate competing hypotheses — If evaluating a research question or design, load references/competing-hypotheses.md and generate 3-5…
  3. Run the debate — Use the multi-turn debate protocol below (default) or single-shot mode for quick checks
  4. Deliver the verdict — Synthesize surviving critiques with severity ratings

What it can do on your machine

Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Devils Advocate loads about 1.8k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 738 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 738 words, ~1,758 tokens.

Download SKILL.mdSave it as .claude/skills/devils-advocate/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
devils-advocate
description
Adversarially challenge research assumptions, mechanisms, and arguments in writing. Use when stress-testing a claim or design before committing to it. Not for an interactive oral drill or a full referee report; use $grill-me or a review agent.
argument-hint
[paper-or-argument-description]

Devil's Advocate Skill

Challenge research assumptions and identify weaknesses in your arguments.

Purpose

Based on Scott Cunningham's Part 3: "Creating Devil's Advocate Agents for Tough Problems" - addressing the "LLM thing of over-confidence in diagnosing a problem."

For formal code audits with replication scripts and referee reports, use the Referee 2 agent instead (.claude/agents/referee2-reviewer.md). This skill is for quick adversarial feedback on arguments, not systematic audits.

When to Use

  • Before submitting a paper
  • When stuck on a research problem
  • When you want to stress-test an argument
  • During paper revision planning

When NOT to Use

  • Code audits — use the Referee 2 agent instead
  • Replication verification — use the Referee 2 agent instead
  • Quick proofreading — just ask for a read-through
  • When you want validation — this skill is designed to challenge, not affirm

Workflow

  1. Understand the claim — Read the paper/argument being evaluated
  2. Generate competing hypotheses — If evaluating a research question or design, load references/competing-hypotheses.md and generate 3-5 rival explanations before critiquing
  3. Run the debate — Use the multi-turn debate protocol below (default) or single-shot mode for quick checks
  4. Deliver the verdict — Synthesize surviving critiques with severity ratings

Multi-Turn Debate Protocol (Default)

Inspired by the simulated scientific debates in Google's AI Co-Scientist. A one-shot critique is easy for an LLM to produce but often superficial. Multi-turn debates force each critique to survive a defense, filtering out weak objections and sharpening the strong ones.

Round 1: Adversarial Critic

Adopt the persona of a hostile but competent reviewer. Challenge on:

  1. Theoretical foundations — Are the assumptions justified?
  2. Methodology — Limitations? Alternative approaches?
  3. Data — Selection bias? Measurement issues? External validity?
  4. Causal claims — Alternative explanations? Confounders?
  5. Contribution — Novel enough? Does it matter?

Produce numbered critiques (aim for 5-8), each with a concrete statement of the problem.

Round 2: Defense

Switch persona to the paper's author. For each numbered critique, provide the strongest possible defense:

  • Cite evidence from the paper that addresses the concern
  • Explain design choices that mitigate the issue
  • Acknowledge limitations honestly where the defense is weak
  • Propose concrete fixes where the critique has merit
Round 3: Adjudication

Switch to an impartial senior reviewer. For each critique-defense pair, rule:

  • Critique stands — the defense is insufficient; this is a real weakness
  • Critique partially addressed — defense has merit but issue remains
  • Critique resolved — the defense adequately addresses the concern
Final Synthesis

Produce a structured report with only the surviving critiques (stands + partially addressed), ranked by severity:

markdown
## Devil's Advocate Report

### Critical (must fix before submission)
1. [Critique] — [Why the defense failed] — [Suggested fix]

### Major (reviewers will likely raise)
2. [Critique] — [What remains after defense] — [Suggested fix]

### Minor (worth acknowledging)
3. [Critique] — [Residual concern] — [How to preempt]

### Dismissed
- [Critiques that were resolved in Round 2, listed briefly for transparency]
Show full SKILL.md (365 more words)Show less

Output Path & Stamping

When run on a paper in a research project (a paper-*/ directory exists), persist the report and stamp it into the review log — the same wiring as proofread and bib-validate, so review-recap renders it as a first-class review (not a manual slot).

  1. Write the report to reviews/<scope>/devils-advocate/<YYYY-MM-DD-HHMM>.md, where <scope> is the in-scope paper slug (e.g. paper-prima) or _project for a project-level argument. Create the dir first (mkdir -p reviews/<scope>/devils-advocate/). Never overwrite — each run is timestamped to the minute. Per rules/review-artefact-routing.md, never write to the project root.
  2. Stamp reviews/INDEX.md:
    bash
    bash <skills-root>/_shared/review-state-log.sh \
      --check devils-advocate \
      --paper "<scope>" \
      --verdict "<PASS|ISSUES FOUND>" \
      --open-issues "<surviving-critiques>/<total-critiques-raised>" \
      --report "reviews/<scope>/devils-advocate/<YYYY-MM-DD-HHMM>.md" \
      --notes "<one line: e.g. '2 Critical, 1 Major survive; identification strategy weakest'>" \
      [--trigger "review-cluster|pre-submission-report"]
    • Verdict: PASS if no critiques survive adjudication (all dismissed); ISSUES FOUND otherwise.
    • Open issues: surviving critiques (Critical+Major+Minor) over total raised in Round 1.
    • Trigger: pass an orchestrator name only if invoked via review-cluster or pre-submission-report; otherwise omit.

Skip stamping only for non-paper use (challenging a bare argument with no project context) or single-shot mode on a paragraph — then the report stays inline. Schema: the installed shared resource shared/review-state-schema.md.

Single-Shot Mode

For quick checks (e.g., "just poke holes in this argument"), skip the multi-turn protocol and produce a direct critique. Use when the user says "quick", "just challenge this", or the input is a paragraph rather than a full paper.

Example Use

"Play devil's advocate on my research paper about preference drift - specifically challenge my identification strategy and the assumptions about utility functions."


Council Mode (Optional)

For the highest-stakes arguments, run the debate across multiple LLM providers — different models have genuinely different reasoning patterns, producing adversarial tension a single model cannot replicate. Each model plays Adversarial Critic, cross-reviews the others, and a chairman ranks surviving critiques by cross-model agreement. Trigger: "council devil's advocate" / "thorough challenge". Full orchestration + invocation: ../shared/council-protocol.md.

Value: High — the multi-turn debate becomes genuinely adversarial when different models play different roles. A critique that survives cross-model scrutiny is almost certainly a real weakness.

Cross-References

SkillWhen to use instead/alongside
interview-meTo develop the idea further through structured interview
multi-perspectiveFor multi-perspective analysis with disciplinary diversity
proofreadFor language/formatting review rather than argument critique

© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/devils-advocate of flonat/flonat-research.

  • SKILL.md
  • references/competing-hypotheses.md

Open the folder on GitHubat commit da27600

Compare with similar skills

Devils Advocate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Devils Advocate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Devils Advocate this skillflonat/flonat-research145—~1.8kAutomated safety check: PassMIT
Grillingbestofjs/bestofjs3.1k30 repos~464Automated safety check: PassMIT
Grill Menateherkai/AIS-OS1.6k—~1.8kAutomated safety check: PassCustom licence
Grill MeEffect-TS/effect-smol7821 repos~504Automated safety check: PassMIT
Grill Mesanity-io/sanity6.4k40 repos~151Automated safety check: PassMIT
Deep Divebyungjunjang/jangpm-meta-skills121—~2.5kAutomated safety check: PassNone

Similar skills

  • Grilling

    bestofjs/bestofjs

    Grill the user relentlessly about a plan, decision, or idea.

    3.1k GitHub starsUsed in 30 repos~464 tokens
    Agent WorkflowsAuto-check passed
  • Grill Me

    nateherkai/AIS-OS

    Interview the user relentlessly about a plan, design, or topic, checkpointing every answer to a brainstorm file so nothing is lost.

    1.6k GitHub stars~1.8k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Grill Me

    Effect-TS/effect-smol

    Interview the user about a plan or design until reaching shared understanding, resolving each branch of the decision tree.

    782 GitHub starsUsed in 1 repo~504 tokens
    Agent WorkflowsAuto-check passed
  • Grill Me

    sanity-io/sanity

    Official

    Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree.

    6.4k GitHub starsUsed in 40 repos~151 tokens
    Agent WorkflowsAuto-check passed
  • Deep Dive

    byungjunjang/jangpm-meta-skills

    Socratic interview skill to deepen a spec or refine an existing agent blueprint.

    121 GitHub stars~2.5k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Grilling

    opencrvs/opencrvs-core

    Grill the user relentlessly about a plan or design. An agent skill from opencrvs/opencrvs-core.

    120 GitHub starsUsed in 2 repos~205 tokens
    Agent WorkflowsAuto-check passed

More from flonat/flonat-research

All 83 skills in this repo
  • Latex Posters

    flonat/flonat-research

    Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.

    145 GitHub stars~1.5k tokensUpdated 8 days ago
    Auto-check: notes
  • Skill Creator

    flonat/flonat-research

    Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.

    145 GitHub stars~4.4k tokensUpdated 8 days ago
    Auto-check passed
  • DOCX

    flonat/flonat-research

    Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.

    145 GitHub stars~1.2k tokensUpdated 8 days ago
    Auto-check passed
  • PDF

    flonat/flonat-research

    Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.

    145 GitHub stars~488 tokensUpdated 8 days ago
    Auto-check passed
  • Init Project Orchestration

    flonat/flonat-research

    Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.

    145 GitHub stars~1.6k tokensUpdated 8 days ago
    Auto-check passed
  • Pre Commit Audit

    flonat/flonat-research

    Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

    145 GitHub stars~2.8k tokensUpdated 8 days ago
    Auto-check: notes

Categories

Questions about Devils Advocate

What does Devils Advocate do?

Adversarially challenge research assumptions, mechanisms, and arguments in writing. Devils Advocate is an agent skill from flonat/flonat-research. Adversarially challenge research assumptions, mechanisms, and arguments in writing.

When should I use Devils Advocate?

Devils Advocate fits situations like: stress-testing a claim; design before committing to it.

How do I install Devils Advocate in Claude Code?

Run `npx skills add flonat/flonat-research --skill devils-advocate -a claude-code`. Or copy the skill folder (skills/devils-advocate in flonat/flonat-research) into .claude/skills/devils-advocate in your project. Claude Code loads it when a task matches its description.

How do I install Devils Advocate in Codex?

Run `npx skills add flonat/flonat-research --skill devils-advocate -a codex`. Or copy the skill folder (skills/devils-advocate in flonat/flonat-research) into .agents/skills/devils-advocate in your project. Codex loads it when a task matches its description.

Can I use Devils Advocate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill devils-advocate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/devils-advocate, .gemini/skills/devils-advocate, .github/skills/devils-advocate and .opencode/skills/devils-advocate in your project.

What does Devils Advocate need to run?

Going by SKILL.md and its folder, Devils Advocate needs the command-line tools its instructions call (bash).

Does Devils Advocate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Devils Advocate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Devils Advocate use?

Devils Advocate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Devils Advocate use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Devils Advocate?

Skills that share tags, products or a category with Devils Advocate: Grilling (bestofjs/bestofjs, 3.1k stars), Grill Me (nateherkai/AIS-OS, 1.6k stars), Grill Me (Effect-TS/effect-smol, 782 stars) and Grill Me (sanity-io/sanity, 6.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Devils Advocate?

flonat (a GitHub user) maintains it in flonat/flonat-research, which has 145 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.

Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.