Agent skill

Benchmark Due Diligence

by daymade in daymade/claude-code-skills

Runs adversarial due-diligence on a benchmark the user envies — a founder, KOL, company, or product whose success looks inflated — splitting marketing bubble from real signal, then mapping the…

MITAuto-check passedResearch & Science

Install Benchmark Due Diligence

skills CLI
$ npx skills add daymade/claude-code-skills --skill benchmark-due-diligence -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install daymade/claude-code-skills benchmark-due-diligence --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/daymade/claude-code-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/daymade-financial/benchmark-due-diligence .claude/skills/benchmark-due-diligence && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-due-diligence
GitHub stars
1.4k
Token cost
~2.3k tokens
SKILL.md length
1,089 words
Files
5 (incl. references)
Skills in repo
103
Repo updated
First seen
Licence
MIT

At a glance

Runs adversarial due-diligence on a benchmark the user envies — a founder, KOL, company, or product whose success looks inflated — splitting marketing bubble from real signal, then mapping the…

  • Works in 2 steps: Inferring relationships between entities… → Treating the commissioner's client as…
  • 尽调/对标/拆解 a competitor
  • SKILL.md covers CRITICAL: run inline, never…, The one rule that protects the…, Phase 0 — nail the foundation… and The four-phase orchestration, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Benchmark Due Diligence is an agent skill from daymade/claude-code-skills. Runs adversarial due-diligence on a benchmark the user envies — a founder, KOL, company, or product whose success looks inflated — splitting marketing bubble from real signal, then mapping the validated playbook onto the user's own resources. Use for 尽调/对标/拆解 a competitor, 抄/偷师 their playbook, or suspecting 水分/泡沫 in claims. Prefer over deep-research when debunking inflated claims, not a neutral briefing.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/attribution_and_resource_mapping.md`, `references/evidence_discipline_traps.md` and `references/evidence_grading_rubric.md`).

It sits in Research & Science, covering Fundraising and pitch decks and Deep research. The repository describes itself as: Professional Claude Code skills marketplace featuring production-ready skills for enhanced development workflows. The licence is MIT.

When your agent uses it

  • 尽调/对标/拆解 a competitor
  • 抄/偷师 their playbook
  • Suspecting 水分/泡沫 in claims

Example prompts

  • “/benchmark-due-diligence”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Inferring relationships between entities from names/domains. "Their content lives at academy.example.com, and they're the founder, so they…
  2. Treating the commissioner's client as the commissioner's asset. If the commissioner does service work for an accelerator/brand, that…

What it can do on your machine

Read from SKILL.md and the folder at commit 872127b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Due Diligence loads about 2.3k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 108 tokens; SKILL.md has 1,089 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from daymade/claude-code-skills at commit 872127b, republished under its MIT licence (© daymade). 1,089 words, ~2,343 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-due-diligence/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
benchmark-due-diligence
description
Runs adversarial due-diligence on a benchmark the user envies — a founder, KOL, company, or product whose success looks inflated — splitting marketing bubble from real signal, then mapping the validated playbook onto the user's own resources. Use for 尽调/对标/拆解 a competitor, 抄/偷师 their playbook, or suspecting 水分/泡沫 in claims. Prefer over deep-research when debunking inflated claims, not a neutral briefing.

Benchmark Due Diligence

Take a benchmark the user envies — a founder, KOL, company, or product whose success looks suspiciously shiny — and produce a teardown that ends in "what this means for ME", not a neutral report. The deliverable answers three questions a balanced briefing never does: How much of this success is real vs marketing bubble? How much is replicable method vs luck/timing? And what, specifically, can the commissioner do with it?

This is the adversarial, decision-oriented cousin of deep-research. Where deep-research builds a trustworthy picture of the world, this skill assumes the picture is inflated until proven otherwise and converts the survivors into the commissioner's own moves.

CRITICAL: run inline, never context: fork

This skill is an orchestrator — it spawns parallel collection + verification agents (via the Workflow tool, or Task agents) and may invoke other skills (deep-research, osint-investigate, qcc). Subagents cannot spawn subagents or call skills. Setting context: fork would silently break the entire fan-out. Do not add a context field. (Same constraint osint-investigate documents — it's a hard runtime rule, not a preference.)

The one rule that protects the commissioner: two injection channels

Everything the agents see flows through exactly two channels. Keeping them separate is the single most important discipline in this skill:

ChannelContentInjected into
FACTSAlready-verified public facts about the benchmark (relationships, who-owns-what, the headline claim flagged ⚠️ to-verify)Every agent — collection, verification, synthesis
COMMISSIONER_CONTEXTThe commissioner's private reality — real resources, client names, strategic intent, what they can actually leverageOnly the final mapping agent (Phase 4)

Why this split is non-negotiable: collection and verification agents take their input and run external WebSearch on it. If the commissioner's client names or strategy leak into those prompts, they get searched on the open web — a privacy breach. The mapping phase genuinely needs "who is the commissioner"; the collection phase must never see it. Encode this in the orchestration (see references/workflow_orchestration_template.md), don't rely on remembering it mid-run.

Phase 0 — nail the foundation by evidence, not appearance (do this BEFORE any agent)

The fastest way to waste a 12-agent fan-out is to build it on a foundation you inferred from appearances. Two failure modes recur and both have burned real runs:

  1. Inferring relationships between entities from names/domains. "Their content lives at academy.example.com, and they're the founder, so they must own that community" — when in reality they were just an invited guest. A shared domain, a similar name, or co-occurrence is an observation, not ownership. Verify with an authoritative source before treating any A↔B relationship as fact.
  2. Treating the commissioner's client as the commissioner's asset. If the commissioner does service work for an accelerator/brand, that accelerator is the client's asset — the commissioner can't leverage its audience or capital. Mapping the benchmark's playbook onto resources the commissioner doesn't actually control produces castles in the air.

So before fanning out, establish by evidence (not vibes):

  • The benchmark's real entity graph — who owns whom, who merely partners/guests. Don't reason from names.
  • The headline-claim attribution — the benchmark's whole narrative usually rests on one trophy stat ("took product X from 0 → 1M users"). Are they the founder, or the departed growth lead? This is the #1 to-verify target; write it into FACTS with a ⚠️.
  • What the commissioner truly controls — separate owned assets from client/partner assets.

Write the results into FACTS (public half) and COMMISSIONER_CONTEXT (private half). A shaky foundation makes every downstream agent confidently wrong.

Show full SKILL.md (527 more words)Show less

The four-phase orchestration

Use the Workflow tool (preferred — deterministic fan-out, see the ready-to-fill template in references/workflow_orchestration_template.md) or Task agents. Scale agent count to how thorough the user wants (a few dimensions for a quick read, 6+ with multi-vote verification for a deep audit).

Phase 1 + 2 — collect → verify, per dimension, as a pipeline (each dimension verifies the moment its collection finishes; no global barrier):

  • Collection agent — objective stance. Every finding carries a source URL and a source_kind (对象自述/营销 vs 第三方独立信源 vs 混合). Anything not found goes in gaps — never filled by guessing.
  • Verification agent — adversarial, default-skeptical stance. Grade every claim L1–L4 and rule 坐实 / 大体可信 / 存疑 / 证伪-水分. The job is to actively hunt falsifying evidence, especially for the headline claims (the trophy stat, "#1 ranking", funding amount, user counts). bubble_summary names the biggest water in that dimension.

Grading rubric, source_kind, verdicts, and both JSON schemas → references/evidence_grading_rubric.md.

Typical dimensions (tailor to the benchmark type — person / company / product):

  1. Subject background + headline-claim attribution (the #1 bubble target)
  2. Corporate base — entity, founding, funding/valuation
  3. Core product/business real metrics — user counts, revenue, rankings, awards, cross-verified against third parties
  4. Playbook teardown — platform matrix, persona, content types, how they borrow other people's audiences, how personal IP funnels to the product
  5. Comparison sample — a structurally-similar peer or parallel path
  6. Sector + how this class of playbook usually wins and usually fails

Phase 3 — synthesis: due-diligence conclusion (single agent, consumes all verdicts):

  1. Real relationship map (correcting the common misreadings from Phase 0)
  2. Bubble-busting table — claim | evidence level | verdict | one-line basis, sorted by most-water-first
  3. Playbook teardown — concrete, copyable actions
  4. Attribution breakdown (the core) — what share of the success is product vs market-timing vs personal-IP-marketing vs operations? Give % ranges with reasons, and explicitly split replicable method from luck / timing / non-transferable endowment.

Phase 4 — synthesis: what this means for the commissioner (single agent; consumes Phase 3 + COMMISSIONER_CONTEXT):

  1. Resource-mapping table — benchmark's playbook elements × the commissioner's real resources; tag each cell ✅ borrow-able / ⚠️ not-replicable (luck/timing) / 🔄 already-doing / 🚫 bubble-don't-copy, one line each
  2. Landing points — exactly how the commissioner uses it (their to-B service / their own IP / their tooling)
  3. Action list + open questions (what's still unconfirmed)

Attribution weighting and the four-tag mapping framework → references/attribution_and_resource_mapping.md.

Don't rebuild what already exists

This skill's edge is the adversarial bubble-busting + attribution + commissioner-mapping layers. The plumbing underneath is not novel — reuse it:

  • Fan-out collection / source governance — borrow the lead-agent + subagent pattern from deep-research. (What's unique here is the skeptical verification stance and the L1–L4 bubble grading, not the parallelism.)
  • Person-subject identity / footprint checks — invoke osint-investigate (ACH hypothesis matrix, Bellingcat-style pivots) rather than re-deriving identity attribution.
  • Mainland-China corporate registration / funding — invoke the qcc family of skills for 工商 data.
  • Social-platform playbook data — the agent-reach CLI covers B站/小红书/抖音/YouTube/X.

Read before you run

  • references/evidence_discipline_traps.md — the recurring traps (inferring relationships from appearances, headline-claim attribution, client-vs-asset, foundation-before-fan-out, grade-don't-binary, privacy leak) with real teardown war-stories. Read this first; it's where runs actually break.
  • references/evidence_grading_rubric.md — L1–L4, source_kind, verdicts, collection/verification schemas.
  • references/attribution_and_resource_mapping.md — attribution weighting + four-tag mapping + landing-point framework.
  • references/workflow_orchestration_template.md — a ready-to-fill Workflow script with the FACTS / COMMISSIONER_CONTEXT injection split already wired in.

Next Step

After the due-diligence conclusion is ready, suggest the natural follow-on (opt-in, never auto-run):

Due-diligence teardown is done.

Options:
A) Render it as a shareable PDF report — pdf-creator (Recommended if this goes to a partner/team)
B) One dimension needs deeper neutral background — deep-research on that sub-topic
C) No thanks — the markdown teardown is enough

© daymade, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in daymade-financial/benchmark-due-diligence of daymade/claude-code-skills.

  • SKILL.md
  • references/attribution_and_resource_mapping.md
  • references/evidence_discipline_traps.md
  • references/evidence_grading_rubric.md
  • references/workflow_orchestration_template.md

Open the folder on GitHubat commit 872127b

Compare with similar skills

Benchmark Due Diligence next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Due Diligence compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Due Diligence this skilldaymade/claude-code-skills1.4k—~2.3kAutomated safety check: PassMIT
Deep Researchsanjay3290/ai-skills4319 repos~683Automated safety check: NotesApache-2.0
Consulting Analysisbytedance/deer-flow84k4 repos~8.4kAutomated safety check: PassMIT
Interceptor ResearchHacker-Valley-Media/Interceptor519—~3.8kAutomated safety check: PassCustom licence
Researcherunderstudy-ai/understudy462—~1.2kAutomated safety check: PassMIT
Deep Scrapedavidondrej/skills4.1k—~2.3kAutomated safety check: PassMIT

Similar skills

  • Deep Research

    sanjay3290/ai-skills

    Execute autonomous multi-step research using Google Gemini Deep Research Agent.

    431 GitHub starsUsed in 9 repos~683 tokens
    Research & ScienceAuto-check: notes
  • Consulting Analysis

    bytedance/deer-flow

    A skill your agent uses when the user requests to generate, create, or write professional research reports including but not limited to market analysis, consumer insights, brand analysis, financial…

    84k GitHub starsUsed in 4 repos~8.4k tokens
    Marketing & SEOAuto-check passed
  • Interceptor Research

    Hacker-Valley-Media/Interceptor

    Deep web-research methodology for the interceptor browser surface — investigate a topic the way researchers, intelligence analysts, investigative journalists, private investigators, and OSINT…

    519 GitHub stars~3.8k tokensUpdated 7 days ago
    Research & ScienceAuto-check passed
  • Researcher

    understudy-ai/understudy

    Research current topics with multiple sources and produce a structured brief, comparison, recommendation, or fact-check.

    462 GitHub stars~1.2k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Deep Scrape

    davidondrej/skills

    Build sourced JSON dossiers on people, companies, or topics with DeepAPI.

    4.1k GitHub stars~2.3k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Vc Industry Research

    zebbern/claude-code-guide

    Generate professional primary market / venture capital industry research reports, including sector deep-dives, investment memos, and market analysis.

    4.7k GitHub stars~1.6k tokensUpdated today
    Documents & OfficeAuto-check passed

More from daymade/claude-code-skills

All 103 skills in this repo
  • Video Comparer

    daymade/claude-code-skills

    This skill should be used when comparing two videos to analyze compression results or quality differences.

    1.4k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check: notes
  • CLI Demo Generator

    daymade/claude-code-skills

    Generates professional animated CLI demos as GIFs using VHS terminal recordings.

    1.4k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Doc To Markdown

    daymade/claude-code-skills

    Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.

    1.4k GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Interaction Design Board

    daymade/claude-code-skills

    Generates several distinct, clickable HTML interaction prototypes for one product surface into a Design Board and collects selection/remix feedback before implementation.

    1.4k GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Auto Repo Setup

    daymade/claude-code-skills

    Diagnoses and repairs repository setup and guarded Git workflows for Claude Code or Codex — environment repair, startup sync, hook auditing, collaborator handoff.

    1.4k GitHub stars~2.8k tokensUpdated today
    Auto-check: notes
  • Bigdata Skill

    daymade/claude-code-skills

    Pulls Bigdata.com (RavenPack) financial and news data via the official bigdata-client SDK and /v1/ REST endpoints — structured financials, prices, analyst estimates, entity-sentiment series…

    1.4k GitHub stars~3.7k tokensUpdated today
    Auto-check passed

Questions about Benchmark Due Diligence

What does Benchmark Due Diligence do?

Runs adversarial due-diligence on a benchmark the user envies — a founder, KOL, company, or product whose success looks inflated — splitting marketing bubble from real signal, then mapping the…. Benchmark Due Diligence is an agent skill from daymade/claude-code-skills. Runs adversarial due-diligence on a benchmark the user envies — a founder, KOL, company, or product whose success looks inflated — splitting marketing bubble from real signal, then mapping the validated playbook onto the user's own resources.

When should I use Benchmark Due Diligence?

Benchmark Due Diligence fits situations like: 尽调/对标/拆解 a competitor; 抄/偷师 their playbook; suspecting 水分/泡沫 in claims.

How do I install Benchmark Due Diligence in Claude Code?

Run `npx skills add daymade/claude-code-skills --skill benchmark-due-diligence -a claude-code`. Or copy the skill folder (daymade-financial/benchmark-due-diligence in daymade/claude-code-skills) into .claude/skills/benchmark-due-diligence in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Due Diligence in Codex?

Run `npx skills add daymade/claude-code-skills --skill benchmark-due-diligence -a codex`. Or copy the skill folder (daymade-financial/benchmark-due-diligence in daymade/claude-code-skills) into .agents/skills/benchmark-due-diligence in your project. Codex loads it when a task matches its description.

Can I use Benchmark Due Diligence in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add daymade/claude-code-skills --skill benchmark-due-diligence -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-due-diligence, .gemini/skills/benchmark-due-diligence, .github/skills/benchmark-due-diligence and .opencode/skills/benchmark-due-diligence in your project.

What does Benchmark Due Diligence need to run?

SKILL.md names no scripts, command-line tools or credentials: Benchmark Due Diligence is instructions for the agent only.

Does Benchmark Due Diligence access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark Due Diligence safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Due Diligence use?

Benchmark Due Diligence is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Due Diligence use?

About 2.3k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.2k tokens, read only when the agent opens those files.

What are the alternatives to Benchmark Due Diligence?

Skills that share tags, products or a category with Benchmark Due Diligence: Deep Research (sanjay3290/ai-skills, 431 stars), Consulting Analysis (bytedance/deer-flow, 84k stars), Interceptor Research (Hacker-Valley-Media/Interceptor, 519 stars) and Researcher (understudy-ai/understudy, 462 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Due Diligence?

daymade (a GitHub user) maintains it in daymade/claude-code-skills, which has 1,447 GitHub stars. The repository holds 103 skills in this directory. The repository was last updated on October 9, 2026.

Source: daymade/claude-code-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.