Agent skill

Evidence Grading

by mohitagw15856 in mohitagw15856/pm-claude-skills

Grade the evidence behind a claim before betting on it — the hierarchy for business evidence (experiments usage data surveys interviews anecdotes opinion), the fit-for-decision test, and the…

MITAuto-check passed

Install Evidence Grading

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill evidence-grading -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills evidence-grading --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evidence-grading .claude/skills/evidence-grading && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evidence-grading
GitHub stars
1.4k
Token cost
~1.5k tokens
SKILL.md length
703 words
Files
1
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Grade the evidence behind a claim before betting on it — the hierarchy for business evidence (experiments usage data surveys interviews anecdotes opinion), the fit-for-decision test, and the…

  • Works in 5 steps: The hierarchy, applied without… → Direction and independence both count:… → Anecdotes prove existence, never… → …
  • Asked how strong is our evidence for this
  • SKILL.md covers What This Skill Produces, Required Inputs, Framework: The Grading Rules and Output Format, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Evidence Grading is an agent skill from mohitagw15856/pm-claude-skills. Grade the evidence behind a claim before betting on it — the hierarchy for business evidence (experiments usage data surveys interviews anecdotes opinion), the fit-for-decision test, and the mixed-evidence verdicts that real questions produce. Use when asked how strong is our evidence for this, grade what we know before the decision, is this enough to bet on, or we have three anecdotes and a survey — now what. Produces the evidence inventory with grades, the sufficiency verdict against the decision's stakes, and…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked how strong is our evidence for this
  • Grade what we know before the decision
  • Is this enough to bet on
  • We have three anecdotes and a survey — now what

Example prompts

  • “/evidence-grading”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. The hierarchy, applied without sentiment: experiments/A-B tests (causal) > usage/behavioral data (what people do, correlational) > surveys…
  2. Direction and independence both count: the inventory includes contradicting evidence at its own grade, and echo-checks the supporting pile…
  3. Anecdotes prove existence, never prevalence: "a customer asked for X" establishes that the need exists somewhere — it cannot establish how…
  4. Sufficiency is relative to stakes: the verdict tests weight against the decision — reversible-and-cheap decisions legitimately run on…
  5. The upgrade path is the constructive ending: when insufficient, name the cheapest test that would change the verdict — the holdout, the…

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evidence Grading loads about 1.5k tokens when it runs. Until then it costs about 143 tokens; SKILL.md has 703 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~143
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 703 words, ~1,458 tokens.

Download SKILL.mdSave it as .claude/skills/evidence-grading/SKILL.md (or your agent's skills folder).
name
evidence-grading
description
Grade the evidence behind a claim before betting on it — the hierarchy for business evidence (experiments > usage data > surveys > interviews > anecdotes > opinion), the fit-for-decision test, and the mixed-evidence verdicts that real questions produce. Use when asked how strong is our evidence for this, grade what we know before the decision, is this enough to bet on, or we have three anecdotes and a survey — now what. Produces the evidence inventory with grades, the sufficiency verdict against the decision's stakes, and the cheapest-upgrade path.

Evidence Grading Skill

"We have evidence" covers everything from a randomized experiment to the CEO's seatmate on a flight — and decisions made on ungraded evidence inherit the confusion. The grading discipline: inventory what actually supports the claim, place each item on the business-evidence hierarchy (what it is, not how confident it feels), test the graded weight against the decision's stakes (a reversible pilot needs less than a one-way rebrand — decision-journal logic), and when evidence falls short, name the cheapest upgrade — because the answer to weak evidence is usually a better test, not a braver bet.

What This Skill Produces

  • The evidence inventory — everything supporting (and contradicting) the claim, each item graded on the hierarchy
  • The weight assessment — what the graded pile actually supports, including the mixed-signals honest read
  • The sufficiency verdict — enough for this decision's stakes / not yet — with the stakes analysis shown
  • The upgrade path — the cheapest next evidence that would move the verdict ("a 2-week holdout test settles this for $0")

Required Inputs

Ask for these if not provided:

  • The claim and the decision riding on it — "users want X" feeding a backlog item vs. feeding a repositioning are different sufficiency bars; the decision's reversibility and cost set the bar
  • The evidence, itemized — every piece: the data pull, the survey, the five customer quotes, the competitor's move, the expert's opinion — including the inconvenient items (an inventory that omits contradicting evidence is advocacy)
  • The evidence's provenance — n, selection, dates, who collected it and with what incentive (source-triangulation supplies the externals; internal evidence has incentives too)

Framework: The Grading Rules

  1. The hierarchy, applied without sentiment: experiments/A-B tests (causal) > usage/behavioral data (what people do, correlational) > surveys (what they say, at scale — survey-design-basics quality adjusts the grade) > interviews (rich, small-n — interview-synthesis counts matter) > anecdotes (existence proofs only) > expert opinion (informed priors) > internal conviction (a hypothesis, not evidence). Each item gets its rung and its quality-within-rung — a leading survey grades below honest interviews.
  2. Direction and independence both count: the inventory includes contradicting evidence at its own grade, and echo-checks the supporting pile (three anecdotes traceable to one loud customer are one anecdote). The weight is the net, honestly netted.
  3. Anecdotes prove existence, never prevalence: "a customer asked for X" establishes that the need exists somewhere — it cannot establish how common; the classic grading error is prevalence conclusions from existence evidence, and it's the error this skill most often catches.
  4. Sufficiency is relative to stakes: the verdict tests weight against the decision — reversible-and-cheap decisions legitimately run on interview-grade evidence (the pilot is the experiment); irreversible-and-expensive ones demand behavioral or experimental grade. "Weak evidence" isn't a verdict; "weak for this bet" is.
  5. The upgrade path is the constructive ending: when insufficient, name the cheapest test that would change the verdict — the holdout, the fake-door, the survey that sizes the interview theme, the pilot-with-metrics. Ranked by cost-to-confidence ratio; the skill's product is often not "no" but "this $0 two-week test first."
Show full SKILL.md (213 more words)Show less

Output Format

Evidence Grade: "[the claim]" — feeding [the decision]

The Inventory

EvidenceRungQuality notes (n, selection, date, independence)Direction

The Weight

[What the graded net actually supports, in one honest paragraph — existence vs. prevalence vs. causation explicitly]

The Sufficiency Verdict

[The decision's stakes (reversibility × cost) · enough / not yet · the reasoning]

The Upgrade Path

[The cheapest evidence that moves the verdict · cost and time · the second option]

Quality Checks

  • Every item has a rung and within-rung quality notes
  • Contradicting evidence is in the inventory at its own grade
  • Echoes were collapsed before weighing
  • The verdict is stakes-relative, not absolute
  • The insufficient branch ends in a priced upgrade, not just a no

Anti-Patterns

  • Do not grade by vividness — the memorable anecdote outshines the boring dataset in every meeting; the hierarchy exists to resist exactly that
  • Do not conclude prevalence from existence — the most common grading felony
  • Do not omit the inconvenient items — an advocacy inventory grades the author, not the claim
  • Do not demand experimental grade for reversible bets — over-evidencing cheap decisions is its own waste
  • Do not end at "insufficient" — the upgrade path is the difference between rigor and obstruction

Example Trigger Phrases

  • "How strong is our evidence for this?"
  • "Grade what we know before the decision."
  • "Is this enough to bet on?"

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/evidence-grading of mohitagw15856/pm-claude-skills.

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Evidence Grading next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evidence Grading compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evidence Grading this skillmohitagw15856/pm-claude-skills1.4k—~1.5kAutomated safety check: PassMIT
Claimsruvnet/ruflo74k2 repos~1.1kAutomated safety check: PassMIT
Heading Hierarchythedaviddias/Front-End-Checklist74k—~516Automated safety check: PassMIT
Grade Iteratealirezarezvani/claude-skills28k—~1kAutomated safety check: PassMIT
Dos Verify Done Claimssickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
Claim Chartanthropics/claude-for-legal9.6k2 repos~10kAutomated safety check: PassApache-2.0

Similar skills

  • Claims

    ruvnet/ruflo

    Claims-based authorization for agents and operations. An agent skill from ruvnet/ruflo.

    74k GitHub starsUsed in 2 repos~1.1k tokens
    Backend & APIsAuto-check passed
  • Heading Hierarchy

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing rendered HTML, interactive components, or design-system patterns related to Use logical heading hierarchy.

    74k GitHub stars~516 tokensUpdated 4 days ago
    Frontend & DesignAuto-check passed
  • Grade Iterate

    alirezarezvani/claude-skills

    Phase 3 of building a Claude Managed Agent — the bounded grade→iterate loop.

    28k GitHub stars~1k tokensUpdated 1 mo ago
    EducationAuto-check passed
  • Dos Verify Done Claims

    sickn33/agentic-awesome-skills

    Before accepting an agent's 'done / shipped / fixed' claim, verify it against ground truth (git ancestry + the commit's own diff) using the DOS kernel's dos verify and dos commit-audit — never the…

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Media & CreativeAuto-check passed
  • Claim Chart

    anthropics/claude-for-legal

    Official

    Build or review an element chart — a patent claim chart (infringement, invalidity, or review) or a civil element chart for any cause of action or defense — with every cell pin-cited and gap…

    9.6k GitHub starsUsed in 2 repos~10k tokens
    Legal & ComplianceAuto-check passed
  • Paper Claim Audit

    wanshuiyin/Auto-claude-code-research-in-sleep

    Zero-context verification that every number, comparison, and scope claim in the paper matches raw result files.

    17k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check: notes

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Questions about Evidence Grading

What does Evidence Grading do?

Grade the evidence behind a claim before betting on it — the hierarchy for business evidence (experiments usage data surveys interviews anecdotes opinion), the fit-for-decision test, and the…. Evidence Grading is an agent skill from mohitagw15856/pm-claude-skills. Grade the evidence behind a claim before betting on it — the hierarchy for business evidence (experiments usage data surveys interviews anecdotes opinion), the fit-for-decision test, and the mixed-evidence verdicts that real questions produce.

When should I use Evidence Grading?

Evidence Grading fits situations like: asked how strong is our evidence for this; grade what we know before the decision; is this enough to bet on; we have three anecdotes and a survey — now what.

How do I install Evidence Grading in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill evidence-grading -a claude-code`. Or copy the skill folder (skills/evidence-grading in mohitagw15856/pm-claude-skills) into .claude/skills/evidence-grading in your project. Claude Code loads it when a task matches its description.

How do I install Evidence Grading in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill evidence-grading -a codex`. Or copy the skill folder (skills/evidence-grading in mohitagw15856/pm-claude-skills) into .agents/skills/evidence-grading in your project. Codex loads it when a task matches its description.

Can I use Evidence Grading in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill evidence-grading -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evidence-grading, .gemini/skills/evidence-grading, .github/skills/evidence-grading and .opencode/skills/evidence-grading in your project.

What does Evidence Grading need to run?

SKILL.md names no scripts, command-line tools or credentials: Evidence Grading is instructions for the agent only.

Does Evidence Grading access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evidence Grading safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evidence Grading use?

Evidence Grading is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evidence Grading use?

About 1.5k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evidence Grading?

Skills that share tags, products or a category with Evidence Grading: Claims (ruvnet/ruflo, 74k stars), Heading Hierarchy (thedaviddias/Front-End-Checklist, 74k stars), Grade Iterate (alirezarezvani/claude-skills, 28k stars) and Dos Verify Done Claims (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evidence Grading?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,434 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 9, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.