Agent skill

Idea Scoring

by MaxKmet in MaxKmet/idea-validation-agents

Aggregates all dimension scores into a final idea score (0–100) and issues a verdict.

MITAuto-check passed

Install Idea Scoring

skills CLI
$ npx skills add MaxKmet/idea-validation-agents --skill idea-scoring -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install MaxKmet/idea-validation-agents idea-scoring --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/MaxKmet/idea-validation-agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/idea-scoring .claude/skills/idea-scoring && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
idea-scoring
GitHub stars
474
Token cost
~3.2k tokens
SKILL.md length
1,269 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Aggregates all dimension scores into a final idea score (0–100) and issues a verdict.

  • Works in 6 steps: Compute dimension sub-scores → Apply floor penalty (the "Killer… → Compute weighted base score → …
  • SKILL.md covers Purpose, Input, Scoring Dimensions and Dimension Score Mapping, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Idea Scoring is an agent skill from MaxKmet/idea-validation-agents. Aggregates all dimension scores into a final idea score (0–100) and issues a verdict. Implements a multiplicative-floor algorithm with Riskiest Assumption Test (RAT). The final output of every validation workflow.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: AI agents that act as your personal venture analyst - from startup idea brainstorming to full validation and go-to-market strategy. Built for developers who'd rather validate in… The licence is MIT.

Example prompts

  • “Use the idea-scoring skill to aggregate all dimension scores into a final idea score (0–100) and issues a verdict”
  • “/idea-scoring”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Compute dimension sub-scores
  2. Apply floor penalty (the "Killer Dimension" rule)
  3. Compute weighted base score
  4. Apply floor penalty and missing-input discount
  5. Determine confidence level
  6. Issue verdict

What it can do on your machine

Read from SKILL.md and the folder at commit 3a4c800. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Idea Scoring loads about 3.2k tokens when it runs. Until then it costs about 57 tokens; SKILL.md has 1,269 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from MaxKmet/idea-validation-agents at commit 3a4c800, republished under its MIT licence (© MaxKmet). 1,269 words, ~3,169 tokens.

Download SKILL.mdSave it as .claude/skills/idea-scoring/SKILL.md (or your agent's skills folder).
name
idea-scoring
description
Aggregates all dimension scores into a final idea score (0–100) and issues a verdict. Implements a multiplicative-floor algorithm with Riskiest Assumption Test (RAT). The final output of every validation workflow.
<!-- version: 0.2.0 | outputs: memory/ideas/<slug>/scores.json -->

Skill: idea-scoring

Purpose

Produce a single, defensible verdict on an idea by aggregating all available dimension scores using a multiplicative-floor algorithm — a single catastrophic weakness kills the score, just like it kills a real startup. Includes a Riskiest Assumption Test (RAT) to convert the verdict into a concrete next action.

Input

  • Idea slug
  • One or more of the following (uses whatever is available):
    • memory/ideas/<slug>/idea.md
    • memory/ideas/<slug>/desire_scores.json
    • memory/ideas/<slug>/competitors.json
    • memory/ideas/<slug>/pricing.json
    • memory/ideas/<slug>/cac.json
    • memory/ideas/<slug>/market_size.json
    • memory/ideas/<slug>/distribution.json
    • memory/ideas/<slug>/retention.json
    • memory/ideas/<slug>/complexity.json
    • memory/ideas/<slug>/weighted_signals.json
  • Optional: memory/user_profile.md (for founder-market fit)
Minimum Viable Input

At least 3 of 7 dimensions must have source data. If fewer are available, refuse to score and list what's missing. Two dimensions are mandatory — Demand and Distribution. Without evidence of a real problem and a path to reach users, scoring is meaningless.

Scoring Dimensions

DimensionWeightSourceWhat it measures
Demand20%desire_scores.json + weighted_signals.json + idea.mdReal human desire + validated market signals
Competition10%competitors.jsonPositioning gaps and defensibility
Monetization20%pricing.json + cac.json + market_size.jsonUnit economics viability (LTV:CAC, WTP, market size)
Distribution20%distribution.jsonOrganic reach, paid viability, founder edge
Retention15%retention.jsonHabit formation, churn risk, usage frequency
Founder-Market Fit15%user_profile.md + domain overlap with ideaBuilder's edge, domain expertise, distribution advantage

Dimension Score Mapping

Each dimension maps source data to a 0–100 sub-score using the rubrics below. When source data uses qualitative labels, apply these conversions.

Demand (0–100)
ConditionScore range
desire_strength_label = "strong" AND trend_velocity = "rising-fast"80–100
desire_strength_label = "strong" OR trend_velocity = "rising"60–79
desire_strength_label = "moderate" AND some signal validation40–59
desire_strength_label = "weak" OR trend_velocity = "declining"15–39
No signal data, speculation only0–14

Adjust within range: +10 if cross-platform resonance confirmed, +5 if monetization_validated is true in idea.md.

Competition (0–100)

Higher = more favorable competitive landscape (counterintuitive — think of it as "opportunity score").

ConditionScore range
market_saturation = "low", clear positioning gaps, no dominant incumbent75–100
market_saturation = "medium", 1–2 positioning gaps identified50–74
market_saturation = "high" but differentiation opportunities exist25–49
market_saturation = "high", no differentiation, dominant incumbents0–24

Adjust: +10 if top competitor complaints reveal an unserved pain point. -15 if a FAANG-class player owns the category.

Monetization (0–100)
ConditionScore range
LTV:CAC ≥ 3:1 on at least 2 channels, WTP target ≥ $5/mo, viable SOM80–100
LTV:CAC ≥ 3:1 on 1 channel, WTP target ≥ $3/mo60–79
LTV:CAC ≥ 2:1, WTP target $1–$3/mo, marginal unit economics35–59
LTV:CAC < 2:1 OR viability_verdict = "not-viable"10–34
No pricing data or CAC data0–9 (flag as missing)

Adjust: +10 if freemium_conversion_estimate > 5%. +5 if market_size_verdict = "large".

Distribution (0–100)
ConditionScore range
distribution_verdict = "strong", viral_loop_exists = true80–100
distribution_verdict = "strong" OR (organic_reach = "high" + creator_economy_fit = "high")60–79
distribution_verdict = "moderate", at least one viable organic channel40–59
distribution_verdict = "weak", paid-only path15–39
No viable channel identified0–14

Adjust: +10 if founder has existing audience or distribution edge (from user_profile.md).

Retention (0–100)
ConditionScore range
retention_verdict = "sticky", D30 ≥ 20%, habit_formation_score ≥ 480–100
retention_verdict = "sticky" OR D30 ≥ 15%60–79
retention_verdict = "moderate", D30 ≥ 8%40–59
retention_verdict = "disposable" OR D30 < 8%15–39
churn_risk = "high" AND no habit loop0–14

Adjust: +5 if natural_usage_frequency is daily. -10 if weekly-or-less with no external trigger.

Founder-Market Fit (0–100)
ConditionScore range
Strong domain match in strong_domains, distribution advantages align, builder/growth tier75–100
Moderate domain overlap OR builder tier with adjacent experience50–74
Beginner tier but high motivation and time commitment (≥ 20 hrs/wk)30–49
No domain overlap, beginner tier, low time commitment0–29

If user_profile.md is unavailable, default to 50 (neutral) and flag as missing.

Scoring Algorithm

Step 1 — Compute dimension sub-scores

Apply the mapping rubrics above. Record each as d_i (0–100).

Step 2 — Apply floor penalty (the "Killer Dimension" rule)

Any dimension scoring below 25 is a potential startup killer. Apply this penalty:

floor_penalty = 1.0
for each dimension d_i:
    if d_i < 25:
        floor_penalty *= (d_i / 25)

This multiplicative penalty means a single catastrophic weakness (score 0–10) can halve or destroy the final score, regardless of how strong other dimensions are. This reflects startup reality: brilliant distribution cannot save a product nobody wants.

Step 3 — Compute weighted base score
weights = {
    demand: 0.20,
    competition: 0.10,
    monetization: 0.20,
    distribution: 0.20,
    retention: 0.15,
    founder_market_fit: 0.15
}

base_score = sum(d_i * w_i for each dimension)
Step 4 — Apply floor penalty and missing-input discount
missing_discount = available_dimensions / total_dimensions
adjusted_score = base_score * floor_penalty * missing_discount
final_score = round(clamp(adjusted_score, 0, 100))

The missing-input discount ensures that ideas scored on only 3 of 7 dimensions can never reach the "pursue" tier without completing more analysis.

Step 5 — Determine confidence level
Available dimensionsConfidence
6–7 of 7high
4–5 of 7medium
3 of 7 (minimum)low
Step 6 — Issue verdict
ScoreVerdictMeaning
75–100pursueStrong across dimensions. Build an MVP.
55–74testPromising but unproven. Run the RAT experiment first.
35–54pivotStructural weakness. Use pivot-engine to explore alternatives.
0–34dropFatal flaws. Move to next idea.

Riskiest Assumption Test (RAT)

Every idea rests on assumptions. The RAT identifies the single assumption that, if wrong, kills the idea — and designs the cheapest possible experiment to test it before building anything.

Show full SKILL.md (507 more words)Show less
RAT Identification Process
  1. List all assumptions embedded in the idea (drawn from dimension scores and source data):

    • Demand: "People actually have this problem and will seek a solution"
    • Monetization: "Users will pay $X/mo for this"
    • Distribution: "We can reach users via [channel] at acceptable cost"
    • Retention: "Users will come back [frequency]"
    • Competition: "Our differentiator matters to users"
    • Founder fit: "I can build this with my current skills/resources" (using user_profile.md if available)
  2. Rank by two axes (each 1–5):

    • Criticality: If wrong, how dead is the idea? (5 = instant kill)
    • Uncertainty: How little evidence do we have? (5 = pure speculation)
  3. RAT = assumption with highest (criticality × uncertainty). Ties broken by criticality.

RAT Experiment Design

For the identified RAT, design an experiment following these constraints:

ConstraintRequirement
Time to run≤ 2 weeks
Cost to run≤ $100 (indie budget)
Signal typeBehavioral (what people DO, not what they SAY)
Sample sizeMinimum credible: 30 responses or 100 landing page visitors
Experiment types by assumption category
Assumption categoryExperiment template
Demand existsLanding page with email capture. Pass: ≥ 10% signup rate from ≥ 100 targeted visitors.
WTP is realLanding page with price shown + "buy" button (payment step). Pass: ≥ 3% click-to-buy from ≥ 100 visitors.
Distribution worksRun 1 channel for 7 days (e.g., 5 TikToks, 10 Reddit posts, ASO test). Pass: CAC below modeled threshold.
Retention holdsConcierge MVP or manual-ops version with 10–30 users for 14 days. Pass: ≥ 3 return sessions per user.
Differentiator mattersShow competitor + your concept side-by-side to 30 target users. Pass: ≥ 60% prefer your concept.
Pass/Fail Threshold

Define the threshold before running the experiment. The threshold is written into scores.json so it can be evaluated later. Thresholds must be:

  • Specific: a number, not "good engagement"
  • Time-bound: measured within the experiment window
  • Binary: pass or fail, no "sort of passed"

Process (step by step)

  1. Load all available dimension files from memory/ideas/<slug>/.
  2. Check minimum viable input (≥ 3 dimensions, including Demand and Distribution). If not met, refuse and list missing inputs.
  3. Map each dimension's source data to a 0–100 sub-score using the rubrics.
  4. Compute floor_penalty from any sub-scores below 25.
  5. Compute base_score using weighted sum.
  6. Apply floor_penalty and missing_discount to get final_score.
  7. Determine score_confidence.
  8. Issue verdict from threshold table.
  9. Identify top_strengths (top 2 dimensions) and top_weaknesses (bottom 2 dimensions).
  10. Run RAT identification: list assumptions, score criticality × uncertainty, select the riskiest.
  11. Design RAT experiment with pass/fail threshold.
  12. If this is a pivot re-score, write to pivot_scores.json instead.

Output

Write to memory/ideas/<slug>/scores.json (or pivot_scores.json for re-scores):

json
{
  "dimension_scores": {
    "demand": 0,
    "competition": 0,
    "monetization": 0,
    "distribution": 0,
    "retention": 0,
    "founder_market_fit": 0
  },
  "weights_applied": {
    "demand": 0.20,
    "competition": 0.10,
    "monetization": 0.20,
    "distribution": 0.20,
    "retention": 0.15,
    "founder_market_fit": 0.15
  },
  "floor_penalty": 1.0,
  "base_score": 0,
  "missing_discount": 1.0,
  "final_score": 0,
  "verdict": "pursue | test | pivot | drop",
  "score_confidence": "high | medium | low",
  "missing_inputs": [],
  "top_strengths": [
    { "dimension": "", "score": 0, "reason": "" }
  ],
  "top_weaknesses": [
    { "dimension": "", "score": 0, "reason": "" }
  ],
  "killer_dimensions": [],
  "riskiest_assumption_test": {
    "assumption": "",
    "category": "demand | monetization | distribution | retention | competition | founder_fit",
    "criticality": 0,
    "uncertainty": 0,
    "rat_score": 0,
    "experiment": {
      "type": "",
      "description": "",
      "duration": "",
      "estimated_cost": "",
      "pass_threshold": "",
      "fail_action": "pivot | drop | re-test with different channel"
    },
    "all_assumptions_ranked": [
      { "assumption": "", "criticality": 0, "uncertainty": 0, "rat_score": 0 }
    ]
  }
}

Notes

  • Re-scoring pivots: When scoring a pivot variant from pivot_options.json, write output to memory/ideas/<slug>/pivot_scores.json. Include a pivot_id field referencing the option.
  • Score decay: If source data is older than 90 days (check analyzed_at or file timestamps), apply a 10% confidence penalty and flag stale inputs.
  • Geometric vs. additive: The floor penalty provides multiplicative dynamics (one zero kills the score) while the weighted sum provides interpretable dimension contributions. This hybrid outperforms pure additive (hides fatal flaws) and pure geometric (too punishing for moderate weaknesses).

© MaxKmet, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/idea-scoring of MaxKmet/idea-validation-agents.

Open the folder on GitHubat commit 3a4c800

Compare with similar skills

Idea Scoring next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Idea Scoring compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Idea Scoring this skillMaxKmet/idea-validation-agents474—~3.2kAutomated safety check: PassMIT
Dimensionsthedaviddias/Front-End-Checklist74k—~926Automated safety check: PassMIT
Harness Scoreruvnet/ruflo74k—~605Automated safety check: NotesMIT
Claw Scoreopenclaw/openclaw392k—~2.5kAutomated safety check: PassMIT
Idea Evaluatorsickn33/agentic-awesome-skills47k1 repos~930Automated safety check: PassMIT
Idea Darwinsickn33/agentic-awesome-skills47k2 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Dimensions

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing image assets, markup, and CDN or build transforms related to Set explicit width and height on images.

    74k GitHub stars~926 tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • Harness Score

    ruvnet/ruflo

    5-dimension harness readiness scorecard from metaharness score <path.

    74k GitHub stars~605 tokensUpdated today
    DevelopmentAuto-check: notes
  • Claw Score

    openclaw/openclaw

    Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

    392k GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Idea Evaluator

    sickn33/agentic-awesome-skills

    Evaluates an idea by hosting a multi-turn debate between a Pro and Con agent, delivering a final verdict on whether it's worth pursuing.

    47k GitHub starsUsed in 1 repo~930 tokens
    Auto-check passed
  • Idea Darwin

    sickn33/agentic-awesome-skills

    Darwinian idea evolution engine — toss rough ideas onto an evolution island, let them compete, crossbreed, and mutate through structured rounds to surface your strongest concepts.

    47k GitHub starsUsed in 2 repos~1.1k tokens
    Auto-check passed
  • Idea Refinement

    addyosmani/agent-skills

    Guides a conversation that takes a vague idea through divergent and convergent thinking and ends in a markdown one-pager covering scope and assumptions.

    103k GitHub starsUsed in 6 repos~2k tokens
    Agent WorkflowsAuto-check passed

More from MaxKmet/idea-validation-agents

All 15 skills in this repo
  • Cac Modeler

    MaxKmet/idea-validation-agents

    Models LTV, CAC by channel, LTV:CAC ratios, and payback period for an indie developer.

    474 GitHub stars~3.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Competitor Mapper

    MaxKmet/idea-validation-agents

    Maps the full competitive landscape — direct, indirect, substitute, and emerging competitors — with positioning gap analysis, review mining, and marketinsights-calibrated saturation scoring.

    474 GitHub stars~3.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Decision Memo

    MaxKmet/idea-validation-agents

    Writes a concise, human-readable decision brief summarizing the full validation analysis — including score, verdict, RAT experiment, pre-mortem, and tier-appropriate next actions.

    474 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check passed
  • Desire Evaluator

    MaxKmet/idea-validation-agents

    Scores the strength of core human desire motivations (survival, status, belonging, control, curiosity) for a given app idea to predict user pull and retention potential.

    474 GitHub stars~551 tokensUpdated 3 mo ago
    Auto-check passed
  • Distribution Analysis

    MaxKmet/idea-validation-agents

    Evaluates organic reach potential, paid feasibility, platform distribution advantages, creator economy fit, and founder edge for a B2C app idea.

    474 GitHub stars~3.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Pivot Engine

    MaxKmet/idea-validation-agents

    Generates structured pivot options for a scored idea based on weak dimensions, marketinsights signals, and founder constraints.

    474 GitHub stars~4.7k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Idea Scoring

What does Idea Scoring do?

Aggregates all dimension scores into a final idea score (0–100) and issues a verdict. Idea Scoring is an agent skill from MaxKmet/idea-validation-agents. Aggregates all dimension scores into a final idea score (0–100) and issues a verdict.

How do I install Idea Scoring in Claude Code?

Run `npx skills add MaxKmet/idea-validation-agents --skill idea-scoring -a claude-code`. Or copy the skill folder (skills/idea-scoring in MaxKmet/idea-validation-agents) into .claude/skills/idea-scoring in your project. Claude Code loads it when a task matches its description.

How do I install Idea Scoring in Codex?

Run `npx skills add MaxKmet/idea-validation-agents --skill idea-scoring -a codex`. Or copy the skill folder (skills/idea-scoring in MaxKmet/idea-validation-agents) into .agents/skills/idea-scoring in your project. Codex loads it when a task matches its description.

Can I use Idea Scoring in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add MaxKmet/idea-validation-agents --skill idea-scoring -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/idea-scoring, .gemini/skills/idea-scoring, .github/skills/idea-scoring and .opencode/skills/idea-scoring in your project.

What does Idea Scoring need to run?

SKILL.md names no scripts, command-line tools or credentials: Idea Scoring is instructions for the agent only.

Does Idea Scoring access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Idea Scoring safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Idea Scoring use?

Idea Scoring is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Idea Scoring use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Idea Scoring?

Skills that share tags, products or a category with Idea Scoring: Dimensions (thedaviddias/Front-End-Checklist, 74k stars), Harness Score (ruvnet/ruflo, 74k stars), Claw Score (openclaw/openclaw, 392k stars) and Idea Evaluator (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Idea Scoring?

MaxKmet (a GitHub user) maintains it in MaxKmet/idea-validation-agents, which has 474 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on June 16, 2026.

Source: MaxKmet/idea-validation-agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.