Agent skill

Tool Design Sprint Test And Score

by product-on-purpose in product-on-purpose/pm-skills

Day 5 (Friday) sprint-closing move of a Design Sprint that produces the bundled Friday artifact covering per-customer interview observations, best quotes, scorecard grid (sprint questions by…

Apache-2.0Auto-check passedProduct & Project Management

Install Tool Design Sprint Test And Score

skills CLI
$ npx skills add product-on-purpose/pm-skills --skill tool-design-sprint-test-and-score -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install product-on-purpose/pm-skills tool-design-sprint-test-and-score --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/product-on-purpose/pm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tool-design-sprint-test-and-score .claude/skills/tool-design-sprint-test-and-score && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tool-design-sprint-test-and-score
GitHub stars
716
Token cost
~2.6k tokens
SKILL.md length
1,232 words
Files
4 (incl. references)
Skills in repo
68
Repo updated
First seen
Licence
Apache-2.0

At a glance

Day 5 (Friday) sprint-closing move of a Design Sprint that produces the bundled Friday artifact covering per-customer interview observations, best quotes, scorecard grid (sprint questions by…

  • Works in 6 steps: Per-customer interview observation… → Best quotes: 5-15 verbatim customer… → Scorecard grid: rows are the sprint… → …
  • Tasks that involve User research
  • SKILL.md covers When to Use, When NOT to Use, What This Skill Produces and Friday Time Structure, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tool Design Sprint Test And Score is an agent skill from product-on-purpose/pm-skills. Day 5 (Friday) sprint-closing move of a Design Sprint that produces the bundled Friday artifact covering per-customer interview observations, best quotes, scorecard grid (sprint questions by customers), observed patterns, hot takes from each team member, and the Decider summary (build, iterate, pivot, or stop, plus highest-confidence learning, most important revision, and next artifact). Use Friday after Thursday's prototype passes trial run and during/after the 5 customer interviews. The sprint's payoff artifact.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `HISTORY.md`, `references/EXAMPLE.md` and `references/TEMPLATE.md`).

It sits in Product & Project Management, covering User research. The repository describes itself as: 68 plug-and-play, best-practice product management skills for AI agents: 30 Triple Diamond phase + 11 foundation + 12 utility + 15 tool (Foundation Sprint + Design Sprint). Plus… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve User research

Example prompts

  • “/tool-design-sprint-test-and-score”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Per-customer interview observation notes: one section per customer; covers Context (Act 2) reactions, Tasks (Act 4) behavior with…
  2. Best quotes: 5-15 verbatim customer quotes the team flags as most signal-bearing. Used in the Decider summary and in any downstream pitch…
  3. Scorecard grid: rows are the sprint questions (from Monday); columns are the 5 customers; each cell is Y / N / partial / unclear with a…
  4. Observed patterns: 4 buckets (worked, hesitated, broke trust, unexpected) with 2-4 patterns per bucket. Each pattern names how many…
  5. Hot takes: one short paragraph per team member capturing their personal read on Friday before group synthesis biases the read. Written…
  6. Decider summary: the Decider's call (build / iterate / pivot / stop / reframe) plus the highest-confidence learning, the most important…

What it can do on your machine

Read from SKILL.md and the folder at commit 1cef1a9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • gv.com
    • character.vc
    • designsprintkit.withgoogle.com
    • nngroup.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tool Design Sprint Test And Score loads about 2.6k tokens when it runs, and up to ~8.6k if it reads all its reference files. Until then it costs about 138 tokens; SKILL.md has 1,232 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~138
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from product-on-purpose/pm-skills at commit 1cef1a9, republished under its Apache-2.0 licence (© product-on-purpose). 1,232 words, ~2,605 tokens.

Download SKILL.mdSave it as .claude/skills/tool-design-sprint-test-and-score/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
tool-design-sprint-test-and-score
description
Day 5 (Friday) sprint-closing move of a Design Sprint that produces the bundled Friday artifact covering per-customer interview observations, best quotes, scorecard grid (sprint questions by customers), observed patterns, hot takes from each team member, and the Decider summary (build, iterate, pivot, or stop, plus highest-confidence learning, most important revision, and next artifact). Use Friday after Thursday's prototype passes trial run and during/after the 5 customer interviews. The sprint's payoff artifact.
license
Apache-2.0
metadata.classification
tool
metadata.version
1.0.0
metadata.updated
2026-07-04
metadata.tool
design-sprint
metadata.move
test-and-score
metadata.category
discovery
metadata.frameworks
design-sprint, sprint
metadata.timebox_minutes
270
metadata.roles
facilitator, researcher, decider, pm, design, engineering
<!-- PM-Skills | https://github.com/product-on-purpose/pm-skills | Apache 2.0 -->

Design Sprint Test and Score (Friday)

Friday is the sprint's payoff. 5 target-profile customers run the prototype while the team observes; the team synthesizes observations into a scorecard against the sprint questions; the Decider makes the build / iterate / pivot / stop call by end-of-day. The week's 35-40 person-days plus customer recruiting cost converts into one actionable decision.

Family contract: docs/reference/skill-families/design-sprint-skills-contract.md. This skill is a member of design-sprint-skills.

When to Use

  • It is Day 5 of the Design Sprint and Thursday's prototype passed trial run.
  • 5 confirmed participants are scheduled (canonical; or 4 if 1 cancelled-and-no-buffer; pause if below 4).
  • The team can observe interviews live (in-person or via Zoom breakout room) and synthesize during the day.
  • The Decider is present Friday PM for the post-interview review (canonically 14:00-18:00 PT window covering observation of slots 4-5 plus Decider review by 17:30 PT).

When NOT to Use

  • Thursday prototype did not pass trial run. Re-run trial; if still failing at 19:00 PT Thursday, postpone Friday.
  • Fewer than 3 customers confirmed. Per Ratified Decision 3, the canonical guidance is 5 customers; 3-4 or 6-7 gets a documented warning; below 3 or above 7 should trigger a re-decision (postpone or split testing). Note: the v0.1.0 family validator does NOT mechanically enforce these thresholds (cohort count is in the EXAMPLE artifact, not in frontmatter); enforcement is a v2.16 validator-expansion candidate.
  • Decider unavailable for the post-interview review window. Without Decider, the day produces observations without a call.
  • The team plans to use this skill to write the executive memo. Per Ratified Decision 4: exec memo authoring is delegated to foundation-stakeholder-update (existing pm-skills foundation skill); this skill produces the Decider summary only.

What This Skill Produces

A single bundled artifact with six sections:

  1. Per-customer interview observation notes: one section per customer; covers Context (Act 2) reactions, Tasks (Act 4) behavior with timestamps, Debrief (Act 5) reactions including pricing. Captured live during the day's interviews.
  2. Best quotes: 5-15 verbatim customer quotes the team flags as most signal-bearing. Used in the Decider summary and in any downstream pitch or planning artifact.
  3. Scorecard grid: rows are the sprint questions (from Monday); columns are the 5 customers; each cell is Y / N / partial / unclear with a one-line note; rightmost column is the team's day-end decision per question (Validated / Invalidated / Inconclusive).
  4. Observed patterns: 4 buckets (worked, hesitated, broke trust, unexpected) with 2-4 patterns per bucket. Each pattern names how many customers showed it.
  5. Hot takes: one short paragraph per team member capturing their personal read on Friday before group synthesis biases the read. Written silently in parallel.
  6. Decider summary: the Decider's call (build / iterate / pivot / stop / reframe) plus the highest-confidence learning, the most important revision the team would make to the prototype direction, and the next artifact the team will produce (the post-sprint deliverable).

See references/TEMPLATE.md for the canonical structure and references/EXAMPLE.md for the Brainshelf book-catalog Friday artifact.

Friday Time Structure

Friday is the longest day: customer interviews start early (canonically 09:00 PT) and the Decider review concludes the day (canonically 17:30 PT).

  • 09:00-16:30: 5 customer interviews of 50-60 minutes each at 09:00 / 10:30 / 12:00 / 14:00 / 15:30. Each slot: 10 min setup + 50-55 min interview + 5 min team huddle to capture observations before next customer.
  • 13:00-14:00: Lunch (slot 3 wraps ~13:00; lunch overlaps the slot 3 to slot 4 buffer)
  • 16:30-16:45: Last-customer wrap; observation note tidy
  • 16:45-17:00: Team writes hot takes silently in parallel
  • 17:00-17:30: Decider reviews scorecard + hot takes; makes the call
  • 17:30-18:00: Decider summary captured; team begins post-sprint disposition (next-step calendar, downstream deliverable assignment)

This skill's 270-minute timebox covers the synthesis sections (scorecard, patterns, hot takes, Decider summary). The 5 interviews themselves (~5 hours of interview time) run in parallel with continuous observation capture.

Scorecard Mechanic

The scorecard is a 2-D grid. Rows are sprint questions from Monday's map-and-target (typically 3-7). Columns are the 5 customers (anonymized IDs). Each cell answers: did this customer's interview validate, invalidate, or leave inconclusive the row's question?

C1C2C3C4C5Day-end decision
Q1YYNYpartialValidated (4 of 5)
Q2NYunclearNNInvalidated (3-of-5 N, 1 of 5 Y)
.....................

Day-end decision rules:

  • Validated: 4 or 5 of 5 Y (strong signal); 3 of 5 Y with no N (directional). For 4-customer cohorts: 4 Y is Validated; 3 Y with no N is directional.
  • Invalidated: 4 or 5 of 5 N. For 4-customer cohorts: 4 N is Invalidated; 3 N with no Y is directional.
  • Inconclusive: all other patterns. Inconclusive questions get scheduled for follow-up (a smaller test, a quant experiment, or a second Design Sprint).

The Decider can override day-end decisions but should record reasoning.

Show full SKILL.md (458 more words)Show less

Common Pitfalls

  • Observation notes too narrative, not behavioral. "Customer seemed confused" is a narrative; "Customer hovered on the capture button for 4 seconds without tapping, then tapped twice in rapid succession" is behavior. Behavior is data; narrative is interpretation.
  • Scorecard cells filled in by consensus. Each observer writes their cell read; differences are surfaced, not averaged. If C1's read on Q1 is split 2 Y vs 2 N across the team, the cell is "split" with an explanatory note.
  • Hot takes written after group synthesis. Hot takes are written SILENTLY and in PARALLEL before group synthesis. Writing them after a group debrief produces consensus, not signal.
  • Decider hesitating on the call because "we want more data." Friday's job is to produce a call with the data you have. If the call truly cannot be made, the call is "iterate" (re-sprint with adjustments). "Defer" is not an answer.
  • Skipping the Decider summary because "we'll write it up Monday." The summary is captured Friday before the team leaves. Monday is too late; context decays fast.
  • Treating "5 customers" as a soft target. Per the canonical research, 5 is where confidence about patterns crosses an inflection point. Fewer than 4 produces noisy signal; more than 7 produces synthesis-overload without much marginal signal.

Cross-Skill Usage

Prerequisites: tool-design-sprint-prototype-plan. Friday consumes the prototype and the interview script from Thursday. Without a working prototype that passed trial run, Friday cannot run.

This skill does NOT invoke tool-note-and-vote. Friday has no voting moment; the scorecard cells are individual reads and the Decider summary is the Decider's call.

This skill does NOT author an executive memo (per Ratified Decision 4). If the team wants an exec memo or stakeholder update, the next invocation is foundation-stakeholder-update, which consumes the Decider summary as input.

Downstream invocations after the sprint closes: deliver-prd (if Decider call is "build"); measure-experiment-design (if "iterate" requires a smaller follow-on experiment); iterate-pivot-decision (if "pivot" requires documenting the pivot rationale); foundation-stakeholder-update (if any of the above need stakeholder communication).

Canonical Sources

Decider Checkpoint

This skill ends with a Decider Checkpoint in references/TEMPLATE.md. The Decider's call (build / iterate / pivot / stop / reframe) IS the checkpoint; the sprint cannot close without it. The checkpoint also captures the next artifact the team owns producing (a PRD, a smaller experiment, a pivot memo, or a stakeholder update), which is what triggers Monday's post-sprint work to begin clean.

© product-on-purpose, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/tool-design-sprint-test-and-score of product-on-purpose/pm-skills.

  • SKILL.md
  • HISTORY.md
  • references/EXAMPLE.md
  • references/TEMPLATE.md

Open the folder on GitHubat commit 1cef1a9

Compare with similar skills

Tool Design Sprint Test And Score next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tool Design Sprint Test And Score compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tool Design Sprint Test And Score this skillproduct-on-purpose/pm-skills716—~2.6kAutomated safety check: PassApache-2.0
User Research Cookiycookiy-ai/user-research-skill1.6k—~954Automated safety check: PassMIT
Fable DomainSahir619/fable-method2.3k—~2.6kAutomated safety check: PassMIT
Produck Feedback To Buildtryproduck/produck-skills511—~1kAutomated safety check: PassApache-2.0
Customer InterviewsRefoundAI/lenny-skills1.4k—~1.7kAutomated safety check: PassMIT
Product Discovery Brief Builderopen-mercato/skills231—~3kAutomated safety check: PassMIT

Similar skills

  • User Research Cookiy

    cookiy-ai/user-research-skill

    End-to-end user research assistant — qualitative and quantitative.

    1.6k GitHub stars~954 tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed
  • Fable Domain

    Sahir619/fable-method

    Discuss a domain with the user, research it from real sources, then generate a trusted skill bundle for it - a step-by-step workflow with a flowchart, a domain adapter, a trap fixture, and a smoke…

    2.3k GitHub stars~2.6k tokensUpdated 8 days ago
    Product & Project ManagementAuto-check passed
  • Produck Feedback To Build

    tryproduck/produck-skills

    Pulls full in-context user feedback tickets through the Produck MCP server and turns them into an aligned product change instead of a guess.

    511 GitHub stars~1k tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed
  • Customer Interviews

    RefoundAI/lenny-skills

    Help users conduct high-impact customer interviews that move beyond surface-level feature requests to identify root emotional frustrations and specific causal triggers.

    1.4k GitHub stars~1.7k tokensUpdated 2 mo ago
    Product & Project ManagementAuto-check passed
  • Guides a product discovery conversation and writes product-brief.md with the problem, evidence, scope, decisions and the next open question, for existing, client or own ideas.

    231 GitHub stars~3k tokensUpdated 5 days ago
    Product & Project ManagementAuto-check passed
  • Discovery Synthesis

    andreaskelm/pm-brain

    Plan customer discovery, turn interview snapshots into synthesis and evidence-based opportunities, build or update an Opportunity Solution Tree, map jobs and segments, and design RAT tests for the…

    234 GitHub stars~2.3k tokensUpdated yesterday
    Product & Project ManagementAuto-check passed

More from product-on-purpose/pm-skills

All 68 skills in this repo
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    716 GitHub stars~966 tokensUpdated 3 days ago
    Auto-check passed
  • Define Jtbd Canvas

    product-on-purpose/pm-skills

    Creates a Jobs to be Done canvas capturing the functional, emotional, and social dimensions of a customer job.

    716 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed
  • Define Opportunity Tree

    product-on-purpose/pm-skills

    Creates an opportunity solution tree connecting a desired outcome to customer opportunities and candidate solutions, preventing solution-first jumps in continuous discovery.

    716 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed
  • Define Problem Statement

    product-on-purpose/pm-skills

    Creates a clear problem framing document with user impact, business context, and success criteria.

    716 GitHub stars~932 tokensUpdated 3 days ago
    Auto-check passed
  • Deliver Acceptance Criteria

    product-on-purpose/pm-skills

    Generates structured Given/When/Then acceptance criteria for a user story or feature slice, covering the happy path, key failure scenarios, and non-functional expectations in testable form.

    716 GitHub stars~1k tokensUpdated 3 days ago
    Auto-check passed
  • Deliver Launch Checklist

    product-on-purpose/pm-skills

    Creates a cross-functional pre-launch checklist covering engineering, design, marketing, support, legal, and operations readiness, with owners, dates, and go/no-go criteria so nothing is missed…

    716 GitHub stars~970 tokensUpdated 3 days ago
    Auto-check passed

Questions about Tool Design Sprint Test And Score

What does Tool Design Sprint Test And Score do?

Day 5 (Friday) sprint-closing move of a Design Sprint that produces the bundled Friday artifact covering per-customer interview observations, best quotes, scorecard grid (sprint questions by…. Tool Design Sprint Test And Score is an agent skill from product-on-purpose/pm-skills. Day 5 (Friday) sprint-closing move of a Design Sprint that produces the bundled Friday artifact covering per-customer interview observations, best quotes, scorecard grid (sprint questions by customers), observed patterns, hot takes from each team member, and the Decider summary (build, iterate, pivot, or stop, plus highest-confidence learning, most important revision, and next artifact).

When should I use Tool Design Sprint Test And Score?

Tool Design Sprint Test And Score fits situations like: tasks that involve User research.

How do I install Tool Design Sprint Test And Score in Claude Code?

Run `npx skills add product-on-purpose/pm-skills --skill tool-design-sprint-test-and-score -a claude-code`. Or copy the skill folder (skills/tool-design-sprint-test-and-score in product-on-purpose/pm-skills) into .claude/skills/tool-design-sprint-test-and-score in your project. Claude Code loads it when a task matches its description.

How do I install Tool Design Sprint Test And Score in Codex?

Run `npx skills add product-on-purpose/pm-skills --skill tool-design-sprint-test-and-score -a codex`. Or copy the skill folder (skills/tool-design-sprint-test-and-score in product-on-purpose/pm-skills) into .agents/skills/tool-design-sprint-test-and-score in your project. Codex loads it when a task matches its description.

Can I use Tool Design Sprint Test And Score in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add product-on-purpose/pm-skills --skill tool-design-sprint-test-and-score -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tool-design-sprint-test-and-score, .gemini/skills/tool-design-sprint-test-and-score, .github/skills/tool-design-sprint-test-and-score and .opencode/skills/tool-design-sprint-test-and-score in your project.

What does Tool Design Sprint Test And Score need to run?

SKILL.md names no scripts, command-line tools or credentials: Tool Design Sprint Test And Score is instructions for the agent only.

Does Tool Design Sprint Test And Score access the network?

SKILL.md names 4 domains. As links in the text: gv.com, character.vc, designsprintkit.withgoogle.com and nngroup.com. This is read from the text; nothing was executed.

Is Tool Design Sprint Test And Score safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tool Design Sprint Test And Score use?

Tool Design Sprint Test And Score is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tool Design Sprint Test And Score use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6k tokens, read only when the agent opens those files.

What are the alternatives to Tool Design Sprint Test And Score?

Skills that share tags, products or a category with Tool Design Sprint Test And Score: User Research Cookiy (cookiy-ai/user-research-skill, 1.6k stars), Fable Domain (Sahir619/fable-method, 2.3k stars), Produck Feedback To Build (tryproduck/produck-skills, 511 stars) and Customer Interviews (RefoundAI/lenny-skills, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tool Design Sprint Test And Score?

product-on-purpose (a GitHub organization) maintains it in product-on-purpose/pm-skills, which has 716 GitHub stars. The repository holds 68 skills in this directory. The repository was last updated on October 8, 2026.

Source: product-on-purpose/pm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.