Draft real-life test SCENARIOS (not smoke tests) from a change/feature spec.

MITAuto-check passedTesting & QA

Install Scenario Design

skills CLI
$ npx skills add BlackBeltTechnology/pi-agent-dashboard --skill scenario-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BlackBeltTechnology/pi-agent-dashboard scenario-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BlackBeltTechnology/pi-agent-dashboard.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/eng-disciplines/.pi/skills/scenario-design .claude/skills/scenario-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scenario-design
GitHub stars
315
Token cost
~2.8k tokens
SKILL.md length
1,218 words
Files
3 (incl. references)
Skills in repo
70
Repo updated
First seen
Licence
MIT

At a glance

Draft real-life test SCENARIOS (not smoke tests) from a change/feature spec.

  • Works in 5 steps: Read the spec, classify requirements → Generate scenarios via the Triple (or a… → The clarification Gate (configurable) → …
  • Tasks that involve Test generation
  • SKILL.md covers Core mechanism — the Triple, Phase 1 — Read the spec,…, Phase 2 — Generate scenarios… and Phase 3 — The clarification…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Scenario Design is an agent skill from BlackBeltTechnology/pi-agent-dashboard. Draft real-life test SCENARIOS (not smoke tests) from a change/feature spec. Derives edge-case, performance, frontend-quirk and error-handling scenarios with ISTQB techniques, routes each to a test level, and writes test-plan.md, emitting clarification questions on a spec gap. Use on "design test scenarios", "what should we test", "build a test plan", "find edge cases".

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/technique-cheatsheet.md` and `references/test-plan-schema.md`). Compatibility notes: Optional: OpenSpec change spec as input. Reads a change/feature spec; writes test-plan.md.

It sits in Testing & QA, covering Test generation, PRD writing and QA and bug reports. The repository describes itself as: Real-time web dashboard for pi coding-agent sessions. Multi-session view, live chat mirroring, integrated terminal, diff viewer, pi-flows execution, and mobile-first remote… The licence is MIT.

When your agent uses it

  • Tasks that involve Test generation
  • Tasks that involve PRD writing
  • Tasks that involve QA and bug reports

Example prompts

  • “design test scenarios”
  • “what should we test”
  • “build a test plan”
  • “/scenario-design”

Requirements

  • Docker
  • Compatibility (from SKILL.md): Optional: OpenSpec change spec as input. Reads a change/feature spec; writes test-plan.md.

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Read the spec, classify requirements
  2. Generate scenarios via the Triple (or a gap)
  3. The clarification Gate (configurable)
  4. Route each scenario to a test level
  5. Write test-plan.md

What it can do on your machine

Read from SKILL.md and the folder at commit 7a2d171. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Optional: OpenSpec change spec as input. Reads a change/feature spec; writes test-plan.md.

    From compatibility in the SKILL.md frontmatter.

Context cost

Scenario Design loads about 2.8k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 97 tokens; SKILL.md has 1,218 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from BlackBeltTechnology/pi-agent-dashboard at commit 7a2d171, republished under its MIT licence (© BlackBeltTechnology). 1,218 words, ~2,833 tokens.

Download SKILL.mdSave it as .claude/skills/scenario-design/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
scenario-design
description
Draft real-life test SCENARIOS (not smoke tests) from a change/feature spec. Derives edge-case, performance, frontend-quirk and error-handling scenarios with ISTQB techniques, routes each to a test level, and writes test-plan.md, emitting clarification questions on a spec gap. Use on "design test scenarios", "what should we test", "build a test plan", "find edge cases".
compatibility
Optional: OpenSpec change spec as input. Reads a change/feature spec; writes test-plan.md.
license
MIT
metadata.author
robson
metadata.version
1.0

Turn a change spec into adversarial, real-life test scenarios — designed to break the system, not confirm it works. A scenario is only as good as it is concrete and executable. If the spec can't supply the concrete bits, that is a spec defect, surfaced as a clarification — not a guess.

Input: A change/feature spec. --change <name> (OpenSpec), or infer from context / point the skill at any spec doc. Mode: --stage proposal|design|apply (default: infer — see Gate). Output: a standalone test-plan.md catalog written to your change/spec's test-plan location (OpenSpec: openspec/changes/<name>/test-plan.md).


Core mechanism — the Triple

Every scenario MUST resolve three concrete slots:

text
   ┌─────────────┬──────────────────────┬────────────────────────────┐
   │  INPUT      │  TRIGGER             │  EXPECTED OBSERVABLE OUTCOME │
   │  concrete   │  the condition /     │  a measurable, visible fact  │
   │  data /     │  action that fires   │  (status, value, latency,    │
   │  state      │  the behaviour       │   DOM, log line, exit code)  │
   └─────────────┴──────────────────────┴────────────────────────────┘

Rule (from spec-coding edge-case practice): if any slot is a verb without a noun, or an adjective instead of a number, the slot is unfillable → spec gap. "Handles errors gracefully" is not a Triple. "POST /api/restart while server already restarting (input) → second caller (trigger) → receives 409 within 500ms, no second orchestrator spawned (observable)" is.

Stance: falsify, don't confirm. For each requirement, the job is to find the input+trigger that makes the observable wrong. Happy path is table stakes; the scenario value is in the boundaries and failures.


Phase 1 — Read the spec, classify requirements

  1. Read what exists (OpenSpec layout shown; in a non-OpenSpec project read whatever spec/design/task docs the user points at — do not fail on a missing openspec/ dir or CLI):

    • openspec/changes/<name>/proposal.md (always)
    • openspec/changes/<name>/design.md (if present — decisions, invariants)
    • openspec/changes/<name>/specs/**/spec.md (requirement deltas)
    • openspec/changes/<name>/tasks.md (if present — to align section numbers)
  2. Extract every testable requirement (each SHALL/MUST, each scenario block, each acceptance criterion). For each, tag its shape — this picks the technique:

    Requirement shapeTechnique to applyScenario class
    Input range / numeric / size / countEquivalence Partitioning + Boundary Value Analysisedge-case
    Multiple boolean/enum flags combineDecision Tableedge-case
    Lifecycle / status transitions / reconnect / restartState-Transitionfrontend-quirk + error-handling
    Async / WebSocket / polling / optimistic UIState-convergence + invariant assertions (not UI-visibility)frontend-quirk
    Latency / throughput / memory / long-runtail-latency (p95/p99) + soak + thresholdperformance
    Depends on network / disk / subprocess / other servicefault injection (delay + abort)error-handling

    See references/technique-cheatsheet.md for how to apply each.


Phase 2 — Generate scenarios via the Triple (or a gap)

For each requirement, walk its technique and try to emit one or more Triples.

  • EP+BVA: emit min, just-below-min (invalid), nominal, just-below-max, max, just-above-max (invalid). Six Triples from one numeric requirement.
  • Decision table: one Triple per reachable flag combination; mark impossible combos.
  • State-transition: one Triple per legal edge AND per illegal edge (event fired in a state that shouldn't accept it).
  • Async/convergence: assert the eventual invariant and the intermediate states, never "element is visible after N ms".
  • Performance: state the workload, the metric (p95/p99/RSS), the threshold, and the measurement window. No threshold in spec → gap.
  • Fault injection: for each dependency, a delay Triple and an abort Triple; assert retry/timeout/degradation behaviour.

When a slot won't fill → STOP generating that scenario. Record a gap with the unfillable slot named (see Gate). Do not invent the missing value.


Phase 3 — The clarification Gate (configurable)

Whether an unfillable Triple blocks or just annotates depends on stage:

text
   stage = proposal | design   →  HARD gate
   stage = apply               →  SOFT gate
   (no --stage)                →  infer: tasks.md absent ⇒ proposal/design (hard)
                                          tasks.md present ⇒ apply (soft)
  • HARD gate: collect all gaps, then call ask_user with decision-forcing questions and STOP. Do not write test-plan.md until answered. The spec is not yet testable; clarify before locking scenarios.
  • SOFT gate: write the scenario row with a [NEEDS CLARIFICATION: <slot> — <question>] marker, continue, and list all markers in a banner at the top of test-plan.md.

Decision-forcing question rules (from ambiguity-detection practice):

  • Name the missing slot and why it blocks a scenario.
  • Offer concrete candidate answers, never propose a solution/implementation.
  • One question per genuine decision; do not pad.

Example: "Restart quiesce window: tasks say bridges 'suppress auto-start for the quiesce window'. To test the boundary I need the exact value — is it 5s (restart) / 60s (shutdown) per AGENTS.md, or spec-defined elsewhere? Without a number I cannot write the just-after-window re-spawn scenario."


Show full SKILL.md (594 more words)Show less

Phase 4 — Route each scenario to a test level

Every scenario carries a level tag fixing where it would be authored, and a disposition (automated | manual-only). Map each scenario's nature to one of your project's actual test levels — the routing method is fixed; the level names and paths are yours to fill. Do not assume a level/harness the project lacks.

Scenario natureRoute to the project level that is…
pure logic / boundary / decision table / pure statethe fast in-process unit tier
process / install / spawn / multi-OS runtimethe process/CLI smoke tier (NO rendered-UI asserts)
rendered UI / WS-driven view / convergence / quirkthe browser/e2e tier
micro perf (fn-level)the unit tier, timed
process/load perf, soakthe smoke tier (or a dedicated perf harness)
aesthetics / hardware / "feels right" / subjectivemanual-only → no fold, no test task (disposition=manual-only, level —)

Keep the rendered-UI-vs-smoke boundary sacred: a UI-visible assertion never lives in a process/CLI smoke row.

Example — pi-agent-dashboard levels (this repo's concrete routing; other projects substitute their own). Honour the AGENTS.md hard rule: rendered-UI assertions are Playwright only; qa/ stays CLI/process smoke.

text
   ┌────────────────────────────┬──────────────────────────────────────────┐
   │ Scenario nature            │ Level → location                          │
   ├────────────────────────────┼──────────────────────────────────────────┤
   │ pure logic / boundary /    │ L1 unit  → packages/*/src/**/__tests__/   │
   │ decision table / state pure│            *.test.ts (vitest)             │
   │ process / install / spawn  │ L2 smoke → qa/tests/*.sh|*.ps1            │
   │  / multi-OS runtime        │            (NO rendered-UI asserts)       │
   │ rendered UI / WS-driven    │ L3 e2e   → tests/e2e/*.spec.ts            │
   │  view / convergence / quirk│   (Playwright vs docker harness port †)   │
   │ micro perf (fn-level)      │ L1 unit (timed)                           │
   │ process/load perf, soak    │ L2 smoke (or dedicated harness)           │
   │ aesthetics / hardware /    │ manual-only → no fold, no test task       │
   │  "feels right" / subjective │   (disposition=manual-only, level —)      │
   └────────────────────────────┴──────────────────────────────────────────┘

† The docker e2e harness port is NOT a fixed :18000 — docker/test-up.sh hash-derives a free port per worktree and records it in .pi-test-harness.json (dashboardPort). An L3 scenario's observable is read against that derived port; never hardcode :18000.

manual-only routing outcome (additive to L1/L2/L3): a scenario whose expected observable is a human judgment with no automatable signal — visual aesthetics, a hardware behaviour, "feels right / looks correct", subjective UX — is NOT routed to a test level. Its manifest row records disposition: manual-only (level —), and no test task is folded for it; it is deferred to post-merge manual verification by ship-change. Every routable scenario keeps its L1/L2/L3 level and disposition: automated — this outcome only diverts the truly un-automatable rows; existing L1/L2/L3 logic is unchanged.

If a scenario implies a brand-new level/harness, flag it in the plan's "New infra needed" section rather than silently assuming it exists.


Phase 5 — Write test-plan.md

Write the test-plan.md to your change/spec's test-plan location (OpenSpec: openspec/changes/<name>/test-plan.md) using references/test-plan-schema.md. It is a standalone catalog, separate from tasks.md. Each scenario is a numbered row with: id, class, technique, level, disposition (automated | manual-only), the full Triple, and (soft gate) any clarification marker. The disposition column is mandatory on every row — it is the manifest's source-of-truth signal that the fold step (in plan-proposal) and the defer rule (in ship-change) both read.

End with a short offer (do not auto-act): "Want me to fold these into the ## Tests / ## Validate sections of tasks.md as checklist items?" — folding is a separate, explicit step.


Guardrails

  • Never invent a missing value to make a scenario "work" — that hides the spec gap this skill exists to expose.
  • Never write app/test code here — this skill drafts the catalog. Authoring the actual *.test.ts / *.spec.ts is implementation (use implement / openspec-apply-change).
  • Don't downgrade scenarios to smoke to make them easy. A scenario that only checks "it exists / exit 0" belongs in qa/ smoke already — this skill's output is the layer above that.
  • Honour the level boundary — no rendered-UI assertion in a qa/ smoke row.
  • Offer, don't auto-fold into tasks.md.

References

  • references/technique-cheatsheet.md — how to apply each ISTQB + resilience technique, with project-specific examples.
  • references/test-plan-schema.md — exact test-plan.md layout.

© BlackBeltTechnology, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in packages/eng-disciplines/.pi/skills/scenario-design of BlackBeltTechnology/pi-agent-dashboard.

  • SKILL.md
  • references/technique-cheatsheet.md
  • references/test-plan-schema.md

Open the folder on GitHubat commit 7a2d171

Compare with similar skills

Scenario Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scenario Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scenario Design this skillBlackBeltTechnology/pi-agent-dashboard315—~2.8kAutomated safety check: PassMIT
Test Planquran/quran.com-frontend-next1.9k—~1.4kAutomated safety check: PassNone
Test Strategynurettincoban/ai-prd-workflow298—~1.4kAutomated safety check: PassMIT
Atmos Testscloudposse/atmos1.4k—~1.9kAutomated safety check: PassApache-2.0
Exploratory Testtobihagemann/turbo408—~2kAutomated safety check: PassMIT
Exploratory Testtobihagemann/turbo408—~2kAutomated safety check: PassMIT

Similar skills

  • Test Plan

    quran/quran.com-frontend-next

    Generates a comprehensive testing plan based on the current branch changes or a specific PR.

    1.9k GitHub stars~1.4k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Test Strategy

    nurettincoban/ai-prd-workflow

    Write TEST-STRATEGY.md, a test plan per RFC, before the tests are written.

    298 GitHub stars~1.4k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Atmos Tests

    cloudposse/atmos

    Author Atmos smoke tests, post-deployment stack checks, and integration tests using type: test, HTTP assertions, shell or script/interpreter checks, require gates, parallel dependencies, matrix…

    1.4k GitHub stars~1.9k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Exploratory Test

    tobihagemann/turbo

    Execute multi-level exploratory testing of the app covering basic functionality, complex operations, adversarial testing, and cross-cutting scenarios, plus usability observations through a UX lens…

    408 GitHub stars~2k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Exploratory Test

    tobihagemann/turbo

    Execute multi-level exploratory testing of the app covering basic functionality, complex operations, adversarial testing, and cross-cutting scenarios, plus usability observations through a UX lens…

    408 GitHub stars~2k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Risk Based Testing

    petrkindlmann/qa-skills

    Produce a risk matrix or heatmap that quantifies what could break by business impact × probability, runs failure mode analysis on the top items, and maps test coverage to risk zones.

    170 GitHub stars~5.3k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed

More from BlackBeltTechnology/pi-agent-dashboard

All 70 skills in this repo
  • Browser

    BlackBeltTechnology/pi-agent-dashboard

    Browser automation via the agent-browser CLI. An agent skill from BlackBeltTechnology/pi-agent-dashboard.

    316 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • CI Troubleshoot

    BlackBeltTechnology/pi-agent-dashboard

    Diagnose failed GitHub Actions runs for pi-agent-dashboard: the 11-file workflow taxonomy, affected-test selection, the release pipeline, known failure modes, and how to read gh run logs and…

    316 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Debug Dashboard

    BlackBeltTechnology/pi-agent-dashboard

    Diagnose problems in the running pi-agent-dashboard system: server.log, /api/health, bridge WebSocket connectivity, vitest triage, known-issue FAQ entries.

    316 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Implement

    BlackBeltTechnology/pi-agent-dashboard

    Disciplined implementation in pi-agent-dashboard: the rebuild matrix (extension→reload, server→restart, client→build+restart, openspec-apply→full rebuild) plus the project's code discipline rules.

    316 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Pi Dashboard

    BlackBeltTechnology/pi-agent-dashboard

    Monitor and control the pi-dashboard server. An agent skill from BlackBeltTechnology/pi-agent-dashboard.

    316 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Session To Guideline

    BlackBeltTechnology/pi-agent-dashboard

    Turn a pi session into a Markdown "how-we-did-it" collaboration guideline: reads the session's JSONL transcript and synthesizes a reusable playbook of which prompts worked, what had to be steered…

    316 GitHub stars~3.2k tokensUpdated today
    Auto-check passed

Categories

Questions about Scenario Design

What does Scenario Design do?

Draft real-life test SCENARIOS (not smoke tests) from a change/feature spec. Scenario Design is an agent skill from BlackBeltTechnology/pi-agent-dashboard. Draft real-life test SCENARIOS (not smoke tests) from a change/feature spec.

When should I use Scenario Design?

Scenario Design fits situations like: tasks that involve Test generation; tasks that involve PRD writing; tasks that involve QA and bug reports.

How do I install Scenario Design in Claude Code?

Run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill scenario-design -a claude-code`. Or copy the skill folder (packages/eng-disciplines/.pi/skills/scenario-design in BlackBeltTechnology/pi-agent-dashboard) into .claude/skills/scenario-design in your project. Claude Code loads it when a task matches its description.

How do I install Scenario Design in Codex?

Run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill scenario-design -a codex`. Or copy the skill folder (packages/eng-disciplines/.pi/skills/scenario-design in BlackBeltTechnology/pi-agent-dashboard) into .agents/skills/scenario-design in your project. Codex loads it when a task matches its description.

Can I use Scenario Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill scenario-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario-design, .gemini/skills/scenario-design, .github/skills/scenario-design and .opencode/skills/scenario-design in your project.

What does Scenario Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Scenario Design is instructions for the agent only. Our summary lists: Docker. Compatibility (from SKILL.md): Optional: OpenSpec change spec as input. Reads a change/feature spec; writes test-plan.md..

Does Scenario Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scenario Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scenario Design use?

Scenario Design is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scenario Design use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.

What are the alternatives to Scenario Design?

Skills that share tags, products or a category with Scenario Design: Test Plan (quran/quran.com-frontend-next, 1.9k stars), Test Strategy (nurettincoban/ai-prd-workflow, 298 stars), Atmos Tests (cloudposse/atmos, 1.4k stars) and Exploratory Test (tobihagemann/turbo, 408 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scenario Design?

BlackBeltTechnology (a GitHub organization) maintains it in BlackBeltTechnology/pi-agent-dashboard, which has 315 GitHub stars. The repository holds 70 skills in this directory. The repository was last updated on October 10, 2026.

Source: BlackBeltTechnology/pi-agent-dashboard on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.