Agent skill

Walkthrough

by andrew-yangy in andrew-yangy/gru-ai

Cognitive walkthrough — simulate real user scenarios against the current system to find gaps between ideal and actual.

MITAuto-check passedAgent Workflows

Install Walkthrough

skills CLI
$ npx skills add andrew-yangy/gru-ai --skill walkthrough -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install andrew-yangy/gru-ai walkthrough --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/andrew-yangy/gru-ai.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/walkthrough .claude/skills/walkthrough && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
walkthrough
GitHub stars
155
Token cost
~3.4k tokens
SKILL.md length
679 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Cognitive walkthrough — simulate real user scenarios against the current system to find gaps between ideal and actual.

  • Works in 6 steps: Load Scenarios → Design the Ideal (per scenario) → Trace the Actual (per scenario) → …
  • Agent Workflows work in your project
  • SKILL.md covers Role Resolution, Step 1: Load Scenarios, Step 2: Design the Ideal (per… and Step 3: Trace the Actual (per…, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Walkthrough is an agent skill from andrew-yangy/gru-ai. Cognitive walkthrough — simulate real user scenarios against the current system to find gaps between ideal and actual. Takes an optional scenario name or 'all' to run standing scenarios. Run after major directives or periodically as a reality check.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows. The repository describes itself as: Autonomous AI agent team for one-man companies. Context engineering + harness engineering drive a pipeline that brainstorms, builds, reviews, and ships. The licence is MIT.

When your agent uses it

  • Agent Workflows work in your project

Example prompts

  • “/walkthrough”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Load Scenarios
  2. Design the Ideal (per scenario)
  3. Trace the Actual (per scenario)
  4. Synthesize Gaps
  5. Present to CEO
  6. Save Report

What it can do on your machine

Read from SKILL.md and the folder at commit 8fba479. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Walkthrough loads about 3.4k tokens when it runs. Until then it costs about 65 tokens; SKILL.md has 679 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from andrew-yangy/gru-ai at commit 8fba479, republished under its MIT licence (© andrew-yangy). 679 words, ~3,351 tokens.

Download SKILL.mdSave it as .claude/skills/walkthrough/SKILL.md (or your agent's skills folder).
name
walkthrough
description
Cognitive walkthrough — simulate real user scenarios against the current system to find gaps between ideal and actual. Takes an optional scenario name or 'all' to run standing scenarios. Run after major directives or periodically as a reality check.

Walkthrough — Cognitive Walkthrough

Role Resolution

Read .claude/agent-registry.json to map roles to agent names. Use each agent's id as the subagent_type when spawning. The CPO designs the ideal experience; the CTO traces the actual implementation.


Simulate user scenarios against the current system. Find what's broken, missing, or surprising.

The pattern: For each scenario, design what SHOULD happen (ideal), trace what DOES happen (actual), report the gaps.

Arguments: $ARGUMENTS

  • A specific scenario name (e.g., ceo-runs-directive) → run just that one
  • all → run all standing scenarios
  • A free-text scenario description (e.g., "seller wants to see competitor prices") → ad-hoc walkthrough
  • Empty → list available scenarios and ask which to run

Step 1: Load Scenarios

If $ARGUMENTS is a scenario name or "all":

Read standing scenarios from .context/lessons/scenarios.md.

Each scenario has:

  • Name: slug identifier
  • Actor: who is performing the action (CEO, seller, shopper, developer)
  • Trigger: what starts the flow ("CEO types /directive improve-security")
  • Goal: what the actor wants to achieve
  • Critical path: the steps that MUST work for the scenario to succeed

If all, load all scenarios. If a specific name, load just that one.

If $ARGUMENTS is free text:

Treat it as an ad-hoc scenario. Spawn the CPO to formalize it:

You are the CPO. The CEO described a user scenario informally:

"{$ARGUMENTS}"

Formalize it into this structure:
{
  "name": "slug-name",
  "actor": "who is doing this",
  "trigger": "what starts the flow",
  "goal": "what the actor wants to achieve",
  "critical_path": [
    "Step 1: what should happen first",
    "Step 2: what should happen next",
    ...
  ],
  "success_criteria": "how do you know the scenario succeeded"
}

Think from the ACTOR's perspective, not the system's. What does the actor expect at each step? What would surprise or frustrate them?

CRITICAL OUTPUT FORMAT: First character must be `{`, last must be `}`. JSON only.
If $ARGUMENTS is empty:

Read .context/lessons/scenarios.md and list available scenarios:

Available scenarios:
1. ceo-runs-directive — CEO issues a directive, wants it handled without blocking
2. ceo-morning-review — CEO opens dashboard, wants to know what happened overnight
3. ...

Which scenario to walk through? (or describe a new one)

Use AskUserQuestion with the scenario names as options.

Step 2: Design the Ideal (per scenario)

For each scenario, spawn the CPO to design the ideal experience — what SHOULD happen if everything worked perfectly.

The CPO receives:

  • Their personality file
  • The scenario definition
  • .context/vision.md — so the ideal aligns with the north star
  • .context/preferences.md — CEO expectations
You are the CPO. You are designing the IDEAL user experience for this scenario. Don't look at the current implementation — design from scratch what the perfect flow would be.

SCENARIO:
- Actor: {actor}
- Trigger: {trigger}
- Goal: {goal}

For each step of the critical path, describe:
1. What the actor does
2. What the system should do in response
3. What the actor sees/experiences
4. How long it should take (instant / seconds / minutes / async)
5. What would frustrate the actor at this step

Then describe the END STATE: what does "success" look like from the actor's perspective?

Think like a product designer, not an engineer. The actor doesn't care about checkpoints, worktrees, or JSON schemas. They care about: did the thing work? Was it fast? Did I have to babysit it?

CRITICAL OUTPUT FORMAT: First character must be `{`, last must be `}`. JSON only.

{
  "scenario": "{name}",
  "ideal_flow": [
    {
      "step": 1,
      "actor_action": "what the actor does",
      "system_response": "what should happen",
      "actor_experience": "what they see/feel",
      "timing": "instant | seconds | minutes | async",
      "frustration_risk": "what could annoy the actor here"
    }
  ],
  "end_state": "what success looks like",
  "key_expectations": ["the non-negotiable things the actor expects"]
}

Step 3: Trace the Actual (per scenario)

For each scenario, spawn the CTO to trace what ACTUALLY happens in the current system. The CTO reads code, config, and skill files to follow the real execution path.

The CTO receives:

  • Their personality file
  • The scenario definition
  • The CPO's ideal flow (from Step 2)
  • .context/lessons/ topic files — known issues
  • .context/preferences.md
You are the CTO. You are tracing what ACTUALLY happens in the current system for this scenario. Read the real code and config — don't guess.

SCENARIO:
- Actor: {actor}
- Trigger: {trigger}
- Goal: {goal}
- Critical path: {steps}

IDEAL FLOW (from the CPO):
{CPO's ideal_flow JSON}

For each step of the ideal flow, trace what the current system actually does:
1. Read the relevant files (SKILL.md, agent files, code)
2. Follow the execution path step by step
3. Note where reality matches the ideal
4. Note where reality DIVERGES from the ideal
5. Note where the system does NOTHING (missing functionality)

Be thorough. Grep for entry points, read the actual instructions, trace the branching logic. Don't assume — verify.

### Doc Consistency Checks

After tracing the execution flow, check the pipeline's internal
documentation for consistency. These checks catch drift between docs
that causes real pipeline failures — fields referenced in one file but
undefined in another, prompt templates injecting stale field names,
validation scripts that don't enforce what the docs promise.

Run these three checks by reading actual file contents (use Grep and
Read). Do NOT guess from memory.

**A. Cross-reference verification.** For each pipeline step doc in
`.claude/skills/directive/docs/pipeline/`, check that any
directive.json or project.json field it references actually exists in
the corresponding schema doc under
`.claude/skills/directive/docs/reference/schemas/` (especially
`directive-json.md` and `plan-schema.md`). Example: if
`09-execute-projects.md` reads `directive.planning.coo_plan`, confirm
`directive-json.md` defines `planning.coo_plan`.

**B. Schema-to-template alignment.** For each prompt template in
`.claude/skills/directive/docs/reference/templates/`, check that the
fields it injects (placeholders like `{field_name}` or references to
JSON paths) match the schema definitions in
`.claude/skills/directive/docs/reference/schemas/`. Example: if
`planner-prompt.md` injects `{audit.risk_areas}`, confirm
`audit-output.md` defines `risk_areas`.

**C. Validation script coverage.** For each validation script in
`.claude/hooks/validate-*.sh`, check that the fields it validates
match what the pipeline docs and schemas claim are enforced. Example:
if `07-plan-approval.md` says "the gate validates `projects` array
exists", confirm `validate-gate.sh` actually checks for that field.
Also check the inverse — if a doc says a field is "required" or
"enforced", a validation script should check it.

Record every discrepancy. Omit checks that pass — only report
mismatches.

CRITICAL OUTPUT FORMAT: First character must be `{`, last must be `}`. JSON only.

{
  "scenario": "{name}",
  "actual_flow": [
    {
      "ideal_step": 1,
      "ideal_expectation": "what the CPO said should happen",
      "actual_behavior": "what the system actually does",
      "status": "match | diverge | missing | broken",
      "evidence": "file:line or config entry that proves this",
      "notes": "explanation of the gap, if any"
    }
  ],
  "gaps_found": [
    {
      "id": "gap-slug",
      "severity": "critical | major | minor | cosmetic",
      "type": "missing | broken | wrong | slow | confusing",
      "description": "what's wrong",
      "ideal": "what should happen",
      "actual": "what does happen",
      "evidence": "file:line",
      "suggested_fix": "how to close the gap"
    }
  ],
  "doc_consistency": {
    "cross_ref_issues": [
      {
        "source_file": "pipeline doc that references the field",
        "references": "the field or path referenced",
        "expected_in": "schema doc where it should be defined",
        "issue": "missing | renamed | wrong_path"
      }
    ],
    "schema_drift": [
      {
        "template_file": "prompt template that injects the field",
        "injects": "field name or path the template uses",
        "schema_file": "schema doc that should define it",
        "issue": "field missing from schema | field renamed | type mismatch"
      }
    ],
    "validation_gaps": [
      {
        "script": "validate-*.sh script name",
        "doc_claims": "what the pipeline doc says is enforced",
        "doc_source": "pipeline doc making the claim",
        "issue": "script does not check this | script checks stale field name"
      }
    ]
  },
  "working_well": ["things that match the ideal — acknowledge what's good"]
}

Step 4: Synthesize Gaps

After all scenarios are traced, consolidate the findings:

  1. Deduplicate — the same gap may appear in multiple scenarios
  2. Prioritize — critical gaps that block the actor's goal come first
  3. Cross-reference — gaps that appear in 2+ scenarios are systemic
  4. Classify effort — quick fix (< 1 hour), medium (half day), large (1+ days)
  5. Consolidate doc_consistency — merge doc_consistency findings across all scenario traces. Deduplicate cross_ref_issues, schema_drift, and validation_gaps that appear in multiple traces (same source_file + references pair, same template_file + injects pair, or same script + doc_claims pair). Flag any issue that appears in 3+ traces as systemic drift. Keep one canonical entry per unique issue with a found_in list of scenario names.
Show full SKILL.md (248 more words)Show less

Step 5: Present to CEO

# Walkthrough Report — {date}

## Scenarios Walked: {count}

### {Scenario Name}
**Actor**: {actor} | **Goal**: {goal}

**Ideal vs Actual:**
| Step | Ideal | Actual | Status |
|------|-------|--------|--------|
| 1 | {ideal} | {actual} | ✅ match / ⚠️ diverge / ❌ missing |
| 2 | ... | ... | ... |

**Gaps Found: {count}**
- [{severity}] **{description}** — {type}
  Ideal: {what should happen}
  Actual: {what does happen}
  Fix: {suggested fix} ({effort})

**Working Well:**
- {things that matched the ideal}

(repeat per scenario)

## Systemic Gaps (appear in 2+ scenarios)
- **{gap}** — found in: {scenario list}

(if doc_consistency findings exist across any scenario trace, include this section)

## Doc Consistency
Issues found by cross-checking pipeline docs, schemas, templates, and
validation scripts. These cause silent pipeline failures when docs drift
out of sync.

### Cross-Reference Mismatches
| Source File | References | Expected In | Issue |
|-------------|-----------|-------------|-------|
| {source_file} | {references} | {expected_in} | {issue} |

### Schema-Template Drift
| Template File | Injects | Schema File | Issue |
|---------------|---------|-------------|-------|
| {template_file} | {injects} | {schema_file} | {issue} |

### Validation Coverage Gaps
| Script | Doc Claims | Doc Source | Issue |
|--------|-----------|------------|-------|
| {script} | {doc_claims} | {doc_source} | {issue} |

(omit any subsection whose table would be empty)

## Summary
- Total gaps: {count} ({critical}, {major}, {minor})
- Scenarios fully passing: {count}/{total}
- Top 3 fixes by impact: {list}

Then ask the CEO:

  • "Create directive from gaps" — bundle gaps into a directive in directives/
  • "Add to backlog" — write gaps to the relevant goal's backlog
  • "Note only" — just keep the report

Step 6: Save Report

Write the full report to .context/reports/walkthrough-{date}.md

If gaps were approved as a directive, create it in .context/directives/.

Standing Scenarios File

If .context/lessons/scenarios.md doesn't exist, create it with starter scenarios on first run. The CEO and team add scenarios over time as new flows become important.

Failure Handling

SituationAction
The CPO can't formalize ad-hoc scenarioAsk CEO to clarify the scenario
The CTO can't find the entry point for a stepMark as "missing — no implementation found"
A scenario has no gapsReport it as passing — this is good news
scenarios.md doesn't existCreate it with starter scenarios, then run

Rules

NEVER
  • Skip the ideal design (Step 2) — the whole point is comparing ideal vs actual
  • Have the same agent design ideal AND trace actual — separate perspectives prevent bias
  • Mark a gap as "minor" if it blocks the actor's goal — that's critical by definition
  • Trace the actual by reading docs/comments — read the real code/config
ALWAYS
  • Design ideal BEFORE tracing actual — don't let current state constrain the ideal
  • Include evidence (file:line) for every gap — no hand-waving
  • Acknowledge what's working well — not just gaps
  • Save the report even if no gaps found (it's a health signal)
  • Use the CPO for ideal (product thinking) and the CTO for actual (technical tracing)

© andrew-yangy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/walkthrough of andrew-yangy/gru-ai.

Open the folder on GitHubat commit 8fba479

Compare with similar skills

Walkthrough next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Walkthrough compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Walkthrough this skillandrew-yangy/gru-ai155—~3.4kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
Hook Development for Claude Code Pluginsanthropics/claude-plugins-official38k10 repos~4.1kAutomated safety check: NotesApache-2.0
Using Superpowersfarm-fe/farm5.6k35 repos~1.4kAutomated safety check: PassMIT
Executing Plans Inlineobra/superpowers297k2 repos~5.1kAutomated safety check: PassMIT
Skill CreatorAzure/azqr79589 repos~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Hook Development for Claude Code Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.

    38k GitHub starsUsed in 10 repos~4.1k tokens
    Agent WorkflowsAuto-check: notes
  • Using Superpowers

    farm-fe/farm

    A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions

    5.6k GitHub starsUsed in 35 repos~1.4k tokens
    Agent WorkflowsAuto-check passed
  • Executing Plans Inline

    obra/superpowers

    Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.

    297k GitHub starsUsed in 2 repos~5.1k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    795 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed

More from andrew-yangy/gru-ai

All 9 skills in this repo
  • Code Review Excellence

    andrew-yangy/gru-ai

    Provides comprehensive code review guidance for React 19, Vue 3, Rust, TypeScript, Java, Python, and C/C++.

    155 GitHub stars~1.7k tokensUpdated 7 mo ago
    Auto-check: notes
  • SEO Audit

    andrew-yangy/gru-ai

    Full website SEO audit with parallel subagent delegation. An agent skill from andrew-yangy/gru-ai.

    155 GitHub stars~731 tokensUpdated 7 mo ago
    Auto-check passed
  • Brainstorm

    andrew-yangy/gru-ai

    Structured brainstorm — from quick Socratic refinement to full C-suite strategy sessions.

    155 GitHub stars~3.2k tokensUpdated 7 mo ago
    Auto-check passed
  • Healthcheck

    andrew-yangy/gru-ai

    Internal codebase and operations health check — the CTO scans technical health, the COO checks operational health.

    155 GitHub stars~2.4k tokensUpdated 7 mo ago
    Auto-check passed
  • Report

    andrew-yangy/gru-ai

    CEO dashboard with progressive disclosure — 3 tiers: headline (5 lines, default), summary (per-goal detail), deep (full weekly analysis).

    155 GitHub stars~4k tokensUpdated 7 mo ago
    Auto-check passed
  • Smoke Test

    andrew-yangy/gru-ai

    Pipeline end-to-end smoke test -- creates a trivial directive, runs it through /directive, validates every pipeline step, and reports pass/fail with evidence.

    155 GitHub stars~941 tokensUpdated 7 mo ago
    Auto-check passed

Categories

Questions about Walkthrough

What does Walkthrough do?

Cognitive walkthrough — simulate real user scenarios against the current system to find gaps between ideal and actual. Walkthrough is an agent skill from andrew-yangy/gru-ai. Cognitive walkthrough — simulate real user scenarios against the current system to find gaps between ideal and actual.

When should I use Walkthrough?

Walkthrough fits situations like: agent Workflows work in your project.

How do I install Walkthrough in Claude Code?

Run `npx skills add andrew-yangy/gru-ai --skill walkthrough -a claude-code`. Or copy the skill folder (.claude/skills/walkthrough in andrew-yangy/gru-ai) into .claude/skills/walkthrough in your project. Claude Code loads it when a task matches its description.

How do I install Walkthrough in Codex?

Run `npx skills add andrew-yangy/gru-ai --skill walkthrough -a codex`. Or copy the skill folder (.claude/skills/walkthrough in andrew-yangy/gru-ai) into .agents/skills/walkthrough in your project. Codex loads it when a task matches its description.

Can I use Walkthrough in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add andrew-yangy/gru-ai --skill walkthrough -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/walkthrough, .gemini/skills/walkthrough, .github/skills/walkthrough and .opencode/skills/walkthrough in your project.

What does Walkthrough need to run?

SKILL.md names no scripts, command-line tools or credentials: Walkthrough is instructions for the agent only.

Does Walkthrough access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Walkthrough safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Walkthrough use?

Walkthrough is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Walkthrough use?

About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Walkthrough?

Skills that share tags, products or a category with Walkthrough: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 297k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Walkthrough?

andrew-yangy (a GitHub user) maintains it in andrew-yangy/gru-ai, which has 155 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on March 11, 2026.

Source: andrew-yangy/gru-ai on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.