Agent skill

Spec Dogfood

by leo-kuang-ai in leo-kuang-ai/spec-first

Hands-off, diff-scoped browser QA of the active branch or PR.

MITAuto-check: notesTesting & QA

Install Spec Dogfood

skills CLI
$ npx skills add leo-kuang-ai/spec-first --skill spec-dogfood -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install leo-kuang-ai/spec-first spec-dogfood --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/spec-dogfood .claude/skills/spec-dogfood && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spec-dogfood
GitHub stars
107
Token cost
~6.1k tokens
SKILL.md length
3,069 words
Files
14 (incl. references)
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

Hands-off, diff-scoped browser QA of the active branch or PR.

  • Works in 7 steps: Scope and Get on the Right Branch → Analyze Changes → Map the Flows, Then Build the Matrix → …
  • A branch needs autonomous user-flow dogfooding before review
  • SKILL.md covers Workflow Contract Summary, Use The Browser Execution Owner, Prerequisites and Reusing Spec-First Skills, plus 2 more sections
  • Runs Shell and JavaScript scripts from its folder; calls git, gh and rails

What it does

Spec Dogfood is an agent skill from leo-kuang-ai/spec-first. Hands-off, diff-scoped browser QA of the active branch or PR. Use when a branch needs autonomous user-flow dogfooding before review or shipping: map changed flows, delegate exact-origin browser execution to spec-test-browser, fix small breakages with regression tests, record human-decision blockers, and write a durable report. Do not use for collaborative UI polish, ordinary browser smoke tests, code review, implementation planning, or broad whole-app exploration.

Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files, including reference files (for example `evals/cases/owner-unavailable-stops.yaml`, `evals/cases/polish-request-routes-out.yaml` and `evals/cases/r2-code-review-routes-out.yaml`).

It sits in Testing & QA, covering QA and bug reports, UI design and UX design. The repository describes itself as: 仓库原生 AI Coding Harness —— 把一次性 AI 对话变成可治理、可验证、可沉淀的工程闭环 · spec-first.cn. The licence is MIT.

When your agent uses it

  • A branch needs autonomous user-flow dogfooding before review
  • Shipping: map changed flows
  • Delegate exact-origin browser execution to spec-test-browser
  • Fix small breakages with regression tests

Example prompts

  • “/spec-dogfood”

Requirements

  • Node.js
  • A Bash shell

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Scope and Get on the Right Branch
  2. Analyze Changes
  3. Map the Flows, Then Build the Matrix
  4. Detect Port and Start the Dev Server
  5. Execute the Matrix
  6. Fix Loop (Authorization-Aware)
  7. Write the Report Artifact

What it can do on your machine

Read from SKILL.md and the folder at commit 74655dc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell and JavaScript, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • gh
    • rails
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, gh and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Spec Dogfood loads about 6.1k tokens when it runs, and up to ~8.1k if it reads all its reference files. Until then it costs about 120 tokens; SKILL.md has 3,069 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~6.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:181
    tructions > `package.json` dev script > `.env*` `PORT=` > default `3000`). If a server is already listening on it, reuse

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from leo-kuang-ai/spec-first at commit 74655dc, republished under its MIT licence (© leo-kuang-ai). 3,069 words, ~6,126 tokens.

Download SKILL.mdSave it as .claude/skills/spec-dogfood/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
spec-dogfood
description
Hands-off, diff-scoped browser QA of the active branch or PR. Use when a branch needs autonomous user-flow dogfooding before review or shipping: map changed flows, delegate exact-origin browser execution to spec-test-browser, fix small breakages with regression tests, record human-decision blockers, and write a durable report. Do not use for collaborative UI polish, ordinary browser smoke tests, code review, implementation planning, or broad whole-app exploration.
disable-model-invocation
true
argument-hint
[PR number, branch name, or blank for current branch] [--port PORT]

Dogfood

Act as a QA engineer who dogfoods the active branch end-to-end: understand every change, test every change in a real browser as a user would, and fix small breakages autonomously until the branch has a clear readiness verdict.

This is diff-scoped, not whole-app exploration. You test what this branch introduced or modified versus the trunk.

Workflow Contract Summary

When To Use

Use when a PR, branch, or current non-trunk branch needs autonomous browser dogfooding before review or shipping: changed-flow mapping, persona-aware journey testing, small fixes, regression tests, and a durable report.

When Not To Use

Do not use for collaborative UI polish (spec-polish), ordinary browser smoke tests (spec-test-browser when delegated), static code review (spec-code-review), implementation planning (spec-plan), broad whole-app exploration, or large product/architecture decisions. When routing out for one of these, name the listed destination skill explicitly in your reply — recognizing the mismatch without naming where the request belongs leaves the owner without a route.

Inputs

A PR number, branch name, or current branch; optional --port; git diff against trunk; project dev-server conventions; persona/strategy docs when present; browser observations and test results.

Outputs

Incrementally updated dogfood report under docs/dogfood-reports/, flowcharts, test matrix statuses, explicitly authorized small source fixes with regression evidence, blocked authorization/human-decision items, and a final readiness verdict.

Artifacts

docs/dogfood-reports/<YYYY-MM-DD>-<branch-slug>-dogfood.md, authorized source/test changes, commits only when separately requested, transient screenshots in OS temp, and optional reusable learnings handed to spec-compound.

Failure Modes

Trunk target with no diff, unsafe checkout or dirty working tree, missing spec-test-browser execution owner/capability, missing or failing dev server, external-interaction flows needing human verification, ambiguous fixes requiring human product/architecture decisions, or failing automated suite after browser matrix completion.

Workflow

Resolve the target branch/PR, optionally isolate with spec-worktree, analyze the diff, map changed user flows, build a matrix, start the app, execute each scenario through spec-test-browser, fix only small unambiguous issues, update the report throughout, then run the automated suite and finalize the verdict.

Downstream Consumers

Human reviewers, spec-code-review, spec-work for larger follow-up fixes, spec-compound for reusable learnings, and PR/commit workflows that consume the readiness evidence.

Use The Browser Execution Owner

This workflow never executes a browser CLI directly. Invoke spec-test-browser with mode:pipeline and an explicit exact loopback target-origin:<origin> so its unique wrapper owns capability probing, request-time exact-origin enforcement, action validation, private evidence, and cleanup. Do not use Chrome MCP tools, other browser-control tools, or hand-built browser argv as a second execution path.

Prerequisites

  • A local dev server you can start (bin/dev, rails server, npm run dev, etc.).
  • The internal spec-test-browser Skill is available to own browser execution. Do not probe or execute its private CLI directly; its mode:pipeline call returns the authoritative capability/exact-origin result. If the owner is unavailable, stop with: "Browser execution owner unavailable. Run spec-runtime-setup to inspect browser readiness, then rerun spec-dogfood. This does not block spec-first baseline."

Reusing Spec-First Skills

spec-dogfood is an orchestrator. Prefer delegating to existing Spec-First skills over re-deriving their behavior:

WhenSkillWhy
Phase 0 isolationspec-worktreeRun the dogfood in an isolated worktree so the main checkout stays clean.
A failure's root cause is non-obviousspec-debugSystematic root-cause analysis instead of guess-and-check.
Authorized commit checkpointspec-commitCreate a consistent, well-scoped commit only after separate commit authorization.
A bug reveals a reusable lessonspec-compoundCapture the learning so the team compounds knowledge.

Mutation Authority Boundary

Before isolation/checkout, browser execution, or the first source fix, derive five independent run-local facts from the current user request and any visible upstream handoff:

yaml
branch_mutation_authorization: authorized | missing
browser_effect_authorization: authorized | missing
local_fix_authorization: authorized | missing
commit_authorization: authorized | missing
landing_authorization: authorized | missing
  • A PR/branch argument selects the dogfood target; branch-selection-is-not-authorization. Branch/worktree mutation requires the current user or upstream owner to explicitly request the exact checkout/isolation action, or the user to approve it after disclosure.
  • Classify every planned browser flow by expected effect as read-only | ephemeral-local | durable-local | external | unknown, regardless of whether the triggering action looks like navigation, click, form submit, or a key press. Read-only and ephemeral-local flows may enter the pipeline with synthetic data. Durable-local, external, and unknown flows require separate browser_effect_authorization: authorized; when it is missing, record browser_effect_authorization_missing, do not put the step in a browser test plan, and keep the scenario blocked. Even when such authority exists, use only behavior the spec-test-browser owner admits; its pipeline refusal remains authoritative.
  • A request to inspect, QA, or dogfood does not by itself authorize source fixes. Set local_fix_authorization: authorized only when the current user/upstream explicitly requests applying small fixes; otherwise keep source findings report-only and record fix_authorization_missing.
  • Set commit_authorization: authorized only when commit creation is separately explicit. A verified fix may remain uncommitted with commit_authorization_missing.
  • Set landing_authorization: authorized only for an explicit push/PR request. Without landing authorization, do not push and do not open a PR.
  • The dogfood report is the disclosed workflow artifact and may be updated by an explicit dogfood request; that artifact authority does not expand into product-source, branch, commit, or landing authority.

Workflow

0. Scope        Resolve the target; change checkout only with branch authorization
1. Analyze      Diff branch vs trunk, understand every change
2. Map+Matrix   Map user flows as Mermaid flowcharts, then derive the test matrix as a task list
3. Serve        Detect port and start the caller-owned dev server
4. Execute      Work the matrix through spec-test-browser mode:pipeline
5. Fix loop     On failure: authorized fix -> regression proof -> optional commit -> continue
6. Report       Write durable doc to docs/dogfood-reports/ (flows, matrix, fixes, learnings, verdict)
Phase 0: Scope and Get on the Right Branch

Parse the invocation arguments supplied by the current host: a PR number, a branch name, or blank (use current branch). Preserve quoted paths/tokens while stripping a recognized --port PORT pair if present.

  1. Identify the target — keep PR identity; do not switch the working tree yet.
    • PR number: the target is the PR — carry the number through every later step (trunk check, isolation, checkout). Read its head only for display (gh pr view <number> --json headRefName,isCrossRepository), but do not reduce it to a bare branch name: a fork PR's head can even be named main/master. Do not check out yet.
    • Branch name: the target is that branch.
    • Blank: the target is the current branch.
  2. Refuse to run on the trunk — branch/blank targets only. If a branch-name or blank target resolves to the trunk (main/master/the detected default), stop — there is no diff to dogfood. A PR is always diffable (it has a base), so this check never applies to a PR target; never refuse spec-dogfood <number> just because the PR's head branch happens to be named main.
  3. Decide isolation by what you're testing; let spec-worktree own the worktree mechanics. Do not re-derive worktree detection or creation here — spec-worktree handles existing-isolation detection, attaching to a ref, and the "already checked out" constraint, and reports its decision back. The target remains scope until branch mutation is authorized:
    • Blank / current-branch target: do not isolate — dogfood in place. You are already on the branch under test, any separately authorized fixes belong in this checkout, and git cannot check the same branch out in a second worktree anyway. (If you happen to already be in a worktree, that is fine — you are simply dogfooding here.)
    • A PR or a different named branch: offer the concrete isolation/checkout choices with their side effects. Choosing an option is the authorization source for that exact action. On authorized isolation, invoke spec-worktree existing-ref mode through its bundled script: isolate pr:<number> for PR targets, or isolate <branch> for branch targets. It may return already_checked_out branch=<name> path=<path>; use that existing checkout and never switch the primary checkout. On an explicitly authorized in-place switch, use gh pr checkout <number> or git checkout <branch> after dirty-state safety checks. If no blocking interaction exists or the user declines both mutations, stop with branch_mutation_authorization_missing; do not silently switch the primary checkout.
  4. Resume if a prior run exists. Look for an existing report at docs/dogfood-reports/*-<branch-slug>-dogfood.md (see the branch-slug rule under Resumability). If one is found with unfinished scenarios, ask whether to resume it or start fresh. To resume, re-hydrate the task list from its matrix: Pass/Fixed/Skipped stay done; Pending and in_progress become the remaining auto-runnable work. The three Blocked states are not auto-runnable — Blocked (fix authorization), Blocked (needs human verify), and Blocked (human decision) all wait on a person, so surface them and ask how to proceed rather than silently re-queuing them.
Resumability (stop and return at any point)

This workflow is designed to be interrupted and resumed. Two pieces of state make that safe:

  • The task list (the harness's task tool — TaskCreate/TaskUpdate on Claude Code, update_plan on Codex, or the equivalent elsewhere) is the live to-do — one task per matrix scenario. Mark each in_progress when you start it and completed only when it genuinely passes.
  • The report doc at docs/dogfood-reports/<YYYY-MM-DD>-<branch-slug>-dogfood.md is the durable checkpoint that survives across sessions. <branch-slug> is the branch name lowercased with every run of non-alphanumeric characters (slashes included) collapsed to a single - (e.g. feature/Foo_Bar -> feature-foo-bar). Create it as soon as the matrix exists (end of Phase 2) by instantiating references/dogfood-report-template.md (read that template now if you haven't) so the checkpoint carries the template-owned section shape from the start — then fill in every scenario at Pending, and update it incrementally after each scenario judgment and each verified fix/commit-status change, not only at the end. An interrupted run must leave a template-shaped checkpoint, not a bare matrix.

Because tasks are session-scoped but the report doc is on disk, the report is the source of truth for resuming. Always keep the two in sync so a later run (or a teammate) can pick up exactly where this one stopped.

Phase 1: Analyze Changes

Derive the trunk ref once, then pull the full diff against it and read it. Do not hard-code main — a repo whose default branch is master (or anything else) would fail with fatal: ambiguous argument 'main...HEAD'.

bash
# Resolve the trunk to a ref that actually exists. Start from the detected
# default name (origin/HEAD, then gh), then fall back to common names. For each
# candidate prefer a local branch; else use the remote-tracking ref QUALIFIED as
# origin/<branch> — an unqualified name resolves via refs/remotes/<name>, NOT
# refs/remotes/origin/<name>, so a remote-only trunk would otherwise miss. This
# qualification applies to the detected default too (PR/CI checkouts often have
# only origin/main, no local main).
DEFAULT=$(git symbolic-ref --quiet --short refs/remotes/origin/HEAD 2>/dev/null | sed 's@^origin/@@')
DEFAULT=${DEFAULT:-$(gh repo view --json defaultBranchRef -q .defaultBranchRef.name 2>/dev/null)}
TRUNK=""
for cand in "$DEFAULT" main master; do
  [ -n "$cand" ] || continue
  if git show-ref --verify --quiet "refs/heads/$cand"; then
    TRUNK=$cand; break
  elif git show-ref --verify --quiet "refs/remotes/origin/$cand"; then
    TRUNK="origin/$cand"; break
  fi
done
TRUNK=${TRUNK:-main}

git diff --name-only "$TRUNK...HEAD"   # what changed
git diff "$TRUNK...HEAD"               # how it changed

Build a mental model of every change: new features, modified behavior, new routes/views/components, touched data flows. Note anything that produces user-visible behavior — that is what the matrix must cover.

Ground in the product's personas and vision. Look for persona and vision context so flows can be judged from real users' eyes, not just "does it work." Check, in order: STRATEGY.md (its "Who it's for" section names the primary persona and their job-to-be-done), VISION.md, and any persona docs (e.g. docs/personas/, PERSONAS.md). Capture the 1-3 primary personas and what each cares about. If none exist, infer a reasonable primary persona from the product and the diff, and say so in the report.

Phase 2: Map the Flows, Then Build the Matrix

Do not jump straight to a flat list of pages. First understand the user flows the diff touches, then derive the matrix from them. A matrix built without a flow model tests pages in isolation and misses the journey — the email that "sends" but lands in the wrong thread.

2a. Map the user flows (required)

For every user-visible change, trace the complete journey end to end and draw it. Map each flow as a Mermaid flowchart so the journey is explicit and reviewable before any testing happens — entry point, each user action, branch points (success / validation error / empty / permission-denied), side effects (emails, jobs, notifications), and the true end state.

Email example: it's not enough that "an email sends." Does it go to the right recipient? When the user clicks through, does the app land on and scroll to the right message? Does the content make sense? Does the whole flow align with the product's vision and UX? The flowchart must carry the click-through and its destination, not stop at "email sent."

mermaid
flowchart TD
    A[User opens /threads] --> B[Clicks 'Reply']
    B --> C{Form valid?}
    C -->|No| D[Inline validation error shown]
    C -->|Yes| E[Reply saved]
    E --> F[Notification email sent to thread participants]
    E --> G[UI scrolls to new reply, focus on it]
    F --> H[Recipient clicks email link]
    H --> I{Lands on correct thread + scrolls to the reply?}

Produce one flowchart per distinct journey, scaled to the diff: a one-route or copy-only change gets a single small flowchart, a multi-step feature gets several. Cover the happy path and the branch points (error, empty, boundary, permission). Mapping the flows before the matrix is never skipped — these diagrams ARE the understanding; they become the spine of the matrix and belong in the final report.

Show full SKILL.md (1,184 more words)Show less
2b. Derive the matrix from the flows

Walk each flowchart and turn every node and branch into one or more test scenarios. Read references/test-matrix-taxonomy.md for the full set of dimensions (journeys, functional checks, experiential checks, edge/error/empty states, accessibility, responsiveness). Cover both functional ("does it work?") and experiential ("does it feel right and align with the product?").

Map changed files to concrete routes (views -> their pages, components -> pages rendering them, layouts -> all pages, stylesheets -> visual regression on key pages) and attach those routes to the flows that exercise them.

Load the matrix as a task list (the harness's task tool, as above), one task per scenario, so progress is tracked and nothing is skipped. Order tasks by flow, following the flowcharts, not by file.

Phase 3: Detect Port and Start the Dev Server

Determine the port (priority: explicit --port > a port explicitly stated in your in-context project instructions > package.json dev script > .env* PORT= > default 3000). If a server is already listening on it, reuse it. Otherwise start the project's dev command (bin/dev, rails server, npm run dev, etc.) in the background and poll the port until it accepts connections. This skill is hands-off, so start the server automatically without asking. Freeze a credential-free exact loopback origin such as http://127.0.0.1:${PORT} for the owner call; do not open the browser here, accept redirects as a replacement origin, or infer browser readiness from the listener alone.

Phase 4: Execute the Matrix

Work the task list one item at a time. For each scenario, mark the task in_progress, then:

  1. Document what you're testing (the journey, expected outcome, exact target origin, and expected-effect classification).

  2. Execute through the owner by invoking spec-test-browser with mode:pipeline, the current branch/PR selector, and target-origin:<exact-loopback-origin>. Supply only routes and expected-safe synthetic interactions derived from the flow. If the owner returns target-origin-*, exact-origin/conformance failure, browser-mutation-authorization-required, cleanup failure, or another not_run / not_supported result, preserve that reason and do not fall back to a direct browser command. Keep transient screenshots in owner-private temp evidence; copy one into the report only when intentionally embedding it.

  3. Judge both correctness and experience: right data, right destination, sensible content, no console errors, and does it feel aligned with the product?

  4. Walk it as each persona. Re-run the journey in your head from each primary persona's perspective (from Phase 1) and ask where they'd feel a paper cut — a small friction that wouldn't fail a functional test but degrades the experience: a confusing label, an extra click, an unexpected jump, a slow-feeling step, missing feedback, copy that doesn't match how that persona thinks. A scenario can be functionally Pass yet still carry paper cuts. Note each paper cut, which persona feels it, and its severity.

  5. Record pass/fail plus any paper cuts, with specifics. Mark the task completed only when it genuinely passes. Paper cuts do not block a Pass, but a sharp paper cut (one severe enough to fix now) is routed into the Phase 5 fix loop just like a failure — apply the same auto-fix-vs-escalate judgment to it. Log the rest in the report.

External-interaction flows (OAuth, real email delivery, payments, SMS) can't be fully driven headlessly — pause, ask the user to verify that leg, and mark the scenario Blocked (needs human verify) until they confirm. Then continue.

Phase 5: Fix Loop (Authorization-Aware)

When a scenario fails — or a passing scenario carries a sharp paper cut worth fixing now — first decide whether local fix authority exists and whether the change is small enough for this workflow.

Judge the size of the fix before touching code. A fix is eligible only when separately authorized and small, well-understood, and low-risk: a clear bug with an obvious correct fix, contained to a few files, with no schema/architecture/product trade-off. Even with local fix authority, do not implement a large or ambiguous change that requires an architectural, schema, product, or UX decision; record it for a human instead.

For authorized fixes:

When local_fix_authorization: missing, do not touch product source. Record the issue and fix_authorization_missing, mark the scenario Blocked (fix authorization), and continue with other independent scenarios. Do not relabel an unfixed failure as Pass.

  1. Investigate the root cause. If it's non-obvious, use spec-debug.
  2. Apply the fix in the code.
  3. Add an automated regression test that fails before the fix and passes after, so the bug can't return. This is the default for behavioral and code bugs. When an automated test is genuinely impractical — a pure copy, spacing, or visual fix with no behavioral assertion to make — substitute a documented browser-replay or screenshot check and state in the report why no automated test was meaningful. Do not invent a hollow test just to satisfy the step.
  4. If commit_authorization: authorized, commit the fix with a clear message using spec-commit, one logical fix per commit. Otherwise leave the verified fix uncommitted, record commit_authorization_missing, and use uncommitted in the report's Commit field.
  5. Re-run the failing scenario in the browser to confirm it now passes; then continue the matrix.
  6. If the bug carried a reusable lesson, capture it with spec-compound.

For changes too big to make autonomously: do not implement. Record it in the report's Decisions for a human section with: what's broken, why it's not a safe autonomous fix, the options you see (with trade-offs), and your recommendation. Mark the scenario Blocked (human decision) in the matrix, then continue with the rest. Never make a large, irreversible, or product-altering change just to clear a matrix item.

Keep iterating until every task is completed or in a terminal Blocked state — Blocked (fix authorization), Blocked (human decision), or Blocked (needs human verify). All three wait on a person, so do not re-queue them. Re-test anything a fix might have affected.

Before declaring the branch ready, run the project's automated test suite once (the new regression tests plus everything that already exists). Discover the test command from the project's active instructions and conventions already in your context — do not assume a specific runner. Record the result in the report; a green matrix with a red suite is not "ready."

Phase 6: Write the Report Artifact

The report doc was created at the end of Phase 2 and updated incrementally throughout (see Resumability). When the matrix is green (or every remaining item is explicitly blocked), finalize it at docs/dogfood-reports/<YYYY-MM-DD>-<branch-slug>-dogfood.md in the repo under test, then surface a short summary in chat with the file path.

Finalize against references/dogfood-report-template.md — the same template the Phase 2 checkpoint was instantiated from, which owns the required sections and what each must carry. Confirm every template-owned section is present and complete; do not reconstruct the section list from memory, as that drifts from the template. Carry forward the cross-phase obligations this skill produced: the Mermaid flowcharts from Phase 2a, a matrix row per scenario with its commit SHA or uncommitted, each fix's root cause and the regression test added (or why none was meaningful), authorization blockers, paper cuts attributed by persona, learnings worth feeding to spec-compound, and a final readiness verdict that records the Phase 5 automated-suite result. Without landing authorization, do not push and do not open a PR.

© leo-kuang-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (references) in skills/spec-dogfood of leo-kuang-ai/spec-first.

  • SKILL.md
  • evals/cases/owner-unavailable-stops.yaml
  • evals/cases/polish-request-routes-out.yaml
  • evals/cases/r2-code-review-routes-out.yaml
  • evals/cases/trunk-target-rejected.yaml
  • evals/eval.yaml
  • evals/fixtures/repos/mini-ledger/README.md
  • evals/fixtures/repos/mini-ledger/package.json
  • evals/fixtures/repos/mini-ledger/src/server.js
  • evals/fixtures/scripts/asks-a-question.sh
  • evals/fixtures/scripts/check-owner-unavailable-stop.sh
  • evals/fixtures/scripts/check-trunk-reject.sh
  • references/dogfood-report-template.md
  • … and 1 more

Open the folder on GitHubat commit 74655dc

Compare with similar skills

Spec Dogfood next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Spec Dogfood compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Spec Dogfood this skillleo-kuang-ai/spec-first107—~6.1kAutomated safety check: NotesMIT
Write Fix Briefn1m21n/Infinite271—~2.9kAutomated safety check: PassCustom licence
Feature Verifysd0xdev/sd0x-harness192—~3.2kAutomated safety check: NotesMIT
Pre-Merge Checkjsmastery-pro/skills1.5k—~1.1kAutomated safety check: NotesMIT
Actionbook Web Testactionbook/actionbook1.6k—~9.7kAutomated safety check: PassApache-2.0
Fix The Classjoetawil7/first-pass98—~1.5kAutomated safety check: PassMIT

Similar skills

  • Write Fix Brief

    n1m21n/Infinite

    Turn a bug report, review findings, broken-UI screenshot or new-node idea into a verified, file/line-precise implementation prompt after checking the real code.

    271 GitHub stars~2.9k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Feature Verify

    sd0xdev/sd0x-harness

    Feature verification (READ-ONLY, P0-P5). An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~3.2k tokensUpdated 3 days ago
    Testing & QAAuto-check: notes
  • Pre-Merge Check

    jsmastery-pro/skills

    A gate before merge: verify runs the real app against the spec, and review has a different model do a senior code review, without editing code.

    1.5k GitHub stars~1.1k tokensUpdated 2 mo ago
    DevelopmentAuto-check: notes
  • Actionbook Web Test

    actionbook/actionbook

    Run browser-based web tests against websites using Actionbook CLI.

    1.6k GitHub stars~9.7k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Fix The Class

    joetawil7/first-pass

    Bug-fix routine that fixes the whole class of bug, not just the reported instance.

    98 GitHub stars~1.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Sdd Audit

    madebyaris/spec-kit-command-cursor

    Compare implementation against specifications, identify gaps and issues.

    197 GitHub stars~405 tokensUpdated 3 mo ago
    DevelopmentAuto-check passed

More from leo-kuang-ai/spec-first

All 35 skills in this repo
  • Spec App Consistency Audit

    leo-kuang-ai/spec-first

    Audit mobile App PRD/Figma/local-source consistency across page routes, KMP/Clean Architecture, components, analytics, i18n, engineering quality, and industry lenses before runtime validation; use…

    107 GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Handoff

    leo-kuang-ai/spec-first

    Create a durable cross-session handoff or resume from a user-selected continuity source.

    107 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Pov

    leo-kuang-ai/spec-first

    Give a decisive, project-grounded verdict on an external input — judged against the current project, not in the abstract.

    107 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Resolve PR Feedback

    leo-kuang-ai/spec-first

    Resolve PR review feedback by evaluating validity and fixing issues with conflict-aware resolver dispatch.

    107 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check: notes
  • Spec Riffrec Feedback Analysis

    leo-kuang-ai/spec-first

    Analyze explicit Riffrec product-feedback captures, including riffrec-.zip, the Riffrec session.json + events.json + recording.webm + voice.webm bundle, or media/notes the user identifies as a…

    107 GitHub stars~1.4k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Compound

    leo-kuang-ai/spec-first

    Document a recently solved problem or durable project vocabulary in docs/solutions/ or CONCEPTS.md.

    107 GitHub stars~18k tokensUpdated 2 days ago
    Auto-check passed

Questions about Spec Dogfood

What does Spec Dogfood do?

Hands-off, diff-scoped browser QA of the active branch or PR. Spec Dogfood is an agent skill from leo-kuang-ai/spec-first. Hands-off, diff-scoped browser QA of the active branch or PR.

When should I use Spec Dogfood?

Spec Dogfood fits situations like: A branch needs autonomous user-flow dogfooding before review; shipping: map changed flows; delegate exact-origin browser execution to spec-test-browser; fix small breakages with regression tests.

How do I install Spec Dogfood in Claude Code?

Run `npx skills add leo-kuang-ai/spec-first --skill spec-dogfood -a claude-code`. Or copy the skill folder (skills/spec-dogfood in leo-kuang-ai/spec-first) into .claude/skills/spec-dogfood in your project. Claude Code loads it when a task matches its description.

How do I install Spec Dogfood in Codex?

Run `npx skills add leo-kuang-ai/spec-first --skill spec-dogfood -a codex`. Or copy the skill folder (skills/spec-dogfood in leo-kuang-ai/spec-first) into .agents/skills/spec-dogfood in your project. Codex loads it when a task matches its description.

Can I use Spec Dogfood in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add leo-kuang-ai/spec-first --skill spec-dogfood -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spec-dogfood, .gemini/skills/spec-dogfood, .github/skills/spec-dogfood and .opencode/skills/spec-dogfood in your project.

What does Spec Dogfood need to run?

Going by SKILL.md and its folder, Spec Dogfood needs a shell and JavaScript for the scripts in its folder and the command-line tools its instructions call (git, gh, rails and npm). Our summary lists: Node.js; A Bash shell.

Does Spec Dogfood access the network?

SKILL.md contains no URLs. Its commands use git, gh and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Spec Dogfood safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Spec Dogfood use?

Spec Dogfood is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Spec Dogfood use?

About 6.1k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Spec Dogfood?

Skills that share tags, products or a category with Spec Dogfood: Write Fix Brief (n1m21n/Infinite, 271 stars), Feature Verify (sd0xdev/sd0x-harness, 192 stars), Pre-Merge Check (jsmastery-pro/skills, 1.5k stars) and Actionbook Web Test (actionbook/actionbook, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Spec Dogfood?

leo-kuang-ai (a GitHub user) maintains it in leo-kuang-ai/spec-first, which has 107 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.

Source: leo-kuang-ai/spec-first on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.