Agent skill

Branch Dogfooding QA Agent

by EveryInc in EveryInc/compound-engineering-plugin

Drives a real browser through every user-visible change on the active branch, fixing small breakages with regression tests until it is ready.

MITAuto-check passedTesting & QA

Install Branch Dogfooding QA Agent

skills CLI
$ npx skills add EveryInc/compound-engineering-plugin --skill ce-dogfood -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EveryInc/compound-engineering-plugin ce-dogfood --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EveryInc/compound-engineering-plugin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ce-dogfood .claude/skills/ce-dogfood && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ce-dogfood
GitHub stars
25k
Token cost
~1.9k tokens
SKILL.md length
1,064 words
Files
6 (incl. scripts, references)
Skills in repo
37
Repo updated
First seen
Licence
MIT

At a glance

Drives a real browser through every user-visible change on the active branch, fixing small breakages with regression tests until it is ready.

  • QA-testing every change on a branch before calling it ready
  • SKILL.md covers Boundaries, Prerequisites, Artifact Root and Delegation, plus 2 more sections
  • Runs Python scripts from its folder; calls npx, rails and npm
  • Judging a branch's user experience against product personas

What it does

Scoped to the diff rather than the whole app, it tests only what the current branch introduced or changed against trunk, following detailed phases in a reference file starting from a scope-setting first phase; it is invoked manually rather than running on its own.

Every matrix scenario must end as passed, fixed, skipped or a terminal blocked state, the project's automated test suite must be run once with its result recorded, and a dated dogfood report must be finalized against its template and committed. A clean matrix over a failing suite still finalizes as not ready, since chasing that suite green is not this run's job. The browser is driven only through a dedicated CLI binary, never a browser MCP tool, and screenshots go to OS temp rather than the repo root unless one is embedded in the final report.

When your agent uses it

  • QA-testing every change on a branch before calling it ready
  • Judging a branch's user experience against product personas
  • Producing a dated dogfood report for a pull request

Example prompts

  • “Dogfood this branch against trunk and fix any small breakages you find.”
  • “Run the dogfood pass on PR 482 and write the report.”
  • “Check whether this branch is ready according to the dogfood matrix.”

Requirements

  • The `agent-browser` CLI
  • A diffable branch or PR target against trunk

What it can do on your machine

Read from SKILL.md and the folder at commit 51aa537. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • rails
    • npm
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, npm and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Branch Dogfooding QA Agent loads about 1.9k tokens when it runs, and up to ~9.6k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 1,064 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from EveryInc/compound-engineering-plugin at commit 51aa537, republished under its MIT licence (© EveryInc). 1,064 words, ~1,925 tokens.

Download SKILL.mdSave it as .claude/skills/ce-dogfood/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
ce-dogfood
description
Hands-off, diff-scoped browser QA of the active branch: maps user flows, drives a real browser, autonomously fixes small breakages with regression tests and commits, judges experience against product personas, and writes a durable dogfood report. Manual invocation only.
disable-model-invocation
true
argument-hint
[PR number, branch name, or blank for current branch] [--port PORT]

Dogfood

Act as a QA engineer who dogfoods the active branch end-to-end, autonomously, until it is genuinely ready.

Outcome: every user-visible change this branch introduced has been driven in a real browser along its whole journey, judged for correctness and for how it feels to the product's personas, with small breakages fixed, regression-tested, and committed. Done: every matrix scenario is Pass, Fixed, Skipped, or in a terminal Blocked state; the project's automated suite has been run once and its result recorded; and the report at <root>/dogfood-reports/<YYYY-MM-DD>-<branch-slug>-dogfood.md is finalized against its template and committed. A green matrix over a red suite finalizes as a not-ready verdict rather than a ready one. Chasing that suite green is not this run's job.

This is diff-scoped, not whole-app exploration. You test what this branch introduced or modified versus the trunk.

Read references/phases.md before Phase 0 (Scope) and follow it. It defines every phase in detail; the run cannot be executed correctly from the phase list below.

Boundaries

  • Drive the browser exclusively through the agent-browser CLI — never Chrome MCP tools (mcp__claude-in-chrome__*), another browser MCP, or a built-in browser-control tool, even when the platform offers one. Use the direct binary, never npx agent-browser (the direct binary uses the fast Rust client).
  • Never dogfood the trunk on a branch-name or blank target — there is no diff. A PR target always has a base, so it is always diffable even when its head branch is named main.
  • A numeric target stays a PR identity through isolation and checkout — never collapse it to its head ref, whose name may itself be main.
  • Never switch the primary checkout out from under the user. This skill decides only whether to offer isolation — no for a blank or current-branch target (you are already on it), yes for a PR or another named ref — and ce-worktree handles the mechanics and reports the verdict. On a declined offer, check the target out in place, confirming first if uncommitted changes would be disturbed.
  • Screenshots and other transient artifacts go to OS temp (mktemp -d "${TMPDIR:-/tmp}/ce-dogfood-XXXXXX"), never the repo root; copy one in only to embed it in the report.
  • Auto-fix only what is small, well-understood, and low-risk. A change that needs an architectural or schema decision, alters product behavior or UX intent, spans many files, or has plausible competing solutions is escalated to the report's Decisions for a human section, never implemented to clear a matrix item.

Prerequisites

User-runnable invocation rendering. In prerequisite failures, default to /ce-setup and /ce-dogfood <original arguments>; use $ce-setup and $ce-dogfood <original arguments> only when the active host is Codex or explicitly documents dollar-prefixed skill invocation. On oh-my-pi (omp), use /skill:ce-setup and /skill:ce-dogfood <original arguments>. Render only each invocation as inline code and output one form only.

  • A local dev server you can start (bin/dev, rails server, npm run dev, etc.).

  • agent-browser installed. Check:

    bash
    command -v agent-browser >/dev/null 2>&1 && echo "Ready" || echo "NOT INSTALLED"

    If not installed, stop and tell the user to install agent-browser: print the rendered ce-setup invocation for the current install command, followed by the rendered ce-dogfood <original arguments> invocation to retry. This workflow cannot function without it.

Artifact Root

Reports live under <root>/dogfood-reports/ and personas under <root>/personas/. Resolve <root> the first time you compose any <root>/ path, whether you are reading or writing, and never before. A run that composes none skips it.

<!-- ce-docs-root:start -->

Resolve the CE artifact root <root> before composing any artifact path.

  • Read docs_root from <repo-root>/.compound-engineering/config.yaml only (<repo-root> = git rev-parse --show-toplevel). Do not read it from config.local.yaml. Unset -> <root> is docs, exactly as before.
  • Validate a set value: a repo-relative directory whose real, symlink-resolved path stays inside the repo and is neither the repo root nor under .git/. Otherwise stop with an error naming docs_root and the value -- never fall back to docs.
  • Use <root> as the sole artifact location: create it if absent, compose each path as <root>/<subdir> with this skill's own subdirectory, and never also read docs.
<!-- ce-docs-root:end -->
Show full SKILL.md (420 more words)Show less

Delegation

ce-dogfood is an orchestrator: prefer an existing CE skill over re-deriving its behavior. Isolate a PR or named-branch target with ce-worktree; take a non-obvious root cause to ce-debug; commit each fix with ce-commit; capture a reusable lesson with ce-compound.

Compound Packs

The repo's declared Compound Packs supply personas and criteria. A pack rule that describes a user is a persona the flows are walked as; one that prescribes how the product must look or behave is a criterion each scenario it reaches is judged against, and a contradiction is a failure that enters the fix loop with its (pack: <id>, <path within the pack>) citation. A contradiction the branch intends is a decision for a human about the rule, not a fix. Judgments that generalize go back to the packs through ce-compound; this skill never writes a pack. references/phases.md defines how packs are applied.

Phase order

Scope -> analyze the diff -> map the flows -> derive the matrix -> serve -> execute -> fix loop -> report. The order is the invariant: the flow model precedes the matrix, and the matrix precedes any browser work. Each phase's conditions are in references/phases.md; read it before Phase 0 (Scope) rather than reconstructing a phase from this line. Work one scenario at a time, judged for correctness and for how it feels to each persona. A fix is not done until a regression test fails before it and passes after, or the report says why no automated test was meaningful.

Checkpoint, not a final write. Create the report from references/dogfood-report-template.md as soon as the matrix exists, with every scenario at Pending, and update it after each scenario is judged and each fix is committed. <branch-slug> is the branch name lowercased, with every run of non-alphanumeric characters — slashes included — collapsed to one -. Find a prior run by globbing <root>/dogfood-reports/*-<branch-slug>-dogfood.md. The task list is session-scoped, but the report on disk is what a later run or a teammate resumes from, so an interrupted run must leave a template-shaped checkpoint rather than a bare matrix.

Terminal states. Blocked (needs human verify) (a step that needs an outside interaction, such as OAuth, real email, payments, or SMS, and cannot be driven headlessly) and Blocked (human decision) (a fix too big to make autonomously) both wait on a person, and each ends that scenario, not the run: continue the rest of the matrix, and never silently re-queue a blocked scenario, on this run or on resume. How a person is reached differs per state, and the phase that sets the state says which.

© EveryInc, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/ce-dogfood of EveryInc/compound-engineering-plugin.

  • SKILL.md
  • agents/openai.yaml
  • references/dogfood-report-template.md
  • references/phases.md
  • references/test-matrix-taxonomy.md
  • scripts/packs-resolve.py

Open the folder on GitHubat commit 51aa537

Compare with similar skills

Branch Dogfooding QA Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Branch Dogfooding QA Agent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Branch Dogfooding QA Agent this skillEveryInc/compound-engineering-plugin25k—~1.9kAutomated safety check: PassMIT
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
Write and Verify Playwright Testsappsmithorg/appsmith41k—~2.9kAutomated safety check: NotesApache-2.0
playwright-cli Browser Automationgithub/gh-aw5.4k25 repos~2.8kAutomated safety check: PassMIT
Cucumber and Playwright E2E Testslanggenius/dify158k—~682Automated safety check: PassCustom licence
E2E Testinglangflow-ai/langflow155k—~3.3kAutomated safety check: PassMIT

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.

    41k GitHub stars~2.9k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Official

    Drives a real browser from the command line with playwright-cli to open pages, interact, mock requests, save state and work with Playwright tests.

    5.4k GitHub starsUsed in 25 repos~2.8k tokens
    Testing & QAAuto-check passed
  • Guides changes and reviews of the Cucumber and Playwright end-to-end suite under `e2e/`: feature files, step definitions, support code, tags, locators and assertions.

    158k GitHub stars~682 tokensUpdated today
    Testing & QAAuto-check passed
  • E2E Testing

    langflow-ai/langflow

    Write and review Playwright E2E tests for Langflow. An agent skill from langflow-ai/langflow.

    155k GitHub stars~3.3k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Handsontable Playwright E2E Tests

    handsontable/handsontable

    Guides writing and changing Playwright end-to-end tests for Handsontable using page objects, data-testid hooks and deterministic waits.

    22k GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from EveryInc/compound-engineering-plugin

All 37 skills in this repo
  • Compound Learning Writer

    EveryInc/compound-engineering-plugin

    Records one solved and verified problem as a durable learning in the repository, but only when the reasoning is not already clear from the final code, tests or docs.

    25k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Compound Learnings Refresh

    EveryInc/compound-engineering-plugin

    Audits a repo's stored learnings against the current codebase, fixes stale, overlapping or superseded docs and reports on every document.

    25k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Compound Engineering Prototype

    EveryInc/compound-engineering-plugin

    Builds a throwaway prototype at just the fidelity needed to settle a specific how-it-should-work-or-feel question, before committing to an approach other work will treat as fixed.

    25k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Compound Engineering Setup

    EveryInc/compound-engineering-plugin

    Checks Compound Engineering plugin health and repo-local config, or scaffolds a Compound Pack when you ask for one by id.

    25k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • PR Babysitter

    EveryInc/compound-engineering-plugin

    Watches an open GitHub pull request over time, routing review comments and CI failures to other skills until the PR is ready to merge.

    25k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • CE Brainstorm

    EveryInc/compound-engineering-plugin

    Turns a vague or ambitious feature idea into a requirements-only plan through dialogue with you, sized to the work, before any code is written.

    25k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Categories

Questions about Branch Dogfooding QA Agent

What does Branch Dogfooding QA Agent do?

Drives a real browser through every user-visible change on the active branch, fixing small breakages with regression tests until it is ready. Scoped to the diff rather than the whole app, it tests only what the current branch introduced or changed against trunk, following detailed phases in a reference file starting from a scope-setting first phase; it is invoked manually rather than running on its own.

When should I use Branch Dogfooding QA Agent?

Branch Dogfooding QA Agent fits situations like: QA-testing every change on a branch before calling it ready; judging a branch's user experience against product personas; producing a dated dogfood report for a pull request.

How do I install Branch Dogfooding QA Agent in Claude Code?

Run `npx skills add EveryInc/compound-engineering-plugin --skill ce-dogfood -a claude-code`. Or copy the skill folder (skills/ce-dogfood in EveryInc/compound-engineering-plugin) into .claude/skills/ce-dogfood in your project. Claude Code loads it when a task matches its description.

How do I install Branch Dogfooding QA Agent in Codex?

Run `npx skills add EveryInc/compound-engineering-plugin --skill ce-dogfood -a codex`. Or copy the skill folder (skills/ce-dogfood in EveryInc/compound-engineering-plugin) into .agents/skills/ce-dogfood in your project. Codex loads it when a task matches its description.

Can I use Branch Dogfooding QA Agent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EveryInc/compound-engineering-plugin --skill ce-dogfood -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ce-dogfood, .gemini/skills/ce-dogfood, .github/skills/ce-dogfood and .opencode/skills/ce-dogfood in your project.

What does Branch Dogfooding QA Agent need to run?

Going by SKILL.md and its folder, Branch Dogfooding QA Agent needs Python for the scripts in its folder and the command-line tools its instructions call (npx, rails, npm and git). Our summary lists: The `agent-browser` CLI; A diffable branch or PR target against trunk.

Does Branch Dogfooding QA Agent access the network?

SKILL.md contains no URLs. Its commands use npx, npm and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Branch Dogfooding QA Agent safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Branch Dogfooding QA Agent use?

Branch Dogfooding QA Agent is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Branch Dogfooding QA Agent use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.6k tokens, read only when the agent opens those files.

What are the alternatives to Branch Dogfooding QA Agent?

Skills that share tags, products or a category with Branch Dogfooding QA Agent: Web Application Testing (anthropics/skills, 180k stars), Write and Verify Playwright Tests (appsmithorg/appsmith, 41k stars), playwright-cli Browser Automation (github/gh-aw, 5.4k stars) and Cucumber and Playwright E2E Tests (langgenius/dify, 158k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Branch Dogfooding QA Agent?

EveryInc (a GitHub organization) maintains it in EveryInc/compound-engineering-plugin, which has 25,461 GitHub stars. The repository holds 37 skills in this directory. The repository was last updated on October 10, 2026.

Source: EveryInc/compound-engineering-plugin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.