Agent skill

Tmux Manual QA Worker

by code-yeongyu in code-yeongyu/senpi

Runs one manual QA scenario for the todo continuation feature in the real ./pi-test.sh CLI inside tmux, captures scrollback and checks a deterministic count marker.

MITAuto-check passedTesting & QA

Install Tmux Manual QA Worker

skills CLI
$ npx skills add code-yeongyu/senpi --skill tmux-manual-qa -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install code-yeongyu/senpi tmux-manual-qa --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/code-yeongyu/senpi.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.factory/skills/tmux-manual-qa .claude/skills/tmux-manual-qa && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tmux-manual-qa
GitHub stars
472
Token cost
~1.6k tokens
SKILL.md length
662 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Runs one manual QA scenario for the todo continuation feature in the real ./pi-test.sh CLI inside tmux, captures scrollback and checks a deterministic count marker.

  • Works in 8 steps: Orient → Set up the fixture (if needed) → Launch tmux session and drive the CLI → …
  • Verifying the todo continuation behavior by hand in an interactive TUI
  • SKILL.md covers Context you MUST read before…, Hard rules, Prerequisites and Procedure, plus 2 more sections
  • Calls rg and npm

What it does

This is a worker skill for a single manual QA feature listed in features.json. It drives the real ./pi-test.sh TUI in a tmux session, saves the scrollback to a log in the gitignored local-ignore folder, and treats a count of the SYSTEM DIRECTIVE: SENPI marker, written to a .count file beside the log, as the pass or fail signal instead of how the screen looks.

Before starting, the agent reads the feature spec, the mission document, the validation contract and the library notes on architecture and user testing, then confirms that pi-test.sh runs, the project has been built, and tmux and rg are installed. Scenarios include default-enabled, settings-disable, flag-override and re-entry-guard. Temporary .pi/settings.json fixtures are deleted after capture, real LLM calls are allowed, the auth file is never touched, and source files stay unchanged: bugs go back to the orchestrator.

When your agent uses it

  • Verifying the todo continuation behavior by hand in an interactive TUI
  • Capturing scrollback evidence for a manual-qa milestone feature
  • Testing a settings or flag override scenario and cleaning up its fixture afterwards

Example prompts

  • “Run the settings-disable manual QA scenario for the todo continuation feature.”
  • “Capture the re-entry-guard scenario in tmux and save the count file.”
  • “Check that the flag-override scenario passes and report the marker count.”

Requirements

  • The repository's ./pi-test.sh CLI and a completed npm run build
  • tmux and ripgrep (rg) installed
  • Configured LLM credentials for real model calls

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Orient
  2. Set up the fixture (if needed)
  3. Launch tmux session and drive the CLI
  4. Capture scrollback
  5. Assert the count marker
  6. Tear down
  7. Verify evidence files
  8. Handoff

What it can do on your machine

Read from SKILL.md and the folder at commit ee6ece1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • rg
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tmux Manual QA Worker loads about 1.6k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 662 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from code-yeongyu/senpi at commit ee6ece1, republished under its MIT licence (© code-yeongyu). 662 words, ~1,561 tokens.

Download SKILL.mdSave it as .claude/skills/tmux-manual-qa/SKILL.md (or your agent's skills folder).
name
tmux-manual-qa
description
Run a single manual tmux-based QA scenario for the todo continuation feature against the real CLI (./pi-test.sh) in an interactive TUI. Captures scrollback, asserts deterministic pass/fail count markers, and cleans up test fixtures. Use only for the manual-qa milestone features.

Tmux Manual QA Worker

You are executing ONE manual QA feature from features.json that drives the real ./pi-test.sh CLI inside a tmux session, captures scrollback, and asserts deterministic pass/fail markers.

Context you MUST read before starting

  1. Feature spec: features.json — your assigned feature.
  2. Mission document: mission.md.
  3. Validation contract: the fulfills IDs for your feature in validation-contract.md.
  4. Architecture: .factory/library/architecture.md.
  5. User testing surface: .factory/library/user-testing.md — especially the Manual tmux TUI section.
  6. Mission AGENTS.md: boundaries and git safety rules.

Hard rules

  • Real LLM calls are allowed in this skill (manual QA only). The user's ~/.pi/agent/auth.json is presumed configured. Do NOT touch that file.
  • Capture to local-ignore/ — never commit QA evidence. The local-ignore/ directory is gitignored.
  • Clean up test fixtures. If you create a temporary .pi/settings.json for a scenario, delete it after capture so subsequent tests start from a clean slate.
  • Deterministic evidence: every manual feature includes a rg -c "SYSTEM DIRECTIVE: SENPI" count check. Always save the count to a .count file alongside the .log file. The count is the canonical pass/fail marker, not the visual scrollback.
  • No src/ changes: you are verifying only. If you find a bug, return to orchestrator with details and do NOT fix it yourself — a coding-agent-extension-worker will handle the fix in a follow-up feature.

Prerequisites

Before running any scenario, confirm:

  1. ./pi-test.sh is executable and runs (check ls -la pi-test.sh).
  2. npm run build has been run at least once after the continuation feature was merged (check packages/coding-agent/dist/cli.js exists and contains the continuation code).
  3. tmux is installed (command -v tmux).
  4. rg is installed (command -v rg).
  5. local-ignore/ directory exists at repo root (create if needed).

If any prerequisite is missing, return to orchestrator.

Procedure

Step 1 — Orient

Read your feature's description carefully. Identify:

  • Which scenario you are running (default-enabled / settings-disable / flag-override / re-entry-guard).
  • What the expected scrollback should contain.
  • The expected rg -c count (0 or ≥1).
  • The log filename convention (local-ignore/qa-cross-XXX-*.log).
Step 2 — Set up the fixture (if needed)

If your scenario requires a test .pi/settings.json:

bash
mkdir -p .pi
echo '{ "todotools": { "continuation": { "enabled": false } } }' > .pi/settings.json.qa-backup-test
# (Back up any existing .pi/settings.json first so we can restore.)

Always back up the existing file before writing the test fixture, and restore it after capture.

Step 3 — Launch tmux session and drive the CLI

Start a new tmux session for the scenario:

bash
TMUX_SESSION="pi-qa-${FEATURE_ID}"
tmux kill-session -t "$TMUX_SESSION" 2>/dev/null || true
tmux new-session -d -s "$TMUX_SESSION" "./pi-test.sh${EXTRA_FLAGS}"

Wait for the CLI to initialize (use sleep 3 or poll for the prompt).

Send the scripted prompts that drive the agent to create todos and end its turn. Example:

bash
tmux send-keys -t "$TMUX_SESSION" 'create a 2-item todo list about testing this feature, mark one as in_progress, then pause' Enter
sleep 10   # wait for the agent to respond

Use sleep generously between interactions — real model calls take time. A 10-20 second pause between prompts is reasonable.

Show full SKILL.md (245 more words)Show less
Step 4 — Capture scrollback
bash
tmux capture-pane -p -t "$TMUX_SESSION" -S -10000 > local-ignore/qa-cross-XXX-tmux.log
Step 5 — Assert the count marker
bash
rg -c 'SYSTEM DIRECTIVE: SENPI' local-ignore/qa-cross-XXX-tmux.log > local-ignore/qa-cross-XXX-tmux.count || true
COUNT=$(cat local-ignore/qa-cross-XXX-tmux.count)
echo "Continuation directive count: $COUNT"

Compare against the expected count in your feature's expectedBehavior:

  • manual-tmux-qa-default-enabled: expect ≥1
  • manual-tmux-qa-settings-disable: expect 0
  • manual-tmux-qa-flag-override: flag-on expect 0, flag-off expect ≥1
  • manual-tmux-qa-reentry-guard: expect per-turn count ≤1 (check manually because it needs per-turn framing)
Step 6 — Tear down
bash
tmux kill-session -t "$TMUX_SESSION" 2>/dev/null || true

Restore any backup settings files. Remove any temporary fixtures you created.

Step 7 — Verify evidence files
bash
ls -l local-ignore/qa-cross-XXX-*.log local-ignore/qa-cross-XXX-*.count

Confirm both files exist.

Step 8 — Handoff

Report:

  • successState: "success" if the count matches expected; "failure" if not.
  • evidenceFiles: paths to the captured log and count files.
  • observedCount: the actual number.
  • expectedCount: the expected range.
  • scrollbackSummary: 3-5 lines describing what you saw (agent behavior, any error messages, any unexpected output).
  • discoveredIssues: anything buggy or surprising observed during the run (surface it — do not silently ignore).

Escalation triggers

Return to orchestrator immediately if:

  • ./pi-test.sh fails to start.
  • tmux is not available.
  • The agent hangs for > 60 seconds without a response (may indicate a provider outage or a bug).
  • The count does not match expected and you cannot reproduce the failure deterministically (this is a bug that needs a coding-agent-extension-worker fix).
  • You observe any runtime error, stack trace, or TypeError in the scrollback.
  • Test fixtures (settings.json) cannot be backed up or restored.

Anti-patterns

  • Modifying source code "to make the test pass."
  • Committing log/count files from local-ignore/.
  • Leaving test settings.json fixtures behind after capture.
  • Visual "it looks fine" confirmation without the rg -c count file.
  • Using the user's live settings without a backup/restore cycle.

© code-yeongyu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .factory/skills/tmux-manual-qa of code-yeongyu/senpi.

Open the folder on GitHubat commit ee6ece1

Compare with similar skills

Tmux Manual QA Worker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tmux Manual QA Worker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tmux Manual QA Worker this skillcode-yeongyu/senpi472—~1.6kAutomated safety check: PassMIT
tmux Real User TestingQwenLM/qwen-code28k—~2.3kAutomated safety check: PassApache-2.0
CodexBar Live QAsteipete/CodexBar22k—~1.2kAutomated safety check: PassMIT
Verify Omnigent End-to-Endomnigent-ai/omnigent11k—~1.6kAutomated safety check: PassApache-2.0
OpenHarness End-to-End EvalsHKUDS/OpenHarness16k1 repos~2.1kAutomated safety check: NotesMIT
Acceptance Evidence for Deliverieslobehub/lobehub83k—~9.7kAutomated safety check: PassApache-2.0

Similar skills

  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • CodexBar Live QA

    steipete/CodexBar

    Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely.

    22k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Verify Omnigent End-to-End

    omnigent-ai/omnigent

    Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code.

    11k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes
  • Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

    83k GitHub stars~9.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Codex Plugin QA

    code-yeongyu/oh-my-openagent

    Tests the omo Codex plugin in an isolated CODEX_HOME with a local mock model, proving hooks fired through app-server notifications without touching ~/.codex.

    70k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed

More from code-yeongyu/senpi

  • Senpi Agent QA Harness

    code-yeongyu/senpi

    Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.

    472 GitHub stars~2.7k tokensUpdated today
    Auto-check: notes
  • Merge Upstream into Fork

    code-yeongyu/senpi

    Syncs a fork branch with its upstream remote using a history-preserving merge commit, with no rebase and no force push.

    472 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Bun 1.4 Builtins Guide

    code-yeongyu/senpi

    Points the agent at Bun 1.4 built-in APIs before it installs an npm package, so image, browser, markdown, cron, PTY and test work uses what Bun already ships.

    472 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Worker brief for implementing one pre-assigned feature in the senpi todotools built-in extension, with strict scope, typing, testing and git-safety rules.

    472 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • GPT Image Prompt Guide

    code-yeongyu/senpi

    Prompt-crafting guide for gpt-image-2.5: which image tool to call, which model to pick, and how to write prompts, edit with references and refine over turns.

    472 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Senpi Release Publishing

    code-yeongyu/senpi

    Walks the canonical CalVer release flow for senpi, from a clean main checkout through changelog audit, checks, tag push, GitHub Release and npm publishing.

    472 GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Tmux Manual QA Worker

What does Tmux Manual QA Worker do?

Runs one manual QA scenario for the todo continuation feature in the real ./pi-test.sh CLI inside tmux, captures scrollback and checks a deterministic count marker. json.count file beside the log, as the pass or fail signal instead of how the screen looks.

When should I use Tmux Manual QA Worker?

Tmux Manual QA Worker fits situations like: verifying the todo continuation behavior by hand in an interactive TUI; capturing scrollback evidence for a manual-qa milestone feature; testing a settings or flag override scenario and cleaning up its fixture afterwards.

How do I install Tmux Manual QA Worker in Claude Code?

Run `npx skills add code-yeongyu/senpi --skill tmux-manual-qa -a claude-code`. Or copy the skill folder (.factory/skills/tmux-manual-qa in code-yeongyu/senpi) into .claude/skills/tmux-manual-qa in your project. Claude Code loads it when a task matches its description.

How do I install Tmux Manual QA Worker in Codex?

Run `npx skills add code-yeongyu/senpi --skill tmux-manual-qa -a codex`. Or copy the skill folder (.factory/skills/tmux-manual-qa in code-yeongyu/senpi) into .agents/skills/tmux-manual-qa in your project. Codex loads it when a task matches its description.

Can I use Tmux Manual QA Worker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add code-yeongyu/senpi --skill tmux-manual-qa -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tmux-manual-qa, .gemini/skills/tmux-manual-qa, .github/skills/tmux-manual-qa and .opencode/skills/tmux-manual-qa in your project.

What does Tmux Manual QA Worker need to run?

Going by SKILL.md and its folder, Tmux Manual QA Worker needs the command-line tools its instructions call (rg and npm). Our summary lists: The repository's ./pi-test.sh CLI and a completed npm run build; tmux and ripgrep (rg) installed; Configured LLM credentials for real model calls.

Does Tmux Manual QA Worker access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Tmux Manual QA Worker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tmux Manual QA Worker use?

Tmux Manual QA Worker is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tmux Manual QA Worker use?

About 1.6k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tmux Manual QA Worker?

Skills that share tags, products or a category with Tmux Manual QA Worker: tmux Real User Testing (QwenLM/qwen-code, 28k stars), CodexBar Live QA (steipete/CodexBar, 22k stars), Verify Omnigent End-to-End (omnigent-ai/omnigent, 11k stars) and OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tmux Manual QA Worker?

code-yeongyu (a GitHub user) maintains it in code-yeongyu/senpi, which has 472 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.

Source: code-yeongyu/senpi on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.