Tmux Manual QA Worker
code-yeongyu/senpi
Runs one manual QA scenario for the todo continuation feature in the real ./pi-test.sh CLI inside tmux, captures scrollback and checks a deterministic count marker.
Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.
$ npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install QwenLM/qwen-code tmux-real-user-testing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.qwen/skills/tmux-real-user-testing .claude/skills/tmux-real-user-testing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tmux-real-user-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/tmux-real-user-testing into .claude/skills/tmux-real-user-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tmux-real-user-testing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/tmux-real-user-testingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install QwenLM/qwen-code tmux-real-user-testing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.qwen/skills/tmux-real-user-testing .agents/skills/tmux-real-user-testing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tmux-real-user-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/tmux-real-user-testing into .agents/skills/tmux-real-user-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tmux-real-user-testing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install QwenLM/qwen-code tmux-real-user-testing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.qwen/skills/tmux-real-user-testing .cursor/skills/tmux-real-user-testing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tmux-real-user-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/tmux-real-user-testing into .cursor/skills/tmux-real-user-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tmux-real-user-testing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/QwenLM/qwen-code.git --path .qwen/skills/tmux-real-user-testing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install QwenLM/qwen-code tmux-real-user-testing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.qwen/skills/tmux-real-user-testing .gemini/skills/tmux-real-user-testing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tmux-real-user-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/tmux-real-user-testing into .gemini/skills/tmux-real-user-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tmux-real-user-testing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install QwenLM/qwen-code tmux-real-user-testingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .github/skills && cp -r skills-src/.qwen/skills/tmux-real-user-testing .github/skills/tmux-real-user-testing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tmux-real-user-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/tmux-real-user-testing into .github/skills/tmux-real-user-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tmux-real-user-testing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install QwenLM/qwen-code tmux-real-user-testing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.qwen/skills/tmux-real-user-testing .opencode/skills/tmux-real-user-testing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tmux-real-user-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/tmux-real-user-testing into .opencode/skills/tmux-real-user-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tmux-real-user-testing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tmux-real-user-testingDrives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.
This skill has the agent run the Qwen Code terminal interface inside tmux, using realistic keystrokes to move through dialogs, slash commands and workflows. After every meaningful state change it captures the pane with tmux capture-pane and appends the snapshot to a log, so the result reads as a narrative of what appeared on screen.
The main deliverable is tmux-readable-full.log, kept in a timestamped folder under the project's tmp directory next to a final-screen capture. Raw pipe-pane output is only optional extra evidence, because the control codes from the React Ink interface look garbled as plain text. A helper, scripts/tmux-real-user-log.sh, starts a session, sends keys, waits for text and adds labeled snapshots. Earlier runs are never overwritten.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6386bb2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
bashnodenpmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
tmux Real User Testing loads about 2.3k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 853 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from QwenLM/qwen-code at commit 6386bb2, republished under its Apache-2.0 licence (© QwenLM). 853 words, ~2,250 tokens.
.claude/skills/tmux-real-user-testing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Run Qwen Code in a real tmux TUI session as a user would: navigate dialogs, trigger slash commands, exercise workflows, and save a readable log that maintainers can review. Prefer this workflow when the goal is not just a pass/fail assertion, but a narrative artifact showing what happened on screen.
Use tmux as a real-use harness. Drive the TUI with realistic keyboard actions,
then save a step-by-step readable transcript with tmux capture-pane -p after
each meaningful state change.
Avoid relying on tmux pipe-pane as the primary report. pipe-pane captures raw
ANSI/control streams from React Ink TUI output and often looks like garbled text
when opened as plain text. Use pipe-pane only as an optional forensic artifact. Make
tmux-readable-full.log the main deliverable.
Use this workflow for:
/auth, /model, /manage-models, MCP setup,
permissions, onboarding, or interactive error recovery.Use headless JSON E2E instead when only tool execution or model API behavior needs structured assertions.
Create a timestamped directory under project tmp/:
tmp/<scenario>-tmux-YYYYMMDD-HHMMSS/
├── tmux-readable-full.log # primary report: step-by-step readable snapshots
├── tmux-final-capture.log # final screen only
├── current-pane.txt # latest poll/snapshot scratch file
└── report.md # short summary with result and artifact pointersDo not overwrite previous runs. Preserve complete logs unless the user explicitly asks to sanitize or trim them.
Use scripts/tmux-real-user-log.sh to avoid rewriting shell glue. The script can
start a session, append labeled snapshots, send keys, wait for text, and finish.
The start command outputs export statements — use eval to set the variables
directly in your shell:
eval "$(bash .qwen/skills/tmux-real-user-testing/scripts/tmux-real-user-log.sh \
start <scenario> . npm run dev -- --approval-mode yolo)"
# → $SESSION, $OUTDIR, $LOG are now availableShow the full usage before running a new scenario:
bash .qwen/skills/tmux-real-user-testing/scripts/tmux-real-user-log.sh helpUse a large tmux viewport so dialogs render fully. Wait for the TUI to render before interacting — poll for a known startup string rather than blind sleeping:
TS=$(date +%Y%m%d-%H%M%S)
PROJECT_ROOT="$(pwd)"
OUT="$PROJECT_ROOT/tmp/<scenario>-$TS"
SESSION="<scenario>-$TS"
mkdir -p "$OUT"
tmux new-session -d -s "$SESSION" -x 200 -y 50 \
-c "$PROJECT_ROOT" \
"npm run dev -- --approval-mode yolo"
# Poll until TUI is ready (adjust regex to match your app's startup line)
for i in $(seq 1 30); do
sleep 1
if tmux capture-pane -t "$SESSION" -p -S -100 | grep -q "Ready\|>"; then
break
fi
doneUse node dist/cli.js instead of npm run dev only when verifying a built
bundle. Use the globally installed qwen only when reproducing a user-reported
installed-version bug.
After each meaningful action, append a section header plus capture-pane -p to
the full log:
LOG="$OUT/tmux-readable-full.log"
{
printf '\n===== 01 /auth dialog =====\n'
tmux capture-pane -t "$SESSION" -p -S -240
} >> "$LOG"Increase -S as the session grows (add ~100 lines per section). The important
part is that each section is a rendered frame, not raw ANSI output.
Split typing and Enter to avoid swallowed submissions:
tmux send-keys -t "$SESSION" "/auth"
sleep 0.5
tmux send-keys -t "$SESSION" Enter
sleep 2For navigation:
tmux send-keys -t "$SESSION" Down
tmux send-keys -t "$SESSION" Space
tmux send-keys -t "$SESSION" EscapeFor text input into Ink fields, prefer one key at a time if bulk text is ignored:
tmux send-keys -t "$SESSION" e n a b l e dUse text on the screen as the completion condition. On timeout, dump the current pane so the log shows what was on screen when the wait expired:
for i in $(seq 1 60); do
sleep 2
tmux capture-pane -t "$SESSION" -p -S -400 > "$OUT/current-pane.txt"
if grep -q "Successfully configured\|Error\|failed" \
"$OUT/current-pane.txt"; then
break
fi
done
# Always append the final poll result (match or timeout) to the log
{
printf '\n===== 04 auth result =====\n'
cat "$OUT/current-pane.txt"
} >> "$LOG"Capture the final screen, append it, then kill the session:
tmux capture-pane -t "$SESSION" -p -S -10000 > "$OUT/tmux-final-capture.log"
{
printf '\n===== final capture before cleanup =====\n'
cat "$OUT/tmux-final-capture.log"
} >> "$LOG"
tmux kill-session -t "$SESSION"Write report.md with:
tmux-readable-full.log marked as the primary log.Keep assertions tied to evidence in the log. Prefer phrases like “log section
07 toggle model on shows 16 enabled” over unsupported summaries.
A good scenario is a linear sequence of observable state transitions. Design it as a series of steps where each step produces visible TUI output you can capture:
wait-for polling.For each step, define:
/auth, Down, Enter, etc.)Successfully configured, Error, Saved)HELPER=.qwen/skills/tmux-real-user-testing/scripts/tmux-real-user-log.sh
# Start
eval "$(bash "$HELPER" start auth-test . npm run dev -- --approval-mode yolo)"
# → prints SESSION=... OUTDIR=...
# Trigger /auth, navigate to OAuth provider
bash "$HELPER" type-submit "$SESSION" /auth
bash "$HELPER" snapshot "$SESSION" "$OUTDIR" "01 auth menu"
bash "$HELPER" send "$SESSION" Down Down Enter
bash "$HELPER" snapshot "$SESSION" "$OUTDIR" "02 provider selected"
# Wait for OAuth flow to complete (may involve browser interaction)
bash "$HELPER" wait-for "$SESSION" "$OUTDIR" "Successfully configured|Error|failed"
bash "$HELPER" snapshot "$SESSION" "$OUTDIR" "03 auth result"
# Finish
bash "$HELPER" finish "$SESSION" "$OUTDIR"For flows involving browser OAuth callbacks, the wait-for poll will catch the
result after the user completes the browser step. If the flow requires the LLM to
open the browser itself, note that side effect in the scenario design.
Ask before deleting logs or reverting settings. Do not sanitize by default if the user explicitly requests complete logs. If logs may be shared externally, offer a separate sanitized copy rather than modifying the original.
Mention likely side effects before starting: OAuth may open a browser, write Qwen settings, set API key config, and update model provider entries.
open <file> on macOS producing no terminal output is normal; it launches the
file in the associated app.tmux-final-capture.log contains only the last screen; it is not the full
journey.tmux-readable-full.log is the report-grade artifact.tmux pipe-pane raw logs can contain ANSI control sequences and look garbled.capture-pane records current rendered state, not transient flicker.© QwenLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in .qwen/skills/tmux-real-user-testing of QwenLM/qwen-code.
Open the folder on GitHubat commit 6386bb2
tmux Real User Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| tmux Real User Testing this skillQwenLM/qwen-code | 28k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Tmux Manual QA Workercode-yeongyu/senpi | 474 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Senpi Agent QA Harnesscode-yeongyu/senpi | 474 | — | ~2.7k | Automated safety check: Notes | MIT | |
| CodexBar Live QAsteipete/CodexBar | 22k | — | ~1.2k | Automated safety check: Pass | MIT | |
| Verify Omnigent End-to-Endomnigent-ai/omnigent | 11k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| OpenHarness End-to-End EvalsHKUDS/OpenHarness | 16k | 1 repos | ~2.1k | Automated safety check: Notes | MIT |
code-yeongyu/senpi
Runs one manual QA scenario for the todo continuation feature in the real ./pi-test.sh CLI inside tmux, captures scrollback and checks a deterministic count marker.
code-yeongyu/senpi
Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.
steipete/CodexBar
Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely.
omnigent-ai/omnigent
Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code.
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
lobehub/lobehub
Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.
QwenLM/qwen-code
Reproduces a feature from Codex or Claude Code in Qwen Code by running the reference agent under capture, reading the traces, then implementing matching behavior.
QwenLM/qwen-code
Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.
QwenLM/qwen-code
Scheduled CI skill that scans a repository for small, certain docs, test and code hygiene issues and fixes them on one branch with a commit per finding.
QwenLM/qwen-code
Builds a rebranded Qwen Code desktop package from the Tauri shell using only a brand id and a logo, with sensible derived defaults.
QwenLM/qwen-code
Walks through capturing and comparing V8 heap snapshots to find memory leaks in the Qwen Code Node.js CLI, using tmux and the chrome-devtools CLI.
QwenLM/qwen-code
Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches.
Categories
Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review. This skill has the agent run the Qwen Code terminal interface inside tmux, using realistic keystrokes to move through dialogs, slash commands and workflows. After every meaningful state change it captures the pane with tmux capture-pane and appends the snapshot to a log, so the result reads as a narrative of what appeared on screen.
tmux Real User Testing fits situations like: checking TUI behavior such as dialogs, keyboard navigation and slash commands; testing flows like /auth, /model or MCP setup where the intermediate screens matter; giving maintainers a readable record of a regression test run.
Run `npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a claude-code`. Or copy the skill folder (.qwen/skills/tmux-real-user-testing in QwenLM/qwen-code) into .claude/skills/tmux-real-user-testing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a codex`. Or copy the skill folder (.qwen/skills/tmux-real-user-testing in QwenLM/qwen-code) into .agents/skills/tmux-real-user-testing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tmux-real-user-testing, .gemini/skills/tmux-real-user-testing, .github/skills/tmux-real-user-testing and .opencode/skills/tmux-real-user-testing in your project.
Going by SKILL.md and its folder, tmux Real User Testing needs a shell for the scripts in its folder and the command-line tools its instructions call (bash, node and npm). Our summary lists: tmux; A Qwen Code project that starts with npm run dev.
SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
tmux Real User Testing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with tmux Real User Testing: Tmux Manual QA Worker (code-yeongyu/senpi, 474 stars), Senpi Agent QA Harness (code-yeongyu/senpi, 474 stars), CodexBar Live QA (steipete/CodexBar, 22k stars) and Verify Omnigent End-to-End (omnigent-ai/omnigent, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
QwenLM (a GitHub organization) maintains it in QwenLM/qwen-code, which has 28,397 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 10, 2026.
Source: QwenLM/qwen-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.