Agent skill

tmux Real User Testing

by QwenLM in QwenLM/qwen-code

Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

Apache-2.0Auto-check passedTesting & QA

Install tmux Real User Testing

skills CLI
$ npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install QwenLM/qwen-code tmux-real-user-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.qwen/skills/tmux-real-user-testing .claude/skills/tmux-real-user-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tmux-real-user-testing
GitHub stars
28k
Token cost
~2.3k tokens
SKILL.md length
853 words
Files
2 (incl. scripts)
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

  • Works in 5 steps: Start the TUI → Append labeled readable snapshots → Send keys like a user → …
  • Checking TUI behavior such as dialogs, keyboard navigation and slash commands
  • SKILL.md covers Core principle, When to use, Standard artifact layout and Recommended helper script, plus 5 more sections
  • Runs Shell scripts from its folder; calls bash, node and npm

What it does

This skill has the agent run the Qwen Code terminal interface inside tmux, using realistic keystrokes to move through dialogs, slash commands and workflows. After every meaningful state change it captures the pane with tmux capture-pane and appends the snapshot to a log, so the result reads as a narrative of what appeared on screen.

The main deliverable is tmux-readable-full.log, kept in a timestamped folder under the project's tmp directory next to a final-screen capture. Raw pipe-pane output is only optional extra evidence, because the control codes from the React Ink interface look garbled as plain text. A helper, scripts/tmux-real-user-log.sh, starts a session, sends keys, waits for text and adds labeled snapshots. Earlier runs are never overwritten.

When your agent uses it

  • Checking TUI behavior such as dialogs, keyboard navigation and slash commands
  • Testing flows like /auth, /model or MCP setup where the intermediate screens matter
  • Giving maintainers a readable record of a regression test run

Example prompts

  • “Test the /model command in Qwen Code in tmux and save a readable log of every screen.”
  • “Walk through the /auth flow like a real user and give me the step-by-step transcript.”
  • “Run the onboarding flow in a tmux session and keep the full capture-pane snapshots.”

Requirements

  • tmux
  • A Qwen Code project that starts with npm run dev

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Start the TUI
  2. Append labeled readable snapshots
  3. Send keys like a user
  4. Poll for completion instead of blind sleeping
  5. Finish cleanly

What it can do on your machine

Read from SKILL.md and the folder at commit 6386bb2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • node
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

tmux Real User Testing loads about 2.3k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 853 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from QwenLM/qwen-code at commit 6386bb2, republished under its Apache-2.0 licence (© QwenLM). 853 words, ~2,250 tokens.

Download SKILL.mdSave it as .claude/skills/tmux-real-user-testing/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
tmux-real-user-testing
description
This skill should be used when the user asks to "用 tmux 做真实测试", "保存 tmux 日志", "像真实用户一样测试 Qwen", "生成可复查的 TUI 测试报告", "测试 slash command 交互", or requests a tmux-based real user E2E run with complete readable logs. It guides real TUI usage with step-by-step capture-pane snapshots rather than ANSI raw pipe logs.

tmux Real User Testing

Run Qwen Code in a real tmux TUI session as a user would: navigate dialogs, trigger slash commands, exercise workflows, and save a readable log that maintainers can review. Prefer this workflow when the goal is not just a pass/fail assertion, but a narrative artifact showing what happened on screen.

Core principle

Use tmux as a real-use harness. Drive the TUI with realistic keyboard actions, then save a step-by-step readable transcript with tmux capture-pane -p after each meaningful state change.

Avoid relying on tmux pipe-pane as the primary report. pipe-pane captures raw ANSI/control streams from React Ink TUI output and often looks like garbled text when opened as plain text. Use pipe-pane only as an optional forensic artifact. Make tmux-readable-full.log the main deliverable.

When to use

Use this workflow for:

  • TUI behavior, rendering, dialogs, keyboard navigation, slash commands, or auth flows.
  • Realistic workflows where a maintainer wants to read the journey afterward.
  • Regression testing where final state is insufficient and intermediate screens matter.
  • User-facing flows such as /auth, /model, /manage-models, MCP setup, permissions, onboarding, or interactive error recovery.

Use headless JSON E2E instead when only tool execution or model API behavior needs structured assertions.

Standard artifact layout

Create a timestamped directory under project tmp/:

text
tmp/<scenario>-tmux-YYYYMMDD-HHMMSS/
├── tmux-readable-full.log   # primary report: step-by-step readable snapshots
├── tmux-final-capture.log   # final screen only
├── current-pane.txt         # latest poll/snapshot scratch file
└── report.md                # short summary with result and artifact pointers

Do not overwrite previous runs. Preserve complete logs unless the user explicitly asks to sanitize or trim them.

Use scripts/tmux-real-user-log.sh to avoid rewriting shell glue. The script can start a session, append labeled snapshots, send keys, wait for text, and finish.

The start command outputs export statements — use eval to set the variables directly in your shell:

bash
eval "$(bash .qwen/skills/tmux-real-user-testing/scripts/tmux-real-user-log.sh \
  start <scenario> . npm run dev -- --approval-mode yolo)"
# → $SESSION, $OUTDIR, $LOG are now available

Show the full usage before running a new scenario:

bash
bash .qwen/skills/tmux-real-user-testing/scripts/tmux-real-user-log.sh help

Manual workflow

1. Start the TUI

Use a large tmux viewport so dialogs render fully. Wait for the TUI to render before interacting — poll for a known startup string rather than blind sleeping:

bash
TS=$(date +%Y%m%d-%H%M%S)
PROJECT_ROOT="$(pwd)"
OUT="$PROJECT_ROOT/tmp/<scenario>-$TS"
SESSION="<scenario>-$TS"
mkdir -p "$OUT"
tmux new-session -d -s "$SESSION" -x 200 -y 50 \
  -c "$PROJECT_ROOT" \
  "npm run dev -- --approval-mode yolo"

# Poll until TUI is ready (adjust regex to match your app's startup line)
for i in $(seq 1 30); do
  sleep 1
  if tmux capture-pane -t "$SESSION" -p -S -100 | grep -q "Ready\|>"; then
    break
  fi
done

Use node dist/cli.js instead of npm run dev only when verifying a built bundle. Use the globally installed qwen only when reproducing a user-reported installed-version bug.

2. Append labeled readable snapshots

After each meaningful action, append a section header plus capture-pane -p to the full log:

bash
LOG="$OUT/tmux-readable-full.log"
{
  printf '\n===== 01 /auth dialog =====\n'
  tmux capture-pane -t "$SESSION" -p -S -240
} >> "$LOG"

Increase -S as the session grows (add ~100 lines per section). The important part is that each section is a rendered frame, not raw ANSI output.

3. Send keys like a user

Split typing and Enter to avoid swallowed submissions:

bash
tmux send-keys -t "$SESSION" "/auth"
sleep 0.5
tmux send-keys -t "$SESSION" Enter
sleep 2

For navigation:

bash
tmux send-keys -t "$SESSION" Down
tmux send-keys -t "$SESSION" Space
tmux send-keys -t "$SESSION" Escape

For text input into Ink fields, prefer one key at a time if bulk text is ignored:

bash
tmux send-keys -t "$SESSION" e n a b l e d
4. Poll for completion instead of blind sleeping

Use text on the screen as the completion condition. On timeout, dump the current pane so the log shows what was on screen when the wait expired:

bash
for i in $(seq 1 60); do
  sleep 2
  tmux capture-pane -t "$SESSION" -p -S -400 > "$OUT/current-pane.txt"
  if grep -q "Successfully configured\|Error\|failed" \
    "$OUT/current-pane.txt"; then
    break
  fi
done
# Always append the final poll result (match or timeout) to the log
{
  printf '\n===== 04 auth result =====\n'
  cat "$OUT/current-pane.txt"
} >> "$LOG"
5. Finish cleanly

Capture the final screen, append it, then kill the session:

bash
tmux capture-pane -t "$SESSION" -p -S -10000 > "$OUT/tmux-final-capture.log"
{
  printf '\n===== final capture before cleanup =====\n'
  cat "$OUT/tmux-final-capture.log"
} >> "$LOG"
tmux kill-session -t "$SESSION"

Reporting expectations

Write report.md with:

  • Date, tmux session name, command, workspace.
  • Scenario scope and exact steps tested.
  • PASS/FAIL result.
  • Key screen observations and important state transitions.
  • Artifact list, with tmux-readable-full.log marked as the primary log.
  • Any known side effects, such as settings updates, opened browser windows, or API calls.

Keep assertions tied to evidence in the log. Prefer phrases like “log section 07 toggle model on shows 16 enabled” over unsupported summaries.

Show full SKILL.md (318 more words)Show less

Designing a test scenario

A good scenario is a linear sequence of observable state transitions. Design it as a series of steps where each step produces visible TUI output you can capture:

  1. Entry point — the slash command or action that starts the flow.
  2. Branch points — dialogs or selectors that require navigation keys (Arrow, Space, Enter).
  3. Waiting states — loading screens, auth callbacks, or async operations that need wait-for polling.
  4. Confirmation — success/error text visible on screen that marks completion.
  5. Side effects — external actions the flow triggers (browser open, file writes, config changes) that may affect subsequent runs.

For each step, define:

  • The keys to send (/auth, Down, Enter, etc.)
  • The expected text to wait for (Successfully configured, Error, Saved)
  • When to take a snapshot (before and after every interaction)
Short example: testing /auth → OAuth
bash
HELPER=.qwen/skills/tmux-real-user-testing/scripts/tmux-real-user-log.sh

# Start
eval "$(bash "$HELPER" start auth-test . npm run dev -- --approval-mode yolo)"
# → prints SESSION=... OUTDIR=...

# Trigger /auth, navigate to OAuth provider
bash "$HELPER" type-submit "$SESSION" /auth
bash "$HELPER" snapshot "$SESSION" "$OUTDIR" "01 auth menu"
bash "$HELPER" send "$SESSION" Down Down Enter
bash "$HELPER" snapshot "$SESSION" "$OUTDIR" "02 provider selected"

# Wait for OAuth flow to complete (may involve browser interaction)
bash "$HELPER" wait-for "$SESSION" "$OUTDIR" "Successfully configured|Error|failed"
bash "$HELPER" snapshot "$SESSION" "$OUTDIR" "03 auth result"

# Finish
bash "$HELPER" finish "$SESSION" "$OUTDIR"

For flows involving browser OAuth callbacks, the wait-for poll will catch the result after the user completes the browser step. If the flow requires the LLM to open the browser itself, note that side effect in the scenario design.

Safety and privacy

Ask before deleting logs or reverting settings. Do not sanitize by default if the user explicitly requests complete logs. If logs may be shared externally, offer a separate sanitized copy rather than modifying the original.

Mention likely side effects before starting: OAuth may open a browser, write Qwen settings, set API key config, and update model provider entries.

Common pitfalls

  • open <file> on macOS producing no terminal output is normal; it launches the file in the associated app.
  • tmux-final-capture.log contains only the last screen; it is not the full journey.
  • tmux-readable-full.log is the report-grade artifact.
  • tmux pipe-pane raw logs can contain ANSI control sequences and look garbled.
  • Search boxes and input fields sometimes ignore bulk text; send characters individually.
  • capture-pane records current rendered state, not transient flicker.
  • Use a timestamped output directory for every run to avoid overwriting evidence.

© QwenLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .qwen/skills/tmux-real-user-testing of QwenLM/qwen-code.

  • SKILL.md
  • scripts/tmux-real-user-log.sh

Open the folder on GitHubat commit 6386bb2

Compare with similar skills

tmux Real User Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

tmux Real User Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
tmux Real User Testing this skillQwenLM/qwen-code28k—~2.3kAutomated safety check: PassApache-2.0
Tmux Manual QA Workercode-yeongyu/senpi474—~1.6kAutomated safety check: PassMIT
Senpi Agent QA Harnesscode-yeongyu/senpi474—~2.7kAutomated safety check: NotesMIT
CodexBar Live QAsteipete/CodexBar22k—~1.2kAutomated safety check: PassMIT
Verify Omnigent End-to-Endomnigent-ai/omnigent11k—~1.6kAutomated safety check: PassApache-2.0
OpenHarness End-to-End EvalsHKUDS/OpenHarness16k1 repos~2.1kAutomated safety check: NotesMIT

Similar skills

  • Tmux Manual QA Worker

    code-yeongyu/senpi

    Runs one manual QA scenario for the todo continuation feature in the real ./pi-test.sh CLI inside tmux, captures scrollback and checks a deterministic count marker.

    474 GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Senpi Agent QA Harness

    code-yeongyu/senpi

    Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.

    474 GitHub stars~2.7k tokensUpdated today
    Testing & QAAuto-check: notes
  • CodexBar Live QA

    steipete/CodexBar

    Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely.

    22k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Verify Omnigent End-to-End

    omnigent-ai/omnigent

    Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code.

    11k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes
  • Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

    83k GitHub stars~9.7k tokensUpdated today
    Testing & QAAuto-check passed

More from QwenLM/qwen-code

All 41 skills in this repo
  • Reproduces a feature from Codex or Claude Code in Qwen Code by running the reference agent under capture, reading the traces, then implementing matching behavior.

    28k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Scheduled CI skill that scans a repository for small, certain docs, test and code hygiene issues and fixes them on one branch with a commit per finding.

    28k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Builds a rebranded Qwen Code desktop package from the Tauri shell using only a brand id and a logo, with sensible derived defaults.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through capturing and comparing V8 heap snapshots to find memory leaks in the Qwen Code Node.js CLI, using tmux and the chrome-devtools CLI.

    28k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Agent Reproduce Align

    QwenLM/qwen-code

    Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches.

    28k GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Works with

Questions about tmux Real User Testing

What does tmux Real User Testing do?

Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review. This skill has the agent run the Qwen Code terminal interface inside tmux, using realistic keystrokes to move through dialogs, slash commands and workflows. After every meaningful state change it captures the pane with tmux capture-pane and appends the snapshot to a log, so the result reads as a narrative of what appeared on screen.

When should I use tmux Real User Testing?

tmux Real User Testing fits situations like: checking TUI behavior such as dialogs, keyboard navigation and slash commands; testing flows like /auth, /model or MCP setup where the intermediate screens matter; giving maintainers a readable record of a regression test run.

How do I install tmux Real User Testing in Claude Code?

Run `npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a claude-code`. Or copy the skill folder (.qwen/skills/tmux-real-user-testing in QwenLM/qwen-code) into .claude/skills/tmux-real-user-testing in your project. Claude Code loads it when a task matches its description.

How do I install tmux Real User Testing in Codex?

Run `npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a codex`. Or copy the skill folder (.qwen/skills/tmux-real-user-testing in QwenLM/qwen-code) into .agents/skills/tmux-real-user-testing in your project. Codex loads it when a task matches its description.

Can I use tmux Real User Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QwenLM/qwen-code --skill tmux-real-user-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tmux-real-user-testing, .gemini/skills/tmux-real-user-testing, .github/skills/tmux-real-user-testing and .opencode/skills/tmux-real-user-testing in your project.

What does tmux Real User Testing need to run?

Going by SKILL.md and its folder, tmux Real User Testing needs a shell for the scripts in its folder and the command-line tools its instructions call (bash, node and npm). Our summary lists: tmux; A Qwen Code project that starts with npm run dev.

Does tmux Real User Testing access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is tmux Real User Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does tmux Real User Testing use?

tmux Real User Testing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does tmux Real User Testing use?

About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to tmux Real User Testing?

Skills that share tags, products or a category with tmux Real User Testing: Tmux Manual QA Worker (code-yeongyu/senpi, 474 stars), Senpi Agent QA Harness (code-yeongyu/senpi, 474 stars), CodexBar Live QA (steipete/CodexBar, 22k stars) and Verify Omnigent End-to-End (omnigent-ai/omnigent, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains tmux Real User Testing?

QwenLM (a GitHub organization) maintains it in QwenLM/qwen-code, which has 28,397 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 10, 2026.

Source: QwenLM/qwen-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.