Exploring The Wizard
PostHog/wizard
Drive the PostHog wizard headlessly against a throwaway app through wizard-ci MCP tools, inspect decisions, and capture the real TUI.
Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.
$ npx skills add QwenLM/qwen-code --skill e2e-testing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install QwenLM/qwen-code e2e-testing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.qwen/skills/e2e-testing .claude/skills/e2e-testing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "e2e-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/e2e-testing into .claude/skills/e2e-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-testing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/e2e-testingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add QwenLM/qwen-code --skill e2e-testing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install QwenLM/qwen-code e2e-testing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.qwen/skills/e2e-testing .agents/skills/e2e-testing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "e2e-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/e2e-testing into .agents/skills/e2e-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-testing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add QwenLM/qwen-code --skill e2e-testing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install QwenLM/qwen-code e2e-testing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.qwen/skills/e2e-testing .cursor/skills/e2e-testing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "e2e-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/e2e-testing into .cursor/skills/e2e-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-testing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/QwenLM/qwen-code.git --path .qwen/skills/e2e-testing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add QwenLM/qwen-code --skill e2e-testing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install QwenLM/qwen-code e2e-testing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.qwen/skills/e2e-testing .gemini/skills/e2e-testing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "e2e-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/e2e-testing into .gemini/skills/e2e-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-testing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install QwenLM/qwen-code e2e-testingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add QwenLM/qwen-code --skill e2e-testing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .github/skills && cp -r skills-src/.qwen/skills/e2e-testing .github/skills/e2e-testing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "e2e-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/e2e-testing into .github/skills/e2e-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-testing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add QwenLM/qwen-code --skill e2e-testing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install QwenLM/qwen-code e2e-testing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.qwen/skills/e2e-testing .opencode/skills/e2e-testing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "e2e-testing" agent skill from https://github.com/QwenLM/qwen-code/tree/main/.qwen/skills/e2e-testing into .opencode/skills/e2e-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "e2e-testing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
e2e-testingGuides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.
The skill explains how to run the Qwen Code CLI through the whole pipeline, from model API to tool validation to tool execution, when unit tests are not enough. It tells the agent which binary to use: the globally installed qwen command to reproduce a reported bug, a local build run with node dist/cli.js to verify a fix, or npm run dev for fast runtime-only checks.
Headless runs pass a prompt with --approval-mode yolo and --output-format json, which emits one JSON array to filter with jq, and --auth-type is what switches providers, since --model alone does not. QWEN_RUNTIME_DIR redirects runtime output away from ~/.qwen so repeated tests do not clutter your real history. Helper files include a mock OpenAI server, an MCP test server script and a token statistics script, with references for MCP testing and the mock server.
Read from SKILL.md and the folder at commit fe4d4e3. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (JavaScript and Python), which the agent can run.
Shell commands in SKILL.md call:
npmnodejqpython3bundleFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Qwen Code E2E Testing loads about 2.1k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 108 tokens; SKILL.md has 792 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from QwenLM/qwen-code at commit fe4d4e3, republished under its Apache-2.0 licence (© QwenLM). 792 words, ~2,076 tokens.
.claude/skills/e2e-testing/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.How to run the Qwen Code CLI end-to-end — from building the bundle to inspecting raw API traffic. Use when unit tests aren't enough and you need to verify behavior through the full pipeline (model API → tool validation → tool execution).
qwen command — this matches
what the user ran when they filed the issue.npm run build && npm run bundle), then run
node dist/cli.js — this tests your local changes.npm run dev -- "<prompt>" <flags> — runs TS
source via tsx, no build. Use build && bundle + node dist/cli.js only when the
shipped artifact itself matters. (<qwen> below can be npm run dev --.)Headless auth comes from ~/.qwen. Force a known-good model with --auth-type +
--model:
<qwen> "your prompt" --auth-type openai --model deepseek-v4-flash \
--approval-mode yolo --output-format jsonGotcha: --model alone won't switch providers — --auth-type (openai/anthropic/qwen-oauth/gemini/vertex-ai) does. Omit it and the run falls back to the default provider and dies
on its missing key.
QWEN_RUNTIME_DIR=<dir> redirects qwen's runtime output — tmp/, debug/,
and projects/<sanitized-cwd>/... (chat recordings, auto-memory, history) —
into <dir> instead of ~/.qwen. Config (settings.json, OAuth tokens,
commands/) still reads from ~/.qwen, so real auth and provider config
work without any setup.
Use when repeated test runs would clutter your real chat history or auto-memory. Skip when the bug you're reproducing depends on the user's actual history or runtime state — that is the repro.
QWEN_RUNTIME_DIR=/tmp/test-1/runtime <qwen> "prompt" ...Run the CLI non-interactively with JSON output (<qwen> = qwen or
node dist/cli.js per above):
<qwen> "your prompt here" \
--approval-mode yolo \
--output-format json \
2>/dev/null--output-format json emits one JSON array (all messages, flushed at end of turn) — filter with jq '.[] | …', never a bare jq 'select(…)'. (--output-format stream-json instead emits NDJSON, one object per line.) Element types:
type: "system" — init: tools, mcp_servers, model, permission_modetype: "assistant" — model output: content[].type is text, tool_use, or thinkingtype: "user" — tool results: content[].type is tool_result with is_errortype: "result" — final output with result text and usage statsFilter with jq — lead with .[] to enter the array, e.g. tool-result errors:
... 2>/dev/null | jq '.[] | select(.type=="user") | .message.content[] | select(.is_error)'
Use when you need to verify TUI rendering, test keyboard interactions, or see what the user sees. Headless mode is simpler when you only need structured output.
tmux new-session -d -s test -x 200 -y 50 \
"cd /tmp/test-dir && <qwen> --approval-mode yolo"
sleep 3 # wait for TUI to initializeSplit text and Enter with a short delay — sending them together can cause the TUI to swallow the submit:
tmux send-keys -t test "your prompt here"
sleep 0.5
tmux send-keys -t test EnterPoll for the streaming indicator to disappear instead of blind sleeping. The
footer placeholder Type your message is always rendered — don't grep for
that or the loop exits on iteration 1 while the model is still working. The
status line esc to cancel is present only while the model is producing
output:
for i in $(seq 1 60); do
sleep 2
tmux capture-pane -t test -p | grep -q "esc to cancel" || break
donetmux capture-pane -t test -p -S -100 # -S -100 = 100 lines of scrollbacktmux send-keys cannot reliably send all key combinations.
C-?, C-Shift-*, and function keys with modifiers are unsupported or
unreliable. For these, use the InteractiveSession harness in
integration-tests/interactive/ or test manually.capture-pane captures the final rendered frame, not
intermediate states. Flicker, tearing, or brief blank frames cannot be
detected this way.tmux kill-session -t testWhen debugging model behavior (wrong tool arguments, schema issues), enable API logging to see the exact request/response payloads:
<qwen> "prompt" \
--approval-mode yolo \
--output-format json \
--openai-logging \
--openai-logging-dir /tmp/api-logsEach API call produces a JSON file (can be 80KB+ due to full message history).
The bulk is in request.messages (conversation history). Trimmed structure:
{
"request": {
"model": "coder-model",
"messages": [
{ "role": "system|user|assistant", "content": "...", "tool_calls?": [...] }
],
"tools": [
{
"type": "function",
"function": {
"name": "tool_name",
"description": "...",
"parameters": { ... } // schema sent to the model
}
}
]
},
"response": {
"choices": [
{
"message": {
"role": "assistant",
"content": "...", // text response (may be null)
"tool_calls": [
{
"id": "call_...",
"function": {
"name": "tool_name",
"arguments": "..." // raw JSON string from the model
}
}
]
}
}
]
}
}Structured-output calls (those requesting a JSON schema, e.g. side queries via
BaseLlmClient.generateJson) deliver the schema as a synthetic tool named
respond_in_schema under request.tools[0] — not under response_format,
which is null for OpenAI-compatible providers. The model's structured reply
lands in tool_calls[0].function.arguments instead of message.content.
Text-mode calls have no tools and use message.content.
Use scripts/token-stats.py to summarize token usage across recent API logs:
python3 .qwen/skills/e2e-testing/scripts/token-stats.py 20 # last 20 requestsShows input, cached, and output tokens per request with cache hit rates. Useful for verifying prompt caching behavior or investigating unexpected token counts.
For testing MCP tool behavior end-to-end, read references/mcp-testing.md. It
covers the setup gotchas (config location, git repo requirement) and includes
a reusable zero-dependency test server template in scripts/mcp-test-server.js.
For driving the CLI through scenarios that are hard to provoke against a real
model — specific error codes, malformed tool calls, deterministic multi-turn
loops, controlled usage blocks — read references/mock-openai-server.md.
It covers when to reach for a mock vs --openai-logging, how to point the
CLI at it, and patterns for specializing the zero-dependency template at
scripts/mock-openai-server.js.
--approval-mode default when testing permission rules. yolo bypasses
rule evaluation entirely — it can't test whether a rule matches.© QwenLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in .qwen/skills/e2e-testing of QwenLM/qwen-code.
Open the folder on GitHubat commit fe4d4e3
Qwen Code E2E Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Qwen Code E2E Testing this skillQwenLM/qwen-code | 28k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Exploring The WizardPostHog/wizard | 197 | — | ~2.2k | Automated safety check: Pass | MIT | |
| Suede MCP Release QAJasonColapietro/suede-creator-skills | 127 | — | ~2.1k | Automated safety check: Pass | MIT | |
| Edt MCP TestingDitriXNew/EDT-MCP | 295 | — | ~2.1k | Automated safety check: Pass | AGPL-3.0 | |
| Browserstackalirezarezvani/claude-skills | 28k | 1 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Edt MCP Ready To DeployDitriXNew/EDT-MCP | 295 | — | ~1.6k | Automated safety check: Pass | AGPL-3.0 |
PostHog/wizard
Drive the PostHog wizard headlessly against a throwaway app through wizard-ci MCP tools, inspect decisions, and capture the real TUI.
JasonColapietro/suede-creator-skills
Checks a Suede AI MCP server release against a live process: the full JSON-RPC lifecycle, schemas, annotations, malformed input, catalog agreement and install docs.
DitriXNew/EDT-MCP
How to manually e2e-test each EDT-MCP server tool against a live EDT workbench + TestConfiguration.
alirezarezvani/claude-skills
Run tests on BrowserStack. An agent skill from alirezarezvani/claude-skills.
DitriXNew/EDT-MCP
The final "definition of done" / ready-to-deploy checklist for EDT-MCP — the ordered gate to run when a piece of work is finished, before declaring it done or merging.
DebugBase/glance
Run E2E browser tests on any web application using Glance MCP.
QwenLM/qwen-code
Reproduces a feature from Codex or Claude Code in Qwen Code by running the reference agent under capture, reading the traces, then implementing matching behavior.
QwenLM/qwen-code
Scheduled CI skill that scans a repository for small, certain docs, test and code hygiene issues and fixes them on one branch with a commit per finding.
QwenLM/qwen-code
Builds a rebranded Qwen Code desktop package from the Tauri shell using only a brand id and a logo, with sensible derived defaults.
QwenLM/qwen-code
Walks through capturing and comparing V8 heap snapshots to find memory leaks in the Qwen Code Node.js CLI, using tmux and the chrome-devtools CLI.
QwenLM/qwen-code
Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.
QwenLM/qwen-code
Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches.
Works with
Categories
Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic. The skill explains how to run the Qwen Code CLI through the whole pipeline, from model API to tool validation to tool execution, when unit tests are not enough.js to verify a fix, or npm run dev for fast runtime-only checks.
Qwen Code E2E Testing fits situations like: verifying CLI behavior with real model calls; reproducing a user-reported bug end to end; testing an MCP tool integration with a test server; inspecting raw API request and response payloads.
Run `npx skills add QwenLM/qwen-code --skill e2e-testing -a claude-code`. Or copy the skill folder (.qwen/skills/e2e-testing in QwenLM/qwen-code) into .claude/skills/e2e-testing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add QwenLM/qwen-code --skill e2e-testing -a codex`. Or copy the skill folder (.qwen/skills/e2e-testing in QwenLM/qwen-code) into .agents/skills/e2e-testing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QwenLM/qwen-code --skill e2e-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/e2e-testing, .gemini/skills/e2e-testing, .github/skills/e2e-testing and .opencode/skills/e2e-testing in your project.
Going by SKILL.md and its folder, Qwen Code E2E Testing needs JavaScript and Python for the scripts in its folder and the command-line tools its instructions call (npm, node, jq, python3 and bundle). Our summary lists: Node.js and npm, to build and run the CLI; jq, to filter the JSON output; Model provider credentials in ~/.qwen, for real model calls.
SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Qwen Code E2E Testing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Qwen Code E2E Testing: Exploring The Wizard (PostHog/wizard, 197 stars), Suede MCP Release QA (JasonColapietro/suede-creator-skills, 127 stars), Edt MCP Testing (DitriXNew/EDT-MCP, 295 stars) and Browserstack (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
QwenLM (a GitHub organization) maintains it in QwenLM/qwen-code, which has 28,349 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 8, 2026.
Source: QwenLM/qwen-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.