Agent skill

Qwen Code E2E Testing

by QwenLM in QwenLM/qwen-code

Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

Apache-2.0Auto-check passedTesting & QA

Install Qwen Code E2E Testing

skills CLI
$ npx skills add QwenLM/qwen-code --skill e2e-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install QwenLM/qwen-code e2e-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.qwen/skills/e2e-testing .claude/skills/e2e-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
e2e-testing
GitHub stars
28k
Token cost
~2.1k tokens
SKILL.md length
792 words
Files
6 (incl. scripts, references)
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

  • Verifying CLI behavior with real model calls
  • SKILL.md covers Setup, Run modes, Inspecting and Test harnesses, plus 1 more section
  • Runs JavaScript and Python scripts from its folder; calls npm, node and jq
  • Reproducing a user-reported bug end to end

What it does

The skill explains how to run the Qwen Code CLI through the whole pipeline, from model API to tool validation to tool execution, when unit tests are not enough. It tells the agent which binary to use: the globally installed qwen command to reproduce a reported bug, a local build run with node dist/cli.js to verify a fix, or npm run dev for fast runtime-only checks.

Headless runs pass a prompt with --approval-mode yolo and --output-format json, which emits one JSON array to filter with jq, and --auth-type is what switches providers, since --model alone does not. QWEN_RUNTIME_DIR redirects runtime output away from ~/.qwen so repeated tests do not clutter your real history. Helper files include a mock OpenAI server, an MCP test server script and a token statistics script, with references for MCP testing and the mock server.

When your agent uses it

  • Verifying CLI behavior with real model calls
  • Reproducing a user-reported bug end to end
  • Testing an MCP tool integration with a test server
  • Inspecting raw API request and response payloads

Example prompts

  • “Reproduce the issue from this bug report in headless mode and show me the JSON output.”
  • “Build the bundle and run my fix through the end-to-end flow with a real model.”
  • “Start the MCP test server and check that its tool shows up in the system init message.”
  • “Run the CLI against the mock OpenAI server and show me the raw request payloads.”

Requirements

  • Node.js and npm, to build and run the CLI
  • jq, to filter the JSON output
  • Model provider credentials in ~/.qwen, for real model calls

What it can do on your machine

Read from SKILL.md and the folder at commit fe4d4e3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (JavaScript and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • npm
    • node
    • jq
    • python3
    • bundle

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Qwen Code E2E Testing loads about 2.1k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 108 tokens; SKILL.md has 792 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from QwenLM/qwen-code at commit fe4d4e3, republished under its Apache-2.0 licence (© QwenLM). 792 words, ~2,076 tokens.

Download SKILL.mdSave it as .claude/skills/e2e-testing/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
e2e-testing
description
Guide for running end-to-end tests of the Qwen Code CLI, including headless mode, MCP server testing, and API traffic inspection. Use this skill whenever you need to verify CLI behavior with real model calls, reproduce user-reported bugs end-to-end, test MCP tool integrations, or inspect raw API request/response payloads. Trigger on mentions of E2E testing, headless testing, MCP tool testing, or reproducing issues.

E2E Testing Guide

How to run the Qwen Code CLI end-to-end — from building the bundle to inspecting raw API traffic. Use when unit tests aren't enough and you need to verify behavior through the full pipeline (model API → tool validation → tool execution).

Setup

Which binary to use
  • Reproducing bugs: use the globally installed qwen command — this matches what the user ran when they filed the issue.
  • Verifying fixes: build first (npm run build && npm run bundle), then run node dist/cli.js — this tests your local changes.
  • Runtime-only checks (fastest): npm run dev -- "<prompt>" <flags> — runs TS source via tsx, no build. Use build && bundle + node dist/cli.js only when the shipped artifact itself matters. (<qwen> below can be npm run dev --.)
Running against a real model

Headless auth comes from ~/.qwen. Force a known-good model with --auth-type + --model:

bash
<qwen> "your prompt" --auth-type openai --model deepseek-v4-flash \
  --approval-mode yolo --output-format json

Gotcha: --model alone won't switch providers — --auth-type (openai/anthropic/qwen-oauth/gemini/vertex-ai) does. Omit it and the run falls back to the default provider and dies on its missing key.

Isolating runtime artifacts

QWEN_RUNTIME_DIR=<dir> redirects qwen's runtime output — tmp/, debug/, and projects/<sanitized-cwd>/... (chat recordings, auto-memory, history) — into <dir> instead of ~/.qwen. Config (settings.json, OAuth tokens, commands/) still reads from ~/.qwen, so real auth and provider config work without any setup.

Use when repeated test runs would clutter your real chat history or auto-memory. Skip when the bug you're reproducing depends on the user's actual history or runtime state — that is the repro.

bash
QWEN_RUNTIME_DIR=/tmp/test-1/runtime <qwen> "prompt" ...

Run modes

Headless Mode

Run the CLI non-interactively with JSON output (<qwen> = qwen or node dist/cli.js per above):

bash
<qwen> "your prompt here" \
  --approval-mode yolo \
  --output-format json \
  2>/dev/null

--output-format json emits one JSON array (all messages, flushed at end of turn) — filter with jq '.[] | …', never a bare jq 'select(…)'. (--output-format stream-json instead emits NDJSON, one object per line.) Element types:

  • type: "system" — init: tools, mcp_servers, model, permission_mode
  • type: "assistant" — model output: content[].type is text, tool_use, or thinking
  • type: "user" — tool results: content[].type is tool_result with is_error
  • type: "result" — final output with result text and usage stats

Filter with jq — lead with .[] to enter the array, e.g. tool-result errors: ... 2>/dev/null | jq '.[] | select(.type=="user") | .message.content[] | select(.is_error)'

Interactive Mode (tmux)

Use when you need to verify TUI rendering, test keyboard interactions, or see what the user sees. Headless mode is simpler when you only need structured output.

Launching
bash
tmux new-session -d -s test -x 200 -y 50 \
  "cd /tmp/test-dir && <qwen> --approval-mode yolo"
sleep 3  # wait for TUI to initialize
Sending prompts

Split text and Enter with a short delay — sending them together can cause the TUI to swallow the submit:

bash
tmux send-keys -t test "your prompt here"
sleep 0.5
tmux send-keys -t test Enter
Waiting for completion

Poll for the streaming indicator to disappear instead of blind sleeping. The footer placeholder Type your message is always rendered — don't grep for that or the loop exits on iteration 1 while the model is still working. The status line esc to cancel is present only while the model is producing output:

bash
for i in $(seq 1 60); do
  sleep 2
  tmux capture-pane -t test -p | grep -q "esc to cancel" || break
done
Capturing output
bash
tmux capture-pane -t test -p -S -100   # -S -100 = 100 lines of scrollback
Limitations
  • Key combos: tmux send-keys cannot reliably send all key combinations. C-?, C-Shift-*, and function keys with modifiers are unsupported or unreliable. For these, use the InteractiveSession harness in integration-tests/interactive/ or test manually.
  • Visual artifacts: capture-pane captures the final rendered frame, not intermediate states. Flicker, tearing, or brief blank frames cannot be detected this way.
Show full SKILL.md (291 more words)Show less
Cleanup
bash
tmux kill-session -t test

Inspecting

Inspecting Raw API Traffic

When debugging model behavior (wrong tool arguments, schema issues), enable API logging to see the exact request/response payloads:

bash
<qwen> "prompt" \
  --approval-mode yolo \
  --output-format json \
  --openai-logging \
  --openai-logging-dir /tmp/api-logs

Each API call produces a JSON file (can be 80KB+ due to full message history). The bulk is in request.messages (conversation history). Trimmed structure:

json
{
  "request": {
    "model": "coder-model",
    "messages": [
      { "role": "system|user|assistant", "content": "...", "tool_calls?": [...] }
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "tool_name",
          "description": "...",
          "parameters": { ... }      // schema sent to the model
        }
      }
    ]
  },
  "response": {
    "choices": [
      {
        "message": {
          "role": "assistant",
          "content": "...",          // text response (may be null)
          "tool_calls": [
            {
              "id": "call_...",
              "function": {
                "name": "tool_name",
                "arguments": "..."   // raw JSON string from the model
              }
            }
          ]
        }
      }
    ]
  }
}

Structured-output calls (those requesting a JSON schema, e.g. side queries via BaseLlmClient.generateJson) deliver the schema as a synthetic tool named respond_in_schema under request.tools[0] — not under response_format, which is null for OpenAI-compatible providers. The model's structured reply lands in tool_calls[0].function.arguments instead of message.content. Text-mode calls have no tools and use message.content.

Token Usage Stats

Use scripts/token-stats.py to summarize token usage across recent API logs:

bash
python3 .qwen/skills/e2e-testing/scripts/token-stats.py 20  # last 20 requests

Shows input, cached, and output tokens per request with cache hit rates. Useful for verifying prompt caching behavior or investigating unexpected token counts.

Test harnesses

MCP Server Testing

For testing MCP tool behavior end-to-end, read references/mcp-testing.md. It covers the setup gotchas (config location, git repo requirement) and includes a reusable zero-dependency test server template in scripts/mcp-test-server.js.

Mock OpenAI Server

For driving the CLI through scenarios that are hard to provoke against a real model — specific error codes, malformed tool calls, deterministic multi-turn loops, controlled usage blocks — read references/mock-openai-server.md. It covers when to reach for a mock vs --openai-logging, how to point the CLI at it, and patterns for specializing the zero-dependency template at scripts/mock-openai-server.js.

Tips

  • Use interactive (tmux) mode when the bug involves permission prompts, slash commands, or keyboard interactions. Headless mode has no TUI — these don't exist there.
  • Use interactive (tmux) mode for hang-related issues. Headless mode produces no output when the process stalls, giving you nothing to work with.
  • Use --approval-mode default when testing permission rules. yolo bypasses rule evaluation entirely — it can't test whether a rule matches.

© QwenLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in .qwen/skills/e2e-testing of QwenLM/qwen-code.

  • SKILL.md
  • references/mcp-testing.md
  • references/mock-openai-server.md
  • scripts/mcp-test-server.js
  • scripts/mock-openai-server.js
  • scripts/token-stats.py

Open the folder on GitHubat commit fe4d4e3

Compare with similar skills

Qwen Code E2E Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Qwen Code E2E Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Qwen Code E2E Testing this skillQwenLM/qwen-code28k—~2.1kAutomated safety check: PassApache-2.0
Exploring The WizardPostHog/wizard197—~2.2kAutomated safety check: PassMIT
Suede MCP Release QAJasonColapietro/suede-creator-skills127—~2.1kAutomated safety check: PassMIT
Edt MCP TestingDitriXNew/EDT-MCP295—~2.1kAutomated safety check: PassAGPL-3.0
Browserstackalirezarezvani/claude-skills28k1 repos~1.3kAutomated safety check: PassMIT
Edt MCP Ready To DeployDitriXNew/EDT-MCP295—~1.6kAutomated safety check: PassAGPL-3.0

Similar skills

  • Exploring The Wizard

    PostHog/wizard

    Official

    Drive the PostHog wizard headlessly against a throwaway app through wizard-ci MCP tools, inspect decisions, and capture the real TUI.

    197 GitHub stars~2.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Suede MCP Release QA

    JasonColapietro/suede-creator-skills

    Checks a Suede AI MCP server release against a live process: the full JSON-RPC lifecycle, schemas, annotations, malformed input, catalog agreement and install docs.

    127 GitHub stars~2.1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Edt MCP Testing

    DitriXNew/EDT-MCP

    How to manually e2e-test each EDT-MCP server tool against a live EDT workbench + TestConfiguration.

    295 GitHub stars~2.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Browserstack

    alirezarezvani/claude-skills

    Run tests on BrowserStack. An agent skill from alirezarezvani/claude-skills.

    28k GitHub starsUsed in 1 repo~1.3k tokens
    Testing & QAAuto-check passed
  • Edt MCP Ready To Deploy

    DitriXNew/EDT-MCP

    The final "definition of done" / ready-to-deploy checklist for EDT-MCP — the ordered gate to run when a piece of work is finished, before declaring it done or merging.

    295 GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Glance Test

    DebugBase/glance

    Run E2E browser tests on any web application using Glance MCP.

    156 GitHub stars~827 tokensUpdated 5 mo ago
    Testing & QAAuto-check passed

More from QwenLM/qwen-code

All 41 skills in this repo
  • Reproduces a feature from Codex or Claude Code in Qwen Code by running the reference agent under capture, reading the traces, then implementing matching behavior.

    28k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Scheduled CI skill that scans a repository for small, certain docs, test and code hygiene issues and fixes them on one branch with a commit per finding.

    28k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Builds a rebranded Qwen Code desktop package from the Tauri shell using only a brand id and a logo, with sensible derived defaults.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through capturing and comparing V8 heap snapshots to find memory leaks in the Qwen Code Node.js CLI, using tmux and the chrome-devtools CLI.

    28k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Agent Reproduce Align

    QwenLM/qwen-code

    Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches.

    28k GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Qwen Code E2E Testing

What does Qwen Code E2E Testing do?

Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic. The skill explains how to run the Qwen Code CLI through the whole pipeline, from model API to tool validation to tool execution, when unit tests are not enough.js to verify a fix, or npm run dev for fast runtime-only checks.

When should I use Qwen Code E2E Testing?

Qwen Code E2E Testing fits situations like: verifying CLI behavior with real model calls; reproducing a user-reported bug end to end; testing an MCP tool integration with a test server; inspecting raw API request and response payloads.

How do I install Qwen Code E2E Testing in Claude Code?

Run `npx skills add QwenLM/qwen-code --skill e2e-testing -a claude-code`. Or copy the skill folder (.qwen/skills/e2e-testing in QwenLM/qwen-code) into .claude/skills/e2e-testing in your project. Claude Code loads it when a task matches its description.

How do I install Qwen Code E2E Testing in Codex?

Run `npx skills add QwenLM/qwen-code --skill e2e-testing -a codex`. Or copy the skill folder (.qwen/skills/e2e-testing in QwenLM/qwen-code) into .agents/skills/e2e-testing in your project. Codex loads it when a task matches its description.

Can I use Qwen Code E2E Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QwenLM/qwen-code --skill e2e-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/e2e-testing, .gemini/skills/e2e-testing, .github/skills/e2e-testing and .opencode/skills/e2e-testing in your project.

What does Qwen Code E2E Testing need to run?

Going by SKILL.md and its folder, Qwen Code E2E Testing needs JavaScript and Python for the scripts in its folder and the command-line tools its instructions call (npm, node, jq, python3 and bundle). Our summary lists: Node.js and npm, to build and run the CLI; jq, to filter the JSON output; Model provider credentials in ~/.qwen, for real model calls.

Does Qwen Code E2E Testing access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Qwen Code E2E Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Qwen Code E2E Testing use?

Qwen Code E2E Testing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Qwen Code E2E Testing use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.

What are the alternatives to Qwen Code E2E Testing?

Skills that share tags, products or a category with Qwen Code E2E Testing: Exploring The Wizard (PostHog/wizard, 197 stars), Suede MCP Release QA (JasonColapietro/suede-creator-skills, 127 stars), Edt MCP Testing (DitriXNew/EDT-MCP, 295 stars) and Browserstack (alirezarezvani/claude-skills, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Qwen Code E2E Testing?

QwenLM (a GitHub organization) maintains it in QwenLM/qwen-code, which has 28,349 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 8, 2026.

Source: QwenLM/qwen-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.