Agent skill

Agent Reproduce Align

by QwenLM in QwenLM/qwen-code

Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches.

Apache-2.0Auto-check passedAgent Workflows

Install Agent Reproduce Align

skills CLI
$ npx skills add QwenLM/qwen-code --skill agent-reproduce-align -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install QwenLM/qwen-code agent-reproduce-align --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.qwen/skills/agent-reproduce-align .claude/skills/agent-reproduce-align && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-reproduce-align
GitHub stars
28k
Token cost
~1.1k tokens
SKILL.md length
463 words
Files
5 (incl. scripts, references)
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches.

  • Works in 8 steps: Re-state the parity target → Run the reference agent and Qwen Code in… → Capture the selected reference agent's… → …
  • Checking that a reimplemented Codex or Claude Code feature behaves the same in Qwen Code
  • SKILL.md covers Purpose, Reference Agent Selection, Workflow and Common Commands, plus 2 more sections
  • Runs Python and Shell scripts from its folder

What it does

Meant for the stage after a feature has been implemented in Qwen Code and needs evidence-based parity with a reference agent, either `codex` or `claude-code`. The aim is to match the observable contract that matters for the feature, not byte-for-byte equality. The agent restates the parity target (feature, reference agent, a baseline prompt, acceptable differences and must-match fields) and runs both tools in separate capture directories.

Traces are normalized with `scripts/normalize_trace.py` and compared with `scripts/compare_traces.py`, while `scripts/run_pair_capture.sh` runs a paired shell scenario driven by the `REPRO_REFERENCE_AGENT` variable. Differences are inspected in a set order: reference-agent state changes, missing tool names, schema shape, model settings, prompt role order, then terminal output and exit status. The agent patches Qwen Code, reruns the smallest failing scenario and keeps only redacted minimal fixtures. Read `references/alignment-workflow.md` before the first comparison.

When your agent uses it

  • Checking that a reimplemented Codex or Claude Code feature behaves the same in Qwen Code
  • Comparing request bodies and tool schemas sent by two coding agents
  • Tracking down why a slash command prints different terminal output in two agents

Example prompts

  • “Our /help command now exists in Qwen Code. Run it against Claude Code and show where the traces differ.”
  • “Compare the tool schemas Codex sends with the ones Qwen Code sends for the same prompt.”
  • “Normalize the traces under .repro-runs/reference and .repro-runs/qwen and list the mismatched fields.”

Requirements

  • Qwen Code and the chosen reference agent (Codex or Claude Code)
  • Python and a POSIX shell for the bundled scripts

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Re-state the parity target
  2. Run the reference agent and Qwen Code in separate capture directories with the same scenario.
  3. Capture the selected reference agent's local state before and after the
  4. Normalize traces with scripts/normalize_trace.py.
  5. Compare normalized traces with scripts/compare_traces.py.
  6. Inspect differences in this order
  7. Patch Qwen Code, rerun the smallest failing scenario, and repeat.
  8. Preserve only redacted minimal fixtures in the repo.

What it can do on your machine

Read from SKILL.md and the folder at commit d9c6f8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python and Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Reproduce Align loads about 1.1k tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 463 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from QwenLM/qwen-code at commit d9c6f8c, republished under its Apache-2.0 licence (© QwenLM). 463 words, ~1,075 tokens.

Download SKILL.mdSave it as .claude/skills/agent-reproduce-align/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
agent-reproduce-align
description
Use after a Codex or Claude Code feature has been implemented in Qwen Code to run the selected reference agent and Qwen Code under the same scenario, capture HTTP and terminal traces, compare request bodies, tool/function schemas, outputs, and iterate until the reproduced behavior is close enough.

Agent Reproduce Align

Purpose

Use this skill when Qwen Code already has a candidate implementation and needs evidence-based parity with a selected reference agent: codex or claude-code. The goal is not byte-for-byte equality; it is matching the observable contract that matters for the feature.

Default target repo: the current working directory. Use a user-specified path only when the user explicitly provides one.

Reference Agent Selection

Use the same reference agent selected during $agent-reproduce-feature. If the earlier choice is unavailable, ask once and record the answer in the scenario or run notes.

Workflow

  1. Re-state the parity target:
    • feature name and trigger
    • selected reference agent
    • one baseline prompt or interaction script
    • acceptable differences
    • must-match fields
  2. Run the reference agent and Qwen Code in separate capture directories with the same scenario.
  3. Capture the selected reference agent's local state before and after the reference run when state may affect parity.
  4. Normalize traces with scripts/normalize_trace.py.
  5. Compare normalized traces with scripts/compare_traces.py.
  6. Inspect differences in this order:
    • reference-agent state changes that explain behavior
    • missing tool/function names
    • schema shape and required fields
    • model settings and response mode
    • prompt role/order differences that affect behavior
    • terminal-visible output and exit status
  7. Patch Qwen Code, rerun the smallest failing scenario, and repeat.
  8. Preserve only redacted minimal fixtures in the repo.

Read references/alignment-workflow.md before the first comparison pass.

Common Commands

Normalize:

sh
.qwen/skills/agent-reproduce-align/scripts/normalize_trace.py \
  .repro-runs/reference/http.jsonl \
  > .repro-runs/reference/normalized.json

Compare:

sh
.qwen/skills/agent-reproduce-align/scripts/compare_traces.py \
  .repro-runs/reference/normalized.json \
  .repro-runs/qwen/normalized.json

Run a paired shell scenario:

sh
REPRO_REFERENCE_AGENT=codex \
.qwen/skills/agent-reproduce-align/scripts/run_pair_capture.sh \
  .repro-runs/slash-help \
  "codex exec '/help'" \
  "npm test -- --runInBand"

For Claude Code, set REPRO_REFERENCE_AGENT=claude-code and replace the first command with the discovered Claude Code command. When REPRO_REFERENCE_AGENT is set, the paired runner writes reference/state-before, reference/state-after, and reference/state-diff. Use the paired runner only when shell quoting is simple. For interactive slash commands, run the two captures manually with tmux so each side can receive the same keystrokes. Use REPRO_REFERENCE_STATE_ROOT=/tmp/some-root only for tests or custom state directories.

Show full SKILL.md (165 more words)Show less

Comparison Rules

  • Compare contracts before wording. Exact prompt text is usually implementation detail.
  • Treat absent schemas, wrong required fields, or wrong argument names as high-signal failures.
  • Treat output ordering as significant only when the user-visible workflow depends on it.
  • Do not chase provider-specific endpoints, model names, IDs, timestamps, token counts, or ephemeral headers unless the feature depends on them.
  • Do not chase every local state write. Treat state diffs as explanatory evidence unless the feature contract requires a particular config, memory, or permission side effect.
  • Stop when Qwen Code passes the user-visible scenario and the remaining trace differences are documented as intentional.

Done Criteria

  • Reference-agent and Qwen Code traces for the same scenario exist locally.
  • Reference-agent state diff exists or state capture is documented as irrelevant for the scenario.
  • The normalized comparison has no unexplained must-match differences.
  • Qwen Code tests or smoke commands cover the fixed behavior.
  • Any remaining mismatch is written down in the task notes or Qwen Code docs when it affects users.

© QwenLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in .qwen/skills/agent-reproduce-align of QwenLM/qwen-code.

  • SKILL.md
  • references/alignment-workflow.md
  • scripts/compare_traces.py
  • scripts/normalize_trace.py
  • scripts/run_pair_capture.sh

Open the folder on GitHubat commit d9c6f8c

Compare with similar skills

Agent Reproduce Align next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Reproduce Align compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Reproduce Align this skillQwenLM/qwen-code28k—~1.1kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers297k3 repos~1.7kAutomated safety check: PassMIT
Debug Cuda Crashsgl-project/sglang37k2 repos~4.9kAutomated safety check: PassApache-2.0
Debugging And Error Recoveryskuramatata/my-pi-agent114—~777Automated safety check: PassNone
Debugging And Error Recoveryskuramatata/my-pi-agent114—~829Automated safety check: PassNone
Browser DebuggingMadAppGang/claude-code285—~4.3kAutomated safety check: PassMIT

Similar skills

  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    297k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Debug Cuda Crash

    sgl-project/sglang

    Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging

    37k GitHub starsUsed in 2 repos~4.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Debugging And Error Recovery

    skuramatata/my-pi-agent

    A skill your agent uses when my-pi-agent tests, typecheck, lint, Pi CLI startup, OpenCode/MCP, Feishu channel, prompt rendering, memory/state, web-console/desktop build, or UI behavior fails…

    114 GitHub stars~777 tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Debugging And Error Recovery

    skuramatata/my-pi-agent

    A skill your agent uses when my-pi-agent tests, typecheck, lint, /code workflow, verifier probes, Feishu channel handling, MCP bootstrap, skill install, or task resume behavior fails unexpectedly.

    114 GitHub stars~829 tokensUpdated 3 mo ago
    DevelopmentAuto-check passed
  • Browser Debugging

    MadAppGang/claude-code

    Systematically tests UI functionality, validates design fidelity with AI visual analysis, monitors console output, tracks network requests, and provides debugging reports using Chrome Extension MCP…

    285 GitHub stars~4.3k tokensUpdated 7 mo ago
    DevelopmentAuto-check passed
  • Autoresearch Iteration Loop

    uditgoenka/autoresearch

    Runs an autonomous modify, verify, keep-or-discard loop against any metric, with subcommands for planning, debugging, fixing, security audits, shipping and more.

    6.5k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed

More from QwenLM/qwen-code

All 41 skills in this repo
  • Reproduces a feature from Codex or Claude Code in Qwen Code by running the reference agent under capture, reading the traces, then implementing matching behavior.

    28k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Scheduled CI skill that scans a repository for small, certain docs, test and code hygiene issues and fixes them on one branch with a commit per finding.

    28k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Builds a rebranded Qwen Code desktop package from the Tauri shell using only a brand id and a logo, with sensible derived defaults.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through capturing and comparing V8 heap snapshots to find memory leaks in the Qwen Code Node.js CLI, using tmux and the chrome-devtools CLI.

    28k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Agent Reproduce Align

What does Agent Reproduce Align do?

Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches. Meant for the stage after a feature has been implemented in Qwen Code and needs evidence-based parity with a reference agent, either `codex` or `claude-code`. The aim is to match the observable contract that matters for the feature, not byte-for-byte equality.

When should I use Agent Reproduce Align?

Agent Reproduce Align fits situations like: checking that a reimplemented Codex or Claude Code feature behaves the same in Qwen Code; comparing request bodies and tool schemas sent by two coding agents; tracking down why a slash command prints different terminal output in two agents.

How do I install Agent Reproduce Align in Claude Code?

Run `npx skills add QwenLM/qwen-code --skill agent-reproduce-align -a claude-code`. Or copy the skill folder (.qwen/skills/agent-reproduce-align in QwenLM/qwen-code) into .claude/skills/agent-reproduce-align in your project. Claude Code loads it when a task matches its description.

How do I install Agent Reproduce Align in Codex?

Run `npx skills add QwenLM/qwen-code --skill agent-reproduce-align -a codex`. Or copy the skill folder (.qwen/skills/agent-reproduce-align in QwenLM/qwen-code) into .agents/skills/agent-reproduce-align in your project. Codex loads it when a task matches its description.

Can I use Agent Reproduce Align in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QwenLM/qwen-code --skill agent-reproduce-align -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-reproduce-align, .gemini/skills/agent-reproduce-align, .github/skills/agent-reproduce-align and .opencode/skills/agent-reproduce-align in your project.

What does Agent Reproduce Align need to run?

Going by SKILL.md and its folder, Agent Reproduce Align needs Python and a shell for the scripts in its folder. Our summary lists: Qwen Code and the chosen reference agent (Codex or Claude Code); Python and a POSIX shell for the bundled scripts.

Does Agent Reproduce Align access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Reproduce Align safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Reproduce Align use?

Agent Reproduce Align is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Reproduce Align use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 603 tokens, read only when the agent opens those files.

What are the alternatives to Agent Reproduce Align?

Skills that share tags, products or a category with Agent Reproduce Align: Diagnosing Superpowers Sessions (obra/superpowers, 297k stars), Debug Cuda Crash (sgl-project/sglang, 37k stars), Debugging And Error Recovery (skuramatata/my-pi-agent, 114 stars) and Debugging And Error Recovery (skuramatata/my-pi-agent, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Reproduce Align?

QwenLM (a GitHub organization) maintains it in QwenLM/qwen-code, which has 28,410 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 11, 2026.

Source: QwenLM/qwen-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.