Agent skill

Claude CLI Real E2E Tests

by liaohch3 in liaohch3/claude-tap

Runs real end-to-end tests for claude-tap against the actual Claude CLI, in a pytest mode and an interactive tmux mode, and checks the resulting traces.

MITAuto-check passedTesting & QA

Install Claude CLI Real E2E Tests

skills CLI
$ npx skills add liaohch3/claude-tap --skill real-e2e-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install liaohch3/claude-tap real-e2e-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/liaohch3/claude-tap.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/real-e2e-test .claude/skills/real-e2e-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
real-e2e-test
GitHub stars
3.3k
Token cost
~693 tokens
SKILL.md length
269 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Runs real end-to-end tests for claude-tap against the actual Claude CLI, in a pytest mode and an interactive tmux mode, and checks the resulting traces.

  • Validating claude-tap end to end against a real Claude CLI session
  • SKILL.md covers Prerequisites, Mode 1: Pytest Real E2E (7…, Mode 2: tmux Interactive Real… and Verification Checklist (for…, plus 2 more sections
  • Calls uv and brew; needs SUBMIT_KEY
  • Debugging a single real E2E test with verbose output

What it does

The tests start claude-tap from local source, connect to the real Claude CLI and verify the trace output. They need the claude CLI installed and authenticated, Python dev dependencies installed with uv sync --extra dev, and tmux for the interactive mode.

Mode one is a pytest suite of seven real E2E cases, run with uv run pytest tests/e2e/ and the --run-real-e2e flag, since these tests are skipped by default. A single test or debug output can be selected, and each case starts a fresh proxy server and trace directory. Cases cover single-turn and multi-turn capture, tool use, HTML viewer generation, API key redaction and streaming SSE capture.

Mode two uses scripts/run_real_e2e_tmux.sh to drive the Claude Code TUI interactively when non-print behavior must be validated, with prompts, submit key and other values overridable through environment variables. Both modes finish with a checklist: the newest trace file contains both prompts, at least two requests hit /v1/messages, one response has a tool_use block, and an HTML viewer file is generated. For portability, shell assertions use grep -F instead of rg.

When your agent uses it

  • Validating claude-tap end to end against a real Claude CLI session
  • Debugging a single real E2E test with verbose output
  • Checking interactive Claude Code behavior inside tmux

Example prompts

  • “Run the real E2E pytest suite for claude-tap and report any failures.”
  • “Run only test_single_turn with debug output.”
  • “Run the tmux real E2E script and check that both prompts appear in the latest trace.”

Requirements

  • The claude CLI, installed and authenticated
  • uv with the dev dependencies installed
  • tmux, for the interactive mode

What it can do on your machine

Read from SKILL.md and the folder at commit 4cc867d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SUBMIT_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Claude CLI Real E2E Tests loads about 693 tokens when it runs. Until then it costs about 19 tokens; SKILL.md has 269 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~693

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from liaohch3/claude-tap at commit 4cc867d, republished under its MIT licence (© liaohch3). 269 words, ~693 tokens.

Download SKILL.mdSave it as .claude/skills/real-e2e-test/SKILL.md (or your agent's skills folder).
name
real-e2e-test
description
Run real E2E tests against Claude CLI in pytest and tmux modes
tags
testing, e2e, integration, tmux

Real E2E Test Skill

Run real end-to-end tests that start claude-tap from local source, connect to the real Claude CLI, and verify trace output.

Prerequisites

  • claude CLI installed and authenticated
  • Python dev dependencies installed: uv sync --extra dev
  • tmux installed for interactive mode (brew install tmux)

Mode 1: Pytest Real E2E (7 test cases)

Run all real E2E tests
bash
uv run pytest tests/e2e/ --run-real-e2e --timeout=300 -v
Run a single test
bash
uv run pytest tests/e2e/test_real_proxy.py::TestRealProxy::test_single_turn --run-real-e2e --timeout=180 -v -s
Run with debug output
bash
uv run pytest tests/e2e/ --run-real-e2e --timeout=300 -v -s --tb=long

Mode 2: tmux Interactive Real E2E

Use this when you need to validate non--p interactive behavior in Claude Code TUI.

bash
scripts/run_real_e2e_tmux.sh

Optional overrides:

bash
PROMPT_ONE="Use the shell tool to run command ls in the current directory, then reply with any 5 filenames only." \
PROMPT_TWO="Thank you." \
SUBMIT_KEY="Enter" \
PERMISSION_MODE="bypassPermissions" \
scripts/run_real_e2e_tmux.sh

Important tmux interaction notes:

  • Submit key is Enter for Claude Code TUI in tmux (confirmed working).
  • PROMPT_ONE should intentionally trigger tool use.
  • For portability, use grep -F instead of rg in shell assertions (rg may be unavailable).

Verification Checklist (for both modes)

  • Latest trace .jsonl contains both prompts (PROMPT_ONE, PROMPT_TWO)
  • At least 2 requests hit /v1/messages
  • At least one response content block has "type": "tool_use"
  • HTML viewer file is generated (trace_*.html)

Notes

  • Real E2E tests are skipped by default; --run-real-e2e is required.
  • Each pytest case starts a fresh proxy server and trace directory.
  • Timeouts are intentionally generous because real API calls are involved.
  • tmux mode includes retry logic for prompt submission and post-run JSONL assertions.

Pytest Test Cases

TestTimeoutWhat It Tests
test_single_turn180sBasic prompt/response trace capture
test_multi_turn300sConversation memory with -c flag
test_tool_use180sTool use generates multiple trace records
test_html_viewer_generated180sHTML viewer generated with embedded trace data
test_api_key_redaction180sAPI keys redacted from trace output
test_streaming_sse_capture180sSSE events captured in streaming mode
test_trace_summary180sCLI stdout includes trace summary and API call count

© liaohch3, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/real-e2e-test of liaohch3/claude-tap.

Open the folder on GitHubat commit 4cc867d

Compare with similar skills

Claude CLI Real E2E Tests next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Claude CLI Real E2E Tests compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Claude CLI Real E2E Tests this skillliaohch3/claude-tap3.3k—~693Automated safety check: PassMIT
CodexBar Live QAsteipete/CodexBar22k—~1.2kAutomated safety check: PassMIT
Drive MiMo CodeXiaomiMiMo/MiMo-Code14k—~3.9kAutomated safety check: PassMIT
Senpi Agent QA Harnesscode-yeongyu/senpi470—~2.7kAutomated safety check: NotesMIT
tmux Real User TestingQwenLM/qwen-code28k—~2.3kAutomated safety check: PassApache-2.0
Verify Omnigent End-to-Endomnigent-ai/omnigent11k—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • CodexBar Live QA

    steipete/CodexBar

    Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely.

    22k GitHub stars~1.2k tokensUpdated today
    Testing & QAAuto-check passed
  • Drive MiMo Code

    XiaomiMiMo/MiMo-Code

    Lets one MiMoCode process drive another, headless with JSON events or interactively through tmux, to test behavior and visual regressions with parseable evidence.

    14k GitHub stars~3.9k tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Senpi Agent QA Harness

    code-yeongyu/senpi

    Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.

    470 GitHub stars~2.7k tokensUpdated today
    Testing & QAAuto-check: notes
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Verify Omnigent End-to-End

    omnigent-ai/omnigent

    Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code.

    11k GitHub stars~1.6k tokensUpdated today
    Testing & QAAuto-check passed
  • Designing Tests

    CloudAI-X/claude-workflow-v2

    Designs and implements testing strategies for any codebase. An agent skill from CloudAI-X/claude-workflow-v2.

    1.4k GitHub starsUsed in 1 repo~1.5k tokens
    Testing & QAAuto-check passed

More from liaohch3/claude-tap

All 11 skills in this repo
  • Codex E2E Trace Validation

    liaohch3/claude-tap

    Runs a real Codex CLI session through claude-tap and produces trace evidence and viewer screenshots for pull requests that touch capture, proxying or the viewer.

    3.3k GitHub stars~3k tokensUpdated 15 days ago
    Auto-check passed
  • Terminal Demo Video Recorder

    liaohch3/claude-tap

    Records a real tmux end-to-end run with asciinema, converts it to a GIF and MP4, and captures trace viewer screenshots with Playwright.

    3.3k GitHub stars~545 tokensUpdated 15 days ago
    Auto-check passed
  • JS-in-HTML Testing

    liaohch3/claude-tap

    Tests JavaScript embedded in an HTML file in two layers: pytest checks of the logic ported to Python, and Playwright runs in a real browser for the DOM.

    3.3k GitHub stars~924 tokensUpdated 15 days ago
    Auto-check passed
  • Runs local checks on maintainer docs for stale review dates, missing manifest paths and plan state drift, matching what the CI workflow enforces.

    3.3k GitHub stars~561 tokensUpdated 15 days ago
    Auto-check passed
  • Playwright Screen Recording

    liaohch3/claude-tap

    Records headless Playwright sessions as .webm videos to show a bug fix working or to give pull request reviewers visual evidence.

    3.3k GitHub stars~714 tokensUpdated 15 days ago
    Auto-check passed
  • PR Preflight Check

    liaohch3/claude-tap

    Runs a single merge-readiness check on a pull request: metadata, GitHub Actions status, local lint, format and test gates, and PR body rules, ending in READY or NOT_READY.

    3.3k GitHub stars~615 tokensUpdated 15 days ago
    Auto-check passed

Works with

Categories

Questions about Claude CLI Real E2E Tests

What does Claude CLI Real E2E Tests do?

Runs real end-to-end tests for claude-tap against the actual Claude CLI, in a pytest mode and an interactive tmux mode, and checks the resulting traces. The tests start claude-tap from local source, connect to the real Claude CLI and verify the trace output. They need the claude CLI installed and authenticated, Python dev dependencies installed with uv sync --extra dev, and tmux for the interactive mode.

When should I use Claude CLI Real E2E Tests?

Claude CLI Real E2E Tests fits situations like: validating claude-tap end to end against a real Claude CLI session; debugging a single real E2E test with verbose output; checking interactive Claude Code behavior inside tmux.

How do I install Claude CLI Real E2E Tests in Claude Code?

Run `npx skills add liaohch3/claude-tap --skill real-e2e-test -a claude-code`. Or copy the skill folder (.agents/skills/real-e2e-test in liaohch3/claude-tap) into .claude/skills/real-e2e-test in your project. Claude Code loads it when a task matches its description.

How do I install Claude CLI Real E2E Tests in Codex?

Run `npx skills add liaohch3/claude-tap --skill real-e2e-test -a codex`. Or copy the skill folder (.agents/skills/real-e2e-test in liaohch3/claude-tap) into .agents/skills/real-e2e-test in your project. Codex loads it when a task matches its description.

Can I use Claude CLI Real E2E Tests in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add liaohch3/claude-tap --skill real-e2e-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/real-e2e-test, .gemini/skills/real-e2e-test, .github/skills/real-e2e-test and .opencode/skills/real-e2e-test in your project.

What does Claude CLI Real E2E Tests need to run?

Going by SKILL.md and its folder, Claude CLI Real E2E Tests needs the command-line tools its instructions call (uv and brew) and credentials named SUBMIT_KEY. Our summary lists: The claude CLI, installed and authenticated; uv with the dev dependencies installed; tmux, for the interactive mode.

Does Claude CLI Real E2E Tests access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Claude CLI Real E2E Tests safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Claude CLI Real E2E Tests use?

Claude CLI Real E2E Tests is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Claude CLI Real E2E Tests use?

About 693 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Claude CLI Real E2E Tests?

Skills that share tags, products or a category with Claude CLI Real E2E Tests: CodexBar Live QA (steipete/CodexBar, 22k stars), Drive MiMo Code (XiaomiMiMo/MiMo-Code, 14k stars), Senpi Agent QA Harness (code-yeongyu/senpi, 470 stars) and tmux Real User Testing (QwenLM/qwen-code, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Claude CLI Real E2E Tests?

liaohch3 (a GitHub user) maintains it in liaohch3/claude-tap, which has 3,267 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on September 22, 2026.

Source: liaohch3/claude-tap on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.