tmux Real User Testing
QwenLM/qwen-code
Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.
Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.
$ npx skills add code-yeongyu/senpi --skill senpi-qa -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install code-yeongyu/senpi senpi-qa --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/code-yeongyu/senpi.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/senpi-qa .claude/skills/senpi-qa && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "senpi-qa" agent skill from https://github.com/code-yeongyu/senpi/tree/main/.agents/skills/senpi-qa into .claude/skills/senpi-qa/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "senpi-qa", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/code-yeongyu/senpi/tree/main/.agents/skills/senpi-qaType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add code-yeongyu/senpi --skill senpi-qa -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install code-yeongyu/senpi senpi-qa --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/code-yeongyu/senpi.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/senpi-qa .agents/skills/senpi-qa && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "senpi-qa" agent skill from https://github.com/code-yeongyu/senpi/tree/main/.agents/skills/senpi-qa into .agents/skills/senpi-qa/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "senpi-qa", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add code-yeongyu/senpi --skill senpi-qa -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install code-yeongyu/senpi senpi-qa --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/code-yeongyu/senpi.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/senpi-qa .cursor/skills/senpi-qa && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "senpi-qa" agent skill from https://github.com/code-yeongyu/senpi/tree/main/.agents/skills/senpi-qa into .cursor/skills/senpi-qa/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "senpi-qa", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/code-yeongyu/senpi.git --path .agents/skills/senpi-qa--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add code-yeongyu/senpi --skill senpi-qa -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install code-yeongyu/senpi senpi-qa --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/code-yeongyu/senpi.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/senpi-qa .gemini/skills/senpi-qa && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "senpi-qa" agent skill from https://github.com/code-yeongyu/senpi/tree/main/.agents/skills/senpi-qa into .gemini/skills/senpi-qa/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "senpi-qa", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install code-yeongyu/senpi senpi-qaInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add code-yeongyu/senpi --skill senpi-qa -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/code-yeongyu/senpi.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/senpi-qa .github/skills/senpi-qa && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "senpi-qa" agent skill from https://github.com/code-yeongyu/senpi/tree/main/.agents/skills/senpi-qa into .github/skills/senpi-qa/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "senpi-qa", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add code-yeongyu/senpi --skill senpi-qa -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install code-yeongyu/senpi senpi-qa --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/code-yeongyu/senpi.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/senpi-qa .opencode/skills/senpi-qa && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "senpi-qa" agent skill from https://github.com/code-yeongyu/senpi/tree/main/.agents/skills/senpi-qa into .opencode/skills/senpi-qa/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "senpi-qa", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
senpi-qaChecks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.
Passing typecheck and unit tests do not count as QA for senpi, so this harness runs the actual CLI through `tsx` after changes to the ai, agent, coding-agent or tui packages. Each run points `SENPI_CODING_AGENT_DIR` at a temporary sandbox with offline mode on, never writes to the real `~/.senpi` directory, and hashes the real auth file to confirm it is unchanged at the end.
Four channels exist: remote RPC with JSONL over stdio, TUI smoke tests using node-pty on Windows or tmux on POSIX systems, a mock loop backed by a local fake model server for deterministic runs that spend no tokens, and a CLI smoke check of `--help`, `--print` and `--list-models`. Every helper script has a `--self-test`. Evidence must be saved under `local-ignore/qa-evidence/`, since no artifact means the QA did not happen, and the skill reports bugs rather than editing `src/`.
Read from SKILL.md and the folder at commit 6073042. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (JavaScript, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
nodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Senpi Agent QA Harness loads about 2.7k tokens when it runs, and up to ~8.2k if it reads all its reference files. Until then it costs about 207 tokens; SKILL.md has 969 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
# installs skill deps (node-pty), wires .env.local + .claude/skillsAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from code-yeongyu/senpi at commit 6073042, republished under its MIT licence (© code-yeongyu). 969 words, ~2,721 tokens.
.claude/skills/senpi-qa/SKILL.md (or your agent's skills folder). This skill also uses 158 other files; get the full folder from GitHub.QA the senpi coding agent (packages/{ai,agent,coding-agent,tui}) by driving the
REAL CLI — not by reading code or trusting unit tests. Each channel runs the
agent from source via tsx in an isolated sandbox and asserts observable
behavior, so a passing run is evidence the user-facing surface actually works.
Every helper script ships a --self-test (or --self-check) that asserts its
scenario against this machine. The scripts are therefore both the QA tools and
their own regression checks.
SENPI_CODING_AGENT_DIR / SENPI_CODING_AGENT_SESSION_DIR pointed at a temp
sandbox and PI_OFFLINE=1. QA must never write into the real ~/.senpi.
scripts/lib/common.mjs does this for you — use it.~/.senpi/agent/auth.json is
the user's real key store. Every script snapshots its sha256 and asserts it is
unchanged at the end. If you script a run by hand, do the same
(guardRealAuth() in common.mjs).src/ edits from this skill. It verifies; it does not fix. If QA finds
a bug, report it with the captured evidence and let a follow-up change fix it.local-ignore/qa-evidence/<YYYYMMDD>-<slug>/. No artifact == the QA did not
happen. local-ignore/ is gitignored — never commit evidence.node scripts/devenv-setup.mjs # installs skill deps (node-pty), wires .env.local + .claude/skills
node .agents/skills/senpi-qa/scripts/lib/common.mjs --self-check # confirm the harnesscommon.mjs --self-check confirms the repo resolves, a sandbox is created and
auto-removed, a free port is allocatable, and the real auth file is untouched.
| You changed… | Run this channel | Reference |
|---|---|---|
| Agent loop, tools, sessions, provider/model resolution, RPC | Channel 1 (RPC) — and Channel 3 for a deterministic loop | references/rpc-protocol.md |
| Interactive TUI, keybindings, rendering, composer | Channel 2 (TUI smoke) | references/tui-driving.md |
| Anything where you want a full agent turn with ZERO tokens | Channel 3 (mock loop) | references/mock-loop.md |
CLI flags, --help, --print, model listing | Channel 4 (CLI smoke) | — |
Persistent terminal tools / packages/pty / PTY sessions | Channel 5 (pty-drive) | — |
| Added a provider / auth path | Channel 3 + 4, and update references/env-vars.md | references/credential-injection.md |
When in doubt, run the channel closest to your change AND Channel 3 (mock loop): the mock loop is the cheapest end-to-end proof that the agent still completes a turn.
All commands are run from the repo root.
scripts/rpc-drive.mjs)Drives --mode rpc (JSON lines over stdio). get_state round-trips with no API
call; --prompt drives a real turn and captures the event stream.
node .agents/skills/senpi-qa/scripts/rpc-drive.mjs --self-test
node .agents/skills/senpi-qa/scripts/rpc-drive.mjs --state
node .agents/skills/senpi-qa/scripts/rpc-drive.mjs --prompt "say PONG" --provider mock --model mock-model --evidence rpc-pongscripts/tui-smoke.mjs)Boots the interactive TUI in a real pseudo-terminal, confirms it renders and a keystroke reaches the composer, then tears it down. Uses node-pty (ConPTY on Windows — no WSL) and falls back to tmux on POSIX.
node .agents/skills/senpi-qa/scripts/tui-smoke.mjs --self-test
node .agents/skills/senpi-qa/scripts/tui-smoke.mjs --self-test --driver tmux --evidence tuiTUI smoke proves boot/render/input, not fine-grained output. For behavioral assertions use Channel 1 or 3.
scripts/mock-loop.mjs)Starts a local fake model server, registers it via a baseUrl override in an
isolated models.json, and drives a REAL turn — deterministic, zero tokens.
Covers all three wire formats senpi uses, so baseUrl override is QA'd for both
OpenAI and Anthropic (pick with --api; default openai-completions):
--api | provider overridden | path / auth |
|---|---|---|
openai-completions | mock | /v1/chat/completions · Bearer |
anthropic-messages | anthropic | /v1/messages · x-api-key |
openai-responses | openai | /v1/responses · Bearer |
--self-test (no --api) round-trips all three and exercises the three
error/retry scenarios below. --with-tool proves the full loop (model → bash tool → final text). --with-mcp-tool registers a sandbox
extension that proxies mcp_fx_tool_<n> to the local MCP stdio fixture, then
asserts the fixture call log exists and the model's second request contains the
fixture result. The loop is hermetic: provider key env vars are stripped so only
the inline mock key is ever used.
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --self-test
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --self-test --api anthropic-messages
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --with-tool --api openai-responses
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --with-mcp-tool mcp_fx_tool_1 --tool-args '{"value":"ok"}'
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --with-eval-hard-limit --evidence eval-hard-limit
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --scenario transient-recover
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --scenario budget-exhaust
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --scenario long-retry-after
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --run "summarize this repo" --evidence mock-summaryscripts/cli-smoke.mjs)Fast, no model: --help, --version, offline --list-models, unknown-flag
handling.
node .agents/skills/senpi-qa/scripts/cli-smoke.mjs --self-testscripts/pty-drive.mjs)Drives the real @earendil-works/pi-pty runtime that backs the built-in terminal
tools (bash / bash_output / bash_input / bash_resize / kill_bash) through the
canonical scenarios: background command + monitor-style line watch and peek, stdin
steering, screen snapshot + resize reflow, and registry teardown with no orphans. Uses native PTY when a host
prebuild is present, else the pipe fallback (--force-pipe forces it).
node .agents/skills/senpi-qa/scripts/pty-drive.mjs --self-test --evidence terminal
node .agents/skills/senpi-qa/scripts/pty-drive.mjs --self-test --force-pipe| Script | --self-test / --self-check asserts |
|---|---|
scripts/lib/common.mjs --self-check | repo + tsx resolve; sandbox created and auto-removed; free port; real auth.json unchanged |
scripts/lib/fake-model-server.mjs --self-test | OpenAI SSE contract: scripted text streams back, [DONE] sent, request recorded |
scripts/rpc-drive.mjs --self-test | get_state returns the documented RpcSessionState, no API call, auth unchanged |
scripts/mock-loop.mjs --self-test | scripted marker returns through the real loop via the mock provider; retry error injection proves same-model recovery, retry-budget fallback, and long-retry-after fallback; zero real calls; auth unchanged |
scripts/mock-loop.mjs --with-tool | full loop: two model turns served, bash tool ran, final text returned |
scripts/mock-loop.mjs --with-mcp-tool <tool> | full loop with a registered sandbox MCP stdio fixture proxy; fails if the requested mcp_fx_tool_<n> is not registered, invoked, and fed back to the model |
scripts/mock-loop.mjs --with-eval-hard-limit | full loop where a never-returning eval cell is killed by its wall-clock hard limit and the kill is reported back to the model |
scripts/eval-hard-limit-rpc-qa.mjs --self-test | RPC channel: a DETACHED eval cell killed by the hard limit injects its <system-reminder> kill notice into the model's next request (print mode cannot show this) |
scripts/tui-smoke.mjs --self-test | TUI boots, renders, accepts a keystroke, tears down; auth unchanged |
scripts/cli-smoke.mjs --self-test | --help/--version/--list-models work offline; unknown flag reported; auth unchanged |
scripts/pty-drive.mjs --self-test | PTY runtime backing the terminal tools: background line watch + peek, stdin steering, screen snapshot + resize, registry teardown (no orphans); auth unchanged |
Run the whole suite:
for s in lib/common.mjs:--self-check lib/fake-model-server.mjs:--self-test \
rpc-drive.mjs:--self-test mock-loop.mjs:--self-test \
tui-smoke.mjs:--self-test cli-smoke.mjs:--self-test; do
node ".agents/skills/senpi-qa/scripts/${s%%:*}" "${s##*:}" || echo "FAILED: $s"
doneev="local-ignore/qa-evidence/$(date +%Y%m%d)-senpi-qa-<slug>"; mkdir -p "$ev"
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --self-test | tee "$ev/mock-loop.txt"
node .agents/skills/senpi-qa/scripts/rpc-drive.mjs --prompt "say PONG" \
--provider mock --model mock-model --evidence senpi-qa-<slug>Most channels accept --evidence <slug> and write artifacts to
local-ignore/qa-evidence/<date>-<slug>/ themselves.
references/rpc-protocol.md — RPC command/response catalog, turn completion, examplesreferences/tui-driving.md — node-pty vs tmux, keybindings files, fragility, isolationreferences/mock-loop.md — fake server, custom-provider models.json shape, in-process faux alternativereferences/credential-injection.md — per-harness credential injection + maskingreferences/env-vars.md — provider keys + isolation env vars© code-yeongyu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 158 other files (scripts, references) in .agents/skills/senpi-qa of code-yeongyu/senpi.
Open the folder on GitHubat commit 6073042
Senpi Agent QA Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Senpi Agent QA Harness this skillcode-yeongyu/senpi | 472 | — | ~2.7k | Automated safety check: Notes | MIT | |
| tmux Real User TestingQwenLM/qwen-code | 28k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| CodexBar Live QAsteipete/CodexBar | 22k | — | ~1.2k | Automated safety check: Pass | MIT | |
| Verify Omnigent End-to-Endomnigent-ai/omnigent | 11k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| OpenHarness End-to-End EvalsHKUDS/OpenHarness | 16k | 1 repos | ~2.1k | Automated safety check: Notes | MIT | |
| Acceptance Evidence for Deliverieslobehub/lobehub | 83k | — | ~9.7k | Automated safety check: Pass | Apache-2.0 |
QwenLM/qwen-code
Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.
steipete/CodexBar
Runs live QA for the CodexBar app: provider usage matrix checks through its packaged CLI, config validation and menu checks, with 1Password-backed credentials handled safely.
omnigent-ai/omnigent
Spins up an isolated Omnigent server, runner and mock model to prove a user-facing behavior or bug fix with recorded evidence instead of reasoning from code.
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
lobehub/lobehub
Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.
XiaomiMiMo/MiMo-Code
Lets one MiMoCode process drive another, headless with JSON events or interactively through tmux, to test behavior and visual regressions with parseable evidence.
code-yeongyu/senpi
Syncs a fork branch with its upstream remote using a history-preserving merge commit, with no rebase and no force push.
code-yeongyu/senpi
Points the agent at Bun 1.4 built-in APIs before it installs an npm package, so image, browser, markdown, cron, PTY and test work uses what Bun already ships.
code-yeongyu/senpi
Worker brief for implementing one pre-assigned feature in the senpi todotools built-in extension, with strict scope, typing, testing and git-safety rules.
code-yeongyu/senpi
Prompt-crafting guide for gpt-image-2.5: which image tool to call, which model to pick, and how to write prompts, edit with references and refine over turns.
code-yeongyu/senpi
Walks the canonical CalVer release flow for senpi, from a clean main checkout through changelog audit, checks, tag push, GitHub Release and npm publishing.
code-yeongyu/senpi
Runs one manual QA scenario for the todo continuation feature in the real ./pi-test.sh CLI inside tmux, captures scrollback and checks a deterministic count marker.
Works with
Categories
Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels. Passing typecheck and unit tests do not count as QA for senpi, so this harness runs the actual CLI through `tsx` after changes to the ai, agent, coding-agent or tui packages.senpi` directory, and hashes the real auth file to confirm it is unchanged at the end.
Senpi Agent QA Harness fits situations like: verifying an agent-loop, tool or provider change end to end; checking keybindings or rendering in the interactive TUI after an edit; running a deterministic agent-loop test without spending real tokens; smoke-testing the CLI flags after a refactor.
Run `npx skills add code-yeongyu/senpi --skill senpi-qa -a claude-code`. Or copy the skill folder (.agents/skills/senpi-qa in code-yeongyu/senpi) into .claude/skills/senpi-qa in your project. Claude Code loads it when a task matches its description.
Run `npx skills add code-yeongyu/senpi --skill senpi-qa -a codex`. Or copy the skill folder (.agents/skills/senpi-qa in code-yeongyu/senpi) into .agents/skills/senpi-qa in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add code-yeongyu/senpi --skill senpi-qa -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/senpi-qa, .gemini/skills/senpi-qa, .github/skills/senpi-qa and .opencode/skills/senpi-qa in your project.
Going by SKILL.md and its folder, Senpi Agent QA Harness needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node). Our summary lists: The senpi repository with Node.js and `tsx`; `node-pty` on Windows or `tmux` on POSIX systems; A one-time run of `node scripts/devenv-setup.mjs`.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Senpi Agent QA Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Senpi Agent QA Harness: tmux Real User Testing (QwenLM/qwen-code, 28k stars), CodexBar Live QA (steipete/CodexBar, 22k stars), Verify Omnigent End-to-End (omnigent-ai/omnigent, 11k stars) and OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
code-yeongyu (a GitHub user) maintains it in code-yeongyu/senpi, which has 472 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 8, 2026.
Source: code-yeongyu/senpi on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.