Copilot Session Failure Analysis
dotnet/maui
Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.
Runbook for the five-arm Codex harness comparison on Terminal-Bench 4.0, pointing out four differences from SWE-Marathon that silently produce wrong scores.
SKILL.md written in Chinese; this summary is our English description.
$ npx skills add loopx-project/loopx --skill tb4-five-arm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install loopx-project/loopx tb4-five-arm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/loopx-project/loopx.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmark/swe-marathon/skills/tb4-five-arm .claude/skills/tb4-five-arm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tb4-five-arm" agent skill from https://github.com/loopx-project/loopx/tree/main/benchmark/swe-marathon/skills/tb4-five-arm into .claude/skills/tb4-five-arm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tb4-five-arm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/loopx-project/loopx/tree/main/benchmark/swe-marathon/skills/tb4-five-armType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add loopx-project/loopx --skill tb4-five-arm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install loopx-project/loopx tb4-five-arm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/loopx-project/loopx.git skills-src && mkdir -p .agents/skills && cp -r skills-src/benchmark/swe-marathon/skills/tb4-five-arm .agents/skills/tb4-five-arm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tb4-five-arm" agent skill from https://github.com/loopx-project/loopx/tree/main/benchmark/swe-marathon/skills/tb4-five-arm into .agents/skills/tb4-five-arm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tb4-five-arm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add loopx-project/loopx --skill tb4-five-arm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install loopx-project/loopx tb4-five-arm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/loopx-project/loopx.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/benchmark/swe-marathon/skills/tb4-five-arm .cursor/skills/tb4-five-arm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tb4-five-arm" agent skill from https://github.com/loopx-project/loopx/tree/main/benchmark/swe-marathon/skills/tb4-five-arm into .cursor/skills/tb4-five-arm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tb4-five-arm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/loopx-project/loopx.git --path benchmark/swe-marathon/skills/tb4-five-arm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add loopx-project/loopx --skill tb4-five-arm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install loopx-project/loopx tb4-five-arm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/loopx-project/loopx.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/benchmark/swe-marathon/skills/tb4-five-arm .gemini/skills/tb4-five-arm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tb4-five-arm" agent skill from https://github.com/loopx-project/loopx/tree/main/benchmark/swe-marathon/skills/tb4-five-arm into .gemini/skills/tb4-five-arm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tb4-five-arm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install loopx-project/loopx tb4-five-armInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add loopx-project/loopx --skill tb4-five-arm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/loopx-project/loopx.git skills-src && mkdir -p .github/skills && cp -r skills-src/benchmark/swe-marathon/skills/tb4-five-arm .github/skills/tb4-five-arm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tb4-five-arm" agent skill from https://github.com/loopx-project/loopx/tree/main/benchmark/swe-marathon/skills/tb4-five-arm into .github/skills/tb4-five-arm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tb4-five-arm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add loopx-project/loopx --skill tb4-five-arm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install loopx-project/loopx tb4-five-arm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/loopx-project/loopx.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/benchmark/swe-marathon/skills/tb4-five-arm .opencode/skills/tb4-five-arm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tb4-five-arm" agent skill from https://github.com/loopx-project/loopx/tree/main/benchmark/swe-marathon/skills/tb4-five-arm into .opencode/skills/tb4-five-arm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tb4-five-arm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tb4-five-armRunbook for the five-arm Codex harness comparison on Terminal-Bench 4.0, pointing out four differences from SWE-Marathon that silently produce wrong scores.
This Chinese-language runbook adapts an existing SWE-Marathon five-arm comparison, which pits bare Codex and Codex's native /goal against three LoopX modes, to Terminal-Bench 4.0. It assumes you have read the SWE-Marathon version first and covers only what differs. The comparison is of the harness, not the model: all arms share the same model, effort, tool surface, sandbox and container.
Switching is done by exporting WEN_BENCH=tb4 before sourcing env.sh, because a prefix assignment is reverted after sourcing and leaves the run on the wrong profile and network policy without any error. Benchmark-specific values live in scripts/bench/tb4.sh. The task data is the v4.0.0 tag of terminal-bench, with 66 tasks, used directly because harbor's registry only lists version 2.0; the files are digest-pinned and must not be edited, so environment variables go through harbor's --ve and --ae options.
Four structural differences from SWE-Marathon are tabulated: a uniform agent timeout of 28800 seconds, a separate verifier environment mode, no declared network mode so harbor defaults to public, and two images per task. Scoring is binary, with no continuous score. Three GPU-heavy tasks are excluded, leaving 63, of which 11 are multi-container and one needs a Playwright MCP server. The network policy is pinned to public, since falling back to no-network would quietly run every task offline and lower scores.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 8205c8b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockerbashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Terminal-Bench 4.0 Five-Arm Runbook loads about 1.6k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 450 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from loopx-project/loopx at commit 8205c8b, republished under its Apache-2.0 licence (© loopx-project). 450 words, ~1,576 tokens.
.claude/skills/tb4-five-arm/SKILL.md (or your agent's skills folder).五臂定义、LoopX profile 装配、agent 类、监控工具全部与 [[swe-marathon-five-arm]] 相同,先读那一份。这里只写 TB4 特有的部分。
对照维度仍是 harness 不是模型:五臂的模型、effort、工具面、沙箱、容器一致。
export WEN_BENCH=tb4 # 必须 export,见下面「最毒的一个坑」
source env.sh
./scripts/prebuild_images.sh music-harmony # 应建 2 个镜像,不是 1 个
./scripts/verify_envs.sh music-harmony # install-only,不烧 token
./scripts/canary_timeout.sh # 五臂 × 短死线
./scripts/marathon_all.sh --dry # 63 任务 × 5 臂 = 315 trialbenchmark 相关的量全在 scripts/bench/tb4.sh,驱动脚本里没有任何 TB4 常量。
WEN_BENCH 不设时默认 swe-marathon,行为与引入 bench 层之前逐字一致
(已用 bash -x 逐参数 diff 验证过)。
terminal-bench/ 是 harbor-framework/terminal-bench 的 tag v4.0.0
(commit 452bf305c6),66 个任务,见 terminal-bench/VERSION。
harbor run -d terminal-bench@4.0:harbor 的注册表(Supabase
dataset 表)里只有 terminal-bench@2.0,实测查 4.0 返回 null。tasks/dataset.toml 里每个任务有 sha256 digest,task.toml 与 tests/
一个字节都不能改。要注入环境变量走 harbor 的 --ve / --ae。| 项 | SWE-Marathon | TB4.0 |
|---|---|---|
agent.timeout_sec | 3600–36000 各异 | 全部 28800(统一 8h) |
verifier.environment_mode | 无(shared) | 全部 separate |
network_mode | 三态,逐任务 | 一个都没声明 → harbor 默认 public |
| 每任务镜像数 | 1(environment/) | 2(environment/ + tests/) |
| 连续分 | metrics.json 有 | 没有,只有二值 reward |
排除 3 个要 H100 的任务(fp8-rmsnorm-gemm / jax-speedrun-gpu /
math-eval-grader),本机是 4090D。剩 63 个,其中 11 个是多容器
(environment/docker-compose.yaml),1 个(medical-claims-processing)
还要 MCP server(playwright,sse http://playwright-mcp:3080/sse)。
WEN_BENCH 用前缀赋值WEN_BENCH=tb4 source env.sh # ✗ 错
export WEN_BENCH=tb4; source env.sh # ✓ 对bash 在 source 返回后会把前缀赋值的变量还原成未设置,但 WEN_TASKS_DIR
是 export 的、留了下来。于是后续任何脚本自己 source env.sh 时:WEN_BENCH 空 →
回退 swe-marathon profile,却继承着 TB4 的任务目录 —— 拿 TB4 的任务、套
marathon 的网络策略(from-task 找不到 network_mode 就断网)、
开着判官注入、写进 marathon-full/。退出码 0,全程无警告。
env.sh 里有个一致性闸门专门拦这个(WEN_TASKS_DIR_BENCH 与当前 bench 不符
就丢弃继承值),但闸门只保证状态自洽,不保证是你想要的那个 bench。
开跑前看一眼各脚本打印的 bench: 那一行。
TB4 的 66 个 task.toml 一个都没有声明 network_mode / allow_internet,
而 harbor 0.20.0 的 NetworkPolicy.network_mode 默认是 PUBLIC ——
上游的标定条件是联网。
而 marathon_run.sh 原本的做法是 grep task.toml 的 network_mode、
找不到就回退 no-network。直接套用会把 66 个任务全跑成断网:
退出码 0、有轨迹、有分数,只是分数偏低,看着像"模型不行"。
bench/tb4.sh 里 BENCH_NET_POLICY=public 显式钉死,不走那条 grep。
另注:--allow-agent-host 在 public 下是空操作。 实测 harbor 会打印
UserWarning: Run-specific allowlist host(s) ['<model-gateway>', '<container-gateway>'] are
ignored because the effective network policy is public.模型端点的可达性靠代理环境变量和宿主机路由,不靠这个参数,别以为加了就生效。
再注:本机实测容器不穿代理也能出网(docker run 里直连 pypi.org 得 200)。
代理注入是沿用 marathon 对 public 任务的既有做法、属冗余保险;
no_proxy 已包含模型网关,不会劫持 codex 的调用。
environment_mode = "separate":两个镜像、两个环境全部 66 个任务都是 separate。harbor 0.20.0 的实际时序
(harbor/trial/single_step.py:38-55、trial/trial.py:610-680):
跑 agent → 上传 agent 日志 → 收 artifacts → **停掉 agent 环境**
→ 起一个从 tests/ 构建的独立 verifier 环境 → 验证 → 停三个后果:
每个任务两个镜像(66 × 2 = 132)。prebuild_images.sh 原来只扫
environment/,漏掉 tests/。漏了的话 verifier 镜像会在运行期首次构建,
而运行期没有代理(那是刻意的,代理进运行时容器会破坏隔离),dockerd 自己钉的
<dead-dockerd-proxy> 又是死的 → apt-get 超时 → verifier 起不来记 errored。
症状极像"任务没做出来":agent 阶段完全正常、有轨迹、有 token 消耗。
现在由 BENCH_IMAGE_DIRS=(environment tests) 覆盖,且跳过会计数。
只有声明在 artifacts 里的文件能跨到 verifier(66 个全声明了)。
agent 把活干在别处、没写到声明路径 → verifier 看到空目录 → reward 0,
而 agent 轨迹完全正常。这是 shared 模式下不存在的失败模式。
串行,所以每 trial 的容器/网段峰值不翻倍,但多一次构建、 多一个 compose project。
SWE-Marathon 靠任务自写的 metrics.json(partial_score / pass_rate /
pytest / gates_*)补救;TB4 不写这个文件,只有 harbor 的二值 reward。
bench/tb4.sh 设 BENCH_HAS_PARTIAL=0,于是:
_partial.py 整体跳过并打印原因,不静默兜底_compare.py 整列不渲染 partial,而不是渲染成一列 0.0
(一列 0.0 会被读成"全都没得分",比不显示更坏)_compare.py 的跨跑次对照自动改用 reward 做差,表头跟着变marathon 那轮实测:撞死线的 23 条 trial 无一得分,自己收尾的 14 条中 11 条得分
——决定分数的是"能不能在预算内做完"。TB4 统一 8h 预算而我们的死线远短于此,
这个问题只会更严重。唯一还能分辨的信号是终止原因(自己收工 vs 撞死线),
_receipts.py 有这个数据,报表必须并排给出。
marathon 用的是 build 6× / setup 3× / verifier 4×,那是为未做资源标定的
Dockerfile 定的。TB 4.0 的卖点恰恰是"重新标定了 time/CPU/memory"
(build_timeout_sec 中位数 900、verifier.timeout_sec 中位数 600),
照抄 6× 等于把上游的标定压掉。bench/tb4.sh 用 4 / 3 / 2。
canary_timeout.sh 的 CANARY_TIMEOUT_MULT=0.05 是按 8h 声明预算定的
(8h × 0.05 = 24 分钟 > GOAL_TIMEOUT_SEC 5 分钟,保证我们先到)。
TB4 恰好也统一 8h,所以同一个值仍成立。换到声明预算短的 benchmark 上,
harbor 侧可能反而先到,冒烟就白做了 —— 而那是静默的:
五臂照样出 result.json,只是走的是另一条超时路径。
63 任务 × 5 臂 = 315 trial,外加 126 个镜像构建。
MARATHON_AGENT_TIMEOUT_MULT 沿用 marathon 的 0.3 得 8640s(2.4h)/trial。
先拿冒烟的真实单 trial 墙钟再算总时长,别直接开全量。
网段池是硬约束:默认约 32 个 bridge 网络,机器共用。起跑前确认池子有余量
(marathon_all.sh 有 MARATHON_NET_CAP 闸门),清理容器时连 compose 网络
一起清(docker rm -f 不删网络)。
冒烟只覆盖单容器路径。ctr-optimization cumulative-layout-shift
freight-dispatch-shift heat-pump-warranty intrastat-meldung
kv-live-surgery legacy-utility-triage live-database-cutover
medical-claims-processing nextjs-performance payments-pipeline-fix
容器数与内存压力显著更高;medical-claims-processing 还要
MCP server 起得来。全量前单独验这 11 个。
tb4-full/<task>/<arm>/<stamp>/<arm>/result.json job 级,含 stats
/<trial>/verifier/ reward.txt
/agent/ trajectory.json / goal_receipt.json
tb4-full/.claims/<task>__<arm>.claim
tb4-jobs/ marathon_run.sh 单独跑的落脚点
verify-envs-tb4/ canary-timeout-tb4/成功判据与 marathon 相同(scripts/_is_done.py,驱动与监控共用):
n_completed_trials >= 1 且 n_errored_trials == 0© loopx-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in benchmark/swe-marathon/skills/tb4-five-arm of loopx-project/loopx.
Open the folder on GitHubat commit 8205c8b
Terminal-Bench 4.0 Five-Arm Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Terminal-Bench 4.0 Five-Arm Runbook this skillloopx-project/loopx | 6.2k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Copilot Session Failure Analysisdotnet/maui | 23k | — | ~3.4k | Automated safety check: Pass | MIT | |
| Autocontext for Hermesgreyhaven-ai/autocontext | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Agentic Harness Design and ReviewNateBJones-Projects/OB1 | 4.7k | — | ~1.8k | Automated safety check: Pass | Custom licence | |
| Octocode Graph Eval Loopbgauryy/octocode | 946 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Write Skilldruxt/druxt.js | 114 | — | ~926 | Automated safety check: Pass | MIT |
dotnet/maui
Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.
greyhaven-ai/autocontext
Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.
NateBJones-Projects/OB1
Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans.
bgauryy/octocode
Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.
druxt/druxt.js
Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it.
affaan-m/ECC
Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.
loopx-project/loopx
Tracks a group of pull or merge requests across repositories as durable LoopX state: inventory, reconcile changes, keep priorities and a roadmap, and monitor over time.
loopx-project/loopx
Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level.
loopx-project/loopx
Role playbook for a LoopX worker running an auto-research lane, with execution checklists, artifact contracts and stop conditions.
loopx-project/loopx
Operates or analyzes a LoopX-managed benchmark experiment: launching runs, maintaining the experiment board, qualifying integrity, and writing case insights.
loopx-project/loopx
Registers durable project materials such as design docs, SOPs and research notes in a LoopX project's own registry so future agents can find them without raw URLs or private content.
loopx-project/loopx
Runs an evidence-backed pull request review through the loopx CLI and posts bilingual reviews: a full Chinese review plus one concise English verdict.
Categories
Runbook for the five-arm Codex harness comparison on Terminal-Bench 4.0, pointing out four differences from SWE-Marathon that silently produce wrong scores. 0. It assumes you have read the SWE-Marathon version first and covers only what differs.
Terminal-Bench 4.0 Five-Arm Runbook fits situations like: running the Codex harness comparison on Terminal-Bench 4.0; debugging scores that look low although the run exited cleanly; setting up the task data from the v4.0.0 tag without relying on the harbor registry; choosing the network policy for Terminal-Bench tasks.
Run `npx skills add loopx-project/loopx --skill tb4-five-arm -a claude-code`. Or copy the skill folder (benchmark/swe-marathon/skills/tb4-five-arm in loopx-project/loopx) into .claude/skills/tb4-five-arm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add loopx-project/loopx --skill tb4-five-arm -a codex`. Or copy the skill folder (benchmark/swe-marathon/skills/tb4-five-arm in loopx-project/loopx) into .agents/skills/tb4-five-arm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add loopx-project/loopx --skill tb4-five-arm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tb4-five-arm, .gemini/skills/tb4-five-arm, .github/skills/tb4-five-arm and .opencode/skills/tb4-five-arm in your project.
Going by SKILL.md and its folder, Terminal-Bench 4.0 Five-Arm Runbook needs the command-line tools its instructions call (docker and bash). Our summary lists: The SWE-Marathon five-arm setup with its env.sh driver scripts; The harbor benchmark runner; A local copy of the terminal-bench v4.0.0 tag.
SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Terminal-Bench 4.0 Five-Arm Runbook is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Terminal-Bench 4.0 Five-Arm Runbook: Copilot Session Failure Analysis (dotnet/maui, 23k stars), Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars), Agentic Harness Design and Review (NateBJones-Projects/OB1, 4.7k stars) and Octocode Graph Eval Loop (bgauryy/octocode, 946 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
loopx-project (a GitHub organization) maintains it in loopx-project/loopx, which has 6,167 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 7, 2026.
Source: loopx-project/loopx on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.