Agent skill

Terminal-Bench 4.0 Five-Arm Runbook

by loopx-project in loopx-project/loopx

Runbook for the five-arm Codex harness comparison on Terminal-Bench 4.0, pointing out four differences from SWE-Marathon that silently produce wrong scores.

Apache-2.0Auto-check passedAgent Workflows

SKILL.md written in Chinese; this summary is our English description.

Install Terminal-Bench 4.0 Five-Arm Runbook

skills CLI
$ npx skills add loopx-project/loopx --skill tb4-five-arm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install loopx-project/loopx tb4-five-arm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/loopx-project/loopx.git skills-src && mkdir -p .claude/skills && cp -r skills-src/benchmark/swe-marathon/skills/tb4-five-arm .claude/skills/tb4-five-arm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tb4-five-arm
GitHub stars
6.2k
Token cost
~1.6k tokens
SKILL.md length
450 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runbook for the five-arm Codex harness comparison on Terminal-Bench 4.0, pointing out four differences from SWE-Marathon that silently produce wrong scores.

  • Works in 3 steps: 每个任务两个镜像(66 × 2 =… → 只有声明在 artifacts 里的文件能跨到 verifier(66… → 串行,所以每 trial 的容器/网段峰值不翻倍,但多一次构建、
  • Running the Codex harness comparison on Terminal-Bench 4.0
  • SKILL.md covers 切换方式, 数据来源, TB4 vs SWE-Marathon:四处结构性差异 and 坑(都会静默通过), plus 1 more section
  • Calls docker and bash

What it does

This Chinese-language runbook adapts an existing SWE-Marathon five-arm comparison, which pits bare Codex and Codex's native /goal against three LoopX modes, to Terminal-Bench 4.0. It assumes you have read the SWE-Marathon version first and covers only what differs. The comparison is of the harness, not the model: all arms share the same model, effort, tool surface, sandbox and container.

Switching is done by exporting WEN_BENCH=tb4 before sourcing env.sh, because a prefix assignment is reverted after sourcing and leaves the run on the wrong profile and network policy without any error. Benchmark-specific values live in scripts/bench/tb4.sh. The task data is the v4.0.0 tag of terminal-bench, with 66 tasks, used directly because harbor's registry only lists version 2.0; the files are digest-pinned and must not be edited, so environment variables go through harbor's --ve and --ae options.

Four structural differences from SWE-Marathon are tabulated: a uniform agent timeout of 28800 seconds, a separate verifier environment mode, no declared network mode so harbor defaults to public, and two images per task. Scoring is binary, with no continuous score. Three GPU-heavy tasks are excluded, leaving 63, of which 11 are multi-container and one needs a Playwright MCP server. The network policy is pinned to public, since falling back to no-network would quietly run every task offline and lower scores.

When your agent uses it

  • Running the Codex harness comparison on Terminal-Bench 4.0
  • Debugging scores that look low although the run exited cleanly
  • Setting up the task data from the v4.0.0 tag without relying on the harbor registry
  • Choosing the network policy for Terminal-Bench tasks

Example prompts

  • “Set up the five-arm comparison on Terminal-Bench 4.0 and tell me which settings differ from SWE-Marathon.”
  • “My TB4 run exited with code 0 but the scores are low. Check whether WEN_BENCH was exported correctly.”
  • “Which Terminal-Bench 4.0 tasks should I exclude on a machine without an H100?”
  • “Explain why the network policy has to be pinned to public for this benchmark.”

Requirements

  • The SWE-Marathon five-arm setup with its env.sh driver scripts
  • The harbor benchmark runner
  • A local copy of the terminal-bench v4.0.0 tag

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. 每个任务两个镜像(66 × 2 = 132)。prebuild_images.sh 原来只扫
  2. 只有声明在 artifacts 里的文件能跨到 verifier(66 个全声明了)。
  3. 串行,所以每 trial 的容器/网段峰值不翻倍,但多一次构建、

What it can do on your machine

Read from SKILL.md and the folder at commit 8205c8b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Terminal-Bench 4.0 Five-Arm Runbook loads about 1.6k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 450 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from loopx-project/loopx at commit 8205c8b, republished under its Apache-2.0 licence (© loopx-project). 450 words, ~1,576 tokens.

Download SKILL.mdSave it as .claude/skills/tb4-five-arm/SKILL.md (or your agent's skills folder).
name
tb4-five-arm
description
在 Terminal-Bench 4.0 上做 codex harness 五臂对照(裸 codex / 原生 /goal / LoopX 三模式)。复用 SWE-Marathon 那套驱动,靠 WEN_BENCH 切换。重点是 TB4 相对 SWE-Marathon 的四处结构性差异——每一处踩错都**退出码 0、有轨迹、有分数**,只是分数不对。

Terminal-Bench 4.0 五臂对照

五臂定义、LoopX profile 装配、agent 类、监控工具全部与 [[swe-marathon-five-arm]] 相同,先读那一份。这里只写 TB4 特有的部分。

对照维度仍是 harness 不是模型:五臂的模型、effort、工具面、沙箱、容器一致。

切换方式

bash
export WEN_BENCH=tb4      # 必须 export,见下面「最毒的一个坑」
source env.sh

./scripts/prebuild_images.sh music-harmony   # 应建 2 个镜像,不是 1 个
./scripts/verify_envs.sh music-harmony       # install-only,不烧 token
./scripts/canary_timeout.sh                  # 五臂 × 短死线
./scripts/marathon_all.sh --dry              # 63 任务 × 5 臂 = 315 trial

benchmark 相关的量全在 scripts/bench/tb4.sh,驱动脚本里没有任何 TB4 常量。 WEN_BENCH 不设时默认 swe-marathon,行为与引入 bench 层之前逐字一致 (已用 bash -x 逐参数 diff 验证过)。

数据来源

terminal-bench/ 是 harbor-framework/terminal-bench 的 tag v4.0.0 (commit 452bf305c6),66 个任务,见 terminal-bench/VERSION。

  • 不能用 harbor run -d terminal-bench@4.0:harbor 的注册表(Supabase dataset 表)里只有 terminal-bench@2.0,实测查 4.0 返回 null。
  • 不要跟 main:上游 README 明说 "will be continuously updated", 跟 main 跑出来的结果不可复现。
  • tasks/dataset.toml 里每个任务有 sha256 digest,task.toml 与 tests/ 一个字节都不能改。要注入环境变量走 harbor 的 --ve / --ae。

TB4 vs SWE-Marathon:四处结构性差异

项SWE-MarathonTB4.0
agent.timeout_sec3600–36000 各异全部 28800(统一 8h)
verifier.environment_mode无(shared)全部 separate
network_mode三态,逐任务一个都没声明 → harbor 默认 public
每任务镜像数1(environment/)2(environment/ + tests/)
连续分metrics.json 有没有,只有二值 reward

排除 3 个要 H100 的任务(fp8-rmsnorm-gemm / jax-speedrun-gpu / math-eval-grader),本机是 4090D。剩 63 个,其中 11 个是多容器 (environment/docker-compose.yaml),1 个(medical-claims-processing) 还要 MCP server(playwright,sse http://playwright-mcp:3080/sse)。

坑(都会静默通过)

一、最毒的一个:WEN_BENCH 用前缀赋值
bash
WEN_BENCH=tb4 source env.sh      # ✗ 错
export WEN_BENCH=tb4; source env.sh   # ✓ 对

bash 在 source 返回后会把前缀赋值的变量还原成未设置,但 WEN_TASKS_DIR 是 export 的、留了下来。于是后续任何脚本自己 source env.sh 时:WEN_BENCH 空 → 回退 swe-marathon profile,却继承着 TB4 的任务目录 —— 拿 TB4 的任务、套 marathon 的网络策略(from-task 找不到 network_mode 就断网)、 开着判官注入、写进 marathon-full/。退出码 0,全程无警告。

env.sh 里有个一致性闸门专门拦这个(WEN_TASKS_DIR_BENCH 与当前 bench 不符 就丢弃继承值),但闸门只保证状态自洽,不保证是你想要的那个 bench。 开跑前看一眼各脚本打印的 bench: 那一行。

二、网络策略反了 → 66 个任务全跑成断网

TB4 的 66 个 task.toml 一个都没有声明 network_mode / allow_internet, 而 harbor 0.20.0 的 NetworkPolicy.network_mode 默认是 PUBLIC —— 上游的标定条件是联网。

而 marathon_run.sh 原本的做法是 grep task.toml 的 network_mode、 找不到就回退 no-network。直接套用会把 66 个任务全跑成断网: 退出码 0、有轨迹、有分数,只是分数偏低,看着像"模型不行"。

bench/tb4.sh 里 BENCH_NET_POLICY=public 显式钉死,不走那条 grep。

另注:--allow-agent-host 在 public 下是空操作。 实测 harbor 会打印

UserWarning: Run-specific allowlist host(s) ['<model-gateway>', '<container-gateway>'] are
ignored because the effective network policy is public.

模型端点的可达性靠代理环境变量和宿主机路由,不靠这个参数,别以为加了就生效。

再注:本机实测容器不穿代理也能出网(docker run 里直连 pypi.org 得 200)。 代理注入是沿用 marathon 对 public 任务的既有做法、属冗余保险; no_proxy 已包含模型网关,不会劫持 codex 的调用。

三、environment_mode = "separate":两个镜像、两个环境

全部 66 个任务都是 separate。harbor 0.20.0 的实际时序 (harbor/trial/single_step.py:38-55、trial/trial.py:610-680):

跑 agent → 上传 agent 日志 → 收 artifacts → **停掉 agent 环境**
        → 起一个从 tests/ 构建的独立 verifier 环境 → 验证 → 停

三个后果:

  1. 每个任务两个镜像(66 × 2 = 132)。prebuild_images.sh 原来只扫 environment/,漏掉 tests/。漏了的话 verifier 镜像会在运行期首次构建, 而运行期没有代理(那是刻意的,代理进运行时容器会破坏隔离),dockerd 自己钉的 <dead-dockerd-proxy> 又是死的 → apt-get 超时 → verifier 起不来记 errored。 症状极像"任务没做出来":agent 阶段完全正常、有轨迹、有 token 消耗。 现在由 BENCH_IMAGE_DIRS=(environment tests) 覆盖,且跳过会计数。

  2. 只有声明在 artifacts 里的文件能跨到 verifier(66 个全声明了)。 agent 把活干在别处、没写到声明路径 → verifier 看到空目录 → reward 0, 而 agent 轨迹完全正常。这是 shared 模式下不存在的失败模式。

  3. 串行,所以每 trial 的容器/网段峰值不翻倍,但多一次构建、 多一个 compose project。

Show full SKILL.md (164 more words)Show less
四、二值 reward 在紧预算下没有区分度,而 TB4 没有连续分

SWE-Marathon 靠任务自写的 metrics.json(partial_score / pass_rate / pytest / gates_*)补救;TB4 不写这个文件,只有 harbor 的二值 reward。

bench/tb4.sh 设 BENCH_HAS_PARTIAL=0,于是:

  • _partial.py 整体跳过并打印原因,不静默兜底
  • _compare.py 整列不渲染 partial,而不是渲染成一列 0.0 (一列 0.0 会被读成"全都没得分",比不显示更坏)
  • _compare.py 的跨跑次对照自动改用 reward 做差,表头跟着变

marathon 那轮实测:撞死线的 23 条 trial 无一得分,自己收尾的 14 条中 11 条得分 ——决定分数的是"能不能在预算内做完"。TB4 统一 8h 预算而我们的死线远短于此, 这个问题只会更严重。唯一还能分辨的信号是终止原因(自己收工 vs 撞死线), _receipts.py 有这个数据,报表必须并排给出。

五、超时倍率不要照抄 marathon

marathon 用的是 build 6× / setup 3× / verifier 4×,那是为未做资源标定的 Dockerfile 定的。TB 4.0 的卖点恰恰是"重新标定了 time/CPU/memory" (build_timeout_sec 中位数 900、verifier.timeout_sec 中位数 600), 照抄 6× 等于把上游的标定压掉。bench/tb4.sh 用 4 / 3 / 2。

六、canary 的预算比例要重算

canary_timeout.sh 的 CANARY_TIMEOUT_MULT=0.05 是按 8h 声明预算定的 (8h × 0.05 = 24 分钟 > GOAL_TIMEOUT_SEC 5 分钟,保证我们先到)。 TB4 恰好也统一 8h,所以同一个值仍成立。换到声明预算短的 benchmark 上, harbor 侧可能反而先到,冒烟就白做了 —— 而那是静默的: 五臂照样出 result.json,只是走的是另一条超时路径。

七、全量的量级要先量再定

63 任务 × 5 臂 = 315 trial,外加 126 个镜像构建。 MARATHON_AGENT_TIMEOUT_MULT 沿用 marathon 的 0.3 得 8640s(2.4h)/trial。 先拿冒烟的真实单 trial 墙钟再算总时长,别直接开全量。

网段池是硬约束:默认约 32 个 bridge 网络,机器共用。起跑前确认池子有余量 (marathon_all.sh 有 MARATHON_NET_CAP 闸门),清理容器时连 compose 网络 一起清(docker rm -f 不删网络)。

八、11 个多容器 + 1 个 MCP 任务未验证

冒烟只覆盖单容器路径。ctr-optimization cumulative-layout-shift freight-dispatch-shift heat-pump-warranty intrastat-meldung kv-live-surgery legacy-utility-triage live-database-cutover medical-claims-processing nextjs-performance payments-pipeline-fix 容器数与内存压力显著更高;medical-claims-processing 还要 MCP server 起得来。全量前单独验这 11 个。

产物

tb4-full/<task>/<arm>/<stamp>/<arm>/result.json          job 级,含 stats
                                   /<trial>/verifier/    reward.txt
                                            /agent/      trajectory.json / goal_receipt.json
tb4-full/.claims/<task>__<arm>.claim
tb4-jobs/                                                marathon_run.sh 单独跑的落脚点
verify-envs-tb4/  canary-timeout-tb4/

成功判据与 marathon 相同(scripts/_is_done.py,驱动与监控共用):

n_completed_trials >= 1 且 n_errored_trials == 0

© loopx-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in benchmark/swe-marathon/skills/tb4-five-arm of loopx-project/loopx.

Open the folder on GitHubat commit 8205c8b

Compare with similar skills

Terminal-Bench 4.0 Five-Arm Runbook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Terminal-Bench 4.0 Five-Arm Runbook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Terminal-Bench 4.0 Five-Arm Runbook this skillloopx-project/loopx6.2k—~1.6kAutomated safety check: PassApache-2.0
Copilot Session Failure Analysisdotnet/maui23k—~3.4kAutomated safety check: PassMIT
Autocontext for Hermesgreyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.0
Agentic Harness Design and ReviewNateBJones-Projects/OB14.7k—~1.8kAutomated safety check: PassCustom licence
Octocode Graph Eval Loopbgauryy/octocode946—~1.6kAutomated safety check: PassMIT
Write Skilldruxt/druxt.js114—~926Automated safety check: PassMIT

Similar skills

  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Autocontext for Hermes

    greyhaven-ai/autocontext

    Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

    1.3k GitHub stars~2.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Agentic Harness Design and Review

    NateBJones-Projects/OB1

    Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans.

    4.7k GitHub stars~1.8k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Octocode Graph Eval Loop

    bgauryy/octocode

    Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

    946 GitHub stars~1.6k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed
  • Write Skill

    druxt/druxt.js

    Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it.

    114 GitHub stars~926 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

    274k GitHub stars~1.5k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed

More from loopx-project/loopx

All 12 skills in this repo
  • LoopX PR Program Manager

    loopx-project/loopx

    Tracks a group of pull or merge requests across repositories as durable LoopX state: inventory, reconcile changes, keep priorities and a roadmap, and monitor over time.

    6.2k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • LoopX Self Repair

    loopx-project/loopx

    Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level.

    6.2k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • LoopX Auto-Research Worker

    loopx-project/loopx

    Role playbook for a LoopX worker running an auto-research lane, with execution checklists, artifact contracts and stop conditions.

    6.2k GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • LoopX Benchmark Operator

    loopx-project/loopx

    Operates or analyzes a LoopX-managed benchmark experiment: launching runs, maintaining the experiment board, qualifying integrity, and writing case insights.

    6.2k GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • LoopX Doc Registry

    loopx-project/loopx

    Registers durable project materials such as design docs, SOPs and research notes in a LoopX project's own registry so future agents can find them without raw URLs or private content.

    6.2k GitHub stars~707 tokensUpdated today
    Auto-check passed
  • LoopX PR Review

    loopx-project/loopx

    Runs an evidence-backed pull request review through the loopx CLI and posts bilingual reviews: a full Chinese review plus one concise English verdict.

    6.2k GitHub stars~3.3k tokensUpdated today
    Auto-check passed

Questions about Terminal-Bench 4.0 Five-Arm Runbook

What does Terminal-Bench 4.0 Five-Arm Runbook do?

Runbook for the five-arm Codex harness comparison on Terminal-Bench 4.0, pointing out four differences from SWE-Marathon that silently produce wrong scores. 0. It assumes you have read the SWE-Marathon version first and covers only what differs.

When should I use Terminal-Bench 4.0 Five-Arm Runbook?

Terminal-Bench 4.0 Five-Arm Runbook fits situations like: running the Codex harness comparison on Terminal-Bench 4.0; debugging scores that look low although the run exited cleanly; setting up the task data from the v4.0.0 tag without relying on the harbor registry; choosing the network policy for Terminal-Bench tasks.

How do I install Terminal-Bench 4.0 Five-Arm Runbook in Claude Code?

Run `npx skills add loopx-project/loopx --skill tb4-five-arm -a claude-code`. Or copy the skill folder (benchmark/swe-marathon/skills/tb4-five-arm in loopx-project/loopx) into .claude/skills/tb4-five-arm in your project. Claude Code loads it when a task matches its description.

How do I install Terminal-Bench 4.0 Five-Arm Runbook in Codex?

Run `npx skills add loopx-project/loopx --skill tb4-five-arm -a codex`. Or copy the skill folder (benchmark/swe-marathon/skills/tb4-five-arm in loopx-project/loopx) into .agents/skills/tb4-five-arm in your project. Codex loads it when a task matches its description.

Can I use Terminal-Bench 4.0 Five-Arm Runbook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add loopx-project/loopx --skill tb4-five-arm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tb4-five-arm, .gemini/skills/tb4-five-arm, .github/skills/tb4-five-arm and .opencode/skills/tb4-five-arm in your project.

What does Terminal-Bench 4.0 Five-Arm Runbook need to run?

Going by SKILL.md and its folder, Terminal-Bench 4.0 Five-Arm Runbook needs the command-line tools its instructions call (docker and bash). Our summary lists: The SWE-Marathon five-arm setup with its env.sh driver scripts; The harbor benchmark runner; A local copy of the terminal-bench v4.0.0 tag.

Does Terminal-Bench 4.0 Five-Arm Runbook access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Terminal-Bench 4.0 Five-Arm Runbook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Terminal-Bench 4.0 Five-Arm Runbook use?

Terminal-Bench 4.0 Five-Arm Runbook is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Terminal-Bench 4.0 Five-Arm Runbook use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Terminal-Bench 4.0 Five-Arm Runbook?

Skills that share tags, products or a category with Terminal-Bench 4.0 Five-Arm Runbook: Copilot Session Failure Analysis (dotnet/maui, 23k stars), Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars), Agentic Harness Design and Review (NateBJones-Projects/OB1, 4.7k stars) and Octocode Graph Eval Loop (bgauryy/octocode, 946 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Terminal-Bench 4.0 Five-Arm Runbook?

loopx-project (a GitHub organization) maintains it in loopx-project/loopx, which has 6,167 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 7, 2026.

Source: loopx-project/loopx on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.