Agent skill

cmux Package Test Bisect

by manaflow-ai in manaflow-ai/cmux

Finds which commit broke a failing Swift package suite in the cmux repo by bisecting on CI, then judges per test whether it went stale or the code regressed.

Custom licenceAuto-check passedTesting & QA

Install cmux Package Test Bisect

skills CLI
$ npx skills add manaflow-ai/cmux --skill cmux-test-bisect -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install manaflow-ai/cmux cmux-test-bisect --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/manaflow-ai/cmux.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cmux-test-bisect .claude/skills/cmux-test-bisect && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cmux-test-bisect
GitHub stars
28k
Token cost
~1.5k tokens
SKILL.md length
822 words
Files
2
Skills in repo
22
Repo updated
First seen
Licence
Custom licence

At a glance

Finds which commit broke a failing Swift package suite in the cmux repo by bisecting on CI, then judges per test whether it went stale or the code regressed.

  • Works in 3 steps: Get today's failing set from the red run → Search for an existing fix: gh search… → Work in your own worktree from…
  • A Swift package suite is red on main and nobody knows which PR broke it
  • SKILL.md covers Before you start, Run the bisect, Read the matrix and Stale test or regression, plus 1 more section
  • Calls python3, git and gh

What it does

In the cmux repository, pull request CI runs only selected tests, so a Swift package suite can sit red on main unnoticed. This skill finds the commit behind each failure on CI rather than on a laptop. It covers the SwiftPM suites under `Packages/iOS` and `Packages/Shared` that `test-ios.yml` can run, while failures in the app-host tests have their own automatic bisect. You read the resulting failure matrix and decide per test whether the test went stale or the code regressed.

Before starting, it collects the current failing set from the red run with `gh`, runs the full suite twice so that a test seen only once is treated as a flake candidate, searches open pull requests for an existing fix and works from a worktree on `upstream/main`. The bisect itself uses `scripts/ci/package_bisect.py`: `start` dispatches probes on old commits with today's CI files laid over them, `status --wait` can take up to 45 minutes, and options such as `--filter`, `--paths`, `--patch` and `--bisect` narrow or extend the search.

When your agent uses it

  • A Swift package suite is red on main and nobody knows which PR broke it
  • PR CI ran only filtered tests and the full suite has drifted
  • Someone asks which commit broke a specific test

Example prompts

  • “CmuxMobileShell tests are red on main. Bisect it on CI and tell me which PR broke them.”
  • “Which commit made the CmuxSyncStore pairing test start failing? Rule out a flake first.”
  • “For each failing test in CmuxMobileCore, decide whether the test is stale or the code regressed.”

Requirements

  • The cmux repository with scripts/ci/package_bisect.py
  • GitHub CLI (gh) access to the repository
  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Get today's failing set from the red run
  2. Search for an existing fix: gh search prs --repo manaflow-ai/cmux --state open ''.
  3. Work in your own worktree from upstream/main. If git log stops at a

What it can do on your machine

Read from SKILL.md and the folder at commit bbb3633. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • git
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

cmux Package Test Bisect loads about 1.5k tokens when it runs. Until then it costs about 86 tokens; SKILL.md has 822 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 822 words (~1,485 tokens).

“PR CI runs selected tests, so a package suite can stay red on main for weeks with nobody noticing. This skill finds the commit behind each failure on CI, never on the Mac (no full cmux builds on laptops).”

— opening of SKILL.md by manaflow-ai, Custom licence
name
cmux-test-bisect

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file in skills/cmux-test-bisect of manaflow-ai/cmux.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit bbb3633

Compare with similar skills

cmux Package Test Bisect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

cmux Package Test Bisect compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
cmux Package Test Bisect this skillmanaflow-ai/cmux28k—~1.5kAutomated safety check: PassCustom licence
Diagnose a Red Rundifferent-ai/openwork24k—~779Automated safety check: PassCustom licence
CI Failure Triage and RepairChachamaru127/claude-code-harness3.2k1 repos~1.1kAutomated safety check: NotesMIT
AI Bug Triagepetrkindlmann/qa-skills170—~5.2kAutomated safety check: PassMIT
Regression Root Cause AnalyzerArabelaTso/Skills-4-SE253—~3.4kAutomated safety check: PassApache-2.0
TiDB Test Diff Triagepingcap/tidb41k—~498Automated safety check: PassApache-2.0

Similar skills

  • Diagnose a Red Run

    different-ai/openwork

    Classifies a failing test, typecheck or CI job before any code changes, by recording the failure and running a clean control to show whether it was already broken.

    24k GitHub stars~779 tokensUpdated today
    Testing & QAAuto-check passed
  • CI Failure Triage and Repair

    Chachamaru127/claude-code-harness

    Diagnoses failing CI pipelines and tests, deciding first whether the test or the implementation is at fault, and hands hard cases to a dedicated fixer subagent.

    3.2k GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check: notes
  • AI Bug Triage

    petrkindlmann/qa-skills

    Hybrid fingerprint + LLM pipeline for bug classification, deduplication, and ticket generation.

    170 GitHub stars~5.2k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Regression Root Cause Analyzer

    ArabelaTso/Skills-4-SE

    Locate root causes of failing regression tests by analyzing code changes, error messages, and test dependencies.

    253 GitHub stars~3.4k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Investigates TiDB plan or test-result diffs that the change does not explain, ruling out failpoint setup and merge effects before expected outputs are updated.

    41k GitHub stars~498 tokensUpdated today
    Testing & QAAuto-check passed
  • Diagnoses a failed GreptimeDB fuzz CI job by pulling its GitHub Actions logs and fuzz artifacts, then matching the evidence to the local source code.

    6.7k GitHub stars~4.4k tokensUpdated today
    Testing & QAAuto-check passed

More from manaflow-ai/cmux

All 22 skills in this repo
  • cmux Diagnostics

    manaflow-ai/cmux

    Runs a read-only health check for cmux and explains what it finds, covering the CLI and socket, settings, agent hooks, session restore and notifications.

    28k GitHub starsUsed in 1 repo~819 tokens
    Auto-check passed
  • cmux Settings Editor

    manaflow-ai/cmux

    Reads and changes cmux preferences in `~/.config/cmux/cmux.json` through a helper that validates each change before writing, covering terminal, browser, viewers and shortcuts.

    28k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • cmux Backend Rules

    manaflow-ai/cmux

    Sets the backend TypeScript and Cloud VM rules for cmux: Effect-based services, thin route handlers, Postgres as source of truth, migrations and provider secrets.

    28k GitHub starsUsed in 1 repo~682 tokens
    Auto-check passed
  • Cmux Debugging Guide

    manaflow-ai/cmux

    Covers debug logging, the Debug menu, profiling rules and runtime pitfalls for working on the cmux macOS terminal app.

    28k GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Workflow rules for the Ghostty submodule in cmux: rebuilding GhosttyKit.xcframework, pushing fork changes and updating the parent submodule pointer safely.

    28k GitHub starsUsed in 1 repo~580 tokens
    Auto-check passed
  • Cmux Release Workflow

    manaflow-ai/cmux

    Runs the cmux release process: gathers per-PR changelog lines since the last tag, bumps the version, guards against a bad tag, then tags and pushes the macOS DMG release.

    28k GitHub starsUsed in 1 repo~856 tokens
    Auto-check passed

Works with

Questions about cmux Package Test Bisect

What does cmux Package Test Bisect do?

Finds which commit broke a failing Swift package suite in the cmux repo by bisecting on CI, then judges per test whether it went stale or the code regressed. In the cmux repository, pull request CI runs only selected tests, so a Swift package suite can sit red on main unnoticed. This skill finds the commit behind each failure on CI rather than on a laptop.

When should I use cmux Package Test Bisect?

cmux Package Test Bisect fits situations like: A Swift package suite is red on main and nobody knows which PR broke it; PR CI ran only filtered tests and the full suite has drifted; someone asks which commit broke a specific test.

How do I install cmux Package Test Bisect in Claude Code?

Run `npx skills add manaflow-ai/cmux --skill cmux-test-bisect -a claude-code`. Or copy the skill folder (skills/cmux-test-bisect in manaflow-ai/cmux) into .claude/skills/cmux-test-bisect in your project. Claude Code loads it when a task matches its description.

How do I install cmux Package Test Bisect in Codex?

Run `npx skills add manaflow-ai/cmux --skill cmux-test-bisect -a codex`. Or copy the skill folder (skills/cmux-test-bisect in manaflow-ai/cmux) into .agents/skills/cmux-test-bisect in your project. Codex loads it when a task matches its description.

Can I use cmux Package Test Bisect in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add manaflow-ai/cmux --skill cmux-test-bisect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cmux-test-bisect, .gemini/skills/cmux-test-bisect, .github/skills/cmux-test-bisect and .opencode/skills/cmux-test-bisect in your project.

What does cmux Package Test Bisect need to run?

Going by SKILL.md and its folder, cmux Package Test Bisect needs the command-line tools its instructions call (python3, git and gh). Our summary lists: The cmux repository with scripts/ci/package_bisect.py; GitHub CLI (gh) access to the repository; Python 3.

Does cmux Package Test Bisect access the network?

SKILL.md contains no URLs. Its commands use git and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is cmux Package Test Bisect safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does cmux Package Test Bisect use?

cmux Package Test Bisect has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does cmux Package Test Bisect use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to cmux Package Test Bisect?

Skills that share tags, products or a category with cmux Package Test Bisect: Diagnose a Red Run (different-ai/openwork, 24k stars), CI Failure Triage and Repair (Chachamaru127/claude-code-harness, 3.2k stars), AI Bug Triage (petrkindlmann/qa-skills, 170 stars) and Regression Root Cause Analyzer (ArabelaTso/Skills-4-SE, 253 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains cmux Package Test Bisect?

manaflow-ai (a GitHub organization) maintains it in manaflow-ai/cmux, which has 28,076 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.

Source: manaflow-ai/cmux on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.