Agent skill

Autopilot Batch

by joshukraine in joshukraine/dotfiles

Fan out a batch of autopilot-queued issues to parallel background worktree subagents — each runs /autopilot at the build model from its 'model:' label — with a gating review at Opus 5 or above and…

MITAuto-check passedDevelopment

Install Autopilot Batch

skills CLI
$ npx skills add joshukraine/dotfiles --skill autopilot-batch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install joshukraine/dotfiles autopilot-batch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/joshukraine/dotfiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude/.claude/skills/autopilot-batch .claude/skills/autopilot-batch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autopilot-batch
GitHub stars
429
Token cost
~6.8k tokens
SKILL.md length
3,908 words
Files
1
Skills in repo
24
Repo updated
First seen
Licence
MIT

At a glance

Fan out a batch of autopilot-queued issues to parallel background worktree subagents — each runs /autopilot at the build model from its 'model:' label — with a gating review at Opus 5 or above and…

  • Works in 6 steps: Preconditions → Assemble the batch → Fan out the builds → …
  • Tasks that involve Git worktrees
  • SKILL.md covers Where this runs (read first), Arguments, The build-model policy — the… and Escape hatch (batch-level), plus 3 more sections
  • Calls gh, git and rails

What it does

Autopilot Batch is an agent skill from joshukraine/dotfiles. Fan out a batch of autopilot-queued issues to parallel background worktree subagents — each runs /autopilot at the build model from its 'model:' label — with a gating review at Opus 5 or above and never below the build (Opus reviews Sonnet and Opus builds, Fable reviews Fable builds).

Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Git worktrees and Subagents. The repository describes itself as: :roundpushpin: My dotfiles for macOS using Neovim, Zsh, and Ghostty + Tmux. The licence is MIT.

When your agent uses it

  • Tasks that involve Git worktrees
  • Tasks that involve Subagents

Example prompts

  • “model:”
  • “/autopilot-batch”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Preconditions
  2. Assemble the batch
  3. Fan out the builds
  4. As each build finishes
  5. Gating review (every --to pr build)
  6. Reclaim worktrees

What it can do on your machine

Read from SKILL.md and the folder at commit b59ad5b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • git
    • rails
    • psql

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Autopilot Batch loads about 6.8k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 3,908 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from joshukraine/dotfiles at commit b59ad5b, republished under its MIT licence (© joshukraine). 3,908 words, ~6,802 tokens.

Download SKILL.mdSave it as .claude/skills/autopilot-batch/SKILL.md (or your agent's skills folder).
name
autopilot-batch
description
Fan out a batch of autopilot-queued issues to parallel background worktree subagents — each runs /autopilot at the build model from its 'model:' label — with a gating review at Opus 5 or above and never below the build (Opus reviews Sonnet and Opus builds, Fable reviews Fable builds).
disable-model-invocation
true
argument-hint
[--merge <issue#,…>]

Autopilot Batch

Run the vetted autopilot-queued queue as a parallel batch: one background isolation: worktree subagent per issue, each running /autopilot <n> end-to-end at the build model its model: label calls for, with a gating review at Opus 5 or above and never below the build (Opus for Sonnet- and Opus-built PRs, Fable for Fable-built PRs). This is the run half of the triage → run split; /autopilot-triage is the vet half that fills the queue.

Use this when you have a queue of independent, well-scoped issues and want them all carried to review-ready PRs in one unattended pass. For a single issue, use /autopilot directly.

Where this runs (read first)

Run this from the target application repository — the repo whose issues and app these are (e.g. the Rails app) — on a clean default branch, pulled up to date. Not from dotfiles. isolation: worktree creates each worktree from the orchestrator's current repo, so the cwd decides where the fan-out worktrees land; running from the wrong repo produces worktrees of the wrong tree. If the cwd is the dotfiles repo (or any repo that doesn't own these issues), stop and say so.

Arguments

  • (no args) — run every autopilot-queued issue at tier --to pr. The whole batch stops at review-ready PRs; nothing merges.
  • --merge <issue#,…> — authorize tier --to merge for the listed issues only (e.g. --merge 847,851). Everything else stays --to pr. Merge is still gated per issue by /autopilot's Step 8 narrow-class gate, which degrades any non-qualifying issue back to --to pr. Passing --merge IS your per-issue authorization to merge those issues; it is not a standing capability.

Default tier is --to pr by design — merge is opt-in, per issue, never batch-wide.

The build-model policy — the label decides the build, review runs at Opus or above

Each issue's build subagent runs at a model chosen per issue; the gating review always runs at Opus 5 or above, and never below the build. Ladder, cheapest to most capable: Sonnet 5 → Opus 5 → Fable 5.

Resolving each issue's build model

Read the issue's model: label first. Repos that have adopted the per-issue model convention (~/.claude/docs/model-selection-strategy.md) carry one on every open issue, recording the build tier assigned at triage. Step 1's issue list already returns labels, so no extra call is needed:

  • model: sonnet → build with Sonnet 5; model: opus → Opus 5; model: fable → Fable 5.
  • No model: label → fall back to the class rubric below. Most repos have not adopted the convention and must keep working unchanged; in an adopted repo, unlabeled means Opus by convention — which the rubric already approximates — so the same fallback is right in both worlds.

The fallback rubric, unchanged:

  • Opus 5 — data-model / migrations, auth / security boundaries, inventory / shipment correctness, thin or ambiguous specs, cross-cutting refactors. Also the model whenever you genuinely can't tell the class (safe default — Opus never under-builds).
  • Sonnet 5 — bounded, well-specified, pattern-following work: i18n / copy, views / Tailwind, config, a single test, straightforward CRUD, docs.

State each issue's model at the confirm gate and where it came from — label or rubric (no label) — with a one-line why for every rubric call. An override you make at the gate is worth writing back to the issue's label afterward, so the record stays honest for the next run.

The review tier is derived, never labeled

The gating review (Step 4) is computed from the build tier and is never read from the label. One invariant: review runs at Opus 5 or above, and never below the build.

BuildGating review
Sonnet 5Opus 5
Opus 5Opus 5 (fresh context, adversarial prompt, independent sample)
Fable 5Fable 5 (ceiling)

This is the structural safety net: an under-build is caught and fixed at review, never shipped. That is exactly what makes leaning on the cheaper tier for the bulk close to free on quality while saving real latency and limit headroom at fan-out scale — so a model: sonnet label must never downgrade the review. It sets the build only.

The floor replaced a "one tier above the build" rule on 2026-07-27. That rule was written to insure cheap builds; Opus was never the cheap build, and at fan-out scale it sent nearly every review in the batch to Fable. A Fable review is now reached by escalation rather than by rule — via a model: fable label, via a mid-flight escalation (the table reads the model the build actually ran at), or by simply re-spawning the reviewer at Fable on the PR. See ~/.claude/docs/model-selection-strategy.md § "The review floor, and the trade it accepts" for the reasoning and the trade it knowingly accepts.

Merge issues are always Fable-built

A build subagent runs the whole /autopilot loop at one model (a subagent can't switch mid-run), so whichever model builds a --to merge issue also makes its merge go/no-go inside /autopilot Step 8. The floor is non-negotiable: the merge decision is always Fable — the autonomous merge is the highest-stakes call in the pipeline, so it gets the most capable model. Therefore: any issue opted into --to merge is built by Fable, regardless of what its model: label or the class rubric would otherwise pick — a model: sonnet label on a merge-tier issue does not buy a cheaper merge decision. Sonnet and Opus only ever build --to pr issues, which stop at a review-ready PR and pass the gating review (Step 4) before you see them. This keeps every autonomous merge Fable-decided without the orchestrator having to re-implement the merge gate. The cost — running the rare, narrow merge class at Fable rates — is negligible because --to merge is opt-in and uncommon.

Escape hatch (batch-level)

A single issue stopping must not halt the batch. If a build subagent hits /autopilot's own escape hatch (a complex/ambiguous issue, a guardrail decision, an unresolvable failure), it stops and reports — that issue is left for you, keeps its queue label, and the rest of the batch continues. Surface every stop prominently in the final report. Stop the whole batch only for a systemic problem: the cwd is the wrong repo, auth to the remote is failing for everyone, or the queue is empty.

Your task

Announce the run first: how many issues are queued, the tier split (all --to pr, or which are --to merge), and that you will fan out one background worktree subagent per issue.

Step 0 — Preconditions
  • Confirm the cwd is the target app repo (see "Where this runs"), on a clean default branch, then git pull.
  • Pre-warm one push approval. Run a single git ls-remote origin up front so any SSH-agent signing approval is granted once, before fan-out — the batch then runs unattended. (Per-push resilience during the run is already handled inside /autopilot: transient sign_and_send_pubkey … communication with agent failed errors are retried with backoff.)
Step 1 — Assemble the batch
  • gh issue list --label autopilot-queued --state open --json number,title,labels.
  • If the list is empty: stop — nothing is queued. Recommend /autopilot-triage to vet candidates and fill the queue.
  • Build the run plan: for each issue, its assigned build model, its source (label or rubric (no label), with a one-line why for rubric calls), and its tier (pr, or merge if named in --merge — and remember a --merge issue is forced to a Fable build whatever its label says).
  • Note any queued issue missing a model: label in a repo that has adopted the convention — check with gh label list --json name --jq '[.[].name | select(startswith("model: "))]'; a non-empty result means adopted. Missing labels are a triage gap, not a blocker: the rubric covers the issue, but flag them at the gate so triage can be corrected.

CONFIRM GATE: Present the run plan — issue list, per-issue model + source, per-issue tier — and get a go-ahead before spawning. Membership was already vetted at triage; this is a light confirm that the models and tiers are right and the queue is still current, plus your chance to override a model call. Not a re-litigation of the queue.

Step 2 — Fan out the builds

First, write each issue's isolated-command wrapper. Do this yourself, before spawning anything — one file per issue, at /tmp/autopilot-i<n>, chmod +x:

bash
#!/usr/bin/env bash
# Isolated commands for autopilot issue <n>. Written by /autopilot-batch; not part of the repo.
set -euo pipefail

# The main checkout also has bin/ci, so a wrong-cwd run would silently test the
# wrong tree while still writing to this issue's database. Refuse unless we are
# in a linked worktree. (Outside a git repo both commands fail and compare
# equal, which also refuses — the safe direction.)
if [ "$(git rev-parse --git-dir 2>/dev/null)" = "$(git rev-parse --git-common-dir 2>/dev/null)" ]; then
  echo "autopilot-i<n>: run this from the agent's worktree, not $PWD" >&2
  exit 1
fi

export PARALLEL_WORKERS=1 RAILS_ENV=test DATABASE_URL=postgres:///<app>_test_i<n>

case "${1:-ci}" in
  ci)      exec bin/ci ;;
  test)    shift; exec bin/rails test "$@" ;;
  prepare) exec bin/rails db:test:prepare ;;
  *)       echo "usage: $(basename "$0") [ci|test <args>|prepare]" >&2; exit 2 ;;
esac

This is the whole point of the mechanism: the three variables are typed once, by you, instead of re-threaded onto every command by an agent for the length of a run. The Bash tool does not persist shell state between calls, so an export at the start of an agent's run does nothing and the variables would otherwise have to be re-typed on every test, CI, and db:test:prepare invocation — including the ones initiated from inside /autopilot, which the agent invokes rather than controls. A wrapper cannot be half-remembered.

Three deliberate choices:

  • It covers all three call sites, not just CI. ci for the full pipeline, test for the fast checks /resolve-issue Step 5 and /autopilot Step 4 need mid-loop, prepare for the one-time database setup. A CI-only wrapper would leave the loop with no isolated way to run tests — and an agent that needs one and hasn't got one will reach for bare bin/rails test, which is exactly the unprefixed invocation that leaked twelve worker databases on 2026-07-27. Forbidding a command without supplying its replacement produces the violation rather than preventing it.
  • It lives in /tmp, not in bin/. Build agents commit frequently, and a bin/ci-isolated inside the worktree is an untracked file that git add -A would sweep into the PR.
  • It guards its cwd instead of cd-ing. Not cd-ing is what lets one wrapper serve the build agent and the later review agent in their different worktrees. But "wrong cwd fails loudly" is only true for directories that lack bin/ci — and the main checkout has one, so an invocation from there would have run CI against the wrong tree and signed off on the wrong branch, silently. The worktree assertion closes that.

Then, for each issue, spawn a background subagent:

  • isolation: worktree — its own checkout, so parallel edits across issues can't collide.
  • run_in_background: true — they run concurrently.
  • model: the issue's assigned build model — from its model: label, or the rubric when it has none; Fable for any --to merge issue. Because the subagent is spawned at that model, /autopilot's own announce-time reconciliation normally finds a match and passes straight through. If you spawn an agent below its issue's label — a downgrade you approved at the confirm gate — say so explicitly in the prompt ("building at Opus is a deliberate override of this issue's model: fable label, approved at the batch confirm gate"), or the agent will treat it as an under-build and stop.
  • prompt: if the issue body opens with a model callout (a > 🤖 Recommended model: … blockquote), quote it in the prompt — it names the watch-item or escalation trigger behind the tier choice, and it is wasted if only the orchestrator reads it. Then: run /autopilot <n> --to <tier> to completion, with this instruction stated explicitly — every test, CI, and database command for this entire run goes through /tmp/autopilot-i<n>, run in the foreground from your worktree root: … prepare once before anything else, then … ci for the full pipeline and … test [args] for the fast checks mid-loop. Never bare bin/ci, never bin/rails test. This holds inside /autopilot too — its Step 7 defers to the CI command you were given. (See the foreground-CI note under Important; the wrapper carries the serial and isolated-database settings, so there is nothing to re-type and nothing to remember.) Then report back, as the final message, the PR number and URL, the loop outcome (which steps ran / were n·a), a one-line summary of the change, and, if it stopped at the escape hatch, exactly where and why.

Spawn them together so they run in parallel. Each subagent lands on an auto-named worktree branch (worktree-agent-…); inside it, /autopilot → /resolve-issue creates the proper feat/gh-<n>-… branch, commits, pushes, and opens the PR from it. Nothing to pre-create.

Step 3 — As each build finishes

When a build subagent reports a review-ready PR:

  • Drop the queue label: gh issue edit <n> --remove-label autopilot-queued. The issue is now in-flight, not pending — lifecycle-contract step 3.
  • Remove that issue's build worktree before spawning its review — git worktree remove --force <path>, using git worktree list to find it. The build agent is finished and its work is pushed, so the worktree is dead weight; but leaving it in place is what causes the gh signoff failure described in Step 4. While it exists it still holds the PR branch checked out, so the reviewer's gh pr checkout <PR> cannot take that branch and silently leaves it on a detached HEAD. Removing it first makes the checkout ordinary and the whole failure mode disappear.
  • Kick off its gating review (Step 4) right away — don't wait for the whole batch: Opus for a Sonnet or Opus build, Fable for a Fable build. (A --to merge issue was Fable-built and Fable-reviewed inside its own /autopilot run — no orchestrator review needed.)

A subagent that stopped at the escape hatch keeps its autopilot-queued label (still pending) and is set aside for the final report — do not review or merge it.

Show full SKILL.md (1,760 more words)Show less
Step 4 — Gating review (every --to pr build)

For each review-ready PR, spawn a review subagent per the derivation table above — Opus for a Sonnet-built or Opus-built PR, Fable for a Fable-built PR. Derive this from the model the build actually ran at, not from the issue's model: label: if a build was overridden at the confirm gate or escalated mid-flight, the review must follow the real build tier, not the recorded one. That is also what makes a mid-flight escalation to Fable pull its own review up with it.

  • model: opus (Sonnet or Opus build) or model: fable (Fable build), isolation: worktree.
  • prompt: check out the PR branch (gh pr checkout <PR>), invoke code-review through the Skill tool with an effort level and the PR number (e.g. high <PR>) — high by default, medium for a small view/copy diff, never ultra (billed, cloud-run, user-triggered only). Name the skill explicitly: the built-in review has come back unavailable or shallow inside subagents, so a prompt that leaves the choice open invites a hand-rolled review. Depth also comes from the spawn tier above plus, for a large or security-/data-sensitive diff, fanning out additional adversarial lenses; mutation-test the key claims either way (delete or disable the change and confirm the covering test actually fails — a test that stays green with the feature removed is the miss a read-only review cannot see). Then fix real correctness bugs and clear quality wins, commit and push (this updates the open PR), then re-run CI with /tmp/autopilot-i<n> ci — the same wrapper the build agent used, so the review reuses that issue's database — so the required sign-off attaches to the final commit. Its test mode is available for iterating on a fix without paying for the full pipeline each time. Report the verdict: clean / N fixes applied / a finding that needs a product decision (→ flag it for the human, don't guess).
  • Also state in the prompt: if gh signoff fails with current branch is not tracking a remote branch, you are on a detached HEAD — gh pr checkout could not take the branch. It is the only red step in an otherwise green run and reads like a test failure, which is what makes it expensive. Step 3 removes the build worktree first precisely to prevent this, so if it still happens, say so rather than working around it silently. The recovery is either a throwaway tracking branch (git switch -c signoff-<PR> then push with -u), or — after confirming the tree is clean and HEAD equals origin/<pr-branch> — gh signoff create -f.

This is the review floor: every --to pr build is reviewed at Opus or above — never below itself — before the PR is "review-ready" for you, so a build-model miss is caught and fixed, never shipped. The build loop's own internal review does not count toward it, however clean it reports — that review runs inside the build's context and with whatever review mechanism the agent could reach. Skipping the gating review because the build "already had an Opus review" is exactly how a small view-and-test PR reached the human on 2026-10-06 with a dead locale test and a misplaced note that a medium code-review caught in under a minute. The only exception is a docs-only PR with no runnable surface, and say so in the report when you take it. If a review comes back thin on a diff you are uneasy about, re-spawning the reviewer at Fable on the same PR is the escalation path: one review's worth of headroom, invoked by judgment rather than by rule. --to merge issues don't pass through here — they were Fable-built and Fable-reviewed inside their own /autopilot run, and merged there only if the narrow-class gate passed.

Step 5 — Reclaim worktrees

Step 3 already removed the build worktree of every issue that reached a review-ready PR. What remains: the review worktrees, the build worktrees of --to merge issues (they merged inside their own run and never passed through Step 3), and the build worktree of anything that stopped at the escape hatch — that last one deliberately, since its work is unpushed and removing it would discard the run.

Don't work from that list; work from git worktree list, which is authoritative. git worktree remove --force <path> each leftover …/worktrees/agent-* entry except any belonging to a stopped issue, then git worktree prune. The branches live on the remote (pushed as feat/gh-<n>-…), so removing the local worktrees is safe. Delete any signoff-<PR> helper branches a reviewer had to create (see Step 4), and remove the per-issue wrappers: rm -f /tmp/autopilot-i*.

Drop the per-issue test databases too — they are the batch's largest leftover and nothing else reclaims them:

bash
psql -lqt | cut -d'|' -f1 | tr -d ' ' \
  | grep -E "^<app>_test_i[0-9]+([-_][0-9]+)?$" \
  | while read -r db; do dropdb "$db" || echo "still in use, skipped: $db"; done

dropdb fails on a database that still has an open connection — a stray agent or a psql you left open. That is why the loop tolerates a failure and names the database instead of aborting the sweep; re-run it once the connection is gone.

The _i[0-9]+ segment is what keeps this off the project's own <app>_test and its worker databases, which belong to the human's local runs — the pattern matches only databases this batch created. Run the leak check under Important before this, not after: dropping the evidence first would hide a bypass you needed to know about.

Completion report

Post a batch summary:

text
## Autopilot batch — N issues

| Issue | PR | Build | Review | Outcome |
| ----- | -- | ----- | ------ | ------- |
| #847  | #M | Fable (merge tier) | (internal)      | merged + deployed ✓ |
| #851  | #P | Sonnet (label)     | clean (Opus)    | PR-ready |
| #863  | #Q | Opus (rubric)      | 2 fixes (Opus)  | PR-ready |
| #870  | —  | —                  | —               | STOPPED — <reason>, needs you |

- Queue: <k> resolved to PRs, <j> stopped and still labeled `autopilot-queued`.
- Model labels: <none missing | #N, #M had no `model:` label in an adopted repo — triage gap, ran on the rubric>.
- Merged: <list, or none — all held at PR-ready>.
- Your call: review + `/merge-pr` on the PR-ready ones; pick up the stopped issues.

Important

  • Run from the target app repo, never dotfiles — the worktree isolation depends on it.

  • Default is --to pr for the whole batch. Only issues named in --merge can merge, and only through /autopilot's narrow-class gate; a mislabeled one degrades to --to pr inside its own run.

  • Every autonomous merge is Fable-decided — --merge issues are Fable-built, so the merge go/no-go is Fable. Every --to pr build is reviewed at Opus or above, never below itself (Step 4), before it reaches you.

  • The model: label sets the build tier only. The review tier is derived from the build (Opus floor, never below the build) and the merge floor is Fable, so no label can cheapen either — the safety net that catches a cheap build's mistakes is never itself cheapened. An issue with no label runs on the rubric exactly as before, which is what keeps unadopted repos working unchanged.

  • One stop never halts the batch. A stopped issue is set aside with its label intact; the rest continue. A wrong merge is the only truly bad outcome, and the gates above prevent it.

  • Compose, don't re-implement. Build subagents run the real /autopilot; review subagents run code-review <effort> <PR>. This skill only orchestrates, gates the model floor, and manages the queue label and worktrees — so improvements to those skills flow through untouched. code-review is model-invocable from inside a worktree subagent (confirmed 2026-10-06 on two gating reviews; this replaces the 2026-07-27 rule that it was user-only and reviewers had to take the built-in review <PR>). The built-in review remains the fallback only if code-review is genuinely missing from the agent's skill listing. A subagent that finds a skill uninvocable must say so, never substitute a hand-rolled review that reports as the real one.

  • Agents never type the isolation variables — the Step 2 wrapper carries them. /tmp/autopilot-i<n> is the only test/CI/database command an agent runs (prepare, then ci and test), and it is the single enforcement point for both settings below. The build agent runs … prepare as its first action, inside its own worktree; the orchestrator cannot, because the wrapper's worktree guard correctly refuses to run from the main checkout.

  • Why the database must be isolated (RAILS_ENV=test DATABASE_URL=… per agent). Several worktree agents run CI at once; if they all share the project's one default test DB they deadlock against each other — and against any other session using that same DB (a common case: a second Claude working in the primary checkout). PARALLEL_WORKERS=1 does not solve this — it only fixes fork-starvation within one run. Each agent gets a name distinct from the project's default test DB, e.g. <app>_test_i<issue#> (…_test_i847), reused by that issue's gating review. RAILS_ENV=test is not optional: DATABASE_URL only maps to the test connection under RAILS_ENV=test. Omit it and db:test:prepare silently falls through to the shared default test DB (clobbering the other session), while bin/ci's setup runs db:prepare in development — creating the isolated DB stamped development and seeding it, which then breaks data-counting tests and makes db:seed:replant abort with EnvironmentMismatchError. Batch-scoped: standalone /autopilot runs (one worktree, no contention) don't need it.

  • Why CI must run serially (PARALLEL_WORKERS=1). Several worktree subagents run CI at once, and on macOS the parallel Minitest system-test workers fork-starve under that combined load — producing flaky failures and fork-crash storms that aren't real, then costly retries. A single test worker sidesteps it. Batch-scoped: standalone /autopilot runs can stay parallel.

  • Verify the isolation held before you report the batch. After the run, check that no per-worker databases leaked:

    bash
    psql -lqt | cut -d'|' -f1 | tr -d ' ' | grep -E "^<app>_test_i[0-9]+[-_][0-9]+$"

    It should return nothing. A hit means some command ran without PARALLEL_WORKERS=1 — the wrapper was bypassed somewhere. Say so in the report rather than cleaning up quietly; that is the signal the mechanism is leaking. (Observed 2026-07-27: …_test_i367_0 through _i367_11, twelve worker databases from a single missed prefix.) The [-_] is deliberate — Rails has used both <base>-<n> and <base>_<n> for parallel worker databases depending on version, and both are present on this machine today, so a pattern hardcoding one separator silently passes on a project using the other.

  • gh signoff fails on a detached HEAD. It resolves @{push}, so it needs a branch with an upstream. A reviewer that lands detached — because gh pr checkout could not take a branch another worktree still held — gets current branch is not tracking a remote branch as the single red step in an otherwise green run, which reads like a test failure. Step 3's worktree removal is the structural prevention; the recoveries are a throwaway signoff-<PR> tracking branch pushed with -u, or gh signoff create -f once the tree is clean and HEAD matches origin/<pr-branch>.

  • Subagents run bin/ci in the FOREGROUND — never as a background shell task. A subagent that backgrounds bin/ci and ends its turn "to wait for the notification" orphans the process: the shell task dies with the agent's turn, no completion callback ever fires, and the run stalls silently at an open PR with no sign-off (observed 2026-07-17, #957/PR #960 — looked hung for ~1.5h with no CI process alive). Background Bash is safe only in the orchestrator's main session, where task exit re-invokes the conversation. If a build agent does stall this way, recovery is SendMessage to the agent (context intact): tell it the background task is dead and to re-run CI in the foreground.

© joshukraine, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in claude/.claude/skills/autopilot-batch of joshukraine/dotfiles.

Open the folder on GitHubat commit b59ad5b

Compare with similar skills

Autopilot Batch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Autopilot Batch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Autopilot Batch this skilljoshukraine/dotfiles429—~6.8kAutomated safety check: PassMIT
Cursor Composer Task DelegateChachamaru127/claude-code-harness3.2k—~4.4kAutomated safety check: NotesMIT
Work With PR Lifecyclecode-yeongyu/oh-my-openagent70k—~4.6kAutomated safety check: PassCustom licence
Worktree ParallelD-Robotics/moss143—~313Automated safety check: PassMIT
Blitzaiskillstore/marketplace430—~2.8kAutomated safety check: PassNone
Inference Format Optimizera2ui-project/a2ui17k—~985Automated safety check: PassApache-2.0

Similar skills

  • Cursor Composer Task Delegate

    Chachamaru127/claude-code-harness

    Hands one implementation task to Cursor Composer in an isolated git worktree, then reviews its diff and cherry-picks the result into the main branch.

    3.2k GitHub stars~4.4k tokensUpdated 4 days ago
    DevelopmentAuto-check: notes
  • Work With PR Lifecycle

    code-yeongyu/oh-my-openagent

    Takes a task through to a merged pull request in an isolated git worktree, splitting it into small independent PRs and looping on CI and review gates until they pass.

    70k GitHub stars~4.6k tokensUpdated today
    DevelopmentAuto-check passed
  • Worktree Parallel

    D-Robotics/moss

    How to run writable sub-agents in isolated git worktrees and merge their lease patches

    143 GitHub stars~313 tokensUpdated today
    DevelopmentAuto-check passed
  • Blitz

    aiskillstore/marketplace

    This skill should be used when parallelizing multi-issue sprints using git worktrees and parallel Claude agents.

    430 GitHub stars~2.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Iterative benchmarking, evaluation, and algorithmic optimization of alternative A2UI inference formats (such as Express, Atom, and Elemental).

    17k GitHub stars~985 tokensUpdated today
    DevelopmentAuto-check passed
  • Bench Baton

    nooga/let-go

    Coordinate heavy local workloads across worktrees, processes, and subagents — benchmarks and timing-sensitive gates run exclusively on a quiesced machine, while builds, test suites, regeneration…

    568 GitHub stars~821 tokensUpdated today
    DevelopmentAuto-check passed

More from joshukraine/dotfiles

All 24 skills in this repo
  • Todoist CLI

    joshukraine/dotfiles

    Manage Todoist tasks, projects, labels, filters, sections, comments, reminders, and workspaces via the td CLI.

    429 GitHub starsUsed in 1 repo~6.9k tokens
    Auto-check passed
  • Autopilot Triage

    joshukraine/dotfiles

    Vet open issues for autonomous resolution and queue the qualifying ones with the autopilot-queued label — the start-of-day "fill the queue" half of the triage → run split.

    429 GitHub stars~2.1k tokensUpdated 3 days ago
    Auto-check passed
  • Checkpoint

    joshukraine/dotfiles

    Quick 2-minute status update on current phase, completed work, blockers, and health check.

    429 GitHub stars~600 tokensUpdated 3 days ago
    Auto-check passed
  • Create PR

    joshukraine/dotfiles

    Create a pull request with auto-generated description, issue linking, ROADMAP updates, and PR-metadata validation.

    429 GitHub stars~1.3k tokensUpdated 3 days ago
    Auto-check passed
  • Debrief

    joshukraine/dotfiles

    Detailed technical walkthrough covering architecture, test coverage, product tour, and key design decisions.

    429 GitHub stars~2.5k tokensUpdated 3 days ago
    Auto-check passed
  • Drift Check

    joshukraine/dotfiles

    Pre-PR advisory check for deviations from the project spec. An agent skill from joshukraine/dotfiles.

    429 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed

Questions about Autopilot Batch

What does Autopilot Batch do?

Fan out a batch of autopilot-queued issues to parallel background worktree subagents — each runs /autopilot at the build model from its 'model:' label — with a gating review at Opus 5 or above and…. Autopilot Batch is an agent skill from joshukraine/dotfiles. Fan out a batch of autopilot-queued issues to parallel background worktree subagents — each runs /autopilot at the build model from its 'model:' label — with a gating review at Opus 5 or above and never below the build (Opus reviews Sonnet and Opus builds, Fable reviews Fable builds).

When should I use Autopilot Batch?

Autopilot Batch fits situations like: tasks that involve Git worktrees; tasks that involve Subagents.

How do I install Autopilot Batch in Claude Code?

Run `npx skills add joshukraine/dotfiles --skill autopilot-batch -a claude-code`. Or copy the skill folder (claude/.claude/skills/autopilot-batch in joshukraine/dotfiles) into .claude/skills/autopilot-batch in your project. Claude Code loads it when a task matches its description.

How do I install Autopilot Batch in Codex?

Run `npx skills add joshukraine/dotfiles --skill autopilot-batch -a codex`. Or copy the skill folder (claude/.claude/skills/autopilot-batch in joshukraine/dotfiles) into .agents/skills/autopilot-batch in your project. Codex loads it when a task matches its description.

Can I use Autopilot Batch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add joshukraine/dotfiles --skill autopilot-batch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autopilot-batch, .gemini/skills/autopilot-batch, .github/skills/autopilot-batch and .opencode/skills/autopilot-batch in your project.

What does Autopilot Batch need to run?

Going by SKILL.md and its folder, Autopilot Batch needs the command-line tools its instructions call (gh, git, rails and psql).

Does Autopilot Batch access the network?

SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Autopilot Batch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Autopilot Batch use?

Autopilot Batch is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Autopilot Batch use?

About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Autopilot Batch?

Skills that share tags, products or a category with Autopilot Batch: Cursor Composer Task Delegate (Chachamaru127/claude-code-harness, 3.2k stars), Work With PR Lifecycle (code-yeongyu/oh-my-openagent, 70k stars), Worktree Parallel (D-Robotics/moss, 143 stars) and Blitz (aiskillstore/marketplace, 430 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Autopilot Batch?

joshukraine (a GitHub user) maintains it in joshukraine/dotfiles, which has 429 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 6, 2026.

Source: joshukraine/dotfiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.