Agent skill

Sandcastle

by will-ness-ai in will-ness-ai/skills

Run a Sandcastle lane — a sandboxed implement→review loop over a fixed batch of work in its own worktree, ended by an agent-authored PR.

MITAuto-check: notesDevelopment

Install Sandcastle

skills CLI
$ npx skills add will-ness-ai/skills --skill sandcastle -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install will-ness-ai/skills sandcastle --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/will-ness-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sandcastle .claude/skills/sandcastle && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sandcastle
GitHub stars
168
Token cost
~1.4k tokens
SKILL.md length
778 words
Files
4 (incl. references)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Run a Sandcastle lane — a sandboxed implement→review loop over a fixed batch of work in its own worktree, ended by an agent-authored PR.

  • Works in 5 steps: Define the lane → Wire the lane → Launch → …
  • Tasks that involve Git worktrees
  • SKILL.md covers 1. Define the lane, 2. Wire the lane, 3. Launch and 4. Run to dry, plus 3 more sections
  • Calls npx, git and gh; needs GH_TOKEN and CLAUDE_CODE_OAUTH_TOKEN

What it does

Sandcastle is an agent skill from will-ness-ai/skills. Run a Sandcastle lane — a sandboxed implement→review loop over a fixed batch of work in its own worktree, ended by an agent-authored PR.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `agents/openai.yaml`, `references/run-hygiene.md` and `references/ticket-sources.md`).

It sits in Development, covering Git worktrees. It works with Docker. The repository describes itself as: Skills for building high quality software. The licence is MIT.

When your agent uses it

  • Tasks that involve Git worktrees

Example prompts

  • “/sandcastle”

Requirements

  • Node.js
  • Docker
  • A credential in CLAUDE_CODE_OAUTH_TOKEN

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Define the lane
  2. Wire the lane
  3. Launch
  4. Run to dry
  5. The PR

What it can do on your machine

Read from SKILL.md and the folder at commit 71d8909. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • git
    • gh
    • pnpm
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx, git, gh, pnpm and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GH_TOKEN
    • CLAUDE_CODE_OAUTH_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sandcastle loads about 1.4k tokens when it runs, and up to ~2.6k if it reads all its reference files. Until then it costs about 37 tokens; SKILL.md has 778 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:29
    - `.sandcastle/.env` is per-working-dir (untracked, never travels with a branch) — copy the agent token in, refresh `GH_

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from will-ness-ai/skills at commit 71d8909, republished under its MIT licence (© will-ness-ai). 778 words, ~1,410 tokens.

Download SKILL.mdSave it as .claude/skills/sandcastle/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
sandcastle
description
Run a Sandcastle lane — a sandboxed implement→review loop over a fixed batch of work in its own worktree, ended by an agent-authored PR.
disable-model-invocation
true

Hand a batch of tickets to Sandcastle's sandboxed agents and let them work AFK. Each run is a lane: its own worktree, its own branch, its own run config naming exactly its tickets — so concurrent lanes never interact and nothing ever queries a shared pool. The loop implements then reviews one ticket per iteration, folds each onto the lane branch, runs until dry (an iteration lands no commits), and ends with a PR agent that gates the branch and opens the lane's PR.

Sandcastle (@ai-hero/sandcastle) orchestrates coding-agent CLIs inside Docker sandboxes, driven from a repo's tracked .sandcastle/ template. Prereqs: Docker running, plus tokens for the agent (CLAUDE_CODE_OAUTH_TOKEN) and GitHub (GH_TOKEN).

1. Define the lane

The batch is whatever work list you were handed — usually the ticket numbers a /to-tickets run just published (they are already in the conversation), but any tracker query result or literal list works. Read references/ticket-sources.md for how your source expresses tickets in the run config, its commit trailer, and its mark-done action. Name the lane: branch sandcastle/<batch-slug>, worktree ../<repo>-<batch-slug>.

Done when you can enumerate the batch exactly — ids in hand, nothing left implied by a label or query that another session could grow.

2. Wire the lane

  • git worktree add -b <branch> ../<repo>-<batch-slug> <base>

  • Repo has no .sandcastle/ template yet: scaffold with npx sandcastle init (check --help; pick sequential-reviewer), then adapt: main.mts reads the run config and passes the batch in via promptArgs, ends with the PR phase and a process.exit(0); mount ~/.claude/skills read-only so in-sandbox agents can run /implement, /tdd, /code-review; pin the docker imageName (the default derives from the directory name and breaks in a worktree); persist the package-manager store across runs — and anchor every runtime dir you add (pnpm-store/, patches/, plus run.json) in the root gitignore, since re-init regenerates the nested one.

  • In the worktree, write the untracked .sandcastle/run.json:

    json
    { "branch": "sandcastle/<batch-slug>", "base": "main",
      "tickets": [292, 293], "notes": "sandbox has no browser/emulator — verify via unit/jsdom tests" }

    cap is optional and defaults to tickets + 1 — pure dry-stop headroom, so a lane always ends dry, never at a cap. notes reaches the implement prompt verbatim.

  • .sandcastle/.env is per-working-dir (untracked, never travels with a branch) — copy the agent token in, refresh GH_TOKEN=$(gh auth token). Seed the package-manager store warm from another worktree's store dir (cp -Rc), then install.

Done when the run config echoes exactly the step-1 batch and the resolved ticket list prints their full bodies.

3. Launch

Show the lane in the same message that launches it — tickets, branch, base, cap, model — then start the orchestrator (pnpm sandcastle, or npx tsx .sandcastle/main.mts) in the background. Config is read once at launch: a wrong cap or list means relaunch, never an edit under a live run.

Done when the loop is running and iteration 1 shows in .sandcastle/logs/.

Show full SKILL.md (332 more words)Show less

4. Run to dry

One ticket per iteration: implement → review → fold onto the lane branch. The implement prompt has the agent skip tickets blocked by open work and pick the next workable one, so one blocker never stalls the lane. Judge a run by its logs and docker ps, never by the process — when a run looks stalled, hung after its final iteration, or needs killing, read references/run-hygiene.md.

Done when the loop went dry and every ticket in the run config is accounted for: landed, or named with why not.

5. The PR

After dry, the PR phase runs as its own sandbox agent: it re-runs the repo gate — judging any red against <base>, so a failure that reproduces on base is filed as its own issue and named in the PR rather than blocking — authors the body from the run config's ticket list (one line per landed ticket, every unlanded one named), pushes the lane branch, and opens the PR against <base>. Merging stays a human act.

Done when the PR URL exists and its body accounts for every ticket in the run config. Report the URL, then do any host-side mark-done your source needs (ticket-sources.md).

Lanes and each other

Lanes are independent by construction. Two residual rules: lanes over overlapping code merge serially — branch the second after the first merges, so cross-lane conflicts happen at merge time as ordinary git, not mid-loop where a conflict kills an orchestrator; and one lane per worktree — a live orchestrator owns its worktree (run-hygiene.md has the liveness check).

Mechanics you'll rely on

  • Prompt variables: {{TICKETS}} and {{NOTES}} arrive via the run's promptArgs; {{SOURCE_BRANCH}} = the branch the agent works on; {{TARGET_BRANCH}} = the host's branch when the iteration started.
  • Termination: the in-sandbox agent ends its turn with <promise>COMPLETE</promise>; the orchestrator stops once an iteration lands no commits.
  • Config knobs at the top of main.mts: the claudeCode("<model>") model, and the docker({ imageName, mounts }) provider with its onSandboxReady install hook. Everything per-batch lives in run.json, not in main.mts.

© will-ness-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/sandcastle of will-ness-ai/skills.

  • SKILL.md
  • agents/openai.yaml
  • references/run-hygiene.md
  • references/ticket-sources.md

Open the folder on GitHubat commit 71d8909

Compare with similar skills

Sandcastle next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sandcastle compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sandcastle this skillwill-ness-ai/skills168—~1.4kAutomated safety check: NotesMIT
Run Dozzle Dev Instanceamir20/dozzle15k—~747Automated safety check: PassMIT
Build Px4 macOSPX4/PX4-Autopilot13k—~1.1kAutomated safety check: PassBSD-3-Clause
Subwave Worktree Devperminder-klair/subwave1.4k—~2.4kAutomated safety check: NotesMIT
Burla Parallel Dev ClustersBurla-Cloud/burla263—~1.6kAutomated safety check: PassCustom licence
Tnr Dev Serverstudie-tech/TheNinjaRPG104—~2.5kAutomated safety check: NotesNone

Similar skills

  • Starts a Dozzle dev server on a port derived from the current worktree so you can test by hand in a browser, without disturbing instances started elsewhere.

    15k GitHub stars~747 tokensUpdated today
    DevelopmentAuto-check passed
  • Build Px4 macOS

    PX4/PX4-Autopilot

    Build PX4 board firmware on macOS in the px4-dev Docker container, including git worktrees, and stage commit-labeled artifacts without flashing hardware.

    13k GitHub stars~1.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Subwave Worktree Dev

    perminder-klair/subwave

    Stage a SUB/WAVE git worktree so the dev stack can run from it, then start it.

    1.4k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check: notes
  • Sets up an isolated Burla dev cluster per git worktree so several agents can work in parallel, and explains when to use local-dev or remote-dev.

    263 GitHub stars~1.6k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Tnr Dev Server

    studie-tech/TheNinjaRPG

    Run a TheNinjaRPG dev server locally from any git worktree, provision disposable test users, and call tRPC endpoints as those users.

    104 GitHub stars~2.5k tokensUpdated today
    DevelopmentAuto-check: notes
  • Verify

    theopenco/llmgateway

    Build, launch, and drive the LLM Gateway stack in an isolated worktree environment to verify API, gateway, dashboard, playground, or screenshot changes.

    1.7k GitHub stars~1k tokensUpdated today
    DevelopmentAuto-check passed

More from will-ness-ai/skills

  • Cmux

    will-ness-ai/skills

    Drive the local cmux app — workspaces, panes, surfaces, terminal input, agent sessions, and browser surfaces.

    168 GitHub stars~2.3k tokensUpdated 2 days ago
    Auto-check passed
  • Code Story

    will-ness-ai/skills

    Build a wizard-style HTML page that teaches how and why a change works.

    168 GitHub stars~828 tokensUpdated 2 days ago
    Auto-check passed
  • Flashlight

    will-ness-ai/skills

    Shine a light into a wayfinder map's fog — work one direction now, out of frontier order, or redraw the map itself.

    168 GitHub stars~1.4k tokensUpdated 2 days ago
    Auto-check passed
  • Test A Skill

    will-ness-ai/skills

    Field-test a skill by running it cold in parallel agent sessions, then turn what they hit into edits.

    168 GitHub stars~1k tokensUpdated 2 days ago
    Auto-check passed
  • Grill Design

    will-ness-ai/skills

    Converge on a frontend look through rounds of prototypes and grilling verdicts.

    168 GitHub stars~254 tokensUpdated 2 days ago
    Auto-check passed
  • Find Standards

    will-ness-ai/skills

    Find how this problem is already solved: standards we could adopt, and the industry's best practices.

    168 GitHub stars~585 tokensUpdated 2 days ago
    Auto-check passed

Works with

Categories

Questions about Sandcastle

What does Sandcastle do?

Run a Sandcastle lane — a sandboxed implement→review loop over a fixed batch of work in its own worktree, ended by an agent-authored PR. Sandcastle is an agent skill from will-ness-ai/skills. Run a Sandcastle lane — a sandboxed implement→review loop over a fixed batch of work in its own worktree, ended by an agent-authored PR.

When should I use Sandcastle?

Sandcastle fits situations like: tasks that involve Git worktrees.

How do I install Sandcastle in Claude Code?

Run `npx skills add will-ness-ai/skills --skill sandcastle -a claude-code`. Or copy the skill folder (skills/sandcastle in will-ness-ai/skills) into .claude/skills/sandcastle in your project. Claude Code loads it when a task matches its description.

How do I install Sandcastle in Codex?

Run `npx skills add will-ness-ai/skills --skill sandcastle -a codex`. Or copy the skill folder (skills/sandcastle in will-ness-ai/skills) into .agents/skills/sandcastle in your project. Codex loads it when a task matches its description.

Can I use Sandcastle in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add will-ness-ai/skills --skill sandcastle -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sandcastle, .gemini/skills/sandcastle, .github/skills/sandcastle and .opencode/skills/sandcastle in your project.

What does Sandcastle need to run?

Going by SKILL.md and its folder, Sandcastle needs the command-line tools its instructions call (npx, git, gh, pnpm and docker) and credentials named GH_TOKEN and CLAUDE_CODE_OAUTH_TOKEN. Our summary lists: Node.js; Docker; A credential in CLAUDE_CODE_OAUTH_TOKEN.

Does Sandcastle access the network?

SKILL.md contains no URLs. Its commands use npx, git, gh and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Sandcastle safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Sandcastle use?

Sandcastle is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Sandcastle use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Sandcastle?

Skills that share tags, products or a category with Sandcastle: Run Dozzle Dev Instance (amir20/dozzle, 15k stars), Build Px4 macOS (PX4/PX4-Autopilot, 13k stars), Subwave Worktree Dev (perminder-klair/subwave, 1.4k stars) and Burla Parallel Dev Clusters (Burla-Cloud/burla, 263 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sandcastle?

will-ness-ai (a GitHub user) maintains it in will-ness-ai/skills, which has 168 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.

Source: will-ness-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.