Agent skill

Code Create Staged Plan

by open-thoughts in open-thoughts/OpenThoughts-Agent

DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Code Create Staged Plan

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill code-create-staged-plan -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent code-create-staged-plan --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/code-create-staged-plan .claude/skills/code-create-staged-plan && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
code-create-staged-plan
GitHub stars
301
Token cost
~1.5k tokens
SKILL.md length
560 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

  • Works in 8 steps: Header: date · status (scoped —… → Goal: the precise, testable end state… → Mechanism: root cause or design rationale. → …
  • The user says scope/plan this change
  • SKILL.md covers When to use, Where it lives, Parent plan doc — required… and Per-stage scope doc…, plus 2 more sections
  • Calls git and rsync

What it does

Code Create Staged Plan is an agent skill from open-thoughts/OpenThoughts-Agent. DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity requirements, a refactor, a kernel/perf change. Produces a parent plan doc + per-stage scope docs under notes/<codebase/ (each stage = scope + GO/NO-GO validation gate + cost), with global invariants (flag-off byte-identical, parity gates), a borrow-map of code anchors (which drift), and safety considerations…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving, Feature launches and release readiness and Refactoring. It works with vLLM. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • The user says scope/plan this change
  • Design before coding
  • A change is too big/risky for one shot

Example prompts

  • “scope/plan this change”
  • “stage it out”
  • “design before coding”
  • “/code-create-staged-plan”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Header: date · status (scoped — propose-only; no code yet at creation) · target repo + canonical local path + isolated working-copy path…
  2. Goal: the precise, testable end state (e.g. "dcp=N rollout bit-identical to dcp=1: greedy token-ids identical + logprobs allclose atol…
  3. Mechanism: root cause or design rationale.
  4. Stage map: Stage | title | what | feature(s) | layer | cost (CPU / 1-GPU / N-GPU) | gate. Each stage must be independently testable and…
  5. Global invariants (assert in EVERY stage): the flag-off / default-off byte-identical contract (a new feature is a no-op until its flag…
  6. Borrow map (don't reinvent): the exact files/functions/line-anchors you'll touch or copy from — and a standing note that anchors DRIFT…
  7. Safety / reward-hacking (where relevant): policy-invariance for RL reward shaping, ground-truth anchors, "down-weight not zero"…
  8. Validation discipline: per-stage, what proves the gate (flag-off byte-identical first → behavior-on test → GPU smoke on the correct…

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • rsync

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and rsync, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Code Create Staged Plan loads about 1.5k tokens when it runs. Until then it costs about 194 tokens; SKILL.md has 560 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~194
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 560 words, ~1,463 tokens.

Download SKILL.mdSave it as .claude/skills/code-create-staged-plan/SKILL.md (or your agent's skills folder).
name
code-create-staged-plan
description
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity requirements, a refactor, a kernel/perf change. Produces a parent plan doc + per-stage scope docs under notes/<codebase>/ (each stage = scope + GO/NO-GO validation gate + cost), with global invariants (flag-off byte-identical, parity gates), a borrow-map of code anchors (which drift), and safety considerations. Evidence/scoping breadcrumbs go in dated agent_logs/. Use when the user says "scope/plan this change", "stage it out", "design before coding", or a change is too big/risky for one shot. Pairs with code-execute-staged-plan (which runs the plan).

code-create-staged-plan

Turn a substantial change into a dependency-ordered staged plan with a gate at each step. Plans live in notes/<codebase>/; scoping evidence lives in dated agent_logs/.

Read an existing notes/vllm/ plan before writing a new one.

When to use

  • A feature port (upstream → our fork), a multi-step fix (esp. with a parity/regression requirement), a refactor, a kernel/perf change, or any change too big or too risky to land in one commit.
  • NOT for a one-line fix or a mechanical edit — just do those (and log if non-obvious).

Where it lives

  • Plan and scope docs: /Users/benjaminfeuer/Documents/notes/<codebase>/ — one parent (README.md or <change>_plan.md) and one stage<N>_<slug>_scope.md per stage.
  • Evidence: dated /Users/benjaminfeuer/Documents/agent_logs/YYYY-MM-DD_<topic>.md; cite it from the plan.

Parent plan doc — required sections

  1. Header: date · status (scoped — propose-only; no code yet at creation) · target repo + canonical local path + isolated working-copy path (the rsync-clone you'll branch in — see "Isolated working copy" below) + branch (the feature branch you'll cut in the clone, e.g. feuer/<slug>) · links to the evidence agent_logs/.
  2. Goal: the precise, testable end state (e.g. "dcp=N rollout bit-identical to dcp=1: greedy token-ids identical + logprobs allclose atol 1e-2").
  3. Mechanism: root cause or design rationale.
  4. Stage map: Stage | title | what | feature(s) | layer | cost (CPU / 1-GPU / N-GPU) | gate. Each stage must be independently testable and build on a gated predecessor; mark the critical path.
  5. Global invariants (assert in EVERY stage): the flag-off / default-off byte-identical contract (a new feature is a no-op until its flag flips — mirror the EP/CP scaffold no-op tests); the parity gate (the load-bearing equivalence, e.g. G2 bit-identical); regression bounds (don't break MLA / the other arms); minimal diff (no gratuitous API/config churn).
  6. Borrow map (don't reinvent): the exact files/functions/line-anchors you'll touch or copy from — and a standing note that anchors DRIFT (reconfirm at impl time; they're from a dated read).
  7. Safety / reward-hacking (where relevant): policy-invariance for RL reward shaping, ground-truth anchors, "down-weight not zero", parse-real-signals-only.
  8. Validation discipline: per-stage, what proves the gate (flag-off byte-identical first → behavior-on test → GPU smoke on the correct SIF/env). Name the measurement (paired McNemar + pass@k, torch.equal, allclose tol at the bf16 floor — don't loosen a tol silently).
Show full SKILL.md (208 more words)Show less

Per-stage scope doc (stage<N>_<slug>_scope.md) — required sections

  • Header: date · status (scoped GO / blocked / …) · companion = the parent · "no fix yet" if scope-only.
  • Why this is the next step.
  • Change-set: exactly what files change (or "test-only; no <repo>/ source touched this stage").
  • Validation gate (GO/NO-GO): the concrete pass condition + cost. This is what code-execute-staged-plan checks before advancing.
  • Composes with / depends on: the upstream stages it assumes are already green.

Isolated working copy — rsync-clone BEFORE you branch (do NOT branch the canonical clone)

Use this rsync-clone path only for penfever/working self-merge repos: ~/Documents/{OpenThoughts-Agent,vllm}. Marin forks (harbor, MarinSkyRL, evalchemy) use git-worktree→PR→main. Never cut a feature branch on the canonical clone; rsync it to an isolated working directory first:

bash
# 1. rsync-clone the canonical repo (INCLUDING .git; skip heavy build/venv dirs) to an isolated copy
SLUG=<change-slug>; REPO=OpenThoughts-Agent          # or vllm  (marin-forks harbor/MarinSkyRL/evalchemy use worktree→PR→main instead)
SRC=/Users/benjaminfeuer/Documents/$REPO
DST=/Users/benjaminfeuer/Documents/staged-work/$SLUG/$REPO
mkdir -p "$(dirname "$DST")"
rsync -a --exclude='.venv' --exclude='__pycache__' --exclude='*.egg-info' --exclude='wandb/' "$SRC/" "$DST/"
# 2. confirm the canonical clone is clean + on penfever/working, then branch IN THE CLONE
cd "$DST" && git fetch origin && git checkout penfever/working && git pull --ff-only && git checkout -b feuer/$SLUG
  • Keep the canonical clone on penfever/working.
  • Record staged-work/<slug>/<repo> in the plan header; execute, commit, and push from it, then git pull other clones.
  • The rsync clone is not the editable-installed copy; test clusters after push+pull. Use mcp__ide__getDiagnostics on clone files.
  • Use one clone per touched repo; remove staged-work/<slug>/ after merge.

Discipline

  • Local clone is ground truth: branch in an rsync clone; execution commits, pushes, and syncs rather than patching clusters.
  • Default-off: flag-off is byte-identical.
  • Cheapest repro first: stage 0 is usually a unit/CPU harness.
  • At creation, plans are propose-only; hand off to code-execute-staged-plan for execution.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/code-create-staged-plan of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Code Create Staged Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Code Create Staged Plan compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Code Create Staged Plan this skillopen-thoughts/OpenThoughts-Agent301—~1.5kAutomated safety check: PassApache-2.0
Ascend Release Manager for vLLMvllm-project/vllm-ascend2.9k—~7.2kAutomated safety check: PassApache-2.0
Dynamo Kv Replay Parityai-dynamo/dynamo8.2k—~4.9kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
CI Fails Buildkiteguqiong96/Lvllm4642 repos~349Automated safety check: PassApache-2.0

Similar skills

  • Ascend Release Manager for vLLM

    vllm-project/vllm-ascend

    Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

    2.9k GitHub stars~7.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Dynamo Kv Replay Parity

    ai-dynamo/dynamo

    Runs deterministic byte-parity and paired performance campaigns for Dynamo offline KV-aware replay across native-G1 vLLM and SGLang configurations, including forced scheduler-pressure and…

    8.2k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    464 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 9 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 9 days ago
    Auto-check passed
  • Commit

    open-thoughts/OpenThoughts-Agent

    Lint, run the pre-PR checks, commit, push, and author or update the branch's pull request in the required plain-text format.

    301 GitHub stars~2.2k tokensUpdated 9 days ago
    Auto-check: notes

Works with

Questions about Code Create Staged Plan

What does Code Create Staged Plan do?

DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…. Code Create Staged Plan is an agent skill from open-thoughts/OpenThoughts-Agent. DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity requirements, a refactor, a kernel/perf change.

When should I use Code Create Staged Plan?

Code Create Staged Plan fits situations like: the user says scope/plan this change; design before coding; A change is too big/risky for one shot.

How do I install Code Create Staged Plan in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill code-create-staged-plan -a claude-code`. Or copy the skill folder (.agents/skills/code-create-staged-plan in open-thoughts/OpenThoughts-Agent) into .claude/skills/code-create-staged-plan in your project. Claude Code loads it when a task matches its description.

How do I install Code Create Staged Plan in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill code-create-staged-plan -a codex`. Or copy the skill folder (.agents/skills/code-create-staged-plan in open-thoughts/OpenThoughts-Agent) into .agents/skills/code-create-staged-plan in your project. Codex loads it when a task matches its description.

Can I use Code Create Staged Plan in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill code-create-staged-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/code-create-staged-plan, .gemini/skills/code-create-staged-plan, .github/skills/code-create-staged-plan and .opencode/skills/code-create-staged-plan in your project.

What does Code Create Staged Plan need to run?

Going by SKILL.md and its folder, Code Create Staged Plan needs the command-line tools its instructions call (git and rsync).

Does Code Create Staged Plan access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Code Create Staged Plan safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Code Create Staged Plan use?

Code Create Staged Plan is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Code Create Staged Plan use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Code Create Staged Plan?

Skills that share tags, products or a category with Code Create Staged Plan: Ascend Release Manager for vLLM (vllm-project/vllm-ascend, 2.9k stars), Dynamo Kv Replay Parity (ai-dynamo/dynamo, 8.2k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars) and Hugging Face Local Model Evals (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Code Create Staged Plan?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.