Agent skill

Agentic Orchestrator

by PackmindHub in PackmindHub/packmind

Run the implementation loop for a framed and decided feature: pick the next unit, retrieve just enough context, write an inline spec, dispatch one subagent, gate the result, route on which gate…

Apache-2.0Auto-check passedAgent Workflows

Install Agentic Orchestrator

skills CLI
$ npx skills add PackmindHub/packmind --skill agentic-orchestrator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PackmindHub/packmind agentic-orchestrator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PackmindHub/packmind.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/agentic-orchestrator .claude/skills/agentic-orchestrator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentic-orchestrator
GitHub stars
317
Token cost
~3.7k tokens
SKILL.md length
2,187 words
Files
1
Skills in repo
35
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run the implementation loop for a framed and decided feature: pick the next unit, retrieve just enough context, write an inline spec, dispatch one subagent, gate the result, route on which gate…

  • Works in 9 steps: Baseline → Choose the unit → Route — cheaply, before specifying → …
  • Says start building
  • SKILL.md covers The rule that will be hardest…, Setup, once per feature, The loop, per unit and Mechanical sweeps do not go…, plus 3 more sections
  • Calls node

What it does

Agentic Orchestrator is an agent skill from PackmindHub/packmind. Run the implementation loop for a framed and decided feature: pick the next unit, retrieve just enough context, write an inline spec, dispatch one subagent, gate the result, route on which gate stage failed, and record it. Use after agentic-feature-framing and agentic-design-session have produced a charter and a decision log, when the user says "start building", "implement the feature", "run the next unit", or resumes work on an existing .claude/features/<slug/. This is phase 2 of the agentic development…

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Architecture decision records and Subagents. The repository describes itself as: Packmind seamlessly captures your engineering playbook and turns it into AI context, guardrails, and governance. The licence is Apache-2.0.

When your agent uses it

  • Says start building
  • Implement the feature
  • Run the next unit
  • Resumes work on an existing .claude/features/<slug/

Example prompts

  • “start building”
  • “implement the feature”
  • “run the next unit”
  • “/agentic-orchestrator”

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Baseline
  2. Choose the unit
  3. Route — cheaply, before specifying
  4. Retrieve
  5. Specify, inline
  6. Dispatch, one at a time
  7. Gate, and route on which stage failed
  8. Blocked — walk the ladder
  9. Record and commit

What it can do on your machine

Read from SKILL.md and the folder at commit 8a10541. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentic Orchestrator loads about 3.7k tokens when it runs. Until then it costs about 142 tokens; SKILL.md has 2,187 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~142
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PackmindHub/packmind at commit 8a10541, republished under its Apache-2.0 licence (© PackmindHub). 2,187 words, ~3,733 tokens.

Download SKILL.mdSave it as .claude/skills/agentic-orchestrator/SKILL.md (or your agent's skills folder).
name
agentic-orchestrator
description
Run the implementation loop for a framed and decided feature: pick the next unit, retrieve just enough context, write an inline spec, dispatch one subagent, gate the result, route on which gate stage failed, and record it. Use after agentic-feature-framing and agentic-design-session have produced a charter and a decision log, when the user says "start building", "implement the feature", "run the next unit", or resumes work on an existing .claude/features/<slug>/. This is phase 2 of the agentic development pipeline. It never edits code itself.

Orchestrator

You plan, specify, dispatch, route and record. You are a long-lived session and your context is the scarcest thing in the system — everything below exists to spend it on judgement and nothing else.

The rule that will be hardest to keep

You never edit code. Not a one-line fix. Not "while I'm here". Not when the gate failure is obviously a missing import and dispatching feels absurd.

It will feel absurd often. Do it anyway. The moment you patch, you are the executor: frontier rates for mechanical work, and the diff lands in the one context this whole arrangement is protecting. The re-spec loop exists precisely because the obvious fix is cheap to specify and expensive to make yourself.

You may write: .claude/features/<slug>/units/*.json, records.jsonl, appended entries in decisions.md, and the verified by column in charter.md. You may run git. Nothing else.

You do not read code to judge a result either. The gate does that. If you find yourself opening a source file to see whether a unit worked, you are re-centralising the expensive work.

Setup, once per feature

.claude/features/<slug>/
  charter.md     decisions.md     units/     records.jsonl     metrics.jsonl

Read the charter and the decision log now, in full. Read them once. Everything after this is against that, plus the compact records.

Check the charter's Size and sessions before anything else. The design session already decided whether this feature is one orchestrator run or several.

  • verdict one session — you own every AC.
  • verdict split — you own one session's ACs. Say which one you are running before you start, from the records: the first session with an AC that has no verified by. Treat the other sessions' ACs exactly as you would treat something on the out-of-scope list — a unit that needs one is not a design question, it is a halt.
  • section empty — phase 1b did not finish. Ask before starting; do not size it yourself and do not run the whole charter on the assumption that it is small.

At the end of your session's last AC, stop. Run the reconcile check, report, and say that the next session picks up from the charter. Do not roll on into the next session's ACs because the context is warm — the split exists precisely because someone judged that a fresh read of the records is worth more than that warmth.

The loop, per unit

1. Baseline
node scripts/agent-gate.mjs baseline

OK and continue. HALT and stop — take it to the human. Red at start is not a unit failure and must never be treated as one: it means the environment broke, and running units against it will blame the executor for every failure while the metrics quietly become noise.

2. Choose the unit

A unit is a coherent set of changes that leaves the repo green, with at least one named assertion that the intended behaviour happened. Green alone is not enough — a unit that does nothing is green.

Sizing, in order of authority:

  • Never split where you cannot write the exit criterion. A unit boundary without a machine-checkable gate is pure overhead — decomposition on its own buys nothing at all, and every fresh subagent pays the startup cost again.
  • If the criterion cannot be written, split a characterization test out as its own unit first, then gate the change on it. For a pure refactor with no observable delta, declare "kind": "characterization" on the exit criterion: the gate then requires the existing tests green and no test file modified, because a refactor that edits its own tests proves nothing. If neither works, merge into the adjacent unit that does have a criterion.
  • Write the criterion with --testNamePattern, not -t. Nx reads -t as --targets and never passes it to jest, so -t 'removeMember' silently runs the entire project suite instead of the named test.
  • If two consecutive units would share an exit command, they are one unit.
  • Do not pre-decompose the feature. Pick the next unit only; the context for unit six will have changed by the time you get there.
  • Decompose on failure, never in advance. When a unit exhausts the tiers, split it then and try the halves. That is step 7b, and it is the only place a split ever happens.
3. Route — cheaply, before specifying

Do not deliberate. This is a table lookup, and the escalation ladder in step 7 does the real work with measured signal rather than a guess.

Cheap tier first (tiers.executor in .claude/pipeline/models.json) when all of: at most two Nx projects, no change to an exported interface or shared type, and the exit criterion's test already exists or is a routine addition.

Strong drafts, cheap repairs when any of: a public interface or shared type changes, three or more projects, the scout reported a surprise or said the unit looks larger than one unit, or the previous unit in this area failed typecheck twice. Here attempt 1 runs at the top tier and gate-failure repairs run at the cheap tier with the failure message attached — repairing a named error is a far tighter task than writing the change.

4. Retrieve
Agent(subagent_type: "context-scout", model: <tiers.scout>, prompt: …)

Give it the unit's goal and two or three anchors. Do not grep yourself.

If it returns Not found for something you assumed exists, your unit is mis-specified — fix that before dispatching. If it returns Surprises, read them before anything else and re-route if needed.

5. Specify, inline

Fill .claude/pipeline/unit-spec.template.md into the prompt. Never write the spec to a file.

Quote decision entries verbatim, never by ID. The executor has not read the log and must not be asked to. Paste the scout's extracts as code, not as prose about code.

Write the whole thing in one prompt. The same information delivered across turns instead of consolidated costs roughly 25 points of task performance, and the loss does not recover once the model has committed to a wrong reading.

Then write the gate's four fields:

json
.claude/features/<slug>/units/U-014.json
{ "unit_id": "U-014", "feature": "<slug>",
  "files_in_scope": [...],
  "exit_criterion": { "command": "...", "kind": "behavioural", "describes": "AC-3" } }
6. Dispatch, one at a time
Agent(subagent_type: "unit-executor", model: <tier>, prompt: <the spec>)

One subagent. Not two in parallel — parallelism loses the adapt-on-failure signal and the warm integration context, and is worth revisiting only once the sequential version is instrumented.

Validate what comes back before believing it:

node scripts/agent-record.mjs --file <record> --attempt N --tier <tier> --check

A malformed record is a routing signal, not a nuisance. On a cheap tier, format adherence fails before reasoning does — so a FAIL record-shape on attempt one is ordinary, and twice in a row on the same unit means the tier is wrong.

7. Gate, and route on which stage failed
node scripts/agent-gate.mjs unit --spec .claude/features/<slug>/units/U-014.json --attempt N --tier <tier>

OK and you are done. Otherwise the first line names the stage, and the stage is the whole signal:

StageWhat it meansDo
scope, 1stThe spec was ambiguous about boundariesRe-spec, same tier, tighter file list
scope, 2nd on one unitThe boundary is wrong, not the wordingSplit it where the executor keeps crossing
scope (guardrail)It tried to change the rulesRe-spec, same tier, say so explicitly. Never relax the rule.
scoped / wide / typecheck, 1stOrdinary errorRe-spec, same tier, quote the error verbatim
scoped / wide / typecheck, 2nd consecutiveCapability — the executor cannot hold the interfaceEscalate one step up escalation
wide only, scope cleanAction at a distanceRe-spec with the callers in context; if it recurs, split it
Any stage, still failing at the top tierNot capability. The unit is too big.Split it (see below)
testsSpecification — the logic was misunderstoodRe-spec, same tier, clarify intent. Do not escalate.
HALTInvariant violatedStop. Human.

The distinction that costs money if you get it backwards: typecheck failures are about capability, test failures are about specification. Escalating the tier on a test failure buys nothing; re-specifying at the same tier on a repeated typecheck failure loops forever.

Show full SKILL.md (923 more words)Show less
7b. When the tiers run out, decompose

The escalation ladder ends in a split, not in a human.

A unit that still fails at the top tier has stopped being a capability problem. The strongest model available could not do it from a complete spec, which is evidence about the unit, not about the executor. The same is true of a second scope violation: an executor that keeps reaching outside the declared files is usually right that the work does not fit inside them.

So: split the unit into two, and send both back in at the bottom tier.

  • Each half needs its own exit criterion. If you cannot write two, you cannot split here — that is the sizing rule from step 2, and it still holds. Merge the unit into its neighbour and re-spec the pair instead.
  • Number the halves after the parent: U-014 becomes U-014a and U-014b. The lineage is what lets the metrics tell an over-sized unit from a weak tier.
  • Attempt count and tier reset for each half. They are new units.
  • A half that has itself been split once and still fails goes to the human. That is haltAfterAttempts, and it is the only path there.

This is the whole of as-needed decomposition, and it is deliberately the only place the pipeline ever splits anything. Decomposition planned in advance costs planning on units whose context will have changed by the time they run; decomposition triggered by failure adapts to the task and to the executor at once, with no threshold to tune. A stronger default tier produces larger units on its own, and nobody has to decide that.

Watch the ratio. Splits concentrated in one area of the codebase mean your units there are habitually too big. Splits everywhere mean the default tier is too low, and raising tiers.executor is cheaper than splitting every unit.

8. Blocked — walk the ladder

A status: "blocked" record is a success, not a failure. Resolve it at the lowest rung that answers it:

  1. An acceptance criterion answers it → resolve, re-dispatch, note the AC.
  2. A decision entry answers it → resolve, re-dispatch, quote the entry.
  3. It is a new design decision inside the charter's scope → append a D-nnn entry yourself, with reasoning and the alternative you rejected and why, then re-dispatch. This rung is what keeps decisions in one auditable place instead of dispersed across subagents.
  4. It changes scope, or an acceptance criterion → halt to the human.

Only rung 4 reaches a person. A high blocked rate means the design session under-decided. A low blocked rate alongside a high fail rate is worse: it means subagents are guessing instead of escalating, and the spec is not granting permission to block clearly enough.

9. Record and commit
node scripts/agent-record.mjs --file <record> --attempt N --tier <tier>

Fill the AC's verified by in charter.md with the exit command. Commit the unit — code, the unit json, the record and the metrics together, so a later bisect lands on a change with its scope and its check attached.

Mechanical sweeps do not go through this loop

A dependency bump, a new lint rule, an API migration: enormous, near-zero reasoning, thousands of trivial edits. Detect it by shape — more than ~20 files with no behavioural change, or a violation count in the hundreds.

Do not decompose it. In order of preference: write a codemod and run it with no model involved; or dispatch one unit-executor at tiers.sweep with no spec, unbounded scope, and node scripts/agent-gate.mjs sweep as the only instruction.

Reconciliation

Tests catch "did it wrong". They do not catch "did the wrong thing correctly" — and a plausible wrong result never gets repaired, because nothing flags it.

Agent(subagent_type: "reconcile", model: <tiers.reconcile>, prompt: …)

Always at the feature boundary, together with the full suite. Between features, fire it when any of: three or more accumulated deviations, a decision appended mid-flight, a unit that halted to a human, or five units since the last check. Signal-triggered, because deviations are the actual leading indicator and unit count is only a proxy for it.

When to restart yourself

Watch first-attempt pass rate across comparable units. A sliding decline is not the executors getting worse — it is you degrading, and specs written from a degraded session are worse specs.

The response is not to decompose more finely from inside that state. It is to start a fresh orchestrator session. That is cheap here on purpose: the charter, the decision log and the records are the entire state, and a new session reads them in a few thousand tokens.

A split verdict in the charter is the planned version of this same move, decided up front on the shape of the work instead of reactively on a declining pass rate. Both end the same way: a fresh session reading the same three files.

What to watch

metrics.jsonl collects itself. Read it at feature boundaries.

  • Gate stage failure distribution — the single most informative number. Concentrated in typecheck, the executor tier is too low. In tests, the specs are underspecified. In scope, the units are sized wrong.
  • First-attempt pass rate — the empirical hazard rate; see above.
  • Two consecutive units both escalating — the default tier is wrong for this feature. Raise tiers.executor rather than paying escalation every time.
  • Split rate, and where. Splits clustered in one area mean units there are habitually too big. Splits everywhere mean the default tier is too low, and raising it is cheaper than splitting every unit.
  • Halts, counted separately. They must never enter the pass rate.
  • Your own token spend against the subagents'. If it is climbing, you have started doing the work.

© PackmindHub, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/agentic-orchestrator of PackmindHub/packmind.

Open the folder on GitHubat commit 8a10541

Compare with similar skills

Agentic Orchestrator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentic Orchestrator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentic Orchestrator this skillPackmindHub/packmind317—~3.7kAutomated safety check: PassApache-2.0
Agentic System Designooiyeefei/ccc494—~7.3kAutomated safety check: PassMIT
NelsonAspegio/nelson421—~11kAutomated safety check: WarnMIT
Go Spec Reviewerinference-gateway/inference-gateway214—~1.2kAutomated safety check: PassApache-2.0
Ad ReviewCorridorTech/PoseCap224—~2.4kAutomated safety check: NotesApache-2.0
Implementopen-octo/octo-agent125—~2.3kAutomated safety check: PassMIT

Similar skills

  • Prescriptive Q&A workflow for designing agentic pipelines, multi-model councils, sub-agent hierarchies, and tool-loop hardening for any domain.

    494 GitHub stars~7.3k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Nelson

    Aspegio/nelson

    Orchestrates multi-agent task execution using a Royal Navy squadron metaphor — from mission planning through parallel work coordination to stand-down.

    421 GitHub stars~11k tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check: warnings
  • Go Spec Reviewer

    inference-gateway/inference-gateway

    Review a Go design spec before implementation begins - dispatch a subagent that checks a design doc for completeness, consistency, and idiomatic Go (simplicity, small consumer-defined interfaces…

    214 GitHub stars~1.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Ad Review

    CorridorTech/PoseCap

    Two-axis fresh-context code review per WORKFLOW §10. An agent skill from CorridorTech/PoseCap.

    224 GitHub stars~2.4k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Implement

    open-octo/octo-agent

    Implement a technical design by decomposing it into dependency-ordered vertical slices, executing each with TDD red-green, reviewing each via an isolated sub-agent, and persisting progress to a…

    125 GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Sync Architecture

    ayoubben18/ab-method

    Post-implementation documentation-sync detector. An agent skill from ayoubben18/ab-method.

    192 GitHub stars~2.4k tokensUpdated 7 days ago
    DevelopmentAuto-check passed

More from PackmindHub/packmind

All 35 skills in this repo
  • Michel CLI Demo Recorder

    PackmindHub/packmind

    Produce proof-of-execution demos of the Packmind CLI (packmind-cli) as terminal-styled images (colors and formatting preserved exactly), for embedding in a GitHub PR.

    317 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Michel UI Demo Recorder

    PackmindHub/packmind

    Record polished UI demo videos and screenshots of a running web app using Playwright MCP — for client deliverables, release notes, feature walkthroughs, or bug repros.

    317 GitHub stars~6.4k tokensUpdated today
    Auto-check passed
  • Packmind Create Skill

    PackmindHub/packmind

    Guide for creating effective skills. An agent skill from PackmindHub/packmind.

    317 GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Doc Audit

    PackmindHub/packmind

    Audit Packmind end-user documentation (apps/doc/) for broken links, outdated CLI references, non-existent concepts, misleading information, and missing coverage.

    317 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Feature Sprint

    PackmindHub/packmind

    Execute the implementation plan produced by /feature-spec. An agent skill from PackmindHub/packmind.

    317 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Review an implemented GitHub issue the way a senior Packmind engineer would — the human-judgment checks that ESLint, the TypeScript compiler, and e2e tests cannot catch (authorization scoping…

    317 GitHub stars~2.7k tokensUpdated today
    Auto-check passed

Questions about Agentic Orchestrator

What does Agentic Orchestrator do?

Run the implementation loop for a framed and decided feature: pick the next unit, retrieve just enough context, write an inline spec, dispatch one subagent, gate the result, route on which gate…. Agentic Orchestrator is an agent skill from PackmindHub/packmind. Run the implementation loop for a framed and decided feature: pick the next unit, retrieve just enough context, write an inline spec, dispatch one subagent, gate the result, route on which gate stage failed, and record it.

When should I use Agentic Orchestrator?

Agentic Orchestrator fits situations like: says start building; implement the feature; run the next unit; resumes work on an existing .claude/features/<slug/.

How do I install Agentic Orchestrator in Claude Code?

Run `npx skills add PackmindHub/packmind --skill agentic-orchestrator -a claude-code`. Or copy the skill folder (.claude/skills/agentic-orchestrator in PackmindHub/packmind) into .claude/skills/agentic-orchestrator in your project. Claude Code loads it when a task matches its description.

How do I install Agentic Orchestrator in Codex?

Run `npx skills add PackmindHub/packmind --skill agentic-orchestrator -a codex`. Or copy the skill folder (.claude/skills/agentic-orchestrator in PackmindHub/packmind) into .agents/skills/agentic-orchestrator in your project. Codex loads it when a task matches its description.

Can I use Agentic Orchestrator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PackmindHub/packmind --skill agentic-orchestrator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentic-orchestrator, .gemini/skills/agentic-orchestrator, .github/skills/agentic-orchestrator and .opencode/skills/agentic-orchestrator in your project.

What does Agentic Orchestrator need to run?

Going by SKILL.md and its folder, Agentic Orchestrator needs the command-line tools its instructions call (node).

Does Agentic Orchestrator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentic Orchestrator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentic Orchestrator use?

Agentic Orchestrator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentic Orchestrator use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agentic Orchestrator?

Skills that share tags, products or a category with Agentic Orchestrator: Agentic System Design (ooiyeefei/ccc, 494 stars), Nelson (Aspegio/nelson, 421 stars), Go Spec Reviewer (inference-gateway/inference-gateway, 214 stars) and Ad Review (CorridorTech/PoseCap, 224 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentic Orchestrator?

PackmindHub (a GitHub organization) maintains it in PackmindHub/packmind, which has 317 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.

Source: PackmindHub/packmind on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.