Agent skill

Spec Compound Refresh

by leo-kuang-ai in leo-kuang-ai/spec-first

Refresh docs/solutions learnings against the current codebase.

MITAuto-check: warningsDevelopment

Install Spec Compound Refresh

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add leo-kuang-ai/spec-first --skill spec-compound-refresh -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install leo-kuang-ai/spec-first spec-compound-refresh --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/leo-kuang-ai/spec-first.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/spec-compound-refresh .claude/skills/spec-compound-refresh && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spec-compound-refresh
GitHub stars
107
Token cost
~15k tokens
SKILL.md length
8,042 words
Files
16 (incl. scripts, references, assets)
Skills in repo
35
Repo updated
First seen
Licence
MIT

At a glance

Refresh docs/solutions learnings against the current codebase.

  • Works in 9 steps: Assess and Route → Investigate Candidate Learnings → 5: Investigate Pattern Docs → …
  • Drifted learnings
  • SKILL.md covers Workflow Contract Summary, Mode Detection, CONCEPTS.md bootstrap requests and Interaction Principles, plus 10 more sections
  • Runs JavaScript and Shell scripts from its folder; calls git

What it does

Spec Compound Refresh is an agent skill from leo-kuang-ai/spec-first. Refresh docs/solutions learnings against the current codebase. Use when auditing stale, overlapping, superseded, or drifted learnings; avoid general refactor, debugging, or code review unless docs/solutions is explicit.

Its SKILL.md is about 15k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files, including scripts, reference files and assets (for example `assets/resolution-template.md`, `evals/cases/non-solutions-scope-gated.yaml` and `evals/eval.yaml`).

It sits in Development. The repository describes itself as: 仓库原生 AI Coding Harness —— 把一次性 AI 对话变成可治理、可验证、可沉淀的工程闭环 · spec-first.cn. The licence is MIT.

When your agent uses it

  • Drifted learnings
  • Avoid general refactor
  • Code review unless docs/solutions is explicit

Example prompts

  • “/spec-compound-refresh”

Requirements

  • Node.js
  • A Bash shell

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Assess and Route
  2. Investigate Candidate Learnings
  3. 5: Investigate Pattern Docs
  4. 75: Document-Set Analysis
  5. Classify the Right Maintenance Action
  6. Ask for Decisions
  7. Execute the Chosen Action
  8. 5: Vocabulary Capture
  9. Commit Changes

What it can do on your machine

Read from SKILL.md and the folder at commit 74655dc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (JavaScript and Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Spec Compound Refresh loads about 15k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 60 tokens; SKILL.md has 8,042 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~15k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~23k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningTells the agent its actions are pre-authorized / not to stop for confirmationSKILL.md:45
    ommended** in the report and continue — do not stop or ask for permissions.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from leo-kuang-ai/spec-first at commit 74655dc, republished under its MIT licence (© leo-kuang-ai). 8,042 words, ~14,512 tokens.

Download SKILL.mdSave it as .claude/skills/spec-compound-refresh/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
spec-compound-refresh
description
Refresh docs/solutions learnings against the current codebase. Use when auditing stale, overlapping, superseded, or drifted learnings; avoid general refactor, debugging, or code review unless docs/solutions is explicit.
argument-hint
[optional: scope hint — directory, filename, module, or keyword] [mode:headless]

Compound Refresh

Maintain the quality of docs/solutions/ over time. This workflow reviews existing learnings against the current codebase, then refreshes any derived pattern docs that depend on them.

Workflow Contract Summary

  • 输入: docs/solutions/、CONCEPTS.md、可选 scope hint,以及当前 source/test/doc evidence。
  • 输出: Keep/Update/Consolidate/Replace/Delete/Stale 分类、已应用的知识维护变更和完整 Applied/Recommended 报告。
  • 硬出口: source truth、目标 repo、写入范围或语义分类无法确认时不得把猜测写成 current knowledge;headless 只能把歧义标 stale。
  • 权威: 当前代码与验证证据优先于旧 learning;本地 mutation、commit 和 landing 分别需要独立授权,mode:headless 只改变交互方式。
  • Refresh is current-source anchored: re-read the defining source refs before updating a learning, retain observed revision/freshness and limitations, and mark the item stale when the source cannot be confirmed. Historical cache, session transcript, or provider output is advisory and never a substitute for the current source.
  • 消费者: 项目维护者,以及读取 docs/solutions//CONCEPTS.md 的规划、实现、调试和审查 workflow。

Mode Detection

Check whether the invocation arguments supplied by the current host contain the exact token mode:headless. If present, strip only that token (preserving the remainder, quoted paths, and token order as the scope hint) and run in headless mode.

ModeWhenBehavior
Interactive (default)User is present and can answer questionsAsk for decisions on ambiguous cases, confirm actions
Headlessmode:headless in argumentsNo user interaction. Apply all unambiguous actions (Keep, Update, Consolidate, auto-Delete, Replace with sufficient evidence). Mark ambiguous cases as stale. Generate a summary report at the end.
Resolve mutation and landing authority

Record three independent run-local facts before any write or Git action:

yaml
mutation_authorization: authorized | missing
commit_authorization: authorized | missing
landing_authorization: authorized | missing

A direct request to refresh or maintain docs/solutions/ authorizes only the bounded local document mutations described by this workflow. Set commit_authorization only when the current user or visible upstream handoff separately requests a commit. Set landing_authorization only when push or PR creation/update is separately explicit. mode:headless, a feature branch, writable permissions, successful edits, or a clean tree do not grant commit, branch creation, push, or PR authority. Without commit authorization, preserve verified edits as uncommitted work; without landing authorization, do not push or open a PR.

Headless mode rules
  • Skip all user questions. Never pause for input.
  • Process all docs in scope. No scope narrowing questions — if no scope hint was provided, process everything.
  • Attempt all safe actions: Keep (no-op), Update (fix references), Consolidate (merge and delete subsumed doc), auto-Delete (unambiguous criteria met), Replace (when evidence is sufficient). If a write succeeds, record it as applied. If a write fails (e.g., permission denied), record the action as recommended in the report and continue — do not stop or ask for permissions.
  • Mark as stale when uncertain. If classification is genuinely ambiguous (Update vs Replace vs Consolidate vs Delete) or Replace evidence is insufficient, mark as stale with status: stale, stale_reason, and stale_date in the frontmatter. If even the stale-marking write fails, include it as a recommendation.
  • Use conservative confidence. In interactive mode, borderline cases get a user question. In headless mode, borderline cases get marked stale. Err toward stale-marking over incorrect action.
  • Always generate a report. The report is the primary deliverable. It has two sections: Applied (actions that were successfully written) and Recommended (actions that could not be written, with full rationale so a human can apply them or run the skill interactively). The report structure is the same regardless of what permissions were granted — the only difference is which section each action lands in.

CONCEPTS.md bootstrap requests

If invoked specifically to create or bootstrap CONCEPTS.md (e.g., "create a CONCEPTS.md", "build the concept map", "set up shared vocabulary"), the intent is ambiguous between two jobs — building the vocabulary file and running a docs/solutions refresh — so disambiguate before proceeding. Use the platform's blocking question tool: AskUserQuestion in Claude Code (call ToolSearch with select:AskUserQuestion first if its schema isn't loaded), request_user_input in Codex. Fall back to numbered options in chat only when no blocking tool exists in the harness or the call errors (e.g., Codex edit modes) — not because a schema load is required. Never silently skip the question. Two options:

  1. Create CONCEPTS.md (build the concept map) — seed the repo-wide concept map; skip only the docs/solutions classification phases (Phases 0–4). Read references/concepts-vocabulary.md and follow its Seed goal and Scope of a seed (repo-wide) rules: seed the project's core domain nouns from the declared domain model (schema, core types, primary models, top-level domain docs), each meeting the qualifying bar, the codebase setting the count. Write the preamble (see Phase 4.5), cluster per the organization rules, and run the Discoverability Check so AGENTS.md/CLAUDE.md surface the new file. Then enter the authority-aware Phase 5 closeout; without separate commit authorization, leave the verified bootstrap edits uncommitted.
  2. Run a refresh cycle — proceed with the normal refresh flow below; CONCEPTS.md is seeded (if absent) and reconciled as part of Phase 4.5.

In headless mode there is no user to ask: default to the refresh cycle (vocabulary is seeded and reconciled within Phase 4.5 regardless) and note in the report that a standalone repo-wide bootstrap was not run.

Interaction Principles

These principles apply to interactive mode only. In headless mode, skip all user questions and apply the headless mode rules above.

Follow the same interaction style as spec-brainstorm:

  • Ask questions one at a time — use the platform's blocking question tool: AskUserQuestion in Claude Code (call ToolSearch with select:AskUserQuestion first if its schema isn't loaded), request_user_input in Codex. Fall back to numbered options in plain text only when no blocking tool exists in the harness or the call errors (e.g., Codex edit modes) — not because a schema load is required. Never silently skip the question
  • Prefer multiple choice when natural options exist
  • Start with scope and intent, then narrow only when needed
  • Do not ask the user to make decisions before you have evidence
  • Lead with a recommendation and explain it briefly

The goal is not to force the user through a checklist. The goal is to help them make a good maintenance decision with the smallest amount of friction.

Refresh Order

Refresh in this order:

  1. Review the relevant individual learning docs first
  2. Note which learnings stayed valid, were updated, were consolidated, were replaced, or were deleted
  3. Then review any pattern docs that depend on those learnings

Why this order:

  • learning docs are the primary evidence
  • pattern docs are derived from one or more learnings
  • stale learnings can make a pattern look more valid than it really is

If the user starts by naming a pattern doc, you may begin there to understand the concern, but inspect the supporting learning docs before changing the pattern.

Maintenance Model

For each candidate artifact, classify it into one of five outcomes:

OutcomeMeaningDefault action
KeepStill accurate and still usefulNo file edit by default; report that it was reviewed and remains trustworthy
UpdateCore solution is still correct, but references driftedApply evidence-backed in-place edits
ConsolidateTwo or more docs overlap heavily but are both correctMerge unique content into the canonical doc, delete the subsumed doc
ReplaceThe old artifact is now misleading, but there is a known better replacementCreate a trustworthy successor, then delete the old artifact
DeleteNo longer useful, applicable, or distinctDelete the file — git history preserves it if anyone needs to recover it later

Core Rules

  1. Evidence informs judgment. The signals below are inputs, not a mechanical scorecard. Use engineering judgment to decide whether the artifact is still trustworthy.
  2. Prefer no-write Keep. Do not update a doc just to leave a review breadcrumb.
  3. Match docs to reality, not the reverse. When current code differs from a learning, update the learning to reflect the current code. The skill's job is doc accuracy, not code review — do not ask the user whether code changes were "intentional" or "a regression." If the code changed, the doc should match. If the user thinks the code is wrong, that is a separate concern outside this workflow.
  4. Be decisive, minimize questions. When evidence is clear (file renamed, class moved, reference broken), apply the update. In interactive mode, only ask the user when the right action is genuinely ambiguous. In headless mode, mark ambiguous cases as stale instead of asking. The goal is automated maintenance with human oversight on judgment calls, not a question for every finding.
  5. Avoid low-value churn. Do not edit a doc just to fix a typo, polish wording, or make cosmetic changes that do not materially improve accuracy or usability.
  6. Use Update only for meaningful, evidence-backed drift. Paths, module names, related links, category metadata, code snippets, and clearly stale wording are fair game when fixing them materially improves accuracy.
  7. Use Replace only when there is a real replacement. That means either:
    • the current conversation contains a recently solved, verified replacement fix, or
    • the user has provided enough concrete replacement context to document the successor honestly, or
    • the codebase investigation found the current approach and can document it as the successor, or
    • newer docs, pattern docs, PRs, or issues provide strong successor evidence.
  8. Delete when the code is gone, and only after checking for inbound links. If the referenced code, controller, or workflow no longer exists in the codebase and no successor can be found, delete the file — don't default to Keep just because the general advice is still "sound." When in doubt between Keep and Delete, ask the user (in interactive mode) or mark as stale (in headless mode). Inbound links inform classification, not cleanup: cleanup is always mechanical, but decorative citations (principle stated inline) allow Delete, while substantive citations (citing doc relies on the cited doc) signal Replace. The auto-delete case is missing code, no matching successor, and citations absent or decorative.
  9. Evaluate document-set design, not just accuracy. In addition to checking whether each doc is accurate, evaluate whether it is still the right unit of knowledge. If two or more docs overlap heavily, determine whether they should remain separate, be cross-scoped more clearly, or be consolidated into one canonical document. Redundant docs are dangerous because they drift silently — two docs saying the same thing will eventually say different things.
  10. Delete, don't archive. There is no _archived/ directory. When a doc is no longer useful, delete it. Git history preserves every deleted file — that is the archive. A dedicated archive directory creates problems: archived docs accumulate, pollute search results, and nobody reads them. If someone needs a deleted doc, git log --diff-filter=D -- docs/solutions/ will find it.

Scope Selection

Start by discovering learnings and pattern docs under docs/solutions/.

Exclude:

  • README.md
  • docs/solutions/_archived/ (legacy — if this directory exists, flag it for cleanup in the report)

Find all .md files under docs/solutions/, excluding README.md files and anything under _archived/. If an _archived/ directory exists, note it in the report as a legacy artifact that should be cleaned up (files either restored or deleted).

If invocation arguments remain, use them to narrow scope before proceeding. Try these matching strategies in order, stopping at the first that produces results:

  1. Directory match — check if the argument matches a subdirectory name under docs/solutions/ (e.g., performance-issues, database-issues)
  2. Frontmatter match — search module, component, or tags fields in learning frontmatter for the argument
  3. Filename match — match against filenames (partial matches are fine)
  4. Content search — search file contents for the argument as a keyword (useful for feature names or feature areas)

If no matches are found, report that and ask the user to clarify. In headless mode, when a scope hint was provided but matched nothing, report the miss in the summary and exit without widening to all docs — do not silently fall back to processing everything. (The "process everything" rule from Headless mode rules applies only when no scope hint was provided.)

If no candidate docs are found, report:

text
No candidate docs found in docs/solutions/.
Run `spec-compound` after solving problems to start building your knowledge base.

Phase 0: Assess and Route

Before asking the user to classify anything:

  1. Discover candidate artifacts
  2. Estimate scope
  3. Choose the lightest interaction path that fits
Route by Scope
ScopeWhen to use itInteraction style
Focused1-2 likely files or user named a specific docInvestigate directly, then present a recommendation
BatchUp to ~8 mostly independent docsInvestigate first, then present grouped recommendations
Broad9+ docs, ambiguous, or repo-wide stale-doc sweepTriage first, then investigate in batches
Broad Scope Triage

When scope is broad (9+ candidate docs), do a lightweight triage before deep investigation:

  1. Inventory — read frontmatter of all candidate docs, group by module/component/category
  2. Impact clustering — identify areas with the densest clusters of learnings + pattern docs. A cluster of 5 learnings and 2 patterns covering the same module is higher-impact than 5 isolated single-doc areas, because staleness in one doc is likely to affect the others.
  3. Spot-check drift — for each cluster, check whether the primary referenced files still exist. Missing references in a high-impact cluster = strongest signal for where to start.
  4. Recommend a starting area — present the highest-impact cluster with a brief rationale and ask the user to confirm or redirect. In headless mode, skip the question and process all clusters in impact order.

Example:

text
Found 24 learnings across 5 areas.

The auth module has 5 learnings and 2 pattern docs that cross-reference
each other — and 3 of those reference files that no longer exist.
I'd start there.

1. Start with auth (recommended)
2. Pick a different area
3. Review everything

Do not ask action-selection questions yet. First gather evidence.

Phase 1: Investigate Candidate Learnings

For each learning in scope, read it, cross-reference its claims against the current codebase, and form a recommendation.

A learning has several dimensions that can independently go stale. Surface-level checks catch the obvious drift, but staleness often hides deeper:

  • References — do the file paths, class names, and modules it mentions still exist or have they moved?
  • Recommended solution — does the fix still match how the code actually works today? A renamed file with a completely different implementation pattern is not just a path update.
  • Code examples — if the learning includes code snippets, do they still reflect the current implementation?
  • Related docs — are cross-referenced learnings and patterns still present and consistent?
  • Auto memory (Claude Code only) — does the injected auto-memory block in your system prompt contain entries in the same problem domain? Scan that block directly. If the block is absent, skip this dimension. A memory note describing a different approach than what the learning recommends is a supplementary drift signal.
  • Overlap — while investigating, note when another doc in scope covers the same problem domain, references the same files, or recommends a similar solution. For each overlap, record: the two file paths, which dimensions overlap (problem, solution, root cause, files, prevention), and which doc appears broader or more current. These signals feed Phase 1.75 (Document-Set Analysis).
  • Vocabulary — note domain terms the learning cites (entities, named processes, status concepts with project-specific meaning). For each term: does it appear in CONCEPTS.md? If yes, does the definition still match how the code uses the term? If no, flag the term for Phase 4.5 to add or bootstrap. Do not edit CONCEPTS.md during investigation — just collect the signal centrally.

Match investigation depth to the learning's specificity — a learning referencing exact file paths and code snippets needs more verification than one describing a general principle.

Drift Classification: Update vs Replace

The critical distinction is whether the drift is cosmetic (references moved but the solution is the same) or substantive (the solution itself changed):

  • Update territory — file paths moved, classes renamed, links broke, metadata drifted, but the core recommended approach is still how the code works. spec-compound-refresh fixes these directly.
  • Replace territory — the recommended solution conflicts with current code, the architectural approach changed, or the pattern is no longer the preferred way. This means a new learning needs to be written. An authorized replacement subagent may draft the successor following spec-compound's document format, using the investigation evidence already gathered, but it never writes the tracked successor. Without authorized dispatch, the orchestrator composes the replacement inline or serially. In every path the orchestrator is the sole tracked-file writer, and the successor must pass the same source_refs / invalidation_condition promotion gate as a new spec-compound learning.

The boundary: if you find yourself rewriting the solution section or changing what the learning recommends, stop — that is Replace, not Update.

Memory-sourced drift signals are supplementary, not primary. A memory note describing a different approach does not alone justify Replace or Delete. Use memory signals to:

  • Corroborate codebase-sourced drift (strengthens the case for Replace)
  • Prompt deeper investigation when codebase evidence is borderline
  • Add context to the evidence report ("(auto memory [claude]) notes suggest approach X may have changed since this learning was written")

In headless mode, memory-only drift (no codebase corroboration) should result in stale-marking, not action.

Judgment Guidelines

Three guidelines that are easy to get wrong:

  1. Contradiction = strong Replace signal. If the learning's recommendation conflicts with current code patterns or a recently verified fix, that is not a minor drift — the learning is actively misleading. Classify as Replace.
  2. Age alone is not a stale signal. A 2-year-old learning that still matches current code is fine. Only use age as a prompt to inspect more carefully.
  3. Check for successors before deleting. Before recommending Replace or Delete, look for newer learnings, pattern docs, PRs, or issues covering the same problem space. If successor evidence exists, prefer Replace over Delete so readers are directed to the newer guidance.

Phase 1.5: Investigate Pattern Docs

After reviewing the underlying learning docs, investigate any relevant pattern docs under docs/solutions/patterns/.

Pattern docs are high-leverage — a stale pattern is more dangerous than a stale individual learning because future work may treat it as broadly applicable guidance. Evaluate whether the generalized rule still holds given the refreshed state of the learnings it depends on.

A pattern doc with no clear supporting learnings is a stale signal — investigate carefully before keeping it unchanged.

Phase 1.75: Document-Set Analysis

After investigating individual docs, step back and evaluate the document set as a whole. The goal is to catch problems that only become visible when comparing docs to each other — not just to reality.

Overlap Detection

For docs that share the same module, component, tags, or problem domain, compare them across these dimensions:

  • Problem statement — do they describe the same underlying problem?
  • Solution shape — do they recommend the same approach, even if worded differently?
  • Referenced files — do they point to the same code paths?
  • Prevention rules — do they repeat the same prevention bullets?
  • Root cause — do they identify the same root cause?

High overlap across 3+ dimensions is a strong Consolidate signal. The question to ask: "Would a future maintainer need to read both docs to get the current truth, or is one mostly repeating the other?"

Supersession Signals

Detect "older narrow precursor, newer canonical doc" patterns:

  • A newer doc covers the same files, same workflow, and broader runtime behavior than an older doc
  • An older doc describes a specific incident that a newer doc generalizes into a pattern
  • Two docs recommend the same fix but the newer one has better context, examples, or scope

When a newer doc clearly subsumes an older one, the older doc is a consolidation candidate — its unique content (if any) should be merged into the newer doc, and the older doc should be deleted.

Canonical Doc Identification

For each topic cluster (docs sharing a problem domain), identify which doc is the canonical source of truth:

  • Usually the most recent, broadest, most accurate doc in the cluster
  • The one a maintainer should find first when searching for this topic
  • The one that other docs should point to, not duplicate

All other docs in the cluster are either:

  • Distinct — they cover a meaningfully different sub-problem and have independent retrieval value. Keep them separate.
  • Subsumed — their unique content fits as a section in the canonical doc. Consolidate.
  • Redundant — they add nothing the canonical doc doesn't already say. Delete.
Retrieval-Value Test

Before recommending that two docs stay separate, apply this test: "If a maintainer searched for this topic six months from now, would having these as separate docs improve discoverability, or just create drift risk?"

Separate docs earn their keep only when:

  • They cover genuinely different sub-problems that someone might search for independently
  • They target different audiences or contexts (e.g., one is about debugging, another about prevention)
  • Merging them would create an unwieldy doc that is harder to navigate than two focused ones

If none of these apply, prefer consolidation. Two docs covering the same ground will eventually drift apart and contradict each other — that is worse than a slightly longer single doc.

Cross-Doc Conflict Check

Look for outright contradictions between docs in scope:

  • Doc A says "always use approach X" while Doc B says "avoid approach X"
  • Doc A references a file path that Doc B says was deprecated
  • Doc A and Doc B describe different root causes for what appears to be the same problem

Contradictions between docs are more urgent than individual staleness — they actively confuse readers. Flag these for immediate resolution, either through Consolidate (if one is right and the other is a stale version of the same truth) or through targeted Update/Replace.

Subagent Strategy

Before any investigation or replacement dispatch, record:

yaml
worker_dispatch_authorization: authorized | missing
capability_probe: not_applicable | attempted | unavailable
worker_dispatch_capability: available | missing | unknown
worker_context_isolation: isolated | inherited | unknown
worker_model_override: supported | unsupported | unknown
worker_bounded_parallelism: supported | unsupported | unknown

workflow invocation does not authorize dispatch。只有当前用户或可见 upstream handoff 明确请求 subagent、delegated work、persona 或 parallel work 时才可派发;headless mode、scope size、permission settings 或本 Skill 被调用都不是授权。缺授权时不得探测 tool schema,固定为 capability_probe: not_applicable + worker_dispatch_capability: unknown,使用 main-thread/serial fallback 并记录 dispatch_authorization_missing。只有授权后才把 current-session registry/schema 作为 provider_untrusted evidence 检查:确认缺失时记录 subagent_capability_missing;surface 不可用、schema 不完整或候选不唯一时记录 worker_capability_unproven,均使用同一 fallback。隔离、模型覆盖和有界并发只取 live facts;required isolation 未满足时保持依赖 gate 打开,model unknown 时继承,parallelism unknown 时串行。记录 worker_dispatch_outcome。Inline fallback 不得声称 independent investigation coverage。

When dispatch is authorized and available, use subagents for context isolation when investigating multiple artifacts — not just because the task sounds complex. Otherwise choose the matching main-thread or serial approach with the same evidence contract:

ApproachWhen to use
Main thread onlySmall scope, short docs
Sequential subagents1-2 artifacts with many supporting files to read
Parallel subagents3+ truly independent artifacts with low overlap
Batched subagentsBroad sweeps — narrow scope first, then investigate in batches

When spawning an authorized subagent, omit the mode parameter so the user's configured permission settings apply; those settings are execution conditions, not authorization. Include this instruction in its task prompt:

Use dedicated file search and read tools (Glob, Grep, Read) for all investigation. Do NOT use shell commands (ls, find, cat, grep, test, bash) for file operations. This avoids permission prompts and is more reliable.

Also scan the "user's auto-memory" block injected into your system prompt (Claude Code only). Check for notes related to the learning's problem domain. Report any memory-sourced drift signals separately from codebase-sourced evidence, tagged with "(auto memory [claude])" in the evidence section. If the block is not present in your context, skip this check.

There are two subagent roles:

  1. Investigation subagents — read-only. They must not edit files, create successors, or delete anything. Each returns: file path, evidence, recommended action, confidence, and open questions. These can run in parallel when artifacts are independent.
  2. Replacement subagents — draft one candidate successor and return its content or a run-local scratch reference. They never write tracked files, stage, commit, or delete. These run one at a time, sequentially; the orchestrator validates the draft and performs the tracked successor write, deletion, and metadata updates.

The orchestrator merges investigation results, detects contradictions, coordinates replacement subagents, and performs all deletions/metadata edits centrally. In interactive mode, it asks the user questions on ambiguous cases. In headless mode, it marks ambiguous cases as stale instead. If two artifacts overlap or discuss the same root issue, investigate them together rather than parallelizing.

Phase 2: Classify the Right Maintenance Action

After gathering evidence, assign one recommended action.

Keep

The learning is still accurate and useful. Do not edit the file — report that it was reviewed and remains trustworthy. Only add last_refreshed if you are already making a meaningful update for another reason.

Update

The core solution is still valid but references have drifted (paths, class names, links, code snippets, metadata). Apply the fixes directly.

Consolidate

Choose Consolidate when Phase 1.75 identified docs that overlap heavily but are both materially correct. This is different from Update (which fixes drift in a single doc) and Replace (which rewrites misleading guidance). Consolidate handles the "both right, one subsumes the other" case.

When to consolidate:

  • Two docs describe the same problem and recommend the same (or compatible) solution
  • One doc is a narrow precursor and a newer doc covers the same ground more broadly
  • The unique content from the subsumed doc can fit as a section or addendum in the canonical doc
  • Keeping both creates drift risk without meaningful retrieval benefit

When NOT to consolidate (apply the Retrieval-Value Test from Phase 1.75):

  • The docs cover genuinely different sub-problems that someone would search for independently
  • Merging would create an unwieldy doc that harms navigation more than drift risk harms accuracy

Consolidate vs Delete: If the subsumed doc has unique content worth preserving (edge cases, alternative approaches, extra prevention rules), use Consolidate to merge that content first. If the subsumed doc adds nothing the canonical doc doesn't already say, skip straight to Delete.

The Consolidate action is: merge unique content from the subsumed doc into the canonical doc, then delete the subsumed doc. Not archive — delete. Git history preserves it.

Replace

Choose Replace when the learning's core guidance is now misleading — the recommended fix changed materially, the root cause or architecture shifted, or the preferred pattern is different.

The user may have invoked the refresh months after the original learning was written. Do not ask them for replacement context they are unlikely to have — use agent intelligence to investigate the codebase and synthesize the replacement.

Evidence assessment:

By the time you identify a Replace candidate, Phase 1 investigation has already gathered significant evidence: the old learning's claims, what the current code actually does, and where the drift occurred. Assess whether this evidence is sufficient to write a trustworthy replacement:

  • Sufficient evidence — you understand both what the old learning recommended AND what the current approach is. The investigation found the current code patterns, the new file locations, the changed architecture. → Proceed to write the replacement (see Phase 4 Replace Flow).
  • Insufficient evidence — the drift is so fundamental that you cannot confidently document the current approach. The entire subsystem was replaced, or the new architecture is too complex to understand from a file scan alone. → Mark as stale in place:
    • Add status: stale, stale_reason: [what you found], stale_date: YYYY-MM-DD to the frontmatter
    • Report what evidence you found and what is missing
    • Recommend the user run spec-compound after their next encounter with that area, when they have fresh problem-solving context
Delete

Choose Delete when:

  • The code or workflow no longer exists and the problem domain is gone
  • The learning is obsolete and has no modern replacement worth documenting
  • The learning is fully redundant with another doc (use Consolidate if there is unique content to merge first)
  • There is no meaningful successor evidence suggesting it should be replaced instead

Action: delete the file. No archival directory, no metadata — just delete it. Git history preserves every deleted file if recovery is ever needed.

Before deleting: check if the problem domain is still active

When a learning's referenced files are gone, that is strong evidence — but only that the implementation is gone. Before deleting, reason about whether the problem the learning solves is still a concern in the codebase:

  • A learning about session token storage where auth_token.rb is gone — does the application still handle session tokens? If so, the concept persists under a new implementation. That is Replace, not Delete.
  • A learning about a deprecated API endpoint where the entire feature was removed — the problem domain is gone. That is Delete.

Do not search mechanically for keywords from the old learning. Instead, understand what problem the learning addresses, then investigate whether that problem domain still exists in the codebase. The agent understands concepts — use that understanding to look for where the problem lives now, not where the old code used to be.

A doc that other files cite is load-bearing in a way the doc itself does not announce. Before classifying as Delete, search the repo's markdown content (other docs, plans, instruction files, READMEs) for citations of the file — not source code, where citations are rare and only appear in comments. The filename slug is usually unique enough that one query covers all citation sites.

Search efficiently:

  • Prefer the platform's native content-search tool (e.g., Grep in Claude Code) over shell. Drop to shell when materially better for the case.
  • Search the filename slug (without .md); narrow to the full path only if matches are noisy.
  • Read context lines around each match (e.g., Grep's -B/-A), not whole files.

Inbound links inform the classification, not the cleanup. Removing a citation is always mechanical (drop the parenthetical, the bare entry, or the deferring clause). The judgment is upstream: given these citations, is Delete still right, or is Replace closer to right?

Classify each citation by what it does in its citing context:

  • Decorative — principle stated inline, citation is a "see also" pointer or bare attribution. Delete is fine; clean up citations in the same refresh change set.
  • Substantive — citing doc relies on the cited doc to provide content not stated inline (e.g., "see X for details on Y" with no inline Y). Signal Replace — write a successor at the same path, or Keep with narrowed scope if the doc's actual content is broader than its title implies.
  • Mixed or unclear — stale-mark.

In headless mode, Delete + decorative cleanup is fine. Any substantive citation, or any genuine ambiguity, downgrades to stale-marking — writing a Replace successor is judgment-heavy and should not happen unattended.

Auto-delete only when all three hold:

  • The implementation is gone (or fully superseded by a clearly better successor, or the doc is plainly redundant).
  • The problem domain is gone — the app no longer deals with what the learning addresses.
  • Inbound links are absent or unambiguously decorative.

If any condition fails, classify as Replace, Update, Consolidate, or stale-mark per the rules above. Do not delete a learning whose problem domain is still active or whose principles are cited substantively — fill the gap with a replacement instead.

Show full SKILL.md (3,144 more words)Show less

Pattern Guidance

Apply the same five outcomes (Keep, Update, Consolidate, Replace, Delete) to pattern docs, but evaluate them as derived guidance rather than incident-level learnings. Key differences:

  • Keep: the underlying learnings still support the generalized rule and examples remain representative
  • Update: the rule holds but examples, links, scope, or supporting references drifted
  • Consolidate: two pattern docs generalize the same set of learnings or cover the same design concern — merge into one canonical pattern
  • Replace: the generalized rule is now misleading, or the underlying learnings support a different synthesis. Base the replacement on the refreshed learning set — do not invent new rules from guesswork
  • Delete: the pattern is no longer valid, no longer recurring, or fully subsumed by a stronger pattern doc with no unique content remaining

Phase 3: Ask for Decisions

Headless mode

Skip this entire phase. Do not ask any questions. Do not present options. Do not wait for input. Proceed directly to Phase 4 and execute all actions based on the classifications from Phase 2:

  • Unambiguous Keep, Update, Consolidate, auto-Delete, and Replace (with sufficient evidence) → execute directly
  • Ambiguous cases → mark as stale
  • Then generate the report (see Output Format)
Interactive mode

Most Updates and Consolidations should be applied directly without asking. Only ask the user when:

  • The right action is genuinely ambiguous (Update vs Replace vs Consolidate vs Delete)
  • You are about to Delete a document and the evidence is not unambiguous (see auto-delete criteria in Phase 2). When auto-delete criteria are met, proceed without asking.
  • You are about to Consolidate and the choice of canonical doc is not clear-cut
  • You are about to create a successor via Replace

Do not ask questions about whether code changes were intentional, whether the user wants to fix bugs in the code, or other concerns outside doc maintenance. Stay in your lane — doc accuracy.

Question Style

Always present choices using the platform's blocking question tool: AskUserQuestion in Claude Code (call ToolSearch with select:AskUserQuestion first if its schema isn't loaded), request_user_input in Codex. Fall back to numbered options in plain text only when no blocking tool exists in the harness or the call errors (e.g., Codex edit modes) — not because a schema load is required. Never silently skip the question.

Question rules:

  • Ask one question at a time
  • Prefer multiple choice
  • Lead with the recommended option
  • Explain the rationale for the recommendation in one concise sentence
  • Avoid asking the user to choose from actions that are not actually plausible
Focused Scope

For a single artifact, present:

  • file path
  • 2-4 bullets of evidence
  • recommended action

Then ask:

text
This [learning/pattern] looks like a [Keep/Update/Consolidate/Replace/Delete].

Why: [one-sentence rationale based on the evidence]

What would you like to do?

1. [Recommended action]
2. [Second plausible action]
3. Skip for now

Do not list all five actions unless all five are genuinely plausible.

Batch Scope

For several learnings:

  1. Group obvious Keep cases together
  2. Group obvious Update cases together when the fixes are straightforward
  3. Present Consolidate cases together when the canonical doc is clear
  4. Present Replace cases individually or in very small groups
  5. Present Delete cases individually unless they are strong auto-delete candidates

Ask for confirmation in stages:

  1. Confirm grouped Keep/Update recommendations
  2. Then handle Consolidate groups (present the canonical doc and what gets merged)
  3. Then handle Replace one at a time
  4. Then handle Delete one at a time unless the deletion is unambiguous and safe to auto-apply
Broad Scope

If the user asked for a sweeping refresh, keep the interaction incremental:

  1. Narrow scope first
  2. Investigate a manageable batch
  3. Present recommendations
  4. Ask whether to continue to the next batch

Do not front-load the user with a full maintenance queue.

Phase 4: Execute the Chosen Action

For each candidate, execute the flow that matches its classification from Phase 2 (confirmed in Phase 3). Read references/per-action-flows.md and follow the matching section:

  • Keep — no file edit by default; summarize why the learning remains trustworthy.
  • Update — in-place edits when the solution is still substantively correct (path renames, link refreshes, module renames).
  • Consolidate — merge overlapping docs into a canonical doc, apply the same promotion exit to the materially rewritten canonical doc (and every new split successor), then update cross-references and delete subsumed docs. The orchestrator handles consolidation directly.
  • Replace — obtain a successor draft through an authorized subagent or inline/serial fallback, then let the orchestrator write the tracked successor, validate parser safety plus the source_refs / invalidation_condition promotion exit, validate cited claims, and only then delete the old. When evidence is insufficient, mark stale instead.
  • Delete — final inbound-link check, then remove. Reclassify if late-discovered substantive citations surface.

Only one flow runs per candidate; the reference contains the per-action criteria, examples, and step-by-step instructions.

Phase 4.5: Vocabulary Capture

After the per-learning actions execute, aggregate the domain terms flagged across Phase 1's Vocabulary dimension and reconcile them with CONCEPTS.md.

First, read references/concepts-vocabulary.md. This is unconditional. Do not pre-judge from memory which Phase 1 signals qualify — the reference's criteria are non-obvious and a "nothing qualifies" judgment without reading is a shortcut, not a result.

Procedure:

  1. Aggregate. Collect qualifying terms surfaced across the learnings in scope, applying the reference's criteria. If the same term surfaced in multiple learnings with different shades of precision, union the shades into one entry — not three entries, not most-recent-wins.

  2. If CONCEPTS.md exists, add missing terms and refine existing entries when the corpus surfaced new precision. Do not duplicate entries already present. Then reconcile the in-scope core nouns: re-derive the core domain nouns of the area in scope from its declared model (per the Seed goal in the reference) and backfill any that are central but missing. This is the every-run safety net for stable-central terms that friction never surfaces — bounded to the area in scope, defining only terms investigated this run, never a repo-wide sweep.

  3. If CONCEPTS.md does not exist and at least one qualifying term was surfaced, bootstrap it — and seed, don't write a single term. Alongside the surfaced term(s), seed the core domain nouns of the area in scope per the reference's Seed goal, so the file is anchored from creation rather than a lone peripheral entry (and so captured terms don't dangle against undefined siblings). The seed stays scoped to the area in scope — a repo-wide concept map comes only from the explicit bootstrap path above, not from a scoped refresh. At creation, hold the qualifying bar conservatively for borderline terms — a borderline term or a class/table/file name dressed up as an entity defers to a later run; clear core nouns are seeded, borderline ones wait. The conservatism is about quality, not count; updates to an existing file follow normal criteria.

  4. Scope discipline and citation hygiene. Bootstrap, seed, and reconcile reflect only the area in scope — do not expand to other categories, and do not retroactively inject (see CONCEPTS.md) pointers into existing learnings. (The repo-wide bootstrap path above is the deliberate exception — it intentionally covers the whole declared model.) The report should note that additional entries are likely from refresh runs on other scopes.

  5. Initial structure. When bootstrapping, start the file with this preamble under the # Concepts heading:

    Shared domain vocabulary for this project — entities, named processes, and status concepts with project-specific meaning. Seeded with core domain vocabulary, then accretes as spec-compound and spec-compound-refresh process learnings; direct edits are fine. Glossary only, not a spec or catch-all.

    Then add entries. Let term count drive shape: 1-4 terms → flat headings, more → cluster by domain relationship per the rules in references/concepts-vocabulary.md.

  6. Scrub violations. Scan existing entries for content that violates references/concepts-vocabulary.md criteria — implementation specifics (file paths, class names, function signatures, code references), current-config values (thresholds, counts, enum values that will drift), status/owner/date metadata, duplicates of terms covered under a different name, or entries that lean on an undefined project-specific sibling (add the sibling or rephrase). Rewrite or consolidate. The full sweep is appropriate here because refresh is an audit; spec-compound's same-named phase scopes corrections to the coherence neighborhood of entries being touched.

If no Phase 1 signals qualified after applying the reference's criteria, record that outcome explicitly in the report's CONCEPTS.md line (e.g., "scanned, no qualifying terms"). Do not silently skip — the visible scan-and-no-result record is the audit signal that the reference was consulted.

Note: if this run creates CONCEPTS.md from scratch, the Discoverability Check below also surfaces it so future agents can discover it — by editing AGENTS.md/CLAUDE.md in interactive mode (with consent), or, in headless mode, by emitting a "Discoverability recommendation" line in the report rather than editing instruction files (per the headless boundary in step 4c — headless does doc maintenance, not project config). Either way the created file is surfaced or flagged for surfacing; subsequent runs skip this because the instruction file is already current or the recommendation was already reported.

Apply edits silently — no user prompt in any mode. Vocabulary capture is a side effect of refreshing, not a decision the user makes per run.

Output Format

The full report MUST be printed as markdown output. Do not summarize findings internally and then output a one-liner. The report is the deliverable — print every section in full, formatted as readable markdown with headers, tables, and bullet points.

After processing the selected scope, output the following report:

text
Compound Refresh Summary
========================
Scanned: N learnings

Kept: X
Updated: Y
Consolidated: C
Replaced: Z
Deleted: W
Skipped: V
Marked stale: S

CONCEPTS.md: <scanned, no qualifying terms | created with N entries (M seeded) | updated — N added, N refined, N reconciled, N scrubbed | repo-wide map created with N entries>

Then for EVERY file processed, list:

  • The file path
  • The classification (Keep/Update/Consolidate/Replace/Delete/Stale)
  • What evidence was found -- tag any memory-sourced findings with "(auto memory [claude])" to distinguish them from codebase-sourced evidence
  • What action was taken (or recommended)
  • For Consolidate: which doc was canonical, what unique content was merged, what was deleted

For Keep outcomes, list them under a reviewed-without-edits section so the result is visible without creating git churn.

Headless mode report

In headless mode, the report is the sole deliverable — there is no user present to ask follow-up questions, so the report must be self-contained and complete. Print the full report. Do not abbreviate, summarize, or skip sections.

Split actions into two sections:

Applied (writes that succeeded):

  • For each Updated file: the file path, what references were fixed, and why
  • For each Consolidated cluster: the canonical doc, what unique content was merged from each subsumed doc, and the subsumed docs that were deleted
  • For each Replaced file: what the old learning recommended vs what the current code does, and the path to the new successor
  • For each Deleted file: the file path and why it was removed (problem domain gone, fully redundant, etc.)
  • For each Marked stale file: the file path, what evidence was found, and why it was ambiguous

Recommended (actions that could not be written — e.g., permission denied):

  • Same detail as above, but framed as recommendations for a human to apply
  • Include enough context that the user can apply the change manually or re-run the skill interactively

If all writes succeed, the Recommended section is empty. If no writes succeed (e.g., read-only invocation), all actions appear under Recommended — the report becomes a maintenance plan.

Legacy cleanup (if docs/solutions/_archived/ exists):

  • List archived files found and recommend disposition: restore (if still relevant), delete (if truly obsolete), or consolidate (if overlapping with active docs)

Phase 5: Commit Changes

After all actions are executed and the report is generated, close out Git state without widening authority. Skip this phase if no files were modified (all Keep, or all writes failed).

Detect git context

Before any Git action, check:

  1. Which branch is currently checked out (main/master vs feature branch)
  2. Whether the working tree has other uncommitted changes beyond what compound-refresh modified
  3. Recent commit messages to match the repo's commit style
  4. The exact files modified by this run and whether any pre-existing dirty hunks overlap them
Missing commit authorization

When commit_authorization: missing, do not create/switch a branch, stage, commit, push, or open a PR. Leave verified refresh edits uncommitted and include:

  • commit_status: not-created
  • commit_reason: commit_authorization_missing
  • the exact uncommitted paths
  • one logical commit candidate message for a later authorized owner

Headless mode never asks for authority and therefore follows this path unless the visible upstream handoff explicitly supplied commit authorization.

Authorized commit

When commit_authorization: authorized, stage only compound-refresh-owned verified paths. Never stage unrelated dirty paths. A pre-existing overlapping dirty hunk requires an explicit bounded preservation decision; otherwise leave the refresh uncommitted instead of guessing ownership.

On main/master/default branch, commit authorization alone does not authorize branch creation or direct default-branch commit. Require the current user to name the intended branch/default-branch action explicitly; otherwise keep the changes uncommitted.

On an existing feature branch, create one isolated commit only after the refresh checks pass. Report the commit SHA and exact paths.

Landing authority

commit_authorization never implies landing_authorization. Without landing authorization, stop after the local commit and do not push or open/update a PR. With explicit landing authorization, push only the authorized branch and create/update only the named PR target after the workflow's verification and report gates pass.

Interactive authorization acquisition

In interactive mode, if the initial request did not authorize commit, offer only a bounded commit decision after presenting the verified diff: Commit these refresh-owned paths or Leave verified changes uncommitted. A positive answer authorizes the local commit only. Ask a separate question for push/PR only when outward landing is genuinely requested; never bundle commit and landing into one implied choice.

Commit message

Write a descriptive commit message that:

  • Summarizes what was refreshed (e.g., "update 3 stale learnings, consolidate 2 overlapping docs, delete 1 obsolete doc")
  • Follows the repo's existing commit conventions (check recent git log for style)
  • Is succinct — the details are in the changed files themselves

Relationship to spec-compound

  • spec-compound captures a newly solved, verified problem
  • spec-compound-refresh maintains older learnings as the codebase evolves — both their individual accuracy and their collective design as a document set

Use Replace only when the refresh process has enough real evidence to write a trustworthy successor. When evidence is insufficient, mark as stale and recommend spec-compound for when the user next encounters that problem area.

Use Consolidate proactively when the document set has grown organically and redundancy has crept in. Every spec-compound invocation adds a new doc — over time, multiple docs may cover the same problem from slightly different angles. Periodic consolidation keeps the document set lean and authoritative.

Discoverability Check

After the refresh report is generated, check whether the project's instruction files would lead an agent to discover and search docs/solutions/ before starting work in a documented area. This runs every time — the knowledge store only compounds value when agents can find it. If this check produces edits, they stay under the same Phase 5 commit and landing authority facts — see step 6 below.

  1. Identify which root-level instruction files exist (AGENTS.md, CLAUDE.md, or both). Read the file(s) and determine which holds the substantive content — one file may just be a shim that @-includes the other (e.g., CLAUDE.md containing only @AGENTS.md, or vice versa). The substantive file is the assessment and edit target; ignore shims. If neither file exists, skip this check entirely.

  2. Assess whether an agent reading the instruction files would learn three things:

    • That a searchable knowledge store of documented solutions exists
    • Enough about its structure to search effectively (category organization, YAML frontmatter fields like module, tags, problem_type)
    • When to search it (before implementing features, debugging issues, or making decisions in documented areas — learnings may cover bugs, best practices, workflow patterns, or other institutional knowledge)

    This is a semantic assessment, not a string match. The information could be a line in an architecture section, a bullet in a gotchas section, spread across multiple places, or expressed without ever using the exact path docs/solutions/. Use judgment — if an agent would reasonably discover and use the knowledge store after reading the file, the check passes.

  3. If the spirit is already met, no action needed.

  4. If not: a. Based on the file's existing structure, tone, and density, identify where a mention fits naturally. Before creating a new section, check whether the information could be a single line in the closest related section — an architecture tree, a directory listing, a documentation section, or a conventions block. A line added to an existing section is almost always better than a new headed section. Only add a new section as a last resort when the file has clear sectioned structure and nothing is even remotely related. b. Draft the smallest addition that communicates the three things. Match the file's existing style and density. The addition should describe the knowledge store itself, not the plugin.

    Keep the tone informational, not imperative. Express timing as description, not instruction — "relevant when implementing or debugging in documented areas" rather than "check before implementing or debugging." Imperative directives like "always search before implementing" cause redundant reads when a workflow already includes a dedicated search step. The goal is awareness: agents learn the folder exists and what's in it, then use their own judgment about when to consult it.

    Examples of calibration (not templates — adapt to the file):

    When there's an existing directory listing or architecture section — add a line:

    docs/solutions/  # documented solutions to past problems (bugs, best practices, workflow patterns), organized by category with YAML frontmatter (module, tags, problem_type)

    When nothing in the file is a natural fit — a small headed section is appropriate:

    ## Documented Solutions
    
    `docs/solutions/` — documented solutions to past problems (bugs, best practices, workflow patterns), organized by category with YAML frontmatter (`module`, `tags`, `problem_type`). Relevant when implementing or debugging in documented areas.

    c. In interactive mode, explain to the user why this matters — agents working in this repo (including fresh sessions, other tools, or collaborators without the plugin) won't know to check docs/solutions/ unless the instruction file surfaces it. Show the proposed change and where it would go, then use the platform's blocking question tool to get consent before making the edit: AskUserQuestion in Claude Code (call ToolSearch with select:AskUserQuestion first if its schema isn't loaded), request_user_input in Codex. Fall back to presenting the proposal in chat only when no blocking tool exists in the harness or the call errors (e.g., Codex edit modes) — not because a schema load is required. Never silently skip the question. In headless mode, include it as a "Discoverability recommendation" line in the report — do not attempt to edit instruction files (headless scope is doc maintenance, not project config).

  5. If CONCEPTS.md exists at repo root, run a parallel discoverability check for it. Use the same workflow as the docs/solutions/ check above: same target file, same edit-placement judgment, same consent-then-edit interaction shape per mode. Example calibration when a directory listing is present:

    CONCEPTS.md  # shared domain vocabulary — read when orienting to the codebase or before discussing domain concepts

    Skip this step entirely if CONCEPTS.md does not exist — never nag for an artifact the project has not adopted. When skipped, this step produces no output and no edit.

  6. Keep discoverability edits under the same authority facts. If step 4 or step 5 edited an instruction file and an authorized local commit already exists, stage only that run-owned file and amend or create a focused follow-up commit. Without commit authorization, leave it unstaged with the other verified refresh changes. Without landing authorization, do not push either commit; an existing remote branch or PR does not widen authority.

© leo-kuang-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files (scripts, references, assets) in skills/spec-compound-refresh of leo-kuang-ai/spec-first.

  • SKILL.md
  • assets/resolution-template.md
  • evals/cases/non-solutions-scope-gated.yaml
  • evals/eval.yaml
  • evals/fixtures/repos/mini-ledger/README.md
  • evals/fixtures/repos/mini-ledger/package.json
  • evals/fixtures/repos/mini-ledger/src/server.js
  • evals/fixtures/scripts/check-scope-gate.sh
  • references/concepts-vocabulary.md
  • references/per-action-flows.md
  • references/schema.yaml
  • references/yaml-schema.md
  • … and 4 more

Open the folder on GitHubat commit 74655dc

Compare with similar skills

Spec Compound Refresh next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Spec Compound Refresh compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Spec Compound Refresh this skillleo-kuang-ai/spec-first107—~15kAutomated safety check: WarnMIT
Vercel Composition Patternssupabase/supabase111k58 repos~726Automated safety check: PassMIT
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
Typescript Advanced Typesrolling-scopes/rsschool-app10k25 repos~4.2kAutomated safety check: PassMPL-2.0
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k4 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Official

    React composition patterns that scale. An agent skill from supabase/supabase.

    111k GitHub starsUsed in 58 repos~726 tokens
    DevelopmentAuto-check passed
  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Typescript Advanced Types

    rolling-scopes/rsschool-app

    Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.

    10k GitHub starsUsed in 25 repos~4.2k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Greploop

    onyx-dot-app/onyx

    Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.

    32k GitHub starsUsed in 4 repos~3.3k tokens
    DevelopmentAuto-check passed

More from leo-kuang-ai/spec-first

All 35 skills in this repo
  • Spec App Consistency Audit

    leo-kuang-ai/spec-first

    Audit mobile App PRD/Figma/local-source consistency across page routes, KMP/Clean Architecture, components, analytics, i18n, engineering quality, and industry lenses before runtime validation; use…

    107 GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Handoff

    leo-kuang-ai/spec-first

    Create a durable cross-session handoff or resume from a user-selected continuity source.

    107 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Pov

    leo-kuang-ai/spec-first

    Give a decisive, project-grounded verdict on an external input — judged against the current project, not in the abstract.

    107 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Resolve PR Feedback

    leo-kuang-ai/spec-first

    Resolve PR review feedback by evaluating validity and fixing issues with conflict-aware resolver dispatch.

    107 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check: notes
  • Spec Riffrec Feedback Analysis

    leo-kuang-ai/spec-first

    Analyze explicit Riffrec product-feedback captures, including riffrec-.zip, the Riffrec session.json + events.json + recording.webm + voice.webm bundle, or media/notes the user identifies as a…

    107 GitHub stars~1.4k tokensUpdated 2 days ago
    Auto-check passed
  • Spec Compound

    leo-kuang-ai/spec-first

    Document a recently solved problem or durable project vocabulary in docs/solutions/ or CONCEPTS.md.

    107 GitHub stars~18k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Spec Compound Refresh

What does Spec Compound Refresh do?

Refresh docs/solutions learnings against the current codebase. Spec Compound Refresh is an agent skill from leo-kuang-ai/spec-first. Refresh docs/solutions learnings against the current codebase.

When should I use Spec Compound Refresh?

Spec Compound Refresh fits situations like: drifted learnings; avoid general refactor; code review unless docs/solutions is explicit.

How do I install Spec Compound Refresh in Claude Code?

Run `npx skills add leo-kuang-ai/spec-first --skill spec-compound-refresh -a claude-code`. Or copy the skill folder (skills/spec-compound-refresh in leo-kuang-ai/spec-first) into .claude/skills/spec-compound-refresh in your project. Claude Code loads it when a task matches its description.

How do I install Spec Compound Refresh in Codex?

Run `npx skills add leo-kuang-ai/spec-first --skill spec-compound-refresh -a codex`. Or copy the skill folder (skills/spec-compound-refresh in leo-kuang-ai/spec-first) into .agents/skills/spec-compound-refresh in your project. Codex loads it when a task matches its description.

Can I use Spec Compound Refresh in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add leo-kuang-ai/spec-first --skill spec-compound-refresh -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spec-compound-refresh, .gemini/skills/spec-compound-refresh, .github/skills/spec-compound-refresh and .opencode/skills/spec-compound-refresh in your project.

What does Spec Compound Refresh need to run?

Going by SKILL.md and its folder, Spec Compound Refresh needs JavaScript and a shell for the scripts in its folder and the command-line tools its instructions call (git). Our summary lists: Node.js; A Bash shell.

Does Spec Compound Refresh access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Spec Compound Refresh safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Spec Compound Refresh use?

Spec Compound Refresh is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Spec Compound Refresh use?

About 15k tokens (SKILL.md is roughly 58k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.4k tokens, read only when the agent opens those files.

What are the alternatives to Spec Compound Refresh?

Skills that share tags, products or a category with Spec Compound Refresh: Vercel Composition Patterns (supabase/supabase, 111k stars), Finishing a Development Branch (obra/superpowers, 297k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars) and PR Babysitter (openinterpreter/openinterpreter, 69k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Spec Compound Refresh?

leo-kuang-ai (a GitHub user) maintains it in leo-kuang-ai/spec-first, which has 107 GitHub stars. The repository holds 35 skills in this directory. The repository was last updated on October 8, 2026.

Source: leo-kuang-ai/spec-first on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.