Agent skill

Refinement Advisor

by swingerman in swingerman/engineer

A skill your agent uses to decide which refinement and hardening tools are worth running on a set of changes — refine, crap-analyzer, arch-check, introversion scan, mutation testing, TLA+ model…

MITAuto-check passedTesting & QA

Install Refinement Advisor

skills CLI
$ npx skills add swingerman/engineer --skill refinement-advisor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install swingerman/engineer refinement-advisor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/swingerman/engineer.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineer/skills/refinement-advisor .claude/skills/refinement-advisor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
refinement-advisor
GitHub stars
154
Token cost
~2.7k tokens
SKILL.md length
1,351 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses to decide which refinement and hardening tools are worth running on a set of changes — refine, crap-analyzer, arch-check, introversion scan, mutation testing, TLA+ model…

  • Decide which refinement and hardening tools are worth running on a set of changes — refine
  • SKILL.md covers When to use, Inputs, The toolbox and Signals, per tool, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Introversion scan

What it does

Refinement Advisor is an agent skill from swingerman/engineer. Use to decide which refinement and hardening tools are worth running on a set of changes — refine, crap-analyzer, arch-check, introversion scan, mutation testing, TLA+ model checking, Lean proofs — and on which files, with the invariant to check. Triggers — "/engineer.refinement-advisor", "what should we harden", "which checks fit this diff", "is this worth TLA+ / Lean", "should we model-check this", "how do I harden this change". Called by /engineer.harden at CP8; also usable ad hoc on any diff, in or out of the…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test coverage. The repository describes itself as: Disciplined Agentic Engineering — a methodology kit for Claude Code: acceptance-test-first specs, explicit checkpoints, and autonomy you can actually leave running. The engineer… The licence is MIT.

When your agent uses it

  • Decide which refinement and hardening tools are worth running on a set of changes — refine
  • Introversion scan
  • Mutation testing
  • TLA+ model checking

Example prompts

  • “/engineer.refinement-advisor”
  • “what should we harden”
  • “which checks fit this diff”
  • “/refinement-advisor”

What it can do on your machine

Read from SKILL.md and the folder at commit 32947eb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Refinement Advisor loads about 2.7k tokens when it runs. Until then it costs about 137 tokens; SKILL.md has 1,351 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from swingerman/engineer at commit 32947eb, republished under its MIT licence (© swingerman). 1,351 words, ~2,683 tokens.

Download SKILL.mdSave it as .claude/skills/refinement-advisor/SKILL.md (or your agent's skills folder).
name
refinement-advisor
description
Use to decide which refinement and hardening tools are worth running on a set of changes — refine, crap-analyzer, arch-check, introversion scan, mutation testing, TLA+ model checking, Lean proofs — and on which files, with the invariant to check. Triggers — "/engineer.refinement-advisor", "what should we harden", "which checks fit this diff", "is this worth TLA+ / Lean", "should we model-check this", "how do I harden this change". Called by /engineer.harden at CP8; also usable ad hoc on any diff, in or out of the pipeline.

refinement-advisor

Reads a diff and recommends which of the DAE refinement/hardening tools to run, where, and why. It advises; it does not run the tools. The expensive ones (TLA+, Lean) cost real minutes each, and running every tool on every diff buries the one finding that matters under noise from checks that never fit the code.

The advice is only as good as its reading of the code: read the changed functions, not just filenames or the diff stat.

When to use

  • Called by /engineer.harden (CP8) to pick that checkpoint's tools.
  • Ad hoc: "which checks fit this change?", before a risky merge, or when deciding whether a piece of code deserves formal verification.

Not for: running the tools (harden, or the tools' own skills), or reviewing code quality itself (refine).

Inputs

  • Scope — a feature dir (diff = feature branch vs its branch point), a fix record, a PR, or an explicit ref range. Default: current branch vs its merge-base with the default branch.
  • Stage (optional) — refine, verify, harden, or any (default). Limits the recommendations to that stage's tools.
  • Prior results (optional) — if a CP7 handoff exists, read its crap_results block (arch-check records crap-analyzer's output there). CRAP scores tell you where complex, poorly tested code is. Use them; don't redo that analysis.

The toolbox

ToolStageAnswersCost
/engineer.refineCP6Is the changed code clean: reuse, clarity, efficiency?medium
crap-analyzerCP7Where does high complexity meet low test coverage?low
/engineer.arch-checkCP7Does it respect the charter's layering and naming?low
introversion scan (dae_introvert.py)CP8Can a test pass without asserting anything?low
atdd:atdd-mutateCP8Do the tests fail when the code is wrong?medium
/engineer.tlaplusCP8Can some interleaving or sequence of events break an invariant?high
/engineer.leanCP8Does an invariant hold for every input, including unbounded ones?high

Every tool in scope gets a verdict, recommend or skip, with a reason. Nothing runs by default. A cheap tool that can't find anything in this diff is still noise, and an expensive one that fits is worth its cost. Judge each one against the signals below.

Signals, per tool

/engineer.refine: recommend when CP6 hasn't run on the current diff, or code changed since it did. Skip when refine's handoff already covers these commits.

crap-analyzer / arch-check: recommend when CP7 hasn't run on the current diff. Skip when a CP7 handoff covers these commits, and reuse its numbers instead.

Introversion scan: recommend when test files were added or changed. Skip when no tests changed (there is nothing new to scan).

Mutation (atdd:atdd-mutate): recommend when the change adds or alters branching logic (conditions, loops, error paths, boundaries) and tests exercise it. That's where a weak assertion hides. Skip when:

  • the change is config, docs, markup, styling, generated code, or trivial accessors
  • no tests cover the changed code (mutation just reports "all survived", which crap-analyzer already told you; recommend writing tests instead)
  • mutation already ran on the same files and they haven't changed since
  • the risk is temporal (races, ordering). Mutants don't model interleavings, so that is a TLA+ job, not a mutation job.

tlaplus: recommend when the changed code has temporal or concurrent shape:

  • retry and backoff loops, timeouts, circuit breakers, rate limiters
  • locks, queues, workers, schedulers, cron-style jobs that can overlap
  • async/await, promises, callbacks, events, or streams where order matters
  • an explicit state machine or status enum with transitions, such as session phase, order status, or a connection lifecycle
  • multi-step protocols (handshake, cancel/redeliver, two-phase anything)

Races don't raise complexity metrics, so a low CRAP score says nothing here. Recommend TLA+ based on the shape of the code, not on the risk score.

lean: recommend when there is a claim of the form "for all inputs X, P holds", about pure logic:

  • parsers, sanitizers, maskers, encoders (e.g. "the output never contains the password")
  • money, units, rounding, or date arithmetic
  • permission and authorization predicates
  • ordering, dedup, and merge functions
  • a state machine whose invariant must hold over unbounded counts or data. TLC only checks a finite model; Lean proves the general case.

Choose between TLA+ and Lean by the question: if you are asking "what if these happen in a different order", use TLA+. If you are asking "what if the input is weird", use Lean. If both apply, recommend both, each scoped to its own function. Skip both for CRUD, glue, configuration, rendering, or straight-line code with no invariant worth stating, or when a small table-driven test already covers the whole input space.

The invariant is the deliverable

For every TLA+/Lean recommendation, draft the invariant in one plain sentence ("attempt count never exceeds MAX_RETRIES and every path exits"). Both formal skills say the invariant is the one thing that can't be inferred from the code: it states what must never happen. The advisor's draft is a starting point for the human to correct. If you can't state an invariant, don't recommend the tool.

Show full SKILL.md (541 more words)Show less

Output

refinement-advisor — <scope>  (<N> files changed, stage: <stage>)

| Tool | Verdict | Target | Why | Invariant / focus |
|---|---|---|---|---|
| tlaplus | recommend | ResidentialProxyHttpClient::get() | bounded retry loop, 4 exit paths | attempt ≤ 3; every path returns or throws |
| atdd:atdd-mutate | recommend | ResidentialProxyHttpClient.php | new retry/error branches, covered by 4 tests | — |
| lean | recommend | ResidentialProxyHttpClient::getMaskedProxyUrl() | regex masker, unbounded input shapes | output never contains the password substring |
| introversion scan | recommend | tests/Unit/Http/* | 2 new test files | — |
| crap-analyzer | skip | — | CP7 already ran on these commits (max CRAP 6) | — |
| refine | skip | — | CP6 handoff covers these commits | — |

Sort the recommend rows by value (the findings you expect relative to their cost), highest first. Every skip row gives its reason. End with a one-line bottom line, e.g. "TLA+ on the retry loop is the highest-value check; mutation is next."

Who decides: autonomy

Autonomy controls who makes the call, never which checks are sound. Use the effective autonomy from ${CLAUDE_PLUGIN_ROOT}/references/handoff-dispatch.md: the feature's autonomy_level, capped by any manifest.autonomy.path_overrides that match the changed files. For a fix record, which has no autonomy_level, start from manifest.autonomy.default_level and apply the same path caps. A critical or blocks_user: true fix always counts as low.

Effective autonomyAdvisor behaviour
highDecides alone. Every recommend row is selected and every drafted invariant is final. Show the table, then one line naming what will run and why ("Running TLA+ on get() and mutation on 1 file; skipped 4, reasons above"). Don't ask anything.
medium, lowAsks. Offers the recommendations as choices (below). Nothing runs that the human didn't pick.

manifest.harden.required: true removes deselection, not the human. The recommended rows are mandatory, so Q1 isn't asked; list them as settled. Below high the human still confirms each invariant (Q2) and may add skipped checks (Q3).

The choices (below high)

A table the human then has to answer in prose is a dead end. Offer the recommendations as selectable choices, so acting on the advice takes one click. If nothing is recommended and nothing was skipped, say so and don't ask anything.

Use AskUserQuestion:

  • Q1 "Which checks should I run?" (multiSelect: true). Skip Q1 when harden.required: true. Make one option per recommend row, in value order, and mark the first (Recommended).

    • Label: <tool> → <target> (e.g. TLA+ → get() retry loop, Mutation → ResidentialProxyHttpClient.php).
    • Description: the why, the cost in minutes, and for TLA+/Lean the drafted invariant.
    • If there are more than 4 recommendations, fold the cheapest ones into a single option (Introversion + arch-check) so every expensive check keeps its own row.
    • The human can add a skipped tool through "Other" ("also run mutation").
  • Q2 (per recommended TLA+/Lean row) "Is this the right invariant for <target>?" Offer:

    • Use as drafted (Recommended)
    • Narrower: <a weaker variant>
    • Stronger: <a stricter variant>

    The built-in "Other" lets the human type their own.

  • Q3 (only when harden.required: true and something was skipped) "Add any of the skipped checks?" (multiSelect: true). Make one option per skip row, labelled with its skip reason.

Keep to the tool's 4-question limit. If more questions are needed, prioritise Q2 for the most expensive formal rows.

The output is the selected tools, each with its target and, for TLA+/Lean, the final invariant.

No interactive human (subagent, headless): at high, decide as above. Below high, print the same choices as a numbered list with a copy-pasteable reply line (reply: "1,3" or "all") and stop.

When called by harden, also return the table as YAML (advisor_picks:) so harden can record it in harden_results.advisor.

Handoff

Off-pipeline: checkpoint: null. When run standalone on a feature, emit a short handoff per ${CLAUDE_PLUGIN_ROOT}/references/handoff-summary.md with recommended_next set to the highest-value pick. When called from harden, return the picks inline and write no handoff.

References

  • /engineer.harden: the CP8 consumer
  • ${CLAUDE_PLUGIN_ROOT}/references/gauntlet.md: the other quality bar (visual/qualitative); not in this toolbox
  • /engineer.tlaplus, /engineer.lean: the formal-verification skills; each has a "Verifying real code" workflow

© swingerman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in engineer/skills/refinement-advisor of swingerman/engineer.

Open the folder on GitHubat commit 32947eb

Compare with similar skills

Refinement Advisor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Refinement Advisor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Refinement Advisor this skillswingerman/engineer154—~2.7kAutomated safety check: PassMIT
Requirementsrizsotto/Bear6.5k—~2kAutomated safety check: PassGPL-3.0
Crap Analysisardalis/RiverBooks1342 repos~3.4kAutomated safety check: PassNone
Code Coverages3s-project/s3s311—~789Automated safety check: PassApache-2.0
Project Statusbactopia/bactopia522—~787Automated safety check: PassMIT
Check Coverageldayton/Dippy243—~403Automated safety check: PassMIT

Similar skills

  • Requirements

    rizsotto/Bear

    Write, modify, or review a requirement file under docs/requirements -- pick the single owning file, keep the text contract-only, name IDs so they need no explanation, and verify cross-references and…

    6.5k GitHub stars~2k tokensUpdated today
    Testing & QAAuto-check passed
  • Crap Analysis

    ardalis/RiverBooks

    Analyze code coverage and CRAP (Change Risk Anti-Patterns) scores to identify high-risk code.

    134 GitHub starsUsed in 2 repos~3.4k tokens
    Testing & QAAuto-check passed
  • Code Coverage

    s3s-project/s3s

    Measure and grow the line coverage of the s3s crate. An agent skill from s3s-project/s3s.

    311 GitHub stars~789 tokensUpdated today
    Testing & QAAuto-check passed
  • Project Status

    bactopia/bactopia

    Show a live snapshot of the Bactopia project state — component counts, GroovyDoc coverage, nf-test coverage, and structural issues.

    522 GitHub stars~787 tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Check Coverage

    ldayton/Dippy

    Ensure comprehensive test coverage for a CLI handler. An agent skill from ldayton/Dippy.

    243 GitHub stars~403 tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Guarding Destructive Operations

    kajisho5/ffmpeg-skill

    Add and review preconditions on operations that delete, overwrite, rewrite history, or resolve a caller-supplied name to a filesystem path — refusing instead of warning, placing the guard ahead of…

    1.9k GitHub stars~2.6k tokensUpdated 2 days ago
    Testing & QAAuto-check passed

More from swingerman/engineer

All 27 skills in this repo
  • Crap Analyzer

    swingerman/engineer

    A skill your agent uses to produce a risk-based refactor + test plan for recently-changed code on a diff/branch/PR by computing CRAP (complexity × untested) on changed methods.

    154 GitHub stars~1.2k tokensUpdated 14 days ago
    Auto-check passed
  • Atdd Mutate

    swingerman/engineer

    A skill your agent uses to add a third validation layer to the ATDD workflow — after acceptance tests verify WHAT and unit tests verify HOW, mutation testing verifies the tests actually catch bugs.

    154 GitHub stars~2.7k tokensUpdated 14 days ago
    Auto-check passed
  • Fix

    swingerman/engineer

    A skill your agent uses to drive a bug fix from first report through close, with a "why didn't we catch it?" loop at the end.

    154 GitHub stars~3k tokensUpdated 14 days ago
    Auto-check passed
  • Atdd

    swingerman/engineer

    A skill your agent uses to drive feature work through the Acceptance Test Driven Development workflow — Given/When/Then specs before code, a project-specific test pipeline, and two parallel test…

    154 GitHub stars~2.8k tokensUpdated 14 days ago
    Auto-check passed
  • Harden

    swingerman/engineer

    Use after a feature passes Light Verify (CP7), to prove the tests actually catch bugs and, where the code warrants it, to formally check its invariants — Checkpoint 8.

    154 GitHub stars~1.8k tokensUpdated 14 days ago
    Auto-check passed
  • Next

    swingerman/engineer

    Use at the start of a work session, or any time the question is "what should I pick up now" across the whole project.

    154 GitHub stars~3k tokensUpdated 14 days ago
    Auto-check passed

Categories

Questions about Refinement Advisor

What does Refinement Advisor do?

A skill your agent uses to decide which refinement and hardening tools are worth running on a set of changes — refine, crap-analyzer, arch-check, introversion scan, mutation testing, TLA+ model…. Refinement Advisor is an agent skill from swingerman/engineer. Use to decide which refinement and hardening tools are worth running on a set of changes — refine, crap-analyzer, arch-check, introversion scan, mutation testing, TLA+ model checking, Lean proofs — and on which files, with the invariant to check.

When should I use Refinement Advisor?

Refinement Advisor fits situations like: decide which refinement and hardening tools are worth running on a set of changes — refine; introversion scan; mutation testing; TLA+ model checking.

How do I install Refinement Advisor in Claude Code?

Run `npx skills add swingerman/engineer --skill refinement-advisor -a claude-code`. Or copy the skill folder (engineer/skills/refinement-advisor in swingerman/engineer) into .claude/skills/refinement-advisor in your project. Claude Code loads it when a task matches its description.

How do I install Refinement Advisor in Codex?

Run `npx skills add swingerman/engineer --skill refinement-advisor -a codex`. Or copy the skill folder (engineer/skills/refinement-advisor in swingerman/engineer) into .agents/skills/refinement-advisor in your project. Codex loads it when a task matches its description.

Can I use Refinement Advisor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swingerman/engineer --skill refinement-advisor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/refinement-advisor, .gemini/skills/refinement-advisor, .github/skills/refinement-advisor and .opencode/skills/refinement-advisor in your project.

What does Refinement Advisor need to run?

SKILL.md names no scripts, command-line tools or credentials: Refinement Advisor is instructions for the agent only.

Does Refinement Advisor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Refinement Advisor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Refinement Advisor use?

Refinement Advisor is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Refinement Advisor use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Refinement Advisor?

Skills that share tags, products or a category with Refinement Advisor: Requirements (rizsotto/Bear, 6.5k stars), Crap Analysis (ardalis/RiverBooks, 134 stars), Code Coverage (s3s-project/s3s, 311 stars) and Project Status (bactopia/bactopia, 522 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Refinement Advisor?

swingerman (a GitHub user) maintains it in swingerman/engineer, which has 154 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on September 23, 2026.

Source: swingerman/engineer on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.