Agent skill

Harness Stripping

by Archive228 in Archive228/loopkit

Systematically remove one harness component at a time and measure impact, killing scaffolding that no longer earns its complexity.

MITAuto-check passedDevelopment

Install Harness Stripping

skills CLI
$ npx skills add Archive228/loopkit --skill harness-stripping -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Archive228/loopkit harness-stripping --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/harness-stripping .claude/skills/harness-stripping && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-stripping
GitHub stars
755
Token cost
~1.2k tokens
SKILL.md length
663 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Systematically remove one harness component at a time and measure impact, killing scaffolding that no longer earns its complexity.

  • Works in 7 steps: Inventory the components. List every… → Rank by suspicion. Put the components… → Pick a baseline eval. You need a… → …
  • Tasks that involve Project scaffolding
  • SKILL.md covers When to apply, Procedure, Anti-patterns and What to strip first, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Harness Stripping is an agent skill from Archive228/loopkit. Systematically remove one harness component at a time and measure impact, killing scaffolding that no longer earns its complexity.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Project scaffolding. The repository describes itself as: 33 battle-tested skills + minimal .claude harness for any coding agent (Claude Code, Cursor, Codex, Gemini CLI). The licence is MIT.

When your agent uses it

  • Tasks that involve Project scaffolding

Example prompts

  • “/harness-stripping”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Inventory the components. List every distinct piece of scaffolding: prompt sections, tool wrappers, post-hoc validators, retry loops…
  2. Rank by suspicion. Put the components most likely to be obsolete at the top: anything added before the last two model bumps, anything…
  3. Pick a baseline eval. You need a repeatable metric before you touch anything. Reuse an existing eval set if you have one; otherwise pick…
  4. Strip one component. Only one. Comment it out or gate it behind a flag — don't delete yet. Re-run the eval.
  5. Compare against baseline.
  6. Commit the delta. Land the strip (or the restore-with-notes) as its own commit. Do not batch multiple strips into one change — you lose…
  7. Repeat for the next component. Re-establish baseline from the new state each round, not the original. Compounding strips have compounding…

What it can do on your machine

Read from SKILL.md and the folder at commit 5ae033e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harness Stripping loads about 1.2k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 663 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Archive228/loopkit at commit 5ae033e, republished under its MIT licence (© Archive228). 663 words, ~1,213 tokens.

Download SKILL.mdSave it as .claude/skills/harness-stripping/SKILL.md (or your agent's skills folder).
name
harness-stripping
description
Systematically remove one harness component at a time and measure impact, killing scaffolding that no longer earns its complexity.
when_to_use
auditing a harness after a model upgrade to see what workarounds are now obsolete, a component encodes an assumption about model weakness worth re-testing, or…

Harness Stripping

Every harness component was added to compensate for a specific model failure. Models improve. Components don't retire themselves. The scaffolding that saved you on Sonnet 4.5 may be dead weight — or actively harmful — on Opus 4.6. Strip it deliberately, one piece at a time, and let evals tell you what still earns its keep.

Inspired by Prithvi's March 2026 harness post on evaluator-generator separation and the general "re-test your assumptions each model bump" discipline.

When to apply

  • A model upgrade just landed and your harness was tuned for the previous generation.
  • A component's justification is "we added this because the model used to do X" — and you haven't checked whether it still does X.
  • The harness has accreted over months and nobody remembers what half the machinery is for.
  • Cost or latency is climbing and you suspect redundant belt-and-suspenders layers.

Procedure

  1. Inventory the components. List every distinct piece of scaffolding: prompt sections, tool wrappers, post-hoc validators, retry loops, evaluator personas, structured-output enforcers, sandbox rules. One row per component. Note the failure mode each was added to prevent.

  2. Rank by suspicion. Put the components most likely to be obsolete at the top: anything added before the last two model bumps, anything targeting a failure mode you haven't seen recently, anything whose original justification is now folklore.

  3. Pick a baseline eval. You need a repeatable metric before you touch anything. Reuse an existing eval set if you have one; otherwise pick 20–60 tasks representative of production work. Record baseline score, cost, and wall-clock.

  4. Strip one component. Only one. Comment it out or gate it behind a flag — don't delete yet. Re-run the eval.

  5. Compare against baseline.

    • Score within noise, cost/latency down → the component is dead weight. Delete.
    • Score drops measurably → the component still earns its complexity. Restore and note what failure mode returned.
    • Score improves → the component was actively harmful. Delete and investigate why (often: over-constraining a now-capable model).
  6. Commit the delta. Land the strip (or the restore-with-notes) as its own commit. Do not batch multiple strips into one change — you lose the ability to attribute the score movement.

  7. Repeat for the next component. Re-establish baseline from the new state each round, not the original. Compounding strips have compounding effects.

Show full SKILL.md (286 more words)Show less

Anti-patterns

  • Stripping two components at once — you can't tell which one mattered. Halve the signal, double the confusion.
  • Skipping the eval "because it's obviously safe to remove" — the harness accreted for reasons. Some are still real. Measure.
  • Deleting instead of gating on the first pass — you will want to A/B mid-review. Flag first, delete after the eval confirms.
  • Trusting anecdotes over the eval — "it feels better without it" is how load-bearing components get removed. If the eval doesn't show it, it isn't there.
  • Stripping components that guard safety, sandboxing, or cost caps — those aren't compensating for model weakness. Leave them.
  • Doing this on prod traffic — run against an eval set, not real users. The failure modes you're re-probing are exactly the ones that hurt users.

What to strip first

Highest yield in practice:

  • Structured-output enforcers layered on top of models that now emit valid JSON natively.
  • Multi-step "plan then execute" wrappers on tasks the model now one-shots.
  • Retry loops around tool calls that no longer flake.
  • Evaluator personas whose critiques the generator now anticipates on its own.
  • Verbose "remember to do X" prompt sections where X is now default behavior.

When NOT to apply

Don't strip mid-project on a live long-running run — you'll perturb sessions in flight. Do it between projects, or on a forked branch. Also skip if you don't have an eval you trust; stripping without measurement is guessing.

  • [[shift-notes]] — record which components were stripped and when, so the next audit doesn't re-strip and re-restore the same piece.
  • [[adversarial-verify]] — the evaluator-generator pattern that may itself be a strip candidate on newer models.
  • [[broken-window-check]] — if you strip a component and the eval regresses in a specific way, that's your new broken window to hunt.

© Archive228, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/harness-stripping of Archive228/loopkit.

Open the folder on GitHubat commit 5ae033e

Compare with similar skills

Harness Stripping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harness Stripping compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harness Stripping this skillArchive228/loopkit755—~1.2kAutomated safety check: PassMIT
Nx Generatenomcopter/react-mosaic4.8k7 repos~1.9kAutomated safety check: PassCustom licence
PonytailDavidObando/gsharp5657 repos~1.7kAutomated safety check: PassMIT
Run Nx Generatornrwl/nx29k2 repos~592Automated safety check: NotesMIT
Conductor Setupgemini-cli-extensions/conductor3.8k—~4.2kAutomated safety check: PassApache-2.0
Mirage VFS Adapter Authoringstrukto-ai/mirage3.7k—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Nx Generate

    nomcopter/react-mosaic

    Generate code using nx generators. An agent skill from nomcopter/react-mosaic.

    4.8k GitHub starsUsed in 7 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Ponytail

    DavidObando/gsharp

    Forces the laziest solution that actually works, simplest, shortest, most minimal.

    565 GitHub starsUsed in 7 repos~1.7k tokens
    DevelopmentAuto-check passed
  • Run Nx generators with prioritization for workspace-plugin generators.

    29k GitHub starsUsed in 2 repos~592 tokens
    DevelopmentAuto-check: notes
  • Conductor Setup

    gemini-cli-extensions/conductor

    Scaffolds the project and sets up the Conductor environment.

    3.8k GitHub stars~4.2k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Builds or extends a custom Mirage virtual filesystem adapter for an API, database, object store or app data, with a working mount configuration and filesystem tests.

    3.7k GitHub stars~2.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Enforces this repository's TypeScript backend module architecture under server/: feature folders, barrel exports, and where shared types and utilities belong.

    14k GitHub stars~1.2k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from Archive228/loopkit

All 43 skills in this repo
  • Hitl Escalate

    Archive228/loopkit

    Escalate blocked runs to a human via configured channel or fallback to BLOCKED.md and exit the loop.

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Structured Output

    Archive228/loopkit

    Get JSON out of the model reliably. An agent skill from Archive228/loopkit.

    755 GitHub stars~830 tokensUpdated 2 mo ago
    Auto-check passed
  • Using Loopkit

    Archive228/loopkit

    A skill your agent uses when starting any conversation in a loopkit-enabled project - establishes how to find and use loopkit's 49 skills, requiring skill invocation before ANY response including…

    755 GitHub stars~1.4k tokensUpdated 2 mo ago
    Auto-check passed
  • Active Memory Reminder

    Archive228/loopkit

    Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable).

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Eval Harness

    Archive228/loopkit

    Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.

    755 GitHub stars~876 tokensUpdated 2 mo ago
    Auto-check passed
  • Feature List JSON

    Archive228/loopkit

    Enumerate every end-to-end feature as strict JSON entries with passes:false, editable-passes-only discipline, and priority order.

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about Harness Stripping

What does Harness Stripping do?

Systematically remove one harness component at a time and measure impact, killing scaffolding that no longer earns its complexity. Harness Stripping is an agent skill from Archive228/loopkit. Systematically remove one harness component at a time and measure impact, killing scaffolding that no longer earns its complexity.

When should I use Harness Stripping?

Harness Stripping fits situations like: tasks that involve Project scaffolding.

How do I install Harness Stripping in Claude Code?

Run `npx skills add Archive228/loopkit --skill harness-stripping -a claude-code`. Or copy the skill folder (skills/harness-stripping in Archive228/loopkit) into .claude/skills/harness-stripping in your project. Claude Code loads it when a task matches its description.

How do I install Harness Stripping in Codex?

Run `npx skills add Archive228/loopkit --skill harness-stripping -a codex`. Or copy the skill folder (skills/harness-stripping in Archive228/loopkit) into .agents/skills/harness-stripping in your project. Codex loads it when a task matches its description.

Can I use Harness Stripping in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Archive228/loopkit --skill harness-stripping -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-stripping, .gemini/skills/harness-stripping, .github/skills/harness-stripping and .opencode/skills/harness-stripping in your project.

What does Harness Stripping need to run?

SKILL.md names no scripts, command-line tools or credentials: Harness Stripping is instructions for the agent only.

Does Harness Stripping access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Harness Stripping safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Harness Stripping use?

Harness Stripping is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harness Stripping use?

About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Harness Stripping?

Skills that share tags, products or a category with Harness Stripping: Nx Generate (nomcopter/react-mosaic, 4.8k stars), Ponytail (DavidObando/gsharp, 565 stars), Run Nx Generator (nrwl/nx, 29k stars) and Conductor Setup (gemini-cli-extensions/conductor, 3.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harness Stripping?

Archive228 (a GitHub user) maintains it in Archive228/loopkit, which has 755 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on July 14, 2026.

Source: Archive228/loopkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.