Score a project's agent harness across 5 subsystems (Instructions / State / Verification / Scope / Lifecycle), identify the bottleneck, and produce a prioritized improvement plan.

MITAuto-check passedAgent Workflows

Install Harness Audit

skills CLI
$ npx skills add AnastasiyaW/codex-claude-code-config --skill harness-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AnastasiyaW/codex-claude-code-config harness-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AnastasiyaW/codex-claude-code-config.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/operational/harness-audit .claude/skills/harness-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-audit
GitHub stars
154
Token cost
~2.5k tokens
SKILL.md length
1,106 words
Files
3 (incl. references)
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Score a project's agent harness across 5 subsystems (Instructions / State / Verification / Scope / Lifecycle), identify the bottleneck, and produce a prioritized improvement plan.

  • Works in 4 steps: Gather → Score → Identify Bottleneck → …
  • Assessing if a project is ready to graduate to [LONG-RUN] status
  • SKILL.md covers What This Skill Does, The Five Subsystems (Our…, How to Run an Audit and Output Format, plus 5 more sections
  • Calls pytest

What it does

Harness Audit is an agent skill from AnastasiyaW/codex-claude-code-config. Score a project's agent harness across 5 subsystems (Instructions / State / Verification / Scope / Lifecycle), identify the bottleneck, and produce a prioritized improvement plan. Use when assessing if a project is ready to graduate to [LONG-RUN] status, when an agent keeps failing despite good models, or when adopting our stack on a new codebase. Do NOT use to design or build a new harness from scratch — this only scores an existing one; for greenfield harness/agent architecture use harness-design (or…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/checklist-per-subsystem.md` and `references/scoring-rubric.md`).

It sits in Agent Workflows. The repository describes itself as: Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development. The licence is MIT.

When your agent uses it

  • Assessing if a project is ready to graduate to [LONG-RUN] status
  • An agent keeps failing despite good models
  • Adopting our stack on a new codebase
  • Build a new harness from scratch — this only scores an existing one

Example prompts

  • “/harness-audit”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Gather
  2. Score
  3. Identify Bottleneck
  4. Prioritized Improvement Plan

What it can do on your machine

Read from SKILL.md and the folder at commit 67709af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • walkinglabs.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harness Audit loads about 2.5k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 136 tokens; SKILL.md has 1,106 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~136
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AnastasiyaW/codex-claude-code-config at commit 67709af, republished under its MIT licence (© AnastasiyaW). 1,106 words, ~2,550 tokens.

Download SKILL.mdSave it as .claude/skills/harness-audit/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
harness-audit
description
Score a project's agent harness across 5 subsystems (Instructions / State / Verification / Scope / Lifecycle), identify the bottleneck, and produce a prioritized improvement plan. Use when assessing if a project is ready to graduate to [LONG-RUN] status, when an agent keeps failing despite good models, or when adopting our stack on a new codebase. Do NOT use to design or build a new harness from scratch — this only scores an existing one; for greenfield harness/agent architecture use harness-design (or agent-harness-design).
license
MIT

Harness Audit

Trigger on phrases like: "audit my harness", "evaluate my agent setup", "score my CLAUDE.md", "is my project ready for long-run", "5-subsystem assessment", "what's missing from my project setup", "/harness-audit". Run proactively when joining an unfamiliar codebase that has agent artifacts (CLAUDE.md, .claude/, AGENTS.md) but obvious gaps. Skip for single-file scripts and pure exploration.

Score a project's agent harness across five subsystems and tell the user which evidenced bottleneck to address first. Distinguish an artifact's presence from demonstrated behavior; never present metadata alone as runtime proof.

Source: Five-subsystem framework adapted from Learn Harness Engineering (walkinglabs, MIT). Adapted to our concrete stack: CLAUDE.md, .claude/rules/, PROBLEMS.md, feature_list.json, init.sh, hooks, handoffs, chronicles.

What This Skill Does

Given a project directory, produces a scorecard like this. This example assumes project-xyz delivers features across sessions and uses pull requests:

=== Harness Audit: project-xyz ===

Instructions  4/5  ✓ Agent entrypoint and modular rules are documented and used
                   ~ PR review guidance is not found in the inspected evidence
State         2/5  ✓ 3 handoffs preserve some continuation state
                   ✗ No current record locates active scope and deferred work
                   ~ PROBLEMS.md / feature_list.json are suitable conventions, not prerequisites
Verification  3/5  ✓ Documented test command; pytest is configured
                   ~ No current execution receipt supplied
                   ✗ No documented staged validation/proof route
Scope         3/5  ✓ in-scope principle in CLAUDE.md
                   ~ shared-resource policy is unknown; no serialized lane is declared
                   ✗ Definition of Done not explicit
Lifecycle     2/5  ✗ Needed session-boundary entry/stop behavior is not documented or demonstrated
                   ~ Manual cleanup convention exists but not enforced

Bottleneck: State (2/5) — no current locator for feature scope and deferred work

Priority improvement (only when the user asks for recommendations):
- Record an execution receipt for the existing test command   ↗ Verification evidence

For an audit-only request, this skill produces the scorecard without making changes or running probes merely to convert unknown into a pass. If the user also requested correction or implementation, the scorecard is an intermediate result: return the confirmed findings to the owning task and execute its necessary, authorized reversible fixes, verification and delivery. Do not stop at a report or assign agent-owned fixes back to the user. Preserve the original acceptance criteria and real external/irreversible boundaries.


The Five Subsystems (Our Adaptation)

SubsystemConcrete files/conventions in our stack
InstructionsCLAUDE.md (root + ~/.claude/), .claude/rules/*.md (project), ~/.claude/rules/*.md (global), optional REVIEW.md
StatePROBLEMS.md, feature_list.json, .claude/handoffs/, .claude/chronicles/
Verificationa documented command plus current receipt appropriate to the target, tests/config where applicable, Proof Loop usage
Scopeexplicit in-scope/Definition of Done policy and a concurrency policy appropriate to the work
LifecycleSessionStart hooks, Stop hooks (stop-test-gate, check-problems-md), cleanup convention

See references/checklist-per-subsystem.md for per-subsystem concrete checks. See references/scoring-rubric.md for how to interpret 1-5 scores.


How to Run an Audit

Phase 1 — Gather

Read these files in order (skip silently if missing):

  1. CLAUDE.md in project root
  2. AGENTS.md in project root (some projects use this name)
  3. .claude/rules/*.md (project-level rules)
  4. .claude/settings.json and .claude/settings.local.json (hooks config)
  5. PROBLEMS.md in root
  6. feature_list.json in root
  7. init.sh in root (and Makefile / package.json scripts as fallback)
  8. .claude/handoffs/ (count files, check INDEX.md existence)
  9. .claude/chronicles/ (count files)
  10. Sample test config: pytest.ini / package.json test script / Cargo.toml

Use Glob + Read for the harness, then inspect the smallest relevant evidence path: a current test/CI receipt, hook execution trace, or sampled state artifact. This is not a broad code review; absence of behavioral evidence is ~ unknown, not ✓ working.

Phase 2 — Score

For each subsystem, first identify the project delivery model and applicable outcomes in references/checklist-per-subsystem.md. Mark every finding as documented, demonstrated, or unknown; score from evidence rather than file presence alone. The canonical files in the table are useful conventions, not universal prerequisites.

  • 5 = applicable outcomes are documented, demonstrated, and consistently followed
  • 4 = documented and demonstrated with bounded gaps
  • 3 = only documented, only demonstrated, materially partial, or unknown at a relevant boundary
  • 2 = isolated or stale structure does not meet most applicable outcomes
  • 1 = missing or actively harmful

For each subsystem, list:

  • ✓ what's present and working
  • ✗ what's missing or broken
  • ~ partial / unclear
Phase 3 — Identify Bottleneck

The lowest-scoring subsystem is the bottleneck. Even if other subsystems are weaker by absolute count of checks, the lowest score is the one to fix first because it limits the value of the rest.

Tie-breaker (multiple subsystems at same low score): pick the one whose improvement unlocks progress in others. State often wins when the project lacks a durable locator for active scope and unresolved work; do not assume two particular filenames are required.

Phase 4 — Prioritized Improvement Plan

Only if the user requests recommendations, propose the smallest number of independently shippable actions that address the evidenced bottleneck. For each, name the expected evidence and a local template/example if one actually fits. Do not invent effort, score gains, or a fixed number of steps. Do not expand an audit-only request into implementation; when implementation was already requested, continue that owning task after the audit instead of requesting the same authorization again.


Show full SKILL.md (436 more words)Show less

Output Format

Use the visual scorecard format shown at the top of this skill. Sections:

  1. Header: === Harness Audit: <project-name> === (one line)
  2. Scorecard: 5 lines, one per subsystem, with score + ✓/✗ findings
  3. Bottleneck: one line naming the subsystem and score
  4. Priority improvements: only when requested, with expected evidence + pointer
  5. Projected total: optional, only if user asks and the stated evidence supports a bounded projection

Keep the entire output under 50 lines. The user is scanning for next steps, not reading an essay. Detail goes into the per-subsystem checklist file, not the audit output.


What This Skill Is NOT

  • Not a code review — does not look at source code quality
  • Not a security audit — does not check for vulnerabilities (use /security-review instead)
  • Not a broad test runner — does not manufacture a green result from configuration. It may inspect a current CI/test receipt, or run one user-authorized, task-relevant probe when runtime evidence is part of the requested audit
  • Not a standalone implementation methodology — an audit-only request ends with its findings; an audit inside an already authorized repair returns control to that repair, not to user homework
  • Not for short-lived work without a durable handoff need — state why the audit is disproportionate instead of applying arbitrary feature/session thresholds

Honest Tradeoffs

  • The 5-subsystem framework is opinionated. A project can be perfectly functional with 3 of 5 strong and 2 weak (e.g., a research repo with no lifecycle needs).
  • Scoring is subjective at the margins. A 3 vs 4 for "covers basics" is a judgment call. Use the checklist to keep it consistent across audits, not to claim numeric precision.
  • The skill assumes our stack conventions. For projects using completely different tooling (e.g., AGENTS.md without .claude/), translate concepts before scoring — don't fail the project on naming.

  • Principle 27 (feature-tracking) — full framework explanation
  • Principle 01 (harness-design) — Generator-Evaluator pattern, source of "subsystems" thinking
  • Templates templates/long-run-project/ — drop-in files for fixing State + Verification gaps
  • Rule rules/long-run-harness.md — convention this audit checks against

Gotchas

  • A scorecard example must use the same evidence rules as the checklist; a missing filename alone does not prove a missing outcome.
  • Presence of this skill or its rubric is not proof that auditors apply it consistently. Do not cite an example-evaluation file unless it actually exists and was inspected.

Troubleshooting

  • Two auditors give different scores to the same evidence: compare the declared delivery model and applicable outcomes, then resolve the checklist/example contradiction; do not add files just to raise the score.
  • An authorized repair ends at the scorecard: return each confirmed finding to the original task, implement the necessary correction and verify it. An audit-only request remains read-only.

© AnastasiyaW, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/operational/harness-audit of AnastasiyaW/codex-claude-code-config.

  • SKILL.md
  • references/checklist-per-subsystem.md
  • references/scoring-rubric.md

Open the folder on GitHubat commit 67709af

Compare with similar skills

Harness Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harness Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harness Audit this skillAnastasiyaW/codex-claude-code-config154—~2.5kAutomated safety check: PassMIT
Orca CLIstablyai/orca89k2 repos~593Automated safety check: PassMIT
OpenSpec Guided OnboardingFission-AI/OpenSpec72k1 repos~4.5kAutomated safety check: PassMIT
Brainstormingobra/superpowers297k1 repos~2.5kAutomated safety check: PassMIT
Claude Code Plugin Structureanthropics/claude-plugins-official38k10 repos~3.4kAutomated safety check: PassApache-2.0
Neat-Freak Knowledge CloseoutKKKKhazix/khazix-skills21k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Orca CLI

    stablyai/orca

    Operate Orca-managed worktrees, folder contexts, terminals, repos, automations, artifacts, skill sharing, worktree comments, and Orca's embedded browser…

    89k GitHub starsUsed in 2 repos~593 tokens
    Agent WorkflowsAuto-check passed
  • OpenSpec Guided Onboarding

    Fission-AI/OpenSpec

    Walks you through a complete OpenSpec workflow cycle with narration while doing real work in your codebase.

    72k GitHub starsUsed in 1 repo~4.5k tokens
    Agent WorkflowsAuto-check passed
  • Brainstorming

    obra/superpowers

    Makes the agent clarify intent and agree on a design with you before writing any code, scaling the process from a quick spike to a written spec.

    297k GitHub starsUsed in 1 repo~2.5k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Plugin Structure

    anthropics/claude-plugins-official

    Official

    Explains the directory layout, plugin.json manifest and component organization of a Claude Code plugin, including auto-discovery and portable paths.

    38k GitHub starsUsed in 10 repos~3.4k tokens
    Agent WorkflowsAuto-check passed
  • Neat-Freak Knowledge Closeout

    KKKKhazix/khazix-skills

    Brings project docs, agent rule files, authorized memory and leftover workspace files back in line with what the code and runtime actually do at the end of a work session.

    21k GitHub stars~1.9k tokensUpdated 10 days ago
    Agent WorkflowsAuto-check passed
  • O2 Review Loop

    openobserve/openobserve

    Splits a change into planner, coder and independent reviewer roles: you confirm a spec, a subagent implements it, and a separate reviewer checks each round's local WIP commit.

    22k GitHub stars~3.7k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from AnastasiyaW/codex-claude-code-config

All 50 skills in this repo
  • Bug Reproducer

    AnastasiyaW/codex-claude-code-config

    Find likely software bugs in a codebase, rank concrete bug candidates, and prove or reject them with focused regression tests before proposing a fix.

    154 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Motion Framer

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when implementing Motion or Framer Motion in React/JavaScript: interactive UI components, micro-interactions, gestures, layout or page transitions, and scroll-based animation.

    154 GitHub starsUsed in 1 repo~5.2k tokens
    Auto-check passed
  • Proof Verify

    AnastasiyaW/codex-claude-code-config

    Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

    154 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Workflow Orchestration

    AnastasiyaW/codex-claude-code-config

    Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

    154 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Notebooklm Grounded Research

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when: NotebookLM, notebooklm MCP, large documentation sets, courses, books, papers, or citation-backed research are mentioned.

    154 GitHub stars~2.4k tokensUpdated today
    Auto-check: warnings
  • Deepseek Provider Contract

    AnastasiyaW/codex-claude-code-config

    Validate a proposed DeepSeek API integration before any key or project context is sent: check thinking-mode tool-call history, strict-schema assumptions, bounded output, and provider data boundaries.

    154 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Questions about Harness Audit

What does Harness Audit do?

Score a project's agent harness across 5 subsystems (Instructions / State / Verification / Scope / Lifecycle), identify the bottleneck, and produce a prioritized improvement plan. Harness Audit is an agent skill from AnastasiyaW/codex-claude-code-config. Score a project's agent harness across 5 subsystems (Instructions / State / Verification / Scope / Lifecycle), identify the bottleneck, and produce a prioritized improvement plan.

When should I use Harness Audit?

Harness Audit fits situations like: assessing if a project is ready to graduate to [LONG-RUN] status; an agent keeps failing despite good models; adopting our stack on a new codebase; build a new harness from scratch — this only scores an existing one.

How do I install Harness Audit in Claude Code?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill harness-audit -a claude-code`. Or copy the skill folder (skills/operational/harness-audit in AnastasiyaW/codex-claude-code-config) into .claude/skills/harness-audit in your project. Claude Code loads it when a task matches its description.

How do I install Harness Audit in Codex?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill harness-audit -a codex`. Or copy the skill folder (skills/operational/harness-audit in AnastasiyaW/codex-claude-code-config) into .agents/skills/harness-audit in your project. Codex loads it when a task matches its description.

Can I use Harness Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AnastasiyaW/codex-claude-code-config --skill harness-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-audit, .gemini/skills/harness-audit, .github/skills/harness-audit and .opencode/skills/harness-audit in your project.

What does Harness Audit need to run?

Going by SKILL.md and its folder, Harness Audit needs the command-line tools its instructions call (pytest).

Does Harness Audit access the network?

SKILL.md names 1 domain. As links in the text: walkinglabs.github.io. This is read from the text; nothing was executed.

Is Harness Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Harness Audit use?

Harness Audit is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harness Audit use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.

What are the alternatives to Harness Audit?

Skills that share tags, products or a category with Harness Audit: Orca CLI (stablyai/orca, 89k stars), OpenSpec Guided Onboarding (Fission-AI/OpenSpec, 72k stars), Brainstorming (obra/superpowers, 297k stars) and Claude Code Plugin Structure (anthropics/claude-plugins-official, 38k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harness Audit?

AnastasiyaW (a GitHub user) maintains it in AnastasiyaW/codex-claude-code-config, which has 154 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: AnastasiyaW/codex-claude-code-config on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.