Agent skill

Harness Eval

by tech-leads-club in tech-leads-club/agent-skills

Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports.

CC-BY-4.0Auto-check passedAgent Workflows

Install Harness Eval

skills CLI
$ npx skills add tech-leads-club/agent-skills --skill harness-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tech-leads-club/agent-skills harness-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tech-leads-club/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'packages/skills-catalog/skills/(development)/harness-eval' .claude/skills/harness-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-eval
GitHub stars
7k
Token cost
~3.9k tokens
SKILL.md length
1,578 words
Files
12 (incl. scripts, references)
Skills in repo
74
Repo updated
First seen
Licence
CC-BY-4.0

At a glance

Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports.

  • Works in 11 steps: Resolve SKILL_DIR → Inventory + claim deck → Track A (deterministic) — always run → …
  • Auditing AGENTS.md and skill files for broken paths and commands
  • SKILL.md covers User questionnaires (HIGH…, Loading this skill's files, Critical rules and Instructions, plus 2 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Harness Eval audits the files that steer an agent in a repository: AGENTS.md, rules, skills and the references they point to. It looks for broken paths and commands, redundant instructions and how useful each piece is, using a stack-agnostic protocol in which two judges review claims and planted traps check the judges. It always ends at reports and does not edit AGENTS.md or skills unless you ask after reviewing the Ship and Slim recommendations.

The run is gated by questions. After an inventory step it asks whether optional project docs outside the skill folders should be in scope, with ADRs and RFCs excluded from one tier, then asks which tracks to run before any spending. Track A, a deterministic correctness check that costs close to no model tokens, always runs; tracks B and C use model judges and run only if you approve. Verdicts use the labels Ship, Review, Hold, Slim and Keep-core. Python scripts handle inventory, surface extraction, correctness and merging of judge results, and PROTOCOL.md must be read in full first.

When your agent uses it

  • Auditing AGENTS.md and skill files for broken paths and commands
  • Finding redundant or duplicated instructions across an agent harness
  • Judging which skills and rules are actually useful before trimming
  • Producing a Ship, Review, Hold, Slim or Keep-core report for a harness

Example prompts

  • “Run harness eval on this repo and stop at the reports.”
  • “Audit our AGENTS.md and .claude/skills for redundant instructions, with Track A only for now.”
  • “Do a full harness eval with the B and C judges, and show me the cost table first.”

Requirements

  • Python to run the bundled scripts

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. Resolve SKILL_DIR
  2. Inventory + claim deck
  3. Track A (deterministic) — always run
  4. Track B — Judge1
  5. Track B — Judge2 (blind)
  6. Merge Track B agreement
  7. Track C — surface deck
  8. Track C — Usefulness Judge1
  9. Track C — Usefulness Judge2 (blind)
  10. Merge Track C agreement
  11. Present results

What it can do on your machine

Read from SKILL.md and the folder at commit 6df68d5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harness Eval loads about 3.9k tokens when it runs, and up to ~9.4k if it reads all its reference files. Until then it costs about 212 tokens; SKILL.md has 1,578 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~212
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tech-leads-club/agent-skills at commit 6df68d5, republished under its CC-BY-4.0 licence (© tech-leads-club). 1,578 words, ~3,865 tokens.

Download SKILL.mdSave it as .claude/skills/harness-eval/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
harness-eval
description
Evaluate a repo agent harness (AGENTS.md, rules, skills, skill refs) for broken paths/commands, redundant instructions, and usefulness using a stack-agnostic dual-judge protocol with planted traps. HIGH PRIORITY questionnaires at top: Q1 optional docs, Q2 B/C budget before Track A (certainty/tokens). A always runs after Q2; B/C opt-in. ADRs/RFCs excluded from T2. Mixed apply uses 11-mixed-apply.md (KEEP/CUT). Use when the user says harness eval, harness-eval, harness debug, audit AGENTS.md, audit skills/rules, instruction audit, redundancy of agent instructions, usefulness of skills, Ship/Review/Hold/Slim/Keep-core for harness, or wants Track A/B/C harness evaluation. Do NOT use for harness setup or init, feature spec-driven work (tlc-spec-driven), or applying Ship/Slim trims unless the user explicitly asks after the report.
license
CC-BY-4.0
metadata.author
Tech Leads Club - github.com/tech-leads-club
metadata.version
1.8.3

Harness Eval

Run a full, stack-agnostic harness evaluation and stop at reports. Do not auto-edit AGENTS.md or skills unless the user explicitly asks after reviewing Ship/Slim.

User questionnaires (HIGH PRIORITY)

Stop and ask before continuing. Do not skip these gates. Do not silently include optional docs or spawn B/C judges.

Order after inventory: Q1 (if needed) → Q2 → then Track A (A always runs) → B/C only if approved.

Q1 — Optional project docs (after inventory)

When optional-docs-candidates.md lists optional types, ask before Q2 / Track A:

markdown
Inventory found cited project docs outside the agent skill trees.

- **Always in scope:** skill-tree files (`.agents/skills`, `.cursor/skills`, `.claude/skills`)
- **Always excluded:** ADRs / RFCs / decision-record trees (never scored as T2)
- **Optional (default: omit):** see types/paths in `optional-docs-candidates.md`

Include any optional doc types or paths in this run?
Reply with: `none` (default), type ids (e.g. `docs`), and/or specific paths.

Re-run inventory with --include-doc-type / --include-doc only after the user answers. If no optional types, skip Q1.

Q2 — Tracks B and C (before Track A — budget)

Ask before Track A so the user sets spend up front. Track A always runs next (deterministic, ~0 model tokens). B/C run only if approved.

markdown
Choose eval scope for this run (before Track A).

| Track | Question | Certainty | Token consumption |
|-------|----------|-----------|-------------------|
| **A — Correctness** | Cited path/command exists? | **Highest** — script only, no LLM. Prefers false negatives over false BROKEN. | **~0 model tokens** (always runs next) |
| **B — Redundancy** | Would an agent rediscover this cheaply without the harness? | **Medium** — dual LLM + plants; Ship only if trap PASS and both agree. Disagree → Hold. Less model-sensitive than C. | **High** — 2 judges × every claim (~N in this inventory). Each may spot-check the repo. |
| **C — Usefulness** | Does this surface change behavior vs theory/demo/overlap? | **Lowest / most subjective** — dual LLM + plants + fan-in; **model-sensitive**. Slim/Mixed need gates; prefer second-model check before large deletes. | **Highest** — 2 judges × every surface (whole files; often dominates the run). |

Notes: Ship (B) ≠ Slim (C). Rediscoverable ≠ useless. A always runs; B/C are optional.

Reply with one of: `A only`, `B`, `C`, or `B+C`.

Fill claim count from claims.md when known; surface count ≈ T0+T1+T2 markdown after extract (or say “after surfaces_extract” if not run yet).

  • A only: run Track A; present 04; stop (no B/C judges).
  • B: Track A, then Steps 4–6.
  • C: Track A, then Steps 7–10 (C does not need B).
  • B+C: Track A, then Steps 4–11.

If the user already requested B/C/full eval in the triggering message, treat as approval — still show the Q2 table once so costs are visible.

Loading this skill's files

This skill is self-contained. Protocol, scripts, and judge prompts live under this skill directory (the folder that contains this SKILL.md). Resolve SKILL_DIR as that directory — never assume another install path.

Run outputs (not protocol) go to the target repo at .harness-eval/runs/<run-id>/.

Critical rules

  1. Report-only by default. Judgment ≠ remediation.
  2. README out of scope as harness surface and as rediscovery/usefulness evidence.
  3. Stack-agnostic. Never hard-code package managers, DBs, frameworks, or folder layouts in prompts or plants. Discover manifests that exist (JS, Python, Make/Task, Rust, Go, PHP, Ruby/Rails, Java/Gradle/Maven, plus bin/*).
  4. Doc scope. T2 always includes agent skill-tree refs (.agents/skills, .cursor/skills, .claude/skills). ADRs / RFCs (decision-record trees) are always excluded from T2 surfaces. Other cited project docs are optional — default omit; ask via Q1 at the top of this skill, then re-run with --include-doc-type / --include-doc.
  5. Track A always runs after inventory (deterministic, high-precision). Prefer false negatives over false BROKEN. Placeholders (SPEC_FOLDER, {x}, [feature]) are never BROKEN. Never normalize paths with str.lstrip('./').
  6. Tracks B and C require user approval via Q2 before Track A. Do not spawn B/C judges until the user opts in. User may approve B only, C only, both, or A only.
  7. Track B needs dual judges + plants. Judge2 is blind (must not read Judge1 scores or trap-key.json). Ship only if trap gate PASS and dual REDUNDANT with Judge2 cost ≤ 1.
  8. Track C needs dual judges + plants. Blind Judge2 must not read 08-usefulness-j1.md or usefulness-trap-key.json. Slim only if trap PASS, dual SLIM/ROUTING-ONLY, and fan-in PASS (no other harness surface hard-loads the path as SoT — merge enforces this on the full skill tree, not just --seed). Usefulness is model-sensitive — record model: <id> in both score files; prefer same model within a run; re-judge on a second model before large Slim deletes.
  9. KEEP / KEEP-CORE plants must not be verbatim copies of claims/surfaces already in the deck.
  10. Subagents: use an allowlisted non-fast model (prefer the same family as the parent when policy allows). Do not use *-fast models.
  11. Do not equate tracks. Track B Ship ≠ Track C Slim. Rediscoverable ≠ useless; useful ≠ non-redundant.
  12. Slim apply / fan-in. Never stub or delete a Slim path listed under “Slim fan-in blocked” (or when python3 "$SKILL_DIR/scripts/slim_fanin.py" --path <P> reports citers) unless those consumers are updated in the same change.
  13. Mixed/Slim apply stays self-contained. Cutting REPO-DEMONSTRATED / THEORY means delete or compress that bulk in the harness surface. Never replace a fenced teaching snippet (or the contract it carried) with See app/... / lib/... / test/... — that swaps SoT for a code-tree pointer. Judge evidence paths stay in score tables only; if the behavior-changing contract must survive, keep a short in-skill rule or snippet.
  14. Mixed apply is mechanical. Dual MIXED alone is not enough. Merge emits 11-mixed-apply.md with per-ID KEEP (from Keep-core columns) and CUT (from Slim columns). Apply agents must follow that file only — do not re-judge, redesign, or invent a different pattern than KEEP. Empty Keep-core/Slim cells → skip that path (Hold).

Instructions

Step 1: Resolve SKILL_DIR

Set SKILL_DIR to the directory containing this SKILL.md. Verify:

  • $SKILL_DIR/references/PROTOCOL.md
  • $SKILL_DIR/scripts/inventory_extract.py
  • $SKILL_DIR/scripts/track_a_correctness.py
  • $SKILL_DIR/scripts/merge_agreement.py
  • $SKILL_DIR/scripts/surfaces_extract.py
  • $SKILL_DIR/scripts/merge_usefulness.py
  • $SKILL_DIR/scripts/slim_fanin.py
  • $SKILL_DIR/scripts/doc_scope.py

If missing, the skill install is broken — stop.

Step 2: Inventory + claim deck

From the target repo root:

bash
RUN_ID=$(date -u +%Y-%m-%d)-full
python3 "$SKILL_DIR/scripts/inventory_extract.py" --root . --run-id "$RUN_ID"
# Optional scope: AGENTS.md + one-hop related skills only
# python3 "$SKILL_DIR/scripts/inventory_extract.py" --root . --run-id "$RUN_ID" --seed AGENTS.md

Expected under .harness-eval/runs/$RUN_ID/: inventory.json, claims.jsonl, claims.md, trap-key.json, optional-docs-candidates.md (+ .json).

Step 2b: Optional docs — Q1 (see top)

Read optional-docs-candidates.md. If optional types exist, run Q1 from User questionnaires. Re-run inventory only after approval:

bash
python3 "$SKILL_DIR/scripts/inventory_extract.py" --root . --run-id "$RUN_ID" \
  --include-doc-type docs   # and/or --include-doc path
Step 2c: Track budget — Q2 (see top)

Run Q2 from User questionnaires before Track A. Record the answer (A only / B / C / B+C). Do not start Steps 4+ unless B and/or C were approved.

Step 3: Track A (deterministic) — always run
bash
python3 "$SKILL_DIR/scripts/track_a_correctness.py" --root . --run-id "$RUN_ID"

Expected: 04-correctness.md (includes term definitions at top). Spot-check that .agents/... cites resolve (not agents/...).

Summarize Track A (broken count + notable clusters). If Q2 was A only, stop. Otherwise continue to the approved B and/or C steps.

Step 4: Track B — Judge1

Read references/judge-prompts.md (Track B Judge1). Spawn an independent subagent with an allowlisted model. Point it at .harness-eval/runs/$RUN_ID/claims.md. It writes 05-redundancy-j1.md (include model: <id>).

Judge1 may read inventory.json. Must not read trap-key.json.

Show full SKILL.md (627 more words)Show less
Step 5: Track B — Judge2 (blind)

Read references/judge-prompts.md (Track B Judge2). Spawn a second subagent. Writes 06-blind-scores.md.

Forbidden for Judge2: trap-key.json, 05-redundancy-j1.md, 07-agreement.md, prior agreement reports.

Prefer Steps 4 and 5 in parallel.

Step 6: Merge Track B agreement
bash
python3 "$SKILL_DIR/scripts/merge_agreement.py" --run-dir .harness-eval/runs/$RUN_ID

Expected: 07-agreement.md (Ship/Review/Hold + What these words mean). On trap FAIL: fix plants per PROTOCOL, rescore P00x, re-merge — do not Ship.

Step 7: Track C — surface deck
bash
python3 "$SKILL_DIR/scripts/surfaces_extract.py" --root . --run-id "$RUN_ID"

Expected: surfaces.md, surfaces.json, usefulness-trap-key.json.

Step 8: Track C — Usefulness Judge1

Read references/judge-prompts.md (Usefulness Judge1). Spawn subagent with allowlisted model (record same id in header). Writes 08-usefulness-j1.md.

Must not read usefulness-trap-key.json.

Step 9: Track C — Usefulness Judge2 (blind)

Read Usefulness Judge2 prompt. Prefer same model as Step 8 for agreement stability. Writes 09-usefulness-j2.md.

Forbidden: usefulness-trap-key.json, 08-usefulness-j1.md, 10-usefulness-agreement.md, and using Track B 05/06/07 to decide usefulness classes.

Prefer Steps 8 and 9 in parallel.

Step 10: Merge Track C agreement
bash
python3 "$SKILL_DIR/scripts/merge_usefulness.py" --run-dir .harness-eval/runs/$RUN_ID

Expected: 10-usefulness-agreement.md (Slim/Keep-core/Mixed/Hold + What these words mean), 11-mixed-apply.md (KEEP/CUT per Mixed ID), plus slim-fanin.json. On trap FAIL: do not Slim. Surfaces with slim-fanin-blocked are Hold — not Slim apply candidates.

Step 11: Present results

Summarize from the agreement reports (each starts with term definitions):

  • Track A broken count → 04-correctness.md
  • Track B trap + Ship/Review/Hold → 07-agreement.md
  • Track C trap + fan-in + Slim/Keep-core/Mixed/Hold → 10-usefulness-agreement.md
  • Call out 11-mixed-apply.md when Mixed count > 0 (the only Mixed apply path)
  • Call out model ids used for Track C and that Slim is model-sensitive
  • Call out any Slim fan-in blocked rows (consumers outside seed may appear here)

Stop unless the user asks to apply Ship/Slim/Mixed. When applying:

  • Slim: only paths in the Slim table (fan-in PASS); never stub fan-in-blocked paths without updating citers first.
  • Mixed: open 11-mixed-apply.md and execute KEEP/CUT per ID only (rule 12). Never re-judge from the Mixed path list alone. Never add code-tree path pointers as substitutes for cut demos (rule 11).

Examples

Example 1: Full harness eval

User says: "run harness eval on this repo"

Actions: inventory → Q1 if needed → Q2 (B/C budget table) → Track A → if approved, Steps 4–11. Parallel B judges, then C judges. Present agreements (terms are in the files).

Example 2: Usefulness only (existing run)

User says: "run Track C usefulness on the last harness-eval run"

Actions: Steps 7–11 on that RUN_ID (inventory must already exist).

Example 3: Wrong skill

User says: "setup harness" / "init harness" → harness setup (not this skill). User says: "specify feature" → tlc-spec-driven.

Troubleshooting

Trap gate FAIL (Track B or C)

Cause: KEEP/KEEP-CORE plants were deck duplicates, or blind judge mis-family. Solution: use skill’s fixed plant templates; rescore plants; re-merge.

Track A false missing .agents/...

Cause: bad path normalization. Solution: skill script must use normalize_cite (strip ./ only). Re-run Track A from $SKILL_DIR/scripts/.

Subagent blocked

Cause: missing/allowlisted model or *-fast blocked. Solution: re-spawn with an allowlisted non-fast model.

Track C Slim looks wrong after model change

Expected: usefulness is model-sensitive. Re-run C1+C2 on a second model; intersection of Slim bands is the safe delete set.

Mixed apply rewrote conventions / removed modules

Cause: apply agent re-judged from the Mixed path list instead of following KEEP/CUT. Solution: apply only via 11-mixed-apply.md; if that file is missing, re-run merge_usefulness.py; if Keep-core/Slim cells are vague, re-score those IDs before apply.

T2 empty / skill references/ missing from inventory

Cause: path normalize used lstrip("./") and turned .agents/… into agents/…. Solution: doc_scope.normalize_rel must strip only a ./ prefix (same rule as Track A).

ADRs appeared in Track C

Cause: old inventory treated all one-hop docs/** as T2. Solution: v1.7+ excludes decision-record trees; only user-approved optional doc types (never ADR/RFC) can enter T2.

Slim stub broke another skill that loads that file

Cause: content OVERLAP/Slim without fan-in — older runs, or apply skipped the gate. Solution: restore the checklist body; re-merge with merge_usefulness.py (fan-in scans full skill trees). Confirm with slim_fanin.py --path <P>.

Scripts missing

Cause: incomplete skill folder. Solution: restore $SKILL_DIR/scripts/ and references/.

© tech-leads-club, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts, references) in packages/skills-catalog/skills/(development)/harness-eval of tech-leads-club/agent-skills.

  • SKILL.md
  • references/GLOSSARY.md
  • references/PROTOCOL.md
  • references/claims.schema.json
  • references/judge-prompts.md
  • scripts/doc_scope.py
  • scripts/inventory_extract.py
  • scripts/merge_agreement.py
  • scripts/merge_usefulness.py
  • scripts/slim_fanin.py
  • scripts/surfaces_extract.py
  • scripts/track_a_correctness.py

Open the folder on GitHubat commit 6df68d5

Compare with similar skills

Harness Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harness Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harness Eval this skilltech-leads-club/agent-skills7k—~3.9kAutomated safety check: PassCC-BY-4.0
Write Skilldruxt/druxt.js114—~926Automated safety check: PassMIT
Waza Skill Evaluatormicrosoft/waza1.4k—~2kAutomated safety check: PassMIT
Waza Interactivemicrosoft/waza1.4k—~1.3kAutomated safety check: PassMIT
Learn From History AuditKonghaYao/peri226—~3.5kAutomated safety check: PassApache-2.0
Skill CreatorZS520L/HanakoPro103—~7.4kAutomated safety check: PassApache-2.0

Similar skills

  • Write Skill

    druxt/druxt.js

    Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it.

    114 GitHub stars~926 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Waza Skill Evaluator

    microsoft/waza

    Official

    Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.

    1.4k GitHub stars~2k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Waza Interactive

    microsoft/waza

    Official

    Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

    1.4k GitHub stars~1.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Audits recent agent conversation history and turns repeated failures and successes into testable harness improvement proposals that later audits can check.

    226 GitHub stars~3.5k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Skill Creator

    ZS520L/HanakoPro

    Create new skills, modify and improve existing skills, and measure skill performance.

    103 GitHub stars~7.4k tokensUpdated 4 mo ago
    Agent WorkflowsAuto-check passed
  • Skill Creator

    avibebuilder/claude-prime

    A skill your agent uses when the user wants to work on a Claude Code skill file (SKILL.md): writing one from scratch, testing whether an existing one works well, running evals or benchmarks…

    120 GitHub stars~8.2k tokensUpdated 4 mo ago
    Agent WorkflowsAuto-check passed

More from tech-leads-club/agent-skills

All 74 skills in this repo
  • Evolutionary Modular Architecture

    tech-leads-club/agent-skills

    Guides design of modular-monolith platforms with DDD, flat-by-aggregate modules, anti-corruption layers, outbox events and resilience, plus an architecture document with SVG diagrams.

    7k GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Excalidraw Diagram Studio

    tech-leads-club/agent-skills

    Generates Excalidraw diagram files from plain descriptions, choosing among flowcharts, mind maps, architecture, swimlane, class, sequence and ER diagrams.

    7k GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Mermaid Studio

    tech-leads-club/agent-skills

    Creates, validates and renders Mermaid diagrams to SVG, PNG or ASCII, including C4 and AWS architecture-beta, flowcharts, sequence diagrams and ERDs.

    7k GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • AWS Cloud Advisor

    tech-leads-club/agent-skills

    Answers AWS architecture, security and service-selection questions by searching AWS documentation through MCP tools first, then adapting advice to your stack and team.

    7k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • NestJS Modular Monolith Architect

    tech-leads-club/agent-skills

    Designs scalable NestJS modular monoliths with domain-driven design, Clean Architecture layers and optional CQRS, defining bounded contexts and strict module boundaries.

    7k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Skill Architect

    tech-leads-club/agent-skills

    Guides you through designing a new agent skill by conversation, from discovery questions and architecture to drafting SKILL.md, validation and delivery.

    7k GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Harness Eval

What does Harness Eval do?

Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports. md, rules, skills and the references they point to. It looks for broken paths and commands, redundant instructions and how useful each piece is, using a stack-agnostic protocol in which two judges review claims and planted traps check the judges.

When should I use Harness Eval?

Harness Eval fits situations like: auditing AGENTS.md and skill files for broken paths and commands; finding redundant or duplicated instructions across an agent harness; judging which skills and rules are actually useful before trimming; producing a Ship, Review, Hold, Slim or Keep-core report for a harness.

How do I install Harness Eval in Claude Code?

Run `npx skills add tech-leads-club/agent-skills --skill harness-eval -a claude-code`. Or copy the skill folder (packages/skills-catalog/skills/(development)/harness-eval in tech-leads-club/agent-skills) into .claude/skills/harness-eval in your project. Claude Code loads it when a task matches its description.

How do I install Harness Eval in Codex?

Run `npx skills add tech-leads-club/agent-skills --skill harness-eval -a codex`. Or copy the skill folder (packages/skills-catalog/skills/(development)/harness-eval in tech-leads-club/agent-skills) into .agents/skills/harness-eval in your project. Codex loads it when a task matches its description.

Can I use Harness Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tech-leads-club/agent-skills --skill harness-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-eval, .gemini/skills/harness-eval, .github/skills/harness-eval and .opencode/skills/harness-eval in your project.

What does Harness Eval need to run?

Going by SKILL.md and its folder, Harness Eval needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python to run the bundled scripts.

Does Harness Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Harness Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Harness Eval use?

Harness Eval is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harness Eval use?

About 3.9k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.6k tokens, read only when the agent opens those files.

What are the alternatives to Harness Eval?

Skills that share tags, products or a category with Harness Eval: Write Skill (druxt/druxt.js, 114 stars), Waza Skill Evaluator (microsoft/waza, 1.4k stars), Waza Interactive (microsoft/waza, 1.4k stars) and Learn From History Audit (KonghaYao/peri, 226 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harness Eval?

tech-leads-club (a GitHub organization) maintains it in tech-leads-club/agent-skills, which has 7,045 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 9, 2026.

Source: tech-leads-club/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.