Agent skill

Agent Harness

by alirezarezvani in alirezarezvani/claude-skills

Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry…

MITAuto-check passedAgent Workflows

Install Agent Harness

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill agent-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills agent-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/agent-harness/skills/agent-harness .claude/skills/agent-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-harness
GitHub stars
28k
Token cost
~2k tokens
SKILL.md length
708 words
Files
28 (incl. scripts, references, assets)
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry…

  • Works in 8 steps: Never adjudicate your own verification.… → Never modify a gate you are judged by.… → One task at a time, writes serialized.… → …
  • You want an agent
  • SKILL.md covers The contract, Quick start, Hard rules and Forcing questions (ask before…, plus 3 more sections
  • Calls python3

What it does

Agent Harness is an agent skill from alirezarezvani/claude-skills. Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until everything is verified or explicitly waived. Use when you want an agent or subagent to pick up a goal and drive it to a verified close across one of this repo's 18 domains ('run this goal through the engineering harness', 'set up an agentic…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 29 other files, including scripts, reference files and assets (for example `assets/harness_manifest.schema.json`, `assets/harnesses/business-growth.json` and `assets/harnesses/business-operations.json`).

It sits in Agent Workflows, covering Autonomous loops and Subagents. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • You want an agent
  • Set up an agentic loop for marketing work
  • Make the finance domain self-verifying)

Example prompts

  • “s 18 domains (”
  • “set up an agentic loop for marketing work”
  • “make the finance domain self-verifying”
  • “/agent-harness”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Never adjudicate your own verification. verify runs the checks via subprocess;
  2. Never modify a gate you are judged by. Check commands come from the manifest/plan.
  3. One task at a time, writes serialized. Parallelize reading and judging, never two
  4. Retry means a changed approach. Same command + same input = same failure. The retry
  5. Budgets are terminal states, not suggestions. max_attempts_per_task → escalated
  6. Fresh context beats long context. Every next directive is executable by a new
  7. State lives in .agent-harness/ — never in .agenthub/, .autoresearch/, or
  8. Plan and state files are a trust boundary. verify shell-executes each task's

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Harness loads about 2k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 207 tokens; SKILL.md has 708 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~207
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 708 words, ~2,000 tokens.

Download SKILL.mdSave it as .claude/skills/agent-harness/SKILL.md (or your agent's skills folder). This skill also uses 27 other files; get the full folder from GitHub.
name
agent-harness
description
Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until everything is verified or explicitly waived. Use when you want an agent or subagent to pick up a goal and drive it to a verified close across one of this repo's 18 domains ('run this goal through the engineering harness', 'set up an agentic loop for marketing work', 'make the finance domain self-verifying'). NOT for authoring Claude Code Workflow-tool .js scripts (workflow-builder), N-agent tournaments on one task (agenthub), single-file metric optimization (autoresearch-agent), or discovering published loop recipes (loop-library).

Agent Harness

You are a harness operator, not a hero. The loop — not your optimism — decides when work is done. Your job: compile the goal into tasks with checks, execute one task at a time, let the controller adjudicate verification, and stop when the state machine says stop.

The contract

GOAL → goal_compiler → PLAN → loop_controller: [execute → verify]* → CLOSE
                                     ↑______retry (≤ max_attempts, changed approach)
                                     └── ESCALATE on exhausted budgets — never fake success

Three layers, all JSON: a committed per-domain manifest (what skills/tools/checks exist), a per-goal plan (which tasks, which verifications, what "done" means), and a per-run state file (the single source of truth; a fresh session resumes from it alone).

Quick start

bash
# 0. Pick the domain manifest (18 committed under assets/harnesses/, e.g. engineering-team.json)
ls assets/harnesses/

# 1. Compile the goal (refuses vague goals with exit 3 + forcing questions)
python3 scripts/goal_compiler.py \
  --goal "audit the payments service and design an SLO with an error budget" \
  --manifest assets/harnesses/engineering.json --out plan.json

# 2. Initialize the loop state
python3 scripts/loop_controller.py init --plan plan.json --state .agent-harness/state.json

# 3. Drive the loop — repeat until directive is "close" or "escalate"
python3 scripts/loop_controller.py next --state .agent-harness/state.json
#    → {"action": "execute", "task": "T1", ...}: open the task's skill (SKILL.md at
#      skill_path), do the work with its tools, then:
python3 scripts/loop_controller.py record --state .agent-harness/state.json \
  --task T1 --phase execute --exit-code 0
#    → the controller runs the task's checks ITSELF (subprocess, timeout, evidence log):
python3 scripts/loop_controller.py verify --state .agent-harness/state.json --task T1 --cwd <repo-root>

# 4. Close — refused (exit 4) while any task is unverified and unwaived
python3 scripts/loop_controller.py close --state .agent-harness/state.json

Regenerate a manifest after skills change (diff-stable, CI-checkable):

bash
python3 scripts/harness_manifest_builder.py --domain engineering-team \
  --repo-root <repo-root> --out-dir assets/harnesses --no-timestamp

Hard rules

  1. Never adjudicate your own verification. verify runs the checks via subprocess; a passing record --phase verify without --evidence is rejected (exit 6). You do not get to declare a task verified.
  2. Never modify a gate you are judged by. Check commands come from the manifest/plan. Editing a check to make it pass is the reward-hacking failure mode (see references/verification_discipline.md) — same invariant as autoresearch-agent's locked evaluator.
  3. One task at a time, writes serialized. Parallelize reading and judging, never two tasks writing the same artifact (references/agentic_loop_canon.md).
  4. Retry means a changed approach. Same command + same input = same failure. The retry directive says so; honor it.
  5. Budgets are terminal states, not suggestions. max_attempts_per_task → escalated (exit 2); max_loop_iterations → escalate (exit 5). Exhausted budgets are never reported as success — a human waives (close --waive T3 --reason "..."), you don't.
  6. Fresh context beats long context. Every next directive is executable by a new session reading only the plan + state files. Long-running goals: run each iteration as its own session against the durable state.
  7. State lives in .agent-harness/ — never in .agenthub/, .autoresearch/, or docs/TC/ (those belong to sibling skills).
  8. Plan and state files are a trust boundary. verify shell-executes each task's check command; only run the harness on plan/state files you or goal_compiler.py produced, never on files from untrusted input (see references/verification_discipline.md).

Forcing questions (ask before compiling; one per turn, with a recommended answer)

#QuestionRecommended answerWhy (canon)
1What single observable outcome means DONE?A named artifact + a command that exits 0 against itVerifier's law: invest in verifiability first
2Which domain harness applies?The domain whose skills name the deliverable; if two, run two sequential loopsOrchestrator-workers: scoped objectives beat mega-goals
3What must NOT change?List no-touch paths; put them in the goal text so the compiler's plan inherits themBoundaries are part of a subagent spec
4Who reviews escalations, and how fast?A named human; escalations block the loop by designApproval-required is a terminal state, not a nuisance
5What is the iteration budget?Default 12 loop iterations / 3 attempts per task; raise only with a reasonCaps are runtime errors, not advice (OpenAI SDK max_turns)
Show full SKILL.md (245 more words)Show less

Exit codes (branch on these mechanically)

CodeToolMeaning
0allOK / directive emitted
2loop_controllerEscalation required — a human must review the evidence log
3goal_compilerGoal too vague — answer the forcing questions, recompile
4goal_compiler / loop_controllerNo skill matched / close refused (unverified tasks)
5loop_controllerGlobal iteration cap reached
6loop_controllerInvalid transition (recording on verified task, evidence missing, unknown task)

Verifiable success

  • python3 scripts/harness_manifest_builder.py --sample, scripts/goal_compiler.py --sample, and scripts/loop_controller.py --sample all exit 0.
  • A vague goal (--goal "make it better") exits 3 and prints forcing questions.
  • loop_controller.py close on a state with an unverified task exits 4.
  • The demo loop in loop_controller.py --sample shows a verify failure consuming an attempt and the loop still closing only after a passing verify with evidence.
  • workflow-builder: authoring deterministic .js scripts for Claude Code's Workflow tool. NOT for goal-to-close loop state (this skill).
  • agenthub: N parallel agents competing on ONE task in git worktrees. Use it inside a harness task that wants competing attempts.
  • autoresearch-agent: metric optimization of a single file against a locked evaluator. Use it when a task's done_when is "metric improves".
  • tc-tracker: per-code-change lifecycle records. Use for change bookkeeping; the harness state file is per-goal, not per-change.
  • loop-library: discover/audit published loop recipes conversationally. This skill is the executable enforcement of that vocabulary.
  • ship-gate / self-eval / spec-driven-workflow: plug in as close-time checks inside a task's verification[].

See references/domain_harness_design.md for the three-layer architecture, the reuse map, and how to raise a domain's harness quality.

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 27 other files (scripts, references, assets) in engineering/agent-harness/skills/agent-harness of alirezarezvani/claude-skills.

  • SKILL.md
  • assets/harness_manifest.schema.json
  • assets/harnesses/business-growth.json
  • assets/harnesses/business-operations.json
  • assets/harnesses/c-level-advisor.json
  • assets/harnesses/commercial.json
  • assets/harnesses/compliance-os.json
  • assets/harnesses/engineering-team.json
  • assets/harnesses/engineering.json
  • assets/harnesses/finance.json
  • assets/harnesses/loop-library.json
  • assets/harnesses/markdown-html.json
  • assets/harnesses/marketing-skill.json
  • assets/harnesses/marketing.json
  • assets/harnesses/product-team.json
  • assets/harnesses/productivity.json
  • assets/harnesses/project-management.json
  • assets/harnesses/ra-qm-team.json
  • assets/harnesses/research-ops.json
  • … and 9 more

Open the folder on GitHubat commit 19392f7

Compare with similar skills

Agent Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Harness this skillalirezarezvani/claude-skills28k—~2kAutomated safety check: PassMIT
Migrating Claude Agent SDK To Pydantic AIpydantic/pydantic-ai20k—~1.6kAutomated safety check: PassMIT
Cursor Orchestratecursor/plugins10k—~1.1kAutomated safety check: PassNone
Agent-Foreman Feature Runmylukin/agent-foreman250—~1.6kAutomated safety check: NotesNone
Spec-Driven Development v2LichAmnesia/lich-skills234—~3.1kAutomated safety check: PassMIT
Optimizeevo-hq/evo1.5k—~13kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Migrate Python applications from the Claude Agent SDK to Pydantic AI and, only when needed, Pydantic AI Harness.

    20k GitHub stars~1.6k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Cursor Orchestrate

    cursor/plugins

    Official

    Splits a large goal into a tree of parallel Cursor cloud agents, with planners, workers and verifiers coordinated by a script and reporting through structured handoffs.

    10k GitHub stars~1.1k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Agent-Foreman Feature Run

    mylukin/agent-foreman

    Runs the agent-foreman task loop for one task by ID or unattended across every pending task, following next, implement, check, done for each one.

    250 GitHub stars~1.6k tokensUpdated 8 mo ago
    Agent WorkflowsAuto-check: notes
  • Spec-Driven Development v2

    LichAmnesia/lich-skills

    Organizes long-running agent work into a Project, Sprint and Task hierarchy with per-task state files, isolated worktrees, review loops and script-checked rules.

    234 GitHub stars~3.1k tokensUpdated 4 mo ago
    Agent WorkflowsAuto-check passed
  • Optimize

    evo-hq/evo

    Drive structured autoresearch iteration after evo:discover and the baseline commit.

    1.5k GitHub stars~13k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Babysitter Process Runner

    a5c-ai/babysitter

    Execute via @babysitter. Use this skill when asked to babysit a task, do anything that is structured process-driven (even a loop) or whenever it is called…

    1.8k GitHub starsUsed in 1 repo~726 tokens
    Agent WorkflowsAuto-check: notes

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Categories

Questions about Agent Harness

What does Agent Harness do?

Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry…. Agent Harness is an agent skill from alirezarezvani/claude-skills. Turn any domain folder of skills into a bounded agentic loop: compile a goal into a verifiable task plan, execute tasks with the domain's own tools, verify every task with machine-run checks, retry with caps, escalate to a human when budgets exhaust, and refuse to close until everything is verified or explicitly waived.

When should I use Agent Harness?

Agent Harness fits situations like: you want an agent; set up an agentic loop for marketing work; make the finance domain self-verifying).

How do I install Agent Harness in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill agent-harness -a claude-code`. Or copy the skill folder (engineering/agent-harness/skills/agent-harness in alirezarezvani/claude-skills) into .claude/skills/agent-harness in your project. Claude Code loads it when a task matches its description.

How do I install Agent Harness in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill agent-harness -a codex`. Or copy the skill folder (engineering/agent-harness/skills/agent-harness in alirezarezvani/claude-skills) into .agents/skills/agent-harness in your project. Codex loads it when a task matches its description.

Can I use Agent Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill agent-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-harness, .gemini/skills/agent-harness, .github/skills/agent-harness and .opencode/skills/agent-harness in your project.

What does Agent Harness need to run?

Going by SKILL.md and its folder, Agent Harness needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Agent Harness access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Harness use?

Agent Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Harness use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.8k tokens, read only when the agent opens those files.

What are the alternatives to Agent Harness?

Skills that share tags, products or a category with Agent Harness: Migrating Claude Agent SDK To Pydantic AI (pydantic/pydantic-ai, 20k stars), Cursor Orchestrate (cursor/plugins, 10k stars), Agent-Foreman Feature Run (mylukin/agent-foreman, 250 stars) and Spec-Driven Development v2 (LichAmnesia/lich-skills, 234 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Harness?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,829 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.