Agent skill

Harness Engineering

by guanyang in guanyang/open-agent-hub

This skill should be used when designing autonomous agent harnesses: research loops, evaluation scaffolds, locked and editable surfaces, durable logs, novelty gates, pruning, rollback, PR…

MITAuto-check passedAgent Workflows

Install Harness Engineering

skills CLI
$ npx skills add guanyang/open-agent-hub --skill harness-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install guanyang/open-agent-hub harness-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/guanyang/open-agent-hub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/harness-engineering .claude/skills/harness-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-engineering
GitHub stars
975
Used in
1 other repo
Token cost
~2.9k tokens
SKILL.md length
1,421 words
Files
1
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when designing autonomous agent harnesses: research loops, evaluation scaffolds, locked and editable surfaces, durable logs, novelty gates, pruning, rollback, PR…

  • Works in 6 steps: Refresh upstream sources on a schedule. → Require novelty checks before spending… → Preserve rejected attempts to avoid… → …
  • Tasks that involve Autonomous loops
  • SKILL.md covers When to Activate, Core Concepts, Detailed Topics and Practical Guidance, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Harness Engineering is an agent skill from guanyang/open-agent-hub. This skill should be used when designing autonomous agent harnesses: research loops, evaluation scaffolds, locked and editable surfaces, durable logs, novelty gates, pruning, rollback, PR preparation, and human approval boundaries.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Autonomous loops and Project scaffolding. The repository describes itself as: A lightweight, zero-dependency CLI tool to manage and activate capabilities for AI coding assistants (such as Claude Code, Cursor, Trae, etc.). The licence is MIT.

When your agent uses it

  • Tasks that involve Autonomous loops
  • Tasks that involve Project scaffolding

Example prompts

  • “/harness-engineering”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Refresh upstream sources on a schedule.
  2. Require novelty checks before spending large budgets.
  3. Preserve rejected attempts to avoid rediscovery.
  4. Run leave-one-out pruning when a stack has multiple additions.
  5. Reward simplification when quality is equal.
  6. Use separate verification before promotion.

What it can do on your machine

Read from SKILL.md and the folder at commit c32921b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harness Engineering loads about 2.9k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 1,421 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from guanyang/open-agent-hub at commit c32921b, republished under its MIT licence (© guanyang). 1,421 words, ~2,854 tokens.

Download SKILL.mdSave it as .claude/skills/harness-engineering/SKILL.md (or your agent's skills folder).
name
harness-engineering
description
This skill should be used when designing autonomous agent harnesses: research loops, evaluation scaffolds, locked and editable surfaces, durable logs, novelty gates, pruning, rollback, PR preparation, and human approval boundaries.

Harness Engineering

Harness engineering designs the control system around an agent: what it may edit, how it receives feedback, where it writes state, how failures recover, and who can approve irreversible actions. The harness is the difference between a helpful agent session and an autonomous loop that can run for days without corrupting its objective.

When to Activate

Activate this skill when:

  • Building autonomous research or experimentation loops
  • Designing an agent environment with locked metrics and editable code or content
  • Creating PR-producing or background agents
  • Evaluating whether an agent can safely run without frequent human prompts
  • Adding novelty, ablation, pruning, rollback, or durable logging to an agent workflow
  • Preventing agents from gaming benchmarks, weakening rubrics, or losing state across compaction

Do not activate this skill for adjacent work owned by other skills:

  • General quality gates, regression suites, or outcome metrics without autonomous control surfaces: evaluation.
  • Tool schemas, response formats, and recovery errors for harness tools: tool-design.
  • Project-level task-model fit, pipeline shape, and cost planning: project-development.
  • Remote sandbox, warm-pool, and hosted session infrastructure: hosted-agents.

Core Concepts

Harness Boundary

Separate the agent from the environment it operates inside. The agent proposes actions; the harness defines allowed surfaces, feedback, persistence, and promotion rules.

Use four surface classes:

SurfaceExamplesRule
LockedEval metric, rubric, validation script, merge policyAgent may read and propose changes, but cannot score itself with modified rules
EditableSkill draft, experiment file, prompt, config under testAgent may mutate during the loop
Append-onlyResults log, research thread, rejected ideasAgent may append, not rewrite
Human-controlledMerge, production deploy, credentials, destructive operationsRequires explicit human approval
Tight Feedback Loops

Autonomy works when feedback is fast, unambiguous, and hard to game. Karpathy's autoresearch is the minimal pattern: one editable file, one locked evaluation file, fixed wall-clock budget, one scalar metric, git rollback, and a durable results log. The lesson is not that every harness needs one metric; it is that ambiguous feedback creates ambiguous autonomy.

For open-ended research-to-skill work, replace the scalar metric with locked rubrics, deterministic structure checks, source traceability, and human review thresholds.

Durable State

Long-running agents must externalize state. Store plans, source queues, results, failures, and handoffs in files so future agents can resume without relying on chat history. Prime Intellect's autonomous nanoGPT work showed the value of durable scratchpads and THREAD.md-style logs for recovery, monitoring, and audit.

Use append-only logs for:

  • What was tried
  • What improved or failed
  • Why a candidate was kept, discarded, or routed to review
  • Which upstream sources were checked
  • What the next agent should do
Search Discipline

Agents tend to exploit the nearest surface, stack complexity, and under-run pruning. Add explicit search rules:

  1. Refresh upstream sources on a schedule.
  2. Require novelty checks before spending large budgets.
  3. Preserve rejected attempts to avoid rediscovery.
  4. Run leave-one-out pruning when a stack has multiple additions.
  5. Reward simplification when quality is equal.
  6. Use separate verification before promotion.
Mechanism Registry

For research-to-skill systems, track accepted mechanisms separately from prose. A mechanism record should include a stable mechanism_id, owning_skill, status, activation scenario, behavior change, evidence, and failure modes. Novelty gates should compare against this registry before using broader corpus overlap, because keyword overlap catches stale phrasing while mechanism comparison catches real duplication.

Governance

Autonomous agents may prepare PRs, but governance must be explicit. They can draft changes, run checks, and write PR summaries. They should not merge, deploy, or push without human approval unless the user has explicitly granted that permission for the specific action.

Detailed Topics

Autoresearch-Style Loop

Use this pattern when optimizing an artifact against a stable evaluator:

text
read locked context -> choose hypothesis -> edit allowed surface -> commit/checkpoint
-> run evaluator -> log result -> keep if better -> discard or rollback if worse
-> repeat

Required properties:

  • The evaluator is outside the editable surface.
  • The feedback cadence is fixed enough to compare attempts.
  • Failed attempts leave an audit trail.
  • Rollback is cheap.
  • The agent has a policy for crashes and timeouts.
Research-To-Skill Loop

Use this pattern when sources become skill changes:

text
discover -> retrieve -> gate -> score -> extract mechanism
-> map to existing or new skill -> draft proposal -> validate structure
-> prepare PR -> human review

The locked evaluator is a combination of source rubrics, skill-change rubrics, structure checks, and reviewer approval. The editable artifact is the proposed skill delta.

Metric Gaming Resistance

Assume an optimizing agent will learn the harness. Guard against:

  • Editing evaluation code or rubrics and then using the new version for self-approval
  • Adding verbose content that pleases a judge but harms skill activation
  • Citing unretrieved sources
  • Optimizing aggregate scores while failing a critical dimension
  • Avoiding failed results in the log

Mitigation: lock rubrics per run, report per-dimension scores, require source retrieval evidence, preserve rejected attempts, and route governance changes to human review.

Monitoring Agents

Use monitoring agents for long runs, but restrict them to read-only reporting unless explicitly tasked otherwise. Monitoring output should report:

  • Best current candidate
  • Active jobs or drafts
  • Last upstream refresh
  • Failed or stale loops
  • Disagreements between logs and claimed state
  • Next action and blocker

Practical Guidance

Harness Design Checklist
  1. Define the objective in one sentence.
  2. Identify locked, editable, append-only, and human-controlled surfaces.
  3. Choose the feedback mechanism: scalar metric, rubric, deterministic tests, human review, or combination.
  4. Define keep, discard, crash, timeout, and review states.
  5. Create a durable thread log before the loop starts.
  6. Add source refresh, mechanism-registry novelty, and pruning rules for long-running loops.
  7. Define what the agent may do without asking and what requires approval.
  8. Validate the harness on one known good and one known bad artifact.
Show full SKILL.md (542 more words)Show less
File Layout
text
research-run/
  THREAD.md
  sources/
    queue.md
    evaluations/
  proposals/
  logs/
    results.tsv
    rejected.md
  drafts/

Use TSV or JSONL for append-only machine-readable logs. Use Markdown for handoffs and reviewer-facing summaries.

Examples

Example 1: Locked metric

An agent optimizes train.py, but prepare.py owns data loading and evaluation. The agent can edit the model but cannot change the metric. Failed experiments are logged and rolled back.

Example 2: Locked rubric

An agent evaluates a new Anthropic or OpenAI engineering post, but the source curation rubric is locked for the run. If the source passes, the agent drafts a skill proposal. It cannot lower the rubric threshold to admit the source.

Example 3: Auto-PR without auto-merge

An agent prepares a branch and PR body after passing source, skill, and structure checks. The PR states unresolved risks and waits for human merge approval.

Guidelines

  1. Lock evaluators before starting the loop.
  2. Keep editable surfaces narrow enough for reliable diffs.
  3. Write durable logs before context compaction can erase state.
  4. Report per-dimension scores instead of only aggregate scores.
  5. Require source retrieval before citation.
  6. Add novelty gates for broad search and pruning gates for complex stacks.
  7. Prefer simplification when quality is equal.
  8. Separate PR preparation from merge authority.
  9. Revalidate harness changes with old and new evaluators.
  10. Treat stopped autonomous loops as harness failures, not agent personality quirks.

Gotchas

  1. Mutable evaluator: If the agent can edit the metric, it may optimize the benchmark instead of the task. Keep rubrics and eval code locked during the run.
  2. Chat-only memory: Long runs fail after compaction when plans live only in conversation history. Write thread logs and result files from the start.
  3. No discard record: Without rejected-attempt logs, agents repeat failed ideas. Preserve failures with enough detail to avoid rediscovery.
  4. Complexity accretion: Agents stack changes and rarely remove them. Require pruning rounds and reward equal-quality simplification.
  5. Premature novelty claims: Agents label recombinations as novel. Compare against existing repo skills, source queue, and rejected logs before claiming novelty.
  6. Monitor misreporting: Monitoring agents can summarize stale or inconsistent state. Require them to cite the files or logs behind claims.
  7. Human approval ambiguity: "Prepare a PR" is not "merge a PR." Make approval boundaries explicit in the harness.
  8. Volatile source drift: Fast-moving lab claims age quickly. Put dated evidence in references and schedule revalidation.

Integration

This skill connects to:

  • evaluation - Rubrics and quality gates provide the locked feedback surface
  • advanced-evaluation - Pairwise comparison and bias mitigation improve proposal review
  • filesystem-context - Durable logs, scratchpads, and thread files preserve state
  • multi-agent-patterns - Researcher, verifier, monitor, and writer agents need isolated contexts
  • tool-design - Harness tools must expose clear contracts and recovery errors
  • project-development - File-based pipelines and task-model fit analysis keep loops simple
  • hosted-agents - Background execution needs sandbox, snapshot, and approval boundaries

References

Internal references:

  • researcher/README.md - Read when implementing the repo-native research-to-skill operating system
  • researcher/rubrics/harness-change.md - Read when evaluating changes to an agent harness
  • researcher/runbooks/autonomous-research-loop.md - Read when running a source-to-skill loop

External resources:

  • Karpathy autoresearch - Constrained autonomous experiment loop with locked evaluation
  • Prime Intellect autonomous nanoGPT speedrun - Durable scratchpads, handoffs, monitoring, and autonomy failure modes
  • AlphaEvolve and FunSearch - LLM-generated candidates paired with systematic evaluators
  • HELM and LM Evaluation Harness - Transparent, reproducible evaluation infrastructure

Skill Metadata

Created: 2026-05-14 Last Updated: 2026-05-15 Author: Agent Skills for Context Engineering Contributors Version: 1.1.0

© guanyang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/harness-engineering of guanyang/open-agent-hub.

Open the folder on GitHubat commit c32921b

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in guanyang/open-agent-hub, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Harness Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harness Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harness Engineering this skillguanyang/open-agent-hub9751 repos~2.9kAutomated safety check: PassMIT
Loop FactoryJuliusBrussee/skills161—~2kAutomated safety check: PassMIT
Autoresearch Loopjdrhyne/agent-skills240—~2kAutomated safety check: PassMIT
AWS Harnesshoodini/ai-agents-skills281—~4.2kAutomated safety check: NotesNone
Example Harnessruvnet/metaharness690—~735Automated safety check: PassMIT
Autoresearch Task QAbosprimigenious/autoresearch-skills151—~1.6kAutomated safety check: PassMIT

Similar skills

  • Loop Factory

    JuliusBrussee/skills

    Run a spec-driven agent loop where coding tasks live as markdown specs that move through inbox → active → archive, get implemented by Claude Code or Codex, and pass a review gate before they count…

    161 GitHub stars~2k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Autoresearch Loop

    jdrhyne/agent-skills

    Domain-agnostic metric-driven improvement loop, generalizing Karpathy's autoresearch.

    240 GitHub stars~2k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • AWS Harness

    hoodini/ai-agents-skills

    Build a new AI agent on AWS and deploy it easily, OR wrap and deploy an agent you already have, using the Amazon Bedrock AgentCore harness.

    281 GitHub stars~4.2k tokensUpdated 2 mo ago
    Backend & APIsAuto-check: notes
  • Example Harness

    ruvnet/metaharness

    Scaffold a ready-made AI agent harness in one command from the 19 published @metaharness/ example packages — 9 host integrations (Claude Code, Codex, Hermes, pi.dev, OpenClaw, RVM, Copilot…

    690 GitHub stars~735 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Autoresearch Task QA

    bosprimigenious/autoresearch-skills

    对 AutoResearch 目录或 ZIP 做只读质检;先审优化面是否仅调参、Baseline 是否合理、Reference 提升是否充分,再查 21 项实现、严格 Docker 路径、Harbor 接口与成对训练证据。输出方法介绍、完整跑分、专家退回说明及 TXT/Markdown/JSON 报告;不用于求解任务。

    151 GitHub stars~1.6k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed
  • Autoresearch Run Isolation

    bosprimigenious/autoresearch-skills

    为 AutoResearch 的双 Agent 轨迹、付费 GPU 长跑、Docker 执行、可信评测与恢复建立共享协议、成本决策和隔离边界。用于小时/包日选择、启动或恢复 campaign、设计证据与防止题目或轨迹串用;不替代具体任务算法或最终平台 QA。

    151 GitHub stars~553 tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed

More from guanyang/open-agent-hub

All 26 skills in this repo
  • Context Compression

    guanyang/open-agent-hub

    This skill should be used when long-running agent sessions need context compression, structured summarization, compaction, token-per-task optimization, or durable handoff summaries that preserve…

    975 GitHub starsUsed in 2 repos~4.6k tokens
    Auto-check passed
  • Context Fundamentals

    guanyang/open-agent-hub

    This skill should be used to explain or reason about the foundational concepts of context engineering: what context is, the anatomy of a context window, how attention mechanics work, the U-shaped…

    975 GitHub starsUsed in 2 repos~4.2k tokens
    Auto-check passed
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    975 GitHub starsUsed in 2 repos~4.2k tokens
    Auto-check passed
  • Multi Agent Patterns

    guanyang/open-agent-hub

    This skill should be used when designing multi-agent systems that need context isolation, supervisor or swarm coordination, explicit handoffs, parallel execution, or a decision on whether multiple…

    975 GitHub starsUsed in 2 repos~4.6k tokens
    Auto-check passed
  • Project Development

    guanyang/open-agent-hub

    This skill should be used for project-level decisions about LLM-powered systems: whether an LLM is the right primitive for the task at hand, the shape of a multi-stage batch or agent pipeline, token…

    975 GitHub starsUsed in 2 repos~4.7k tokens
    Auto-check passed
  • Tool Design

    guanyang/open-agent-hub

    This skill should be used for the tool-interface layer of an agent system specifically: writing tool descriptions agents can route on, designing tool schemas and response formats, naming…

    975 GitHub starsUsed in 2 repos~5k tokens
    Auto-check passed

Questions about Harness Engineering

What does Harness Engineering do?

This skill should be used when designing autonomous agent harnesses: research loops, evaluation scaffolds, locked and editable surfaces, durable logs, novelty gates, pruning, rollback, PR…. Harness Engineering is an agent skill from guanyang/open-agent-hub. This skill should be used when designing autonomous agent harnesses: research loops, evaluation scaffolds, locked and editable surfaces, durable logs, novelty gates, pruning, rollback, PR preparation, and human approval boundaries.

When should I use Harness Engineering?

Harness Engineering fits situations like: tasks that involve Autonomous loops; tasks that involve Project scaffolding.

How do I install Harness Engineering in Claude Code?

Run `npx skills add guanyang/open-agent-hub --skill harness-engineering -a claude-code`. Or copy the skill folder (skills/harness-engineering in guanyang/open-agent-hub) into .claude/skills/harness-engineering in your project. Claude Code loads it when a task matches its description.

How do I install Harness Engineering in Codex?

Run `npx skills add guanyang/open-agent-hub --skill harness-engineering -a codex`. Or copy the skill folder (skills/harness-engineering in guanyang/open-agent-hub) into .agents/skills/harness-engineering in your project. Codex loads it when a task matches its description.

Can I use Harness Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guanyang/open-agent-hub --skill harness-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-engineering, .gemini/skills/harness-engineering, .github/skills/harness-engineering and .opencode/skills/harness-engineering in your project.

What does Harness Engineering need to run?

SKILL.md names no scripts, command-line tools or credentials: Harness Engineering is instructions for the agent only.

Does Harness Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Harness Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Harness Engineering use?

Harness Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harness Engineering use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Harness Engineering?

Skills that share tags, products or a category with Harness Engineering: Loop Factory (JuliusBrussee/skills, 161 stars), Autoresearch Loop (jdrhyne/agent-skills, 240 stars), AWS Harness (hoodini/ai-agents-skills, 281 stars) and Example Harness (ruvnet/metaharness, 690 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harness Engineering?

guanyang (a GitHub user) maintains it in guanyang/open-agent-hub, which has 975 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on October 7, 2026.

Source: guanyang/open-agent-hub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.