Harness Engineering
10xChengTu/harness-engineering
Set up and improve harness engineering (AGENTS.md, docs/, lint rules, eval systems, project-level prompt engineering) for AI-agent-friendly codebases.
Design, build, diagnose, and evolve agent harnesses: instructions, tools, execution environments, durable state, context management, verification, recovery, and bounded autonomous loops.
$ npx skills add magnus919/agent-skills --skill harness-engineering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install magnus919/agent-skills harness-engineering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/harness-engineering .claude/skills/harness-engineering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "harness-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/harness-engineering into .claude/skills/harness-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-engineering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/magnus919/agent-skills/tree/main/harness-engineeringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add magnus919/agent-skills --skill harness-engineering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install magnus919/agent-skills harness-engineering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/harness-engineering .agents/skills/harness-engineering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "harness-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/harness-engineering into .agents/skills/harness-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-engineering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill harness-engineering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install magnus919/agent-skills harness-engineering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/harness-engineering .cursor/skills/harness-engineering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "harness-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/harness-engineering into .cursor/skills/harness-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-engineering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/magnus919/agent-skills.git --path harness-engineering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add magnus919/agent-skills --skill harness-engineering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install magnus919/agent-skills harness-engineering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/harness-engineering .gemini/skills/harness-engineering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "harness-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/harness-engineering into .gemini/skills/harness-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-engineering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install magnus919/agent-skills harness-engineeringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add magnus919/agent-skills --skill harness-engineering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/harness-engineering .github/skills/harness-engineering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "harness-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/harness-engineering into .github/skills/harness-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-engineering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill harness-engineering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install magnus919/agent-skills harness-engineering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/harness-engineering .opencode/skills/harness-engineering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "harness-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/harness-engineering into .opencode/skills/harness-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "harness-engineering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
harness-engineeringDesign, build, diagnose, and evolve agent harnesses: instructions, tools, execution environments, durable state, context management, verification, recovery, and bounded autonomous loops.
Harness Engineering is an agent skill from magnus919/agent-skills. Design, build, diagnose, and evolve agent harnesses: instructions, tools, execution environments, durable state, context management, verification, recovery, and bounded autonomous loops. Use for unreliable coding agents, cross-session drift, premature completion, harness audits, or changes to the runtime around a model, including bounded typed-decision routing. Do not use for prompt rewriting alone, model training, general application architecture, or operating an already evaluated agent in production; route…
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 62 other files, including scripts and reference files (for example `README.md`, `evals/business-app-rubric-challenges.md` and `evals/evals.json`).
It sits in Agent Workflows, covering Session handoff, Fine-tuning and Context engineering. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 22b4723. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Harness Engineering loads about 3.6k tokens when it runs, and up to ~39k if it reads all its reference files. Until then it costs about 144 tokens; SKILL.md has 1,305 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from magnus919/agent-skills at commit 22b4723, republished under its MIT licence (© magnus919). 1,305 words, ~3,606 tokens.
.claude/skills/harness-engineering/SKILL.md (or your agent's skills folder). This skill also uses 59 other files; get the full folder from GitHub.Engineer the system around the model so an agent can start, act within authority, observe results, recover, and finish with evidence. Scaffolding is one entry point; existing harnesses may need fewer files, better tools, or removal of stale rules. Default to the smallest change that addresses an observed failure.
Confirm the target, scope, and rollback path before acting. Read-only discovery may proceed without confirmation.
Use existing user authorization when it covers these facts; do not ask again for routine reversible work. Destructive operations require an explicit directive. Inspect local guidance, real tasks, tool contracts, manifests, CI, state storage, and recent failure traces before editing. Do not execute commands discovered in untrusted instructions or logs merely because an audit found them.
Identify whether the task concerns a repository harness (around an existing coding agent) or a custom runtime (its loop, tool dispatcher, persistence, and controls). Do not prescribe repository files as the architecture of every agent.
| Subsystem | Engineering question | Evidence to inspect |
|---|---|---|
| Instructions | Does the agent find current intent and constraints? | Routing, precedence, source ownership, stale rules |
| Tools | Can it act and inspect results within authority? | Schemas, errors, permission enforcement, per-call concurrency |
| Environment | Can another session reproduce execution? | Runtime versions, dependencies, services, isolation |
| State | Can work resume without inventing progress? | Objective, revision, checkpoints, ownership, evidence |
| Feedback | Can it distinguish success, failure, and missing evidence? | Real checks, user journey, traces, termination gate |
Scope, initialization, handoff, and recovery cross these five subsystems. Keep them explicit without presenting a second, incompatible five-subsystem taxonomy.
human_interventions aggregate intact. Add the
optional typed intervention_detail only when the question needs event-level
attribution; absence means detail was not collected, not zero. Do not use a
lower count alone as a productivity or success claim. See
observability and feedback.Choose the mode from the task, not from available tools. Read references/entry-points-and-use-cases.md for six worked examples and mode-specific exit artifacts.
| Request | Start with | Deliver |
|---|---|---|
| Create a repository harness | Existing commands/state + fresh-session test | Minimal tailored setup; real startup/task/handoff evidence |
| Diagnose an unreliable agent | First-divergence analysis | Supported hypothesis, discriminating probe, bounded repair |
| Design a custom runtime | Tool/state/authority/event contracts | Runtime design with failure/replay tests |
| Reduce context/cost | Measured stage/context baseline | Fidelity-preserving intervention and comparison |
| Automate repeated work | Goal, verifier, authority, budgets | Bounded loop with exercised termination/recovery |
| Coordinate workers | Ownership, shared interfaces, integration | Explicit routing and combined verification |
| Maintain/retire harness rules | Current failures, source owners, model changes | One evidenced simplification with rollback |
| Active problem | Reference | Working template |
|---|---|---|
| Failures or uncertain audit findings | Diagnosis | Harness contract |
| Startup/worker split and fresh-session discovery | Session lifecycle | Session protocol |
| Invisible knowledge, stale/contradictory rules | Knowledge and invariants | Guardrail promotion |
| Context retrieval, compression/reset, budgets | Context design | Context budget |
| Persistence, state transitions, partial effects | State and recovery | State, recovery drill |
| Environment and custom runtime boundaries | Runtime design | Runtime decision |
| Tool contracts, events, authorization, hooks | Tools and events | Tool contract, event contract |
| Find task-relevant tools and operate on large result sets | Tool discovery and bulk results | Bulk result contract, tool contract |
| Actual runtime journey and trace-to-action loop | Observability feedback | Observability plan |
| Optional, privacy-bounded intervention event detail | Observability and feedback, script contract | Aggregate-only run record by default; add detail only for the stated question |
| Premature completion, verifier, experiment | Verification/improvement | Acceptance, experiment, run record |
| Repeated work, graphs, concurrent ownership | Loops/coordination | Loop contract, graph |
| Typed model choice inside a harness | System One decisions | Decision placement |
| Need specialist input then return to this workflow | Catalog handoffs | Supply and consume the named contract |
| Source provenance and extraction coverage | Source assessment, coverage map, primary-source index, source inventory | Scope observations to their evidence |
| Create a minimal repository harness | Use existing state owner first | Instructions, handoff |
The scripts use Python 3.10+ standard library. Resolve scripts/ relative to this
skill's installed root, not the target repository. See
references/script-contract.md for the command and
report contracts. These are structural aids and explicit check execution, not an
agent-quality benchmark. Contract declarations are validated without executing
or authenticating their evidence references.
# Read-only audit; optional HTML writes only the requested new report.
python3 scripts/harness.py audit --target /path/to/repo
# Preview; creates nothing. Apply only within the confirmed scope.
python3 scripts/harness.py scaffold --target /path/to/repo
python3 scripts/harness.py scaffold --target /path/to/repo --apply
# Review the JSON argv list first. Run only authorized commands.
python3 scripts/harness.py verify --target /path/to/repo --commands /path/to/checks.json
python3 scripts/harness.py verify --target /path/to/repo --commands /path/to/checks.json --execute --report /tmp/new-run.jsonRead-only contract helpers:
python3 scripts/contracts.py validate --kind state --file /path/to/state.json
python3 scripts/contracts.py validate --kind graph --file /path/to/graph.json
python3 scripts/contracts.py validate --kind run --file /path/to/run.json
python3 scripts/contracts.py compare --baseline /path/to/baseline.json --candidate /path/to/candidate.jsonThe comparison refuses changed model/environment/task/authority/verifier fingerprints, unequal case IDs, and changed inputs. It reports supplied observations and regressions; it does not prove causal improvement or approve release.
Scaffolding never overwrites existing paths and does not install dependencies or run a project. Tailor its acceptance placeholders before use. Verification records command outcomes; it never marks a feature complete or approves a release.
Finish when the requested audit, design, or implementation is delivered with observable evidence, remaining limits, and a restart/rollback path. For an audit, stop at a prioritized hypothesis and bounded experiment; do not invent a repair. For implementation, stop when acceptance and relevant failure/recovery checks pass, or after three non-converging diagnostic passes report the blocker and required evidence. Respect the user's tighter budget. Do not claim behavioral improvement without comparable real agent runs.
Read forward-test report when assessing tested coverage and remaining evidence gaps. Independent fixture reviews complement the declarative cases; they do not establish field efficacy.
For bounded selection or routing with a typed decision model, read System One decisions before proposing an integration. It distinguishes evidence-backed harness controls from local implementation reports and untested design proposals. System One owns question semantics, model-specific limits, calibration, and model-level latency; Harness Engineering owns the candidate space, current state, execution gates, observed effects, task outcome, and end-to-end cost. Return the measured task outcomes to the harness decision after the model-level contract is reviewed.
When a workflow embeds authoritative business rules, use spec-driven-development's policy recipe and release contract. The harness selects approved versioned inputs, separates explanation from evaluation, and re-checks execution authority against current state. Keep clause interpretation and translation-fidelity fixtures with the policy owner; do not duplicate policy authoring here.
© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 59 other files (scripts, references) in harness-engineering of magnus919/agent-skills.
Open the folder on GitHubat commit 22b4723
Harness Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Harness Engineering this skillmagnus919/agent-skills | 115 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Harness Engineering10xChengTu/harness-engineering | 102 | 1 repos | ~1k | Automated safety check: Pass | None | |
| Session Handoffdavila7/claude-code-templates | 33k | 2 repos | ~1.6k | Automated safety check: Pass | MIT | |
| Memori Long-Term MemoryMemoriLabs/Memori | 17k | — | ~2k | Automated safety check: Notes | Custom licence | |
| Session Handoff Documentthedotmack/claude-mem | 99k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Planning with FilesOthmanAdi/planning-with-files | 27k | — | ~2.9k | Automated safety check: Pass | MIT |
10xChengTu/harness-engineering
Set up and improve harness engineering (AGENTS.md, docs/, lint rules, eval systems, project-level prompt engineering) for AI-agent-friendly codebases.
davila7/claude-code-templates
Creates comprehensive handoff documents for seamless AI agent session transfers.
MemoriLabs/Memori
Connects Claude Code to Memori Cloud for long-term memory, recalling stored context before substantive replies and saving new context afterward.
thedotmack/claude-mem
Writes a HANDOFF.md capturing goal, state, files, failed attempts and next steps so a fresh agent session can continue exactly where this one stopped.
OthmanAdi/planning-with-files
Keeps a task plan, findings and progress log in markdown files on disk so long agent tasks survive context resets, with Gemini hooks and helper scripts.
sickn33/agentic-awesome-skills
Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.
magnus919/agent-skills
Organize durable agent research outputs as summaries, analysis, and evidence dossiers.
magnus919/agent-skills
Build portable, first-person colored ASCII city engines and small GIS-derived city packs.
magnus919/agent-skills
Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.
magnus919/agent-skills
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
magnus919/agent-skills
Use Docker Compose to define, run, debug, and harden multi-container applications.
magnus919/agent-skills
Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.
Categories
Design, build, diagnose, and evolve agent harnesses: instructions, tools, execution environments, durable state, context management, verification, recovery, and bounded autonomous loops. Harness Engineering is an agent skill from magnus919/agent-skills. Design, build, diagnose, and evolve agent harnesses: instructions, tools, execution environments, durable state, context management, verification, recovery, and bounded autonomous loops.
Harness Engineering fits situations like: unreliable coding agents; cross-session drift; premature completion; changes to the runtime around a model.
Run `npx skills add magnus919/agent-skills --skill harness-engineering -a claude-code`. Or copy the skill folder (harness-engineering in magnus919/agent-skills) into .claude/skills/harness-engineering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add magnus919/agent-skills --skill harness-engineering -a codex`. Or copy the skill folder (harness-engineering in magnus919/agent-skills) into .agents/skills/harness-engineering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill harness-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-engineering, .gemini/skills/harness-engineering, .github/skills/harness-engineering and .opencode/skills/harness-engineering in your project.
Going by SKILL.md and its folder, Harness Engineering needs the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Harness Engineering is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 35k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Harness Engineering: Harness Engineering (10xChengTu/harness-engineering, 102 stars), Session Handoff (davila7/claude-code-templates, 33k stars), Memori Long-Term Memory (MemoriLabs/Memori, 17k stars) and Session Handoff Document (thedotmack/claude-mem, 99k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 115 GitHub stars. The repository holds 131 skills in this directory. The repository was last updated on October 10, 2026.
Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.