Agent skill

Harness Design

by AnastasiyaW in AnastasiyaW/codex-claude-code-config

Design a scoped multi-agent harness for a requested long-running AI workflow.

MITAuto-check passedAgent Workflows

Install Harness Design

skills CLI
$ npx skills add AnastasiyaW/codex-claude-code-config --skill harness-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AnastasiyaW/codex-claude-code-config harness-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AnastasiyaW/codex-claude-code-config.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/architecture/harness-design .claude/skills/harness-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harness-design
GitHub stars
154
Token cost
~2.6k tokens
SKILL.md length
921 words
Files
1
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Design a scoped multi-agent harness for a requested long-running AI workflow.

  • Works in 3 steps: Planner (Планировщик) → Generator (Генератор) → Evaluator (Оценщик)
  • : build a harness
  • SKILL.md covers Когда нужен harness, а когда…, Архитектура: Three-Agent System, Sprint Contract Pattern and Generator-Evaluator: Почему…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Harness Design is an agent skill from AnastasiyaW/codex-claude-code-config. Design a scoped multi-agent harness for a requested long-running AI workflow. Preserve the user-approved product boundary while using Generator-Evaluator separation, testable sprint contracts, context management, and task-appropriate validation. Based on Anthropic Engineering patterns. Use when: "build a harness", "multi-agent architecture", "agent orchestration", "generator-evaluator", "long-running app", "harness design", "agent pipeline", "quality evaluation loop", "sprint contract", "build app with agents"…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Multi-agent orchestration, Load testing and Context engineering. It works with Claude Agent SDK. The repository describes itself as: Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development. The licence is MIT.

When your agent uses it

  • : build a harness
  • Multi-agent architecture
  • Agent orchestration
  • Generator-evaluator

Example prompts

  • “build a harness”
  • “multi-agent architecture”
  • “agent orchestration”
  • “/harness-design”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Planner (Планировщик)
  2. Generator (Генератор)
  3. Evaluator (Оценщик)

What it can do on your machine

Read from SKILL.md and the folder at commit 67709af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harness Design loads about 2.6k tokens when it runs. Until then it costs about 232 tokens; SKILL.md has 921 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~232
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AnastasiyaW/codex-claude-code-config at commit 67709af, republished under its MIT licence (© AnastasiyaW). 921 words, ~2,558 tokens.

Download SKILL.mdSave it as .claude/skills/harness-design/SKILL.md (or your agent's skills folder).
name
harness-design
description
Design a scoped multi-agent harness for a requested long-running AI workflow. Preserve the user-approved product boundary while using Generator-Evaluator separation, testable sprint contracts, context management, and task-appropriate validation. Based on Anthropic Engineering patterns. Use when: "build a harness", "multi-agent architecture", "agent orchestration", "generator-evaluator", "long-running app", "harness design", "agent pipeline", "quality evaluation loop", "sprint contract", "build app with agents", "Claude Agent SDK architecture", or when building complex full-stack apps that need planning → generation → evaluation cycles. Also use when discussing context degradation, self-evaluation bias, or assumption testing in AI workflows. Do NOT use to stress-test or critique an already-written plan document; use plan-swarm-review for that (this skill designs the harness, it does not review plans).
user-invocable
true
model
opus

Multi-Agent Harness Design

Источники:

  • Anthropic Engineering — "Harness design for long-running apps"
  • OpenClaw-RL paper (arxiv 2603.10165) — personal agent verification
  • DenisSergeevitch/repo-task-proof-loop — execution protocol with durable proof

См. также: references/proof-loop-research.md — детали paper + repo mapping

Когда нужен harness, а когда хватит solo agent

СигналSolo agentHarness
ScopeОдна фича, bug fix, refactorFull-stack app, multi-feature product
ДлительностьBounded work with direct evidenceWork that needs durable state, resumes, or independent evaluation
КачествоBaseline достаточноНужен polish, originality, craft
СтоимостьВ рамках явного бюджета задачиБюджет и stop condition фиксируются для этой задачи
ПроверкаПрямой task-appropriate proofНезависимая проверка и runtime proof, если требуются контрактом

Правило: Evaluator оправдан когда задача за пределами reliable solo performance. Не фиксированное yes/no — зависит от complexity tier.


Архитектура: Three-Agent System

1. Planner (Планировщик)
  • Превращает запрос пользователя в детальную, проверяемую спецификацию
  • Сохраняет утверждённый scope, явные non-goals и существующие границы продукта
  • Не добавляет AI-фичи, миграции стека или новые deliverables без запроса пользователя либо подтверждённой причинной необходимости
  • НЕ over-specify реализацию — только what, не how
  • Если AI-фичи уже входят в утверждённый scope, вписывает их в продукт органично
2. Generator (Генератор)
  • Реализует фичи итеративно
  • Включает self-evaluation перед handoff (но она ненадёжна — см. ниже)
  • Работает в рамках Sprint Contract
3. Evaluator (Оценщик)
  • Независимый от генератора — отдельный контекст, отдельный промпт
  • Выбирает проверку по acceptance contract: unit/integration/runtime probe для сервисов, browser/UI проверку только для затронутого пользовательского пути
  • Откалиброван через few-shot примеры
  • Ловит то, что self-evaluation пропускает

Sprint Contract Pattern

Перед каждой итерацией:

1. Planner определяет фичу и user story
2. Generator и Evaluator ДОГОВАРИВАЮТСЯ о:
   - Что значит "done" для этой фичи
   - Конкретные testable success criteria
   - Что НЕ входит в scope
3. Generator реализует
4. Evaluator валидирует по контракту
5. Если конкретный критерий не пройден → конкретный feedback → повтор с п.3 только для этого критерия

Контракт = мост между user stories и implementation. Без него evaluator судит по своим критериям, generator не знает что проверять.

Останавливай цикл, когда все утверждённые criteria имеют требуемое доказательство. Не назначай число раундов заранее; продолжай только при наблюдаемом незакрытом критерии или новом опровергающем evidence.


Generator-Evaluator: Почему раздельно

Self-evaluation bias

Модели уверенно хвалят свою работу — даже когда качество посредственное. Это не баг модели, а свойство: генератор оптимизирован на producing, не на judging.

Решение: Independent evaluator
  • Другой system prompt с calibrated skepticism
  • Few-shot примеры с детальными score breakdowns
  • Проверяет исполнимую поверхность acceptance contract: browser только для затронутого UI journey; native/API/worker задачи — их unit, integration, CLI или runtime probe, а не browser по умолчанию
  • Конкретные failure criteria, а не общие "looks good"
Калибровка оценщика (QA Tuning Loop)
1. Evaluator выдаёт оценку
2. Ты проверяешь: согласен ли с оценкой?
3. Расхождение → обновляешь QA промпт
4. Типичные проблемы:
   - Superficial testing, пропускает edge cases
   - Premature approval посредственной работы
   - Слишком строгие критерии → бесконечные итерации
5. Повторяешь пока evaluator judgment ≈ твой judgment

Quality Criteria Framework (для фронтенда)

4 измерения, каждое 0-10:

1. Design Quality — Целостность

Дизайн ощущается как единое целое, а не коллекция частей?

  • Интеграция color, typography, layout, imagery
  • Consistent visual language

2. Originality — Уникальность

Штраф за:

  • Template layouts, library defaults
  • AI slop patterns: purple gradients over white cards
  • "Telltale signs of AI generation"
  • Cookie-cutter структуры

3. Craft — Техническое мастерство

  • Typography hierarchy
  • Spacing consistency
  • Color harmony, contrast ratios
  • Pixel-perfect alignment

4. Functionality — Работоспособность

Пользователь завершает задачу без угадывания?

  • Все интерактивные элементы работают
  • Нет stub features
  • Error states обработаны
Влияние формулировок на генерацию

Фразы в criteria прямо влияют на вывод генератора:

  • "museum quality" → visual convergence к одному стилю
  • "best designs" → перфекционизм за счёт creativity
  • Тестируй формулировки — они стируют модель ДО оценки

Контекст-менеджмент

Context Degradation

Модели теряют coherence по мере заполнения context window.

Context reset > Compaction:

  • Compaction сохраняет continuity, но не даёт чистый лист
  • Reset + structured handoff artifact = лучший баланс
  • Handoff artifact = документ с state, decisions, progress
Context Anxiety

Модели могут сворачивать работу раньше времени из-за роста контекста.

  • Решение: clean context resets, когда они уменьшают риск потери важных ограничений
Structured Handoff

При context reset передавать:

- Что уже сделано (с конкретными файлами/строками)
- Какие решения приняты и почему
- Что осталось сделать
- Текущие проблемы и blockers
- Sprint contract для текущей итерации

Assumption Testing

"Every component in a harness encodes an assumption about what the model can't do on its own"

Show full SKILL.md (363 more words)Show less
Принцип: предположения устаревают
  • Возможности моделей меняются, поэтому необходимость каждого harness-компонента должна подтверждаться наблюдаемым риском
  • Стратегия: при наличии безопасного измеримого эксперимента убирать компоненты по одному, измерять влияние и откатывать ухудшение
Simplification Loop
1. Текущий harness работает? Да →
2. Убери один компонент (напр. sprint decomposition)
3. Качество упало? Да → верни. Нет →
4. Повтори с другим компонентом
5. Остановись на минимальном harness для текущей задачи

Реальные failure modes (пойманные evaluator'ом)

  • Rectangle fill tool ставит тайлы только на endpoints drag, вместо заполнения области
  • Delete key handler требует два условия, когда нужно одно
  • FastAPI route matching: "reorder" матчится как integer frame_id
  • Audio recording: stub без mic capture
  • Missing clip resize/split operations
  • Effect visualizations как числовые слайдеры вместо графики
  • Display-only features без интерактивности
  • Missing instrument panels
  • Unimplemented recording functionality

Инструментарий

Claude Agent SDK
  • Handles agent orchestration + compaction автоматически
  • Manages context growth across long sessions
  • Рекомендуемый стек для production harnesses
Проверочные инструменты
  • Для UI: browser automation и screenshots проверяют реальный затронутый user journey
  • Для API, очередей и данных: contract/integration tests, health/runtime probes и проверка нужных persistent states
  • Используй инструменты, уже поддержанные проектом; не меняй стек ради соответствия этой skill
Иллюстративные технологии

React/Vite, Nuxt/Vue, FastAPI/Fastify, SQLite/PostgreSQL и Playwright — примеры веб-проектов на момент написания, а не default, требование или инструкция к миграции. Выбирай технологии по существующей архитектуре, пользовательскому запросу и текущей документации.


Gotchas

  • Language shapes output: формулировки в criteria сдвигают генератор ДО обратной связи от оценщика. "Museum quality" → convergence, "experimental" → divergence
  • Не обосновывай продолжение цикла номером итерации: новая итерация нужна только при незакрытом criterion или новом evidence
  • Cost/time figures are dated examples, not gates: если бюджет нужен, согласуй task-specific limit и не понижай acceptance ради формального PASS
  • Evaluator may need tuning: меняй QA prompt только в ответ на конкретный false positive/negative и повторно проверяй затронутый criterion
  • Self-evaluation is seductive: генератор БУДЕТ говорить "всё отлично" — не верь, проверяй через independent evaluator

Troubleshooting

СимптомПричинаРешение
Evaluator всё одобряетПромпт слишком мягкийДобавь few-shot с detailed score breakdowns, конкретные failure criteria
Generator не улучшаетсяFeedback слишком абстрактныйEvaluator должен давать конкретные файлы/строки/проблемы
Бесконечные итерацииCriteria невыполнимы или feedback не связан с нимиПересмотри контракт; сохраняй требуемую safety/runtime планку, а несовместимый scope split только с явным решением пользователя
Context degradationДлинная сессия без resetStructured handoff + clean context reset
Все итерации выглядят одинаковоCriteria слишком узкиеРасширь пространство, убери "museum quality" формулировки
Evaluator ловит мелочи, пропускает крупноеWrong priority в промптеRestructure: critical → high → medium → cosmetic

© AnastasiyaW, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/architecture/harness-design of AnastasiyaW/codex-claude-code-config.

Open the folder on GitHubat commit 67709af

Compare with similar skills

Harness Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harness Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harness Design this skillAnastasiyaW/codex-claude-code-config154—~2.6kAutomated safety check: PassMIT
Project Developmentguanyang/open-agent-hub9772 repos~4.7kAutomated safety check: PassMIT
Harness Engineering10xChengTu/harness-engineering1021 repos~1kAutomated safety check: PassNone
Cozempic Session GuardRuya-AI/cozempic425—~434Automated safety check: PassMIT
Filesystem Contextguanyang/open-agent-hub9771 repos~4kAutomated safety check: PassMIT
Durable Session StateZaxbyHub/opencode-swarm496—~896Automated safety check: PassMIT

Similar skills

  • Project Development

    guanyang/open-agent-hub

    This skill should be used for project-level decisions about LLM-powered systems: whether an LLM is the right primitive for the task at hand, the shape of a multi-stage batch or agent pipeline, token…

    977 GitHub starsUsed in 2 repos~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Harness Engineering

    10xChengTu/harness-engineering

    Set up and improve harness engineering (AGENTS.md, docs/, lint rules, eval systems, project-level prompt engineering) for AI-agent-friendly codebases.

    102 GitHub starsUsed in 1 repo~1k tokens
    Agent WorkflowsAuto-check passed
  • Cozempic Session Guard

    Ruya-AI/cozempic

    Starts a background daemon that watches a Claude Code session's size and prunes it before auto-compaction can discard context or agent-team state.

    425 GitHub stars~434 tokensUpdated 3 mo ago
    Agent WorkflowsAuto-check passed
  • Filesystem Context

    guanyang/open-agent-hub

    This skill should be used when agent work needs file-backed context: durable scratchpads, tool-output offloading, just-in-time discovery, cross-agent handoff files, filesystem memory, or cleanup…

    977 GitHub starsUsed in 1 repo~4k tokens
    Agent WorkflowsAuto-check passed
  • Durable Session State

    ZaxbyHub/opencode-swarm

    Keeps plans, decisions, evidence and reviewer verdicts in small files so long multi-phase tasks survive context compaction and session resumes.

    496 GitHub stars~896 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Gives a fast orientation to OpenRig for an agent that just booted into a seat, covering rigs, topologies, layers and where context and skills come from.

    6.8k GitHub stars~4k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from AnastasiyaW/codex-claude-code-config

All 50 skills in this repo
  • Bug Reproducer

    AnastasiyaW/codex-claude-code-config

    Find likely software bugs in a codebase, rank concrete bug candidates, and prove or reject them with focused regression tests before proposing a fix.

    154 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Motion Framer

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when implementing Motion or Framer Motion in React/JavaScript: interactive UI components, micro-interactions, gestures, layout or page transitions, and scroll-based animation.

    154 GitHub starsUsed in 1 repo~5.2k tokens
    Auto-check passed
  • Proof Verify

    AnastasiyaW/codex-claude-code-config

    Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

    154 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Workflow Orchestration

    AnastasiyaW/codex-claude-code-config

    Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

    154 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Notebooklm Grounded Research

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when: NotebookLM, notebooklm MCP, large documentation sets, courses, books, papers, or citation-backed research are mentioned.

    154 GitHub stars~2.4k tokensUpdated today
    Auto-check: warnings
  • Deepseek Provider Contract

    AnastasiyaW/codex-claude-code-config

    Validate a proposed DeepSeek API integration before any key or project context is sent: check thinking-mode tool-call history, strict-schema assumptions, bounded output, and provider data boundaries.

    154 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Categories

Questions about Harness Design

What does Harness Design do?

Design a scoped multi-agent harness for a requested long-running AI workflow. Harness Design is an agent skill from AnastasiyaW/codex-claude-code-config. Design a scoped multi-agent harness for a requested long-running AI workflow.

When should I use Harness Design?

Harness Design fits situations like: : build a harness; multi-agent architecture; agent orchestration; generator-evaluator.

How do I install Harness Design in Claude Code?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill harness-design -a claude-code`. Or copy the skill folder (skills/architecture/harness-design in AnastasiyaW/codex-claude-code-config) into .claude/skills/harness-design in your project. Claude Code loads it when a task matches its description.

How do I install Harness Design in Codex?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill harness-design -a codex`. Or copy the skill folder (skills/architecture/harness-design in AnastasiyaW/codex-claude-code-config) into .agents/skills/harness-design in your project. Codex loads it when a task matches its description.

Can I use Harness Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AnastasiyaW/codex-claude-code-config --skill harness-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harness-design, .gemini/skills/harness-design, .github/skills/harness-design and .opencode/skills/harness-design in your project.

What does Harness Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Harness Design is instructions for the agent only.

Does Harness Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Harness Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Harness Design use?

Harness Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harness Design use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Harness Design?

Skills that share tags, products or a category with Harness Design: Project Development (guanyang/open-agent-hub, 977 stars), Harness Engineering (10xChengTu/harness-engineering, 102 stars), Cozempic Session Guard (Ruya-AI/cozempic, 425 stars) and Filesystem Context (guanyang/open-agent-hub, 977 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harness Design?

AnastasiyaW (a GitHub user) maintains it in AnastasiyaW/codex-claude-code-config, which has 154 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: AnastasiyaW/codex-claude-code-config on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.