Agent skill

Gan Style Harness

by affaan-m in affaan-m/ECC

GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously.

MITAuto-check passed

Install Gan Style Harness

skills CLI
$ npx skills add affaan-m/ECC --skill gan-style-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC gan-style-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gan-style-harness .claude/skills/gan-style-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gan-style-harness
GitHub stars
277k
Used in
1 other repo
Token cost
~3k tokens
SKILL.md length
930 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously.

  • Works in 6 steps: Planner Agent → Generator Agent → Evaluator Agent → …
  • A feature should be built autonomously through generator and evaluator iteration until it clears a quality bar
  • SKILL.md covers Core Insight, When to Use, When NOT to Use and Architecture, plus 8 more sections
  • Calls claude and npm

What it does

Gan Style Harness is an agent skill from affaan-m/ECC. GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously. Based on Anthropic's March 2026 harness design paper. Use when a feature should be built autonomously through generator and evaluator iteration until it clears a quality bar.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • A feature should be built autonomously through generator and evaluator iteration until it clears a quality bar

Example prompts

  • “/gan-style-harness”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Planner Agent
  2. Generator Agent
  3. Evaluator Agent
  4. Weaker Models (Sonnet-class)
  5. Capable Models (Opus 4.5-class)
  6. Frontier Models (Opus 4.6-class)

What it can do on your machine

Read from SKILL.md and the folder at commit 2d515e4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • anthropic.com
    • epsilla.com
    • martinfowler.com
    • openai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gan Style Harness loads about 3k tokens when it runs. Until then it costs about 73 tokens; SKILL.md has 930 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 2d515e4, republished under its MIT licence (© affaan-m). 930 words, ~2,998 tokens.

Download SKILL.mdSave it as .claude/skills/gan-style-harness/SKILL.md (or your agent's skills folder).
name
gan-style-harness
description
GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously. Based on Anthropic's March 2026 harness design paper. Use when a feature should be built autonomously through generator and evaluator iteration until it clears a quality bar.
metadata.origin
ECC-community
tools
Read, Write, Edit, Bash, Grep, Glob, Task

GAN-Style Harness Skill

Inspired by Anthropic's Harness Design for Long-Running Application Development (March 24, 2026)

A multi-agent harness that separates generation from evaluation, creating an adversarial feedback loop that drives quality far beyond what a single agent can achieve.

Core Insight

When asked to evaluate their own work, agents are pathological optimists — they praise mediocre output and talk themselves out of legitimate issues. But engineering a separate evaluator to be ruthlessly strict is far more tractable than teaching a generator to self-critique.

This is the same dynamic as GANs (Generative Adversarial Networks): the Generator produces, the Evaluator critiques, and that feedback drives the next iteration.

When to Use

  • Building complete applications from a one-line prompt
  • Frontend design tasks requiring high visual quality
  • Full-stack projects that need working features, not just code
  • Any task where "AI slop" aesthetics are unacceptable
  • Projects where you want to invest $50-200 for production-quality output

When NOT to Use

  • Quick single-file fixes (use standard claude -p)
  • Tasks with tight budget constraints (<$10)
  • Simple refactoring (use de-sloppify pattern instead)
  • Tasks that are already well-specified with tests (use TDD workflow)

Architecture

                    ┌─────────────┐
                    │   PLANNER   │
                    │  (Sonnet)   │
                    └──────┬──────┘
                           │ Product Spec
                           │ (features, sprints, design direction)
                           ▼
              ┌────────────────────────┐
              │                        │
              │   GENERATOR-EVALUATOR  │
              │      FEEDBACK LOOP     │
              │                        │
              │  ┌──────────┐          │
              │  │GENERATOR │--build-->│──┐
              │  │ (Sonnet) │          │  │
              │  └────▲─────┘          │  │
              │       │                │  │ live app
              │    feedback             │  │
              │       │                │  │
              │  ┌────┴─────┐          │  │
              │  │EVALUATOR │<-test----│──┘
              │  │ (Sonnet) │          │
              │  │+Playwright│         │
              │  └──────────┘          │
              │                        │
              │   5-15 iterations      │
              └────────────────────────┘

The Three Agents

1. Planner Agent

Role: Product manager — expands a brief prompt into a full product specification.

Key behaviors:

  • Takes a one-line prompt and produces a 16-feature, multi-sprint specification
  • Defines user stories, technical requirements, and visual design direction
  • Is deliberately ambitious — conservative planning leads to underwhelming results
  • Produces evaluation criteria that the Evaluator will use later

Model: Sonnet by default; raise via GAN_PLANNER_MODEL=opus for deeper spec expansion

2. Generator Agent

Role: Developer — implements features according to the spec.

Key behaviors:

  • Works in structured sprints (or continuous mode with newer models)
  • Negotiates a "sprint contract" with the Evaluator before writing code
  • Uses full-stack tooling: React, FastAPI/Express, databases, CSS
  • Manages git for version control between iterations
  • Reads Evaluator feedback and incorporates it in next iteration

Model: Sonnet by default; raise via GAN_GENERATOR_MODEL=opus for maximum coding capability

3. Evaluator Agent

Role: QA engineer — tests the live running application, not just code.

Key behaviors:

  • Uses Playwright MCP to interact with the live application
  • Clicks through features, fills forms, tests API endpoints
  • Scores against four criteria (configurable):
    1. Design Quality — Does it feel like a coherent whole?
    2. Originality — Custom decisions vs. template/AI patterns?
    3. Craft — Typography, spacing, animations, micro-interactions?
    4. Functionality — Do all features actually work?
  • Returns structured feedback with scores and specific issues
  • Is engineered to be ruthlessly strict — never praises mediocre work

Model: Sonnet by default; raise via GAN_EVALUATOR_MODEL=opus for stronger judgment + tool use

Evaluation Criteria

The default four criteria, each scored 1-10:

markdown
## Evaluation Rubric

### Design Quality (weight: 0.3)
- 1-3: Generic, template-like, "AI slop" aesthetics
- 4-6: Competent but unremarkable, follows conventions
- 7-8: Distinctive, cohesive visual identity
- 9-10: Could pass for a professional designer's work

### Originality (weight: 0.2)
- 1-3: Default colors, stock layouts, no personality
- 4-6: Some custom choices, mostly standard patterns
- 7-8: Clear creative vision, unique approach
- 9-10: Surprising, delightful, genuinely novel

### Craft (weight: 0.3)
- 1-3: Broken layouts, missing states, no animations
- 4-6: Works but feels rough, inconsistent spacing
- 7-8: Polished, smooth transitions, responsive
- 9-10: Pixel-perfect, delightful micro-interactions

### Functionality (weight: 0.2)
- 1-3: Core features broken or missing
- 4-6: Happy path works, edge cases fail
- 7-8: All features work, good error handling
- 9-10: Bulletproof, handles every edge case
Scoring
  • Weighted score = sum of (criterion_score * weight)
  • Pass threshold = 7.0 (configurable)
  • Max iterations = 15 (configurable, typically 5-15 sufficient)

Usage

Via Command
bash
# Full three-agent harness
/project:gan-build "Build a project management app with Kanban boards, team collaboration, and dark mode"

# With custom config
/project:gan-build "Build a recipe sharing platform" --max-iterations 10 --pass-threshold 7.5

# Frontend design mode (generator + evaluator only, no planner)
/project:gan-design "Create a landing page for a crypto portfolio tracker"
Via Shell Script
bash
# Basic usage
./scripts/gan-harness.sh "Build a music streaming dashboard"

# With options
GAN_MAX_ITERATIONS=10 \
GAN_PASS_THRESHOLD=7.5 \
GAN_EVAL_CRITERIA="functionality,performance,security" \
./scripts/gan-harness.sh "Build a REST API for task management"
Via Claude Code (Manual)
bash
# Step 1: Plan
claude -p --model sonnet "You are a Product Planner. Read PLANNER_PROMPT.md. Expand this brief into a full product spec: 'Build a Kanban board app'. Write spec to spec.md"

# Step 2: Generate (iteration 1)
claude -p --model sonnet "You are a Generator. Read spec.md. Implement Sprint 1. Start the dev server on port 3000."

# Step 3: Evaluate (iteration 1)
claude -p --model sonnet --allowedTools "Read,Bash,mcp__playwright__*" "You are an Evaluator. Read EVALUATOR_PROMPT.md. Test the live app at http://localhost:3000. Score against the rubric. Write feedback to feedback-001.md"

# Step 4: Generate (iteration 2 — reads feedback)
claude -p --model sonnet "You are a Generator. Read spec.md and feedback-001.md. Address all issues. Improve the scores."

# Repeat steps 3-4 until pass threshold met

Evolution Across Model Capabilities

The harness should simplify as models improve. Following Anthropic's evolution:

Stage 1 — Weaker Models (Sonnet-class)
  • Full sprint decomposition required
  • Context resets between sprints (avoid context anxiety)
  • 2-agent minimum: Initializer + Coding Agent
  • Heavy scaffolding compensates for model limitations
Stage 2 — Capable Models (Opus 4.5-class)
  • Full 3-agent harness: Planner + Generator + Evaluator
  • Sprint contracts before each implementation phase
  • 10-sprint decomposition for complex apps
  • Context resets still useful but less critical
Stage 3 — Frontier Models (Opus 4.6-class)
  • Simplified harness: single planning pass, continuous generation
  • Evaluation reduced to single end-pass (model is smarter)
  • No sprint structure needed
  • Automatic compaction handles context growth

Key principle: Every harness component encodes an assumption about what the model can't do alone. When models improve, re-test those assumptions. Strip away what's no longer needed.

Show full SKILL.md (349 more words)Show less

Configuration

Environment Variables
VariableDefaultDescription
GAN_MAX_ITERATIONS15Maximum generator-evaluator cycles
GAN_PASS_THRESHOLD7.0Weighted score to pass (1-10)
GAN_PLANNER_MODELsonnetModel for planning agent
GAN_GENERATOR_MODELsonnetModel for generator agent
GAN_EVALUATOR_MODELsonnetModel for evaluator agent
GAN_EVAL_CRITERIAdesign,originality,craft,functionalityComma-separated criteria
GAN_DEV_SERVER_PORT3000Port for the live app
GAN_DEV_SERVER_CMDnpm run devCommand to start dev server
GAN_PROJECT_DIR.Project working directory
GAN_SKIP_PLANNERfalseSkip planner, use spec directly
GAN_EVAL_MODEplaywrightplaywright, screenshot, or code-only
Evaluation Modes
ModeToolsBest For
playwrightBrowser MCP + live interactionFull-stack apps with UI
screenshotScreenshot + visual analysisStatic sites, design-only
code-onlyTests + linting + buildAPIs, libraries, CLI tools

Anti-Patterns

  1. Evaluator too lenient — If the evaluator passes everything on iteration 1, your rubric is too generous. Tighten scoring criteria and add explicit penalties for common AI patterns.

  2. Generator ignoring feedback — Ensure feedback is passed as a file, not inline. The generator should read feedback-NNN.md at the start of each iteration.

  3. Infinite loops — Always set GAN_MAX_ITERATIONS. If the generator can't improve past a score plateau after 3 iterations, stop and flag for human review.

  4. Evaluator testing superficially — The evaluator must use Playwright to interact with the live app, not just screenshot it. Click buttons, fill forms, test error states.

  5. Evaluator praising its own fixes — Never let the evaluator suggest fixes and then evaluate those fixes. The evaluator only critiques; the generator fixes.

  6. Context exhaustion — For long sessions, use Claude Agent SDK's automatic compaction or reset context between major phases.

Results: What to Expect

Based on Anthropic's published results:

MetricSolo AgentGAN HarnessImprovement
Time20 min4-6 hours12-18x longer
Cost$9$125-20014-22x more
QualityBarely functionalProduction-readyPhase change
Core featuresBrokenAll workingN/A
DesignGeneric AI slopDistinctive, polishedN/A

The tradeoff is clear: ~20x more time and cost for a qualitative leap in output quality. This is for projects where quality matters.

References

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/gan-style-harness of affaan-m/ECC.

Open the folder on GitHubat commit 2d515e4

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Gan Style Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gan Style Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gan Style Harness this skillaffaan-m/ECC277k1 repos~3kAutomated safety check: PassMIT
Generatealirezarezvani/claude-skills28k1 repos~1.1kAutomated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k1 repos~8.1kAutomated safety check: NotesMIT
Fal Generatenexu-io/open-design100k—~306Automated safety check: PassApache-2.0
Video Generationbytedance/deer-flow84k3 repos~1.4kAutomated safety check: PassMIT
LLM Evaluationdavila7/claude-code-templates33k12 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Generate

    alirezarezvani/claude-skills

    Generate Playwright tests. An agent skill from alirezarezvani/claude-skills.

    28k GitHub starsUsed in 1 repo~1.1k tokens
    Testing & QAAuto-check passed
  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 1 repo~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • Fal Generate

    nexu-io/open-design

    Generate images and videos using fal.ai AI models. An agent skill from nexu-io/open-design.

    100k GitHub stars~306 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    84k GitHub starsUsed in 3 repos~1.4k tokens
    Media & CreativeAuto-check passed
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    33k GitHub starsUsed in 12 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Image Generation

    onyx-dot-app/onyx

    Generate or edit raster images (photos, illustrations, textures, sprites, mockups, logos, infographics) using the workspace's configured image-generation provider via onyx-cli image.

    32k GitHub starsUsed in 1 repo~1.7k tokens
    Media & CreativeAuto-check passed

More from affaan-m/ECC

All 682 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    277k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    277k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    277k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    277k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    277k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Set an ECC-specific frontend design direction for production UI work.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Questions about Gan Style Harness

What does Gan Style Harness do?

GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously. Gan Style Harness is an agent skill from affaan-m/ECC. GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously.

When should I use Gan Style Harness?

Gan Style Harness fits situations like: A feature should be built autonomously through generator and evaluator iteration until it clears a quality bar.

How do I install Gan Style Harness in Claude Code?

Run `npx skills add affaan-m/ECC --skill gan-style-harness -a claude-code`. Or copy the skill folder (skills/gan-style-harness in affaan-m/ECC) into .claude/skills/gan-style-harness in your project. Claude Code loads it when a task matches its description.

How do I install Gan Style Harness in Codex?

Run `npx skills add affaan-m/ECC --skill gan-style-harness -a codex`. Or copy the skill folder (skills/gan-style-harness in affaan-m/ECC) into .agents/skills/gan-style-harness in your project. Codex loads it when a task matches its description.

Can I use Gan Style Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill gan-style-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gan-style-harness, .gemini/skills/gan-style-harness, .github/skills/gan-style-harness and .opencode/skills/gan-style-harness in your project.

What does Gan Style Harness need to run?

Going by SKILL.md and its folder, Gan Style Harness needs the command-line tools its instructions call (claude and npm).

Does Gan Style Harness access the network?

SKILL.md names 4 domains. As links in the text: anthropic.com, epsilla.com, martinfowler.com and openai.com. This is read from the text; nothing was executed.

Is Gan Style Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gan Style Harness use?

Gan Style Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gan Style Harness use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gan Style Harness?

Skills that share tags, products or a category with Gan Style Harness: Generate (alirezarezvani/claude-skills, 28k stars), Arize Evaluator (github/awesome-copilot, 40k stars), Fal Generate (nexu-io/open-design, 100k stars) and Video Generation (bytedance/deer-flow, 84k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gan Style Harness?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,673 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 11, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.