Agent skill

Eval-Driven Development Harness

by affaan-m in affaan-m/ECC

Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

MITAuto-check passedAgent Workflows

Install Eval-Driven Development Harness

skills CLI
$ npx skills add affaan-m/ECC --skill eval-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC eval-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/docs/es/skills/eval-harness .claude/skills/eval-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-harness
GitHub stars
276k
Token cost
~1.5k tokens
SKILL.md length
440 words
Files
1
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

  • Works in 7 steps: Evaluador Basado en Código → Evaluador Basado en Modelo → Evaluador Humano → …
  • Defining pass/fail criteria before building an AI-assisted workflow
  • SKILL.md covers Cuándo Activar, Filosofía, Tipos de Eval and Tipos de Evaluador, plus 7 more sections
  • Calls npm

What it does

This skill treats evals as the unit tests of AI development: define expected behavior before implementation, run evals continuously and track regressions with each change. It describes two eval types, capability evals that test whether Claude can do something new and regression evals that compare against a baseline SHA or checkpoint, each written as a short structured template with a task, success criteria or tests.

Three grader types are covered: code-based graders using deterministic checks such as grep, model-based graders where Claude judges open-ended output, and flags for human review. Reliability is measured with pass@k (at least one success in k attempts) and pass^k (all k trials succeed), with a typical target of pass@3 above 90%. The workflow runs define, implement, evaluate and report, ending in an eval report. The excerpt cuts off before the integration patterns section.

When your agent uses it

  • Defining pass/fail criteria before building an AI-assisted workflow
  • Measuring how reliably an agent completes a task over several attempts
  • Building a regression suite for prompt or agent changes
  • Comparing agent performance across model versions

Example prompts

  • “Write a capability eval for the new auth feature before we implement it.”
  • “Set up regression evals against the current commit so prompt changes can be checked.”
  • “Report pass@3 for the refactoring workflow across these ten test tasks.”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Evaluador Basado en Código
  2. Evaluador Basado en Modelo
  3. Evaluador Humano
  4. Definir (Antes de Codificar)
  5. Implementar
  6. Evaluar
  7. Reportar

What it can do on your machine

Read from SKILL.md and the folder at commit 4eb71d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval-Driven Development Harness loads about 1.5k tokens when it runs. Until then it costs about 34 tokens; SKILL.md has 440 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~34
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 4eb71d9, republished under its MIT licence (© affaan-m). 440 words, ~1,491 tokens.

Download SKILL.mdSave it as .claude/skills/eval-harness/SKILL.md (or your agent's skills folder).
name
eval-harness
description
Framework formal de evaluación para sesiones de Claude Code que implementa principios de desarrollo orientado a evals (EDD)
origin
ECC
tools
Read, Write, Edit, Bash, Grep, Glob

Skill Eval Harness

Un framework formal de evaluación para sesiones de Claude Code, implementando principios de desarrollo orientado a evals (EDD).

Cuándo Activar

  • Configurar desarrollo orientado a evals (EDD) para flujos de trabajo asistidos por IA
  • Definir criterios de pass/fail para la completitud de tareas en Claude Code
  • Medir confiabilidad del agente con métricas pass@k
  • Crear suites de pruebas de regresión para cambios de prompts o agentes
  • Comparar rendimiento del agente entre versiones de modelos

Filosofía

El Desarrollo Orientado a Evals trata los evals como las "pruebas unitarias del desarrollo de IA":

  • Definir el comportamiento esperado ANTES de la implementación
  • Ejecutar evals continuamente durante el desarrollo
  • Rastrear regresiones con cada cambio
  • Usar métricas pass@k para medición de confiabilidad

Tipos de Eval

Evals de Capacidad

Probar si Claude puede hacer algo que antes no podía:

markdown
[CAPABILITY EVAL: feature-name]
Task: Descripción de lo que Claude debe lograr
Success Criteria:
  - [ ] Criterio 1
  - [ ] Criterio 2
  - [ ] Criterio 3
Expected Output: Descripción del resultado esperado
Evals de Regresión

Asegurar que los cambios no rompan la funcionalidad existente:

markdown
[REGRESSION EVAL: feature-name]
Baseline: SHA o nombre del checkpoint
Tests:
  - existing-test-1: PASS/FAIL
  - existing-test-2: PASS/FAIL
  - existing-test-3: PASS/FAIL
Result: X/Y pasaron (anteriormente Y/Y)

Tipos de Evaluador

1. Evaluador Basado en Código

Verificaciones deterministas usando código:

bash
# Verificar si el archivo contiene el patrón esperado
grep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"

# Verificar si las pruebas pasan
npm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"

# Verificar si el build tiene éxito
npm run build && echo "PASS" || echo "FAIL"
2. Evaluador Basado en Modelo

Usar Claude para evaluar salidas de forma abierta:

markdown
[MODEL GRADER PROMPT]
Evalúa el siguiente cambio de código:
1. ¿Resuelve el problema declarado?
2. ¿Está bien estructurado?
3. ¿Se manejan los casos límite?
4. ¿El manejo de errores es apropiado?

Puntuación: 1-5 (1=pobre, 5=excelente)
Razonamiento: [explicación]
3. Evaluador Humano

Marcar para revisión manual:

markdown
[HUMAN REVIEW REQUIRED]
Cambio: Descripción de qué cambió
Razón: Por qué se necesita revisión humana
Nivel de Riesgo: BAJO/MEDIO/ALTO

Métricas

pass@k

"Al menos un éxito en k intentos"

  • pass@1: Tasa de éxito en el primer intento
  • pass@3: Éxito dentro de 3 intentos
  • Objetivo típico: pass@3 > 90%
pass^k

"Todos los k ensayos tienen éxito"

  • Barra más alta para confiabilidad
  • pass^3: 3 éxitos consecutivos
  • Usar para rutas críticas

Flujo de Trabajo de Eval

1. Definir (Antes de Codificar)
markdown
## EVAL DEFINITION: feature-xyz

### Capability Evals
1. Puede crear nueva cuenta de usuario
2. Puede validar formato de email
3. Puede hashear contraseña de forma segura

### Regression Evals
1. El login existente sigue funcionando
2. La gestión de sesiones no cambió
3. El flujo de logout está intacto

### Success Metrics
- pass@3 > 90% para evals de capacidad
- pass^3 = 100% para evals de regresión
2. Implementar

Escribir código para pasar los evals definidos.

3. Evaluar
bash
# Ejecutar evals de capacidad
[Ejecutar cada eval de capacidad, registrar PASS/FAIL]

# Ejecutar evals de regresión
npm test -- --testPathPattern="existing"

# Generar reporte
4. Reportar
markdown
EVAL REPORT: feature-xyz
========================

Capability Evals:
  create-user:     PASS (pass@1)
  validate-email:  PASS (pass@2)
  hash-password:   PASS (pass@1)
  Overall:         3/3 passed

Regression Evals:
  login-flow:      PASS
  session-mgmt:    PASS
  logout-flow:     PASS
  Overall:         3/3 passed

Metrics:
  pass@1: 67% (2/3)
  pass@3: 100% (3/3)

Status: READY FOR REVIEW

Patrones de Integración

Pre-Implementación
/eval define feature-name

Crea el archivo de definición de eval en .claude/evals/feature-name.md

Show full SKILL.md (179 more words)Show less
Durante la Implementación
/eval check feature-name

Ejecuta los evals actuales y reporta el estado

Post-Implementación
/eval report feature-name

Genera el reporte completo de eval

Almacenamiento de Evals

Almacenar evals en el proyecto:

.claude/
  evals/
    feature-xyz.md      # Definición de eval
    feature-xyz.log     # Historial de ejecuciones
    baseline.json       # Líneas base de regresión

Buenas Prácticas

  1. Definir evals ANTES de codificar — Fuerza pensar claramente sobre los criterios de éxito
  2. Ejecutar evals con frecuencia — Detectar regresiones temprano
  3. Rastrear pass@k con el tiempo — Monitorear tendencias de confiabilidad
  4. Usar evaluadores de código cuando sea posible — Determinístico > probabilístico
  5. Revisión humana para seguridad — Nunca automatizar completamente las verificaciones de seguridad
  6. Mantener los evals rápidos — Los evals lentos no se ejecutan
  7. Versionar evals con el código — Los evals son artefactos de primera clase

Guía de pass@k

  • pass@1: confiabilidad directa
  • pass@3: confiabilidad práctica bajo reintentos controlados
  • pass^3: prueba de estabilidad (las 3 ejecuciones deben pasar)

Umbrales recomendados:

  • Evals de capacidad: pass@3 >= 0.90
  • Evals de regresión: pass^3 = 1.00 para rutas críticas de release

Anti-Patrones de Eval

  • Sobreajustar prompts a ejemplos de eval conocidos
  • Medir solo salidas del camino feliz
  • Ignorar deriva de costo y latencia mientras se persiguen tasas de pass
  • Permitir evaluadores inestables en compuertas de release

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in docs/es/skills/eval-harness of affaan-m/ECC.

Open the folder on GitHubat commit 4eb71d9

Compare with similar skills

Eval-Driven Development Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval-Driven Development Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval-Driven Development Harness this skillaffaan-m/ECC276k—~1.5kAutomated safety check: PassMIT
Octocode Graph Eval Loopbgauryy/octocode949—~1.6kAutomated safety check: PassMIT
Write Skilldruxt/druxt.js114—~926Automated safety check: PassMIT
Benchmark Agentsvercel/vercel-plugin301—~3.6kAutomated safety check: PassCustom licence
Autocontextgreyhaven-ai/autocontext1.3k—~892Automated safety check: PassApache-2.0
Harness Evaltech-leads-club/agent-skills7k—~3.9kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Octocode Graph Eval Loop

    bgauryy/octocode

    Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

    949 GitHub stars~1.6k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Write Skill

    druxt/druxt.js

    Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it.

    114 GitHub stars~926 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Benchmark Agents

    vercel/vercel-plugin

    Official

    Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

    301 GitHub stars~3.6k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Autocontext

    greyhaven-ai/autocontext

    Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.

    1.3k GitHub stars~892 tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Harness Eval

    tech-leads-club/agent-skills

    Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports.

    7k GitHub stars~3.9k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote.

    6.2k GitHub stars~2.2k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed

More from affaan-m/ECC

All 683 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    276k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    276k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    276k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    276k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    276k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    276k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed

Questions about Eval-Driven Development Harness

What does Eval-Driven Development Harness do?

Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics. This skill treats evals as the unit tests of AI development: define expected behavior before implementation, run evals continuously and track regressions with each change. It describes two eval types, capability evals that test whether Claude can do something new and regression evals that compare against a baseline SHA or checkpoint, each written as a short structured template with a task, success criteria or tests.

When should I use Eval-Driven Development Harness?

Eval-Driven Development Harness fits situations like: defining pass/fail criteria before building an AI-assisted workflow; measuring how reliably an agent completes a task over several attempts; building a regression suite for prompt or agent changes; comparing agent performance across model versions.

How do I install Eval-Driven Development Harness in Claude Code?

Run `npx skills add affaan-m/ECC --skill eval-harness -a claude-code`. Or copy the skill folder (docs/es/skills/eval-harness in affaan-m/ECC) into .claude/skills/eval-harness in your project. Claude Code loads it when a task matches its description.

How do I install Eval-Driven Development Harness in Codex?

Run `npx skills add affaan-m/ECC --skill eval-harness -a codex`. Or copy the skill folder (docs/es/skills/eval-harness in affaan-m/ECC) into .agents/skills/eval-harness in your project. Codex loads it when a task matches its description.

Can I use Eval-Driven Development Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill eval-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-harness, .gemini/skills/eval-harness, .github/skills/eval-harness and .opencode/skills/eval-harness in your project.

What does Eval-Driven Development Harness need to run?

Going by SKILL.md and its folder, Eval-Driven Development Harness needs the command-line tools its instructions call (npm).

Does Eval-Driven Development Harness access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Eval-Driven Development Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval-Driven Development Harness use?

Eval-Driven Development Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval-Driven Development Harness use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval-Driven Development Harness?

Skills that share tags, products or a category with Eval-Driven Development Harness: Octocode Graph Eval Loop (bgauryy/octocode, 949 stars), Write Skill (druxt/druxt.js, 114 stars), Benchmark Agents (vercel/vercel-plugin, 301 stars) and Autocontext (greyhaven-ai/autocontext, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval-Driven Development Harness?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,111 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 10, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.