Agent skill

Agent Eval Loop

by Consensys in Consensys/c0

Run and improve the c0 agent evaluation loop using the local and remote Cloudflare MCP harness, versioned history notes, and current Cloudflare product updates.

LGPL-3.0Auto-check passedAgent Workflows

Install Agent Eval Loop

skills CLI
$ npx skills add Consensys/c0 --skill agent-eval-loop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Consensys/c0 agent-eval-loop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Consensys/c0.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/agent-eval-loop .claude/skills/agent-eval-loop && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-eval-loop
GitHub stars
105
Token cost
~1.2k tokens
SKILL.md length
534 words
Files
2
Skills in repo
11
Repo updated
First seen
Licence
LGPL-3.0

At a glance

Run and improve the c0 agent evaluation loop using the local and remote Cloudflare MCP harness, versioned history notes, and current Cloudflare product updates.

  • Works in 3 steps: Read every versioned note in… → Extract → Do not repeat failed ideas unless you…
  • Evaluating agent quality
  • SKILL.md covers Purpose, Start Here, Pick An Eval Target and Local Versus Remote, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Eval Loop is an agent skill from Consensys/c0. Run and improve the c0 agent evaluation loop using the local and remote Cloudflare MCP harness, versioned history notes, and current Cloudflare product updates. Use when evaluating agent quality, harness behavior, retrieval quality, model/tool performance, or when the user asks to benchmark, score, improve, or retest the agent workflow.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `reference.md`).

It sits in Agent Workflows, covering Agent evaluation and testing. It works with Cloudflare and Model Context Protocol. The repository describes itself as: c0 is an open-source AI platform built for organizational work, tools, and context. A single deployment to Cloudflare to get started. The licence is LGPL-3.0.

When your agent uses it

  • Evaluating agent quality
  • Harness behavior
  • Retrieval quality
  • Model/tool performance

Example prompts

  • “/agent-eval-loop”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Read every versioned note in notes/agent-eval-history/ before proposing a new improvement.
  2. Extract
  3. Do not repeat failed ideas unless you have new evidence that changes the hypothesis.

What it can do on your machine

Read from SKILL.md and the folder at commit e7a1810. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Eval Loop loads about 1.2k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 534 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Consensys/c0 at commit e7a1810, republished under its LGPL-3.0 licence (© Consensys). 534 words, ~1,198 tokens.

Download SKILL.mdSave it as .claude/skills/agent-eval-loop/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
agent-eval-loop
description
Run and improve the c0 agent evaluation loop using the local and remote Cloudflare MCP harness, versioned history notes, and current Cloudflare product updates. Use when evaluating agent quality, harness behavior, retrieval quality, model/tool performance, or when the user asks to benchmark, score, improve, or retest the agent workflow.

Agent Eval Loop

Purpose

Use this skill to run repeatable agent evaluations, compare outcomes over time, apply targeted improvements, retest, and record versioned results in notes/agent-eval-history/.

Start Here

  1. Read every versioned note in notes/agent-eval-history/ before proposing a new improvement.
  2. Extract:
    • what worked
    • what regressed behavior
    • what has already been tried
    • what still looks promising
  3. Do not repeat failed ideas unless you have new evidence that changes the hypothesis.

Pick An Eval Target

List available eval targets first:

bash
nub run agent:eval:targets

Use --target <id> to inspect one target:

bash
nub run agent:eval:targets --target linea-coordinator-runbooks

Prefer:

  • configured-ai-search for narrow retrieval quality and /mcp behavior
  • ai-search-mcp-regression for server, config, and route regressions
  • agent-investigation-loop for broader multi-step workflows

Local Versus Remote

  • nub run agent:eval:local runs the real Worker and /mcp route with a deterministic AI mock.
  • nub run agent:eval:remote runs the real Worker and live Cloudflare retrieval path.
  • Treat local runs as harness correctness signals.
  • Treat remote runs as live quality signals.
  • Never treat a passing local run as proof that live retrieval or live agent quality is correct.

Required Metrics

Score every eval on a 1-5 scale:

  • relevance
  • helpfulness
  • correctness

Also record:

  • taskDurationMs
  • local run duration when available
  • remote run duration when available

Rubric:

  • 1: unusable or misleading
  • 2: weak, incomplete, or mostly off-target
  • 3: acceptable but missing important value
  • 4: strong and useful with minor issues
  • 5: highly effective, correct, and ready to trust

Web Research

Each eval loop should search the web for recent Cloudflare changes that may improve the harness or implementation:

  • Agents SDK
  • AI Search
  • Vectorize
  • Workers AI
  • MCP- and Workers-related testing improvements

Capture:

  • title
  • URL
  • why it matters for the current eval target

Core Loop

  1. Read prior notes and summarize the important lessons.
  2. Choose a target and define success criteria.
  3. Run the relevant local and/or remote harness commands.
  4. Review outputs and score the metrics.
  5. If the user wants improvement work, make the smallest high-signal change that addresses the current hypothesis.
  6. Retest using the same target.
  7. Compare metrics and raw evidence.
  8. Record a versioned history note.
Show full SKILL.md (195 more words)Show less

Use The Helper Scripts

Create a machine-readable summary first:

bash
nub run agent:eval:metrics --input /tmp/agent-eval-summary.json --output /tmp/agent-eval-summary.json

Preview the next note path:

bash
nub run agent:eval:note:next --recap "bootstrap-agent-eval-workflow"

Write the versioned note and artifact:

bash
nub run agent:eval:note:write --input /tmp/agent-eval-summary.json --recap "bootstrap-agent-eval-workflow"

Summary Payload Expectations

The input JSON for agent:eval:metrics should include:

  • targetId
  • taskSummary
  • workflowType
  • recap
  • local
  • remote
  • metrics
  • baselineObservations
  • changesAttempted
  • worked
  • didntWork
  • nextHypotheses
  • webUpdates
  • rawEvidence
  • existingHistoryReviewed
  • sourceCommands

See reference.md for a complete example payload and note shape.

Output Requirements

For each eval run, produce:

  • a concise summary of findings
  • exact commands run
  • the metric scores
  • whether the result came from local or remote harness coverage
  • the note path written under notes/agent-eval-history/
  • the next best hypothesis

Improvement Guidance

  • Prefer one focused change per loop when validating a hypothesis.
  • Keep the before/after comparison tight.
  • If a change helps locally but not remotely, say so explicitly.
  • If a Cloudflare product update suggests a new approach, reference it in both the write-up and the next hypothesis.

Stop Conditions

Stop the loop when:

  • the user only asked for evaluation, not changes
  • the remote run is blocked by missing credentials or infrastructure access
  • the next change would be speculative without stronger evidence
  • you have already reached a clear conclusion for the current target

Manual Invocation

Use .cursor/commands/run-agent-eval.md for a repeatable chat entrypoint.

© Consensys, LGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/agent-eval-loop of Consensys/c0.

  • SKILL.md
  • reference.md

Open the folder on GitHubat commit e7a1810

Compare with similar skills

Agent Eval Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Eval Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Eval Loop this skillConsensys/c0105—~1.2kAutomated safety check: PassLGPL-3.0
MCP Server Builderanthropics/skills180k62 repos~2.3kAutomated safety check: PassApache-2.0
Agents SDKcloudflare/skills3k2 repos~3kAutomated safety check: PassApache-2.0
Autocontext for Hermesgreyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.0
Turnstile Testingjumodada/Drissionpage-MCP-Server487—~1.7kAutomated safety check: PassCustom licence
Memoraagentic-box/memora731—~813Automated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 62 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Agents SDK

    cloudflare/skills

    Official

    Build, debug, or review Cloudflare Agents SDK applications using the agents package.

    3k GitHub starsUsed in 2 repos~3k tokens
    Agent WorkflowsAuto-check passed
  • Autocontext for Hermes

    greyhaven-ai/autocontext

    Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

    1.3k GitHub stars~2.5k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Turnstile Testing

    jumodada/Drissionpage-MCP-Server

    A skill your agent uses when testing an authorized Cloudflare Turnstile integration or operating an authorized production challenge with drissionpage-mcp.

    487 GitHub stars~1.7k tokensUpdated 23 days ago
    Agent WorkflowsAuto-check passed
  • Memora

    agentic-box/memora

    A skill your agent uses when working with persistent memory across sessions, storing/retrieving knowledge, managing TODOs/issues, or when context from previous sessions would be helpful.

    731 GitHub stars~813 tokensUpdated 9 days ago
    Agent WorkflowsAuto-check passed
  • Skill Forge

    AgriciDaniel/skill-forge

    Ultimate Claude Code skill creator and architect. An agent skill from AgriciDaniel/skill-forge.

    177 GitHub stars~1.9k tokensUpdated 6 mo ago
    Agent WorkflowsAuto-check: notes

More from Consensys/c0

All 11 skills in this repo
  • Debug deployed Cloudflare Workers using the cfobservability MCP, Wrangler, D1/R2 state, repo evidence, and safe live reproduction.

    105 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check: notes
  • Deslop Typescript

    Consensys/c0

    Run a final de-slopping pass on nearly finished JavaScript or TypeScript work before commit or PR.

    105 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • A skill your agent uses when adding, changing, or reviewing Workflow Nodes in the c0 agent repo, including shared node catalog metadata, Workflow Node options and ports, runtime node Adapters…

    105 GitHub stars~943 tokensUpdated 1 mo ago
    Auto-check passed
  • Guidance for safely changing the c0 Workflow runtime ABI, manifest ABI, runtime kernels, manifest migrations, workflow artifact compatibility, and audit/backfill tooling.

    105 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Alchemy Env Types

    Consensys/c0

    Derive Cloudflare worker env types from Alchemy resource Env types instead of hand-writing env shapes.

    105 GitHub stars~707 tokensUpdated 1 mo ago
    Auto-check passed
  • C0 Config

    Consensys/c0

    A skill your agent uses when changing c0 deployment configuration, runtime-editable global settings, auth providers, AI providers, MCP Context Forge, discovered registries, secret references, or…

    105 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check: notes

Categories

Questions about Agent Eval Loop

What does Agent Eval Loop do?

Run and improve the c0 agent evaluation loop using the local and remote Cloudflare MCP harness, versioned history notes, and current Cloudflare product updates. Agent Eval Loop is an agent skill from Consensys/c0. Run and improve the c0 agent evaluation loop using the local and remote Cloudflare MCP harness, versioned history notes, and current Cloudflare product updates.

When should I use Agent Eval Loop?

Agent Eval Loop fits situations like: evaluating agent quality; harness behavior; retrieval quality; model/tool performance.

How do I install Agent Eval Loop in Claude Code?

Run `npx skills add Consensys/c0 --skill agent-eval-loop -a claude-code`. Or copy the skill folder (.agents/skills/agent-eval-loop in Consensys/c0) into .claude/skills/agent-eval-loop in your project. Claude Code loads it when a task matches its description.

How do I install Agent Eval Loop in Codex?

Run `npx skills add Consensys/c0 --skill agent-eval-loop -a codex`. Or copy the skill folder (.agents/skills/agent-eval-loop in Consensys/c0) into .agents/skills/agent-eval-loop in your project. Codex loads it when a task matches its description.

Can I use Agent Eval Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Consensys/c0 --skill agent-eval-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-eval-loop, .gemini/skills/agent-eval-loop, .github/skills/agent-eval-loop and .opencode/skills/agent-eval-loop in your project.

What does Agent Eval Loop need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Eval Loop is instructions for the agent only.

Does Agent Eval Loop access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Eval Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Eval Loop use?

Agent Eval Loop is published under the LGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Eval Loop use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Eval Loop?

Skills that share tags, products or a category with Agent Eval Loop: MCP Server Builder (anthropics/skills, 180k stars), Agents SDK (cloudflare/skills, 3k stars), Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars) and Turnstile Testing (jumodada/Drissionpage-MCP-Server, 487 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Eval Loop?

Consensys (a GitHub organization) maintains it in Consensys/c0, which has 105 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on August 18, 2026.

Source: Consensys/c0 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.