Evaluate and score agent behavior against a golden reference.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Eval

skills CLI
$ npx skills add agentevals-dev/agentevals --skill eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentevals-dev/agentevals eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentevals-dev/agentevals.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/eval .claude/skills/eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval
GitHub stars
162
Token cost
~904 tokens
SKILL.md length
396 words
Files
2
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Evaluate and score agent behavior against a golden reference.

  • Works in 4 steps: Get the file path(s). Check the extension → Ask if they have a golden eval set JSON.… → Call evaluate_traces with the file(s),… → …
  • The user wants to run evaluation
  • SKILL.md covers Determine the input type, Evaluating trace files, Evaluating sessions… and Score interpretation, plus 1 more section
  • Calls uv

What it does

Eval is an agent skill from agentevals-dev/agentevals. Evaluate and score agent behavior against a golden reference. Use this skill whenever the user wants to run evaluation, check pass/fail status, understand metric scores, compare sessions for regressions, validate agent behavior, or score a trace from a file or a live session. Trigger on phrases like "eval this trace", "check my agent output", "did my agent do the right thing", "compare runs", "did my agent regress", "score session X", "evaluate against golden", "run evals". Works with both local trace files and…

Its SKILL.md is about 900 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`).

It sits in AI & LLM Engineering, covering LLM evaluation and Observability. It works with OpenTelemetry. The repository describes itself as: agentevals is a framework-agnostic evaluations solution based on OpenTelemetry traces. The licence is Apache-2.0.

When your agent uses it

  • The user wants to run evaluation
  • Check pass/fail status
  • Understand metric scores
  • Compare sessions for regressions

Example prompts

  • “eval this trace”
  • “check my agent output”
  • “did my agent do the right thing”
  • “/eval”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Get the file path(s). Check the extension
  2. Ask if they have a golden eval set JSON. For tool_trajectory_avg_score (the
  3. Call evaluate_traces with the file(s), format, and eval set.
  4. Present results as a score table (see Score interpretation below) and explain failures.

What it can do on your machine

Read from SKILL.md and the folder at commit 94785bb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval loads about 904 tokens when it runs. Until then it costs about 137 tokens; SKILL.md has 396 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~904

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentevals-dev/agentevals at commit 94785bb, republished under its Apache-2.0 licence (© agentevals-dev). 396 words, ~904 tokens.

Download SKILL.mdSave it as .claude/skills/eval/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
eval
description
Evaluate and score agent behavior against a golden reference. Use this skill whenever the user wants to run evaluation, check pass/fail status, understand metric scores, compare sessions for regressions, validate agent behavior, or score a trace from a file or a live session. Trigger on phrases like "eval this trace", "check my agent output", "did my agent do the right thing", "compare runs", "did my agent regress", "score session X", "evaluate against golden", "run evals". Works with both local trace files and live streaming sessions.

Evaluate agent behavior and explain what the scores mean.

Determine the input type

First, figure out what to evaluate:

  • Trace file(s) — user mentions a .json or .jsonl file path → use evaluate_traces
  • Sessions vs golden — user has multiple live sessions and wants regression testing → use evaluate_sessions
  • Single live session — user wants to score one session against a golden eval set → guide them to use evaluate_sessions with one session as golden

Evaluating trace files

  1. Get the file path(s). Check the extension: .jsonl → trace_format: "otlp-json" | .json → "jaeger-json" (default)

  2. Ask if they have a golden eval set JSON. For tool_trajectory_avg_score (the default metric), an eval set is required — it provides the expected tool call sequence to compare against. If they don't have one yet, explain this and suggest starting with hallucinations_v1, or ask if they want to create a golden set from a reference run first.

  3. Call evaluate_traces with the file(s), format, and eval set.

  4. Present results as a score table (see Score interpretation below) and explain failures.

Evaluating sessions (regression testing)

This workflow requires the server to be running with the --dev flag (which enables WebSocket and session streaming). Plain agentevals serve will not have sessions. If you get a connection error from any tool below, tell the user:

bash
uv run agentevals serve --dev
  1. Call list_sessions to show available sessions.

  2. Help the user identify the "golden" session — the reference run that represents correct behavior. The server derives the eval set from it automatically.

  3. Call evaluate_sessions(golden_session_id=...). This scores all other completed sessions against the golden.

  4. Present a comparison table:

    Session             | Score | Status  | Delta
    session-abc (golden)| 1.00  | —       | baseline
    session-def         | 0.85  | PASSED  | -0.15
    session-ghi         | 0.40  | FAILED  | -0.60 ⚠️
  5. Explain regressions specifically: which tools the golden called that a failing session skipped, or unexpected extra calls. Concrete tool names are more useful than just quoting the score.

Show full SKILL.md (109 more words)Show less

Score interpretation

ScoreMeaning
1.0Exact match — right tools, right order
0.7–0.9Minor deviations (extra call or slightly different args)
0.5–0.7Partial match — some turns correct, others missing or wrong tool calls
0.0–0.5Major divergence — most tool calls don't match golden

Important: evalStatus: PASSED does not mean the agent did well — it only means the score met the configured threshold. Without a configured threshold, every session shows PASSED regardless of score. Focus on the numeric score, not the status label.

After results

If the user wants to understand what the agent did step by step (not just the score), suggest /inspect to get a readable narrative of a session.

© agentevals-dev, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .claude/skills/eval of agentevals-dev/agentevals.

  • SKILL.md
  • evals/evals.json

Open the folder on GitHubat commit 94785bb

Compare with similar skills

Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval this skillagentevals-dev/agentevals162—~904Automated safety check: PassApache-2.0
Dt Obs GenaiDynatrace/dynatrace-for-ai161—~4.5kAutomated safety check: PassApache-2.0
Improve PromptAgentX-ai/AgentX-Trace-Eval106—~2kAutomated safety check: PassCustom licence
Phoenix LLM ObservabilityOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT
RAG Observability Evalssickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: PassMIT
Exploring LLM EvaluationsPostHog/posthog-foss721—~5.7kAutomated safety check: PassMIT

Similar skills

  • Dt Obs Genai

    Dynatrace/dynatrace-for-ai

    Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.

    161 GitHub stars~4.5k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Improve Prompt

    AgentX-ai/AgentX-Trace-Eval

    Propose an improved version of a prompt registered in a self-hosted AgentX (AgentX-trace-eval) instance, using real low-rated evaluation results as evidence, then publish it as a new version once…

    106 GitHub stars~2k tokensUpdated 6 days ago
    DevOps & CloudAuto-check passed
  • Phoenix LLM Observability

    Orchestra-Research/AI-Research-SKILLs

    Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • RAG Observability Evals

    sickn33/agentic-awesome-skills

    Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.

    47k GitHub starsUsed in 2 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Exploring LLM Evaluations

    PostHog/posthog-foss

    Official

    Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment).

    721 GitHub stars~5.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Failproof AI SDK Integration

    FailproofAI/failproofai

    Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

    5.3k GitHub stars~6k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from agentevals-dev/agentevals

  • Inspect

    agentevals-dev/agentevals

    Inspect and debug live streaming agent sessions to understand what the agent did.

    162 GitHub stars~534 tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Eval

What does Eval do?

Evaluate and score agent behavior against a golden reference. Eval is an agent skill from agentevals-dev/agentevals. Evaluate and score agent behavior against a golden reference.

When should I use Eval?

Eval fits situations like: the user wants to run evaluation; check pass/fail status; understand metric scores; compare sessions for regressions.

How do I install Eval in Claude Code?

Run `npx skills add agentevals-dev/agentevals --skill eval -a claude-code`. Or copy the skill folder (.claude/skills/eval in agentevals-dev/agentevals) into .claude/skills/eval in your project. Claude Code loads it when a task matches its description.

How do I install Eval in Codex?

Run `npx skills add agentevals-dev/agentevals --skill eval -a codex`. Or copy the skill folder (.claude/skills/eval in agentevals-dev/agentevals) into .agents/skills/eval in your project. Codex loads it when a task matches its description.

Can I use Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentevals-dev/agentevals --skill eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval, .gemini/skills/eval, .github/skills/eval and .opencode/skills/eval in your project.

What does Eval need to run?

Going by SKILL.md and its folder, Eval needs the command-line tools its instructions call (uv).

Does Eval access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval use?

Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval use?

About 904 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval?

Skills that share tags, products or a category with Eval: Dt Obs Genai (Dynatrace/dynatrace-for-ai, 161 stars), Improve Prompt (AgentX-ai/AgentX-Trace-Eval, 106 stars), Phoenix LLM Observability (Orchestra-Research/AI-Research-SKILLs, 13k stars) and RAG Observability Evals (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval?

agentevals-dev (a GitHub organization) maintains it in agentevals-dev/agentevals, which has 162 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 6, 2026.

Source: agentevals-dev/agentevals on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.