Agent skill

Eval Report

by agentscope-ai in agentscope-ai/OpenJudge

A skill your agent uses when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive…

Apache-2.0Auto-check passedWriting & Content

Install Eval Report

skills CLI
$ npx skills add agentscope-ai/OpenJudge --skill eval-report -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/OpenJudge eval-report --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval_pipeline/04-eval-report .claude/skills/eval-report && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-report
GitHub stars
870
Token cost
~2.4k tokens
SKILL.md length
819 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive…

  • Works in 7 steps: Inventory Scan → Maturity Assessment → Cross-Skill Signal Synthesis → …
  • The user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment
  • SKILL.md covers When to Activate, Checklist, Step 1: Inventory Scan and Step 2: Maturity Assessment, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Report is an agent skill from agentscope-ai/OpenJudge. Use when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive summary. Also use when the user mentions eval health check, evaluation audit, ship readiness, evaluation maturity, or "how good is my evaluation system itself." This is a read-only analysis skill.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Writing & Content, covering Summarization. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.

When your agent uses it

  • The user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment
  • Cross-skill signals
  • Prioritized actions
  • An executive summary

Example prompts

  • “how good is my evaluation system itself.”
  • “/eval-report”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Inventory Scan
  2. Maturity Assessment
  3. Cross-Skill Signal Synthesis
  4. Weakness Diagnosis
  5. Root Cause Classification
  6. Prioritized Recommendations
  7. Executive Summary

What it can do on your machine

Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Report loads about 2.4k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 819 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 819 words, ~2,437 tokens.

Download SKILL.mdSave it as .claude/skills/eval-report/SKILL.md (or your agent's skills folder).
name
eval-report
description
Use when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive summary. Also use when the user mentions eval health check, evaluation audit, ship readiness, evaluation maturity, or "how good is my evaluation system itself." This is a read-only analysis skill.
<HARD-GATE>
NO recommendation WITHOUT statistical evidence backing it.
NO "system ready" declaration WITHOUT all calibrated judges passing AND all production gates green.
NO trend analysis WITHOUT at least 2 data points in history.
</HARD-GATE>

Eval Report

Synthesize everything from your evaluation journey into a comprehensive report. This skill is read-only — it analyzes what exists, doesn't create new graders or datasets.

When to Activate

  • You've run 2+ evaluation skills and want the big picture
  • You need to report evaluation status to non-technical stakeholders
  • You're making a ship/no-ship decision and need evidence
  • The evaluation system has been running for a while — time for a health check

Checklist

You MUST create a task for each item and complete them in order:

  1. Inventory scan — catalog everything in eval-design.md + runs/ history
  2. Maturity assessment — 5 dimensions × 4 levels
  3. Cross-skill signal synthesis — consistent findings + contradictions
  4. Weakness diagnosis — failure concentration, correlations, stratum gaps
  5. Root cause classification — system / metric / data / unclear
  6. Prioritized recommendations — P0/P1/P2 actions with impact estimates
  7. Executive summary — ship readiness + top 3 risks + next actions

Step 1: Inventory Scan

Read eval-design.md and all runs/ directories. Build a timeline:

Timeline:
  2026-04-15  01-eval-design   → 5 failure modes → 3 dimensions from 200 traces
  2026-04-18  02-metric-design → 4 graders configured (2 LLM + 1 rule + 1 executable)
  2026-04-25  (evaluation run) → 90-sample stratified dataset scored
  2026-05-01  03-align-human   → 2 judges Phase 3, 1 Phase 2, 1 Phase 1 (TPR/TNR + kappa)
  2026-05-10  07-redteam       → safety audit not yet run

Report key metrics:

  • Total skills run, total principles, total labels
  • Calibrated judges: X of Y (with TPR/TNR range)
  • Last activity date per skill

Step 2: Maturity Assessment

Rate the evaluation system across 5 dimensions:

DimensionL1 (Initial)L2 (Developing)L3 (Established)L4 (Optimizing)
Failure DiscoveryNo systematic analysisFailure modes identifiedCoverage validated with stratificationContinuous triage from production
Judge Qualityv0 uncalibrated onlySome calibrated (TPR/TNR measured)All calibrated with CICalibrated + aligned with humans
Label Coverage< 50 labels50-200 labels200+ stratified labelsCoverage audit passed, drift monitored
Safety CoverageNo redteamingAd-hoc redteam runSystematic redteam with policy docContinuous redteam with sign-off
Human AlignmentNo alignment dataKappa measured for some judgesKappa ≥ 0.8 for all judgesHuman spot-check only, quarterly audit

Scoring rule: The overall maturity level is the minimum across dimensions (weakest link principle). If 4 dimensions are L3 but Safety is L1, the system is L1.

Step 3: Cross-Skill Signal Synthesis

Consistent Signals (high confidence)

Find themes confirmed by multiple skills. Example:

  • "Factuality is the top risk" — evidence chain:
    • 01-eval-design: #1 failure mode (38% prevalence in traces)
    • 02-metric-design: weighted as the highest-impact dimension
    • 03-align-human: TPR=0.92 TNR=0.88 (confirmed measurable)
Contradictions (needs investigation)

Find where skills disagree. These are the most valuable findings:

  • "02-metric-design weighted hallucination as a top signal, but 03-align-human shows the hallucination judge has TPR=0.74" → Possible explanations: the judge prompt captures surface patterns, not real hallucination. Or the judge prompt needs refinement, or the labels are noisy.
Coverage Gaps (blind spots)

What hasn't been touched by any skill?

  • "01-eval-design coverage shows multilingual input_type n=0, no workflow has addressed non-English queries"

Step 4: Weakness Diagnosis

Failure Concentration

Which principle/grader has the lowest pass rate? Where are failures clustering?

Failure Correlation

Compute Jaccard similarity between principle pairs — when sample A fails on principle X, does it also fail on principle Y? Highly correlated pairs (Jaccard > 0.5) likely share a root cause.

Show full SKILL.md (314 more words)Show less
Per-Stratum Weakness

Which difficulty stratum performs worst across all principles? If boundary stratum TPR < 0.7 for 3 of 4 principles, boundary discrimination is a systemic weakness.

Step 5: Root Cause Classification

For each weakness area, classify the root cause:

TypeDefinitionKey indicator
system_problemThe application itself performs poorlyLow pass rate + high judge-human agreement
metric_problemThe judge/eval is flawedLow pass rate + low judge-human agreement
data_problemThe eval dataset isn't representative01-eval-design coverage shows thin strata OR label drift detected
unclearNot enough evidenceConflicting signals, need more data

This classification is critical — fixing a metric problem by changing the system (or vice versa) wastes effort.

Step 6: Prioritized Recommendations

Generate P0/P1/P2 actions. Each must include: priority, concrete action, current state, target state, expected impact, and the skill to use.

🔴 P0 | Calibrate hallucination judge
       Current: TPR=0.74 (below 0.8 threshold)
       Target:  TPR >= 0.8
       Impact:  Judge becomes usable as production gate
       Use:     03-align-human

🔴 P0 | Add boundary samples for hallucination
       Current: n=6 boundary samples (CI half-width ±18%)
       Target:  n >= 20 (CI narrows to ±10%)
       Impact:  Reliable per-stratum TPR measurement
       Use:     01-eval-design

🟡 P1 | Align tone_consistency judge
       Current: kappa=0.72, bias=-0.15 (lenient)
       Target:  kappa >= 0.8, |bias| < 0.1
       Impact:  Reduce false positive rate ~15%
       Use:     03-align-human

🟢 P2 | Enable auto-gate for factuality judge
       Current: kappa=0.87, TPR=0.92, TNR=0.88
       Impact:  Eliminate 90% of human review for this dimension
       Use:     03-align-human (mark Phase 3)

Step 7: Executive Summary

One page for non-technical stakeholders:

Executive Summary
=================
System: Customer support chatbot for e-commerce
Stakes: Production
Report Date: 2026-05-12

SHIP READINESS: Conditional
  2 items must be resolved before production gate:
  1. Hallucination judge TPR below threshold (0.74 < 0.8)
  2. No safety/redteam evaluation has been run

TOP 3 RISKS:
  1. Hallucination detection unreliable — severity: HIGH
     The judge measuring whether the bot fabricates information itself has poor
     recall (TPR=0.74), meaning ~26% of hallucinations go undetected.
     Mitigation: Calibrate with more boundary labels (2-3 weeks).

  2. Safety coverage missing — severity: MEDIUM
     No jailbreak, injection, or PII leakage testing has been performed.
     Mitigation: Run redteam skill this sprint (1-2 days).

  3. Tone evaluation is biased lenient — severity: LOW
     The tone judge systematically rates responses as better than humans do.
     Mitigation: Refine judge prompt with borderline examples.

EVAL MATURITY: L2 (Developing) → Target L3 in 3-4 weeks
  Strongest: Failure Discovery (L3)
  Weakest:  Safety Coverage (L1), Human Alignment (L2)

NEXT ACTIONS (this sprint):
  [P0] Add boundary labels + recalibrate hallucination judge
  [P0] Run initial redteam evaluation
  [P1] Refine tone judge alignment

Output

FileContent
eval-design.mdeval_report: namespace (maturity, signals, trends, actions)
runs/eval-report/<ts>/report.mdFull analysis report
runs/eval-report/<ts>/executive-summary.mdOne-page stakeholder summary

Each per-metric run you synthesize should be a runs/<skill>/<ts>/results.json row of the form:

python
{"metric": "order_accuracy", "mean": 0.82, "ci_95": [0.78, 0.86], "n": 90,
 "by_stratum": {"easy": 0.95, "boundary": 0.71, "adversarial": 0.60},
 "verdict": "pass | fail | insufficient_evidence"}

Common Mistakes

  • Giving recommendations without root cause classification. "Pass rate is low" doesn't tell you what to fix. Classify as system/metric/data problem first.
  • Maturity scoring by gut feel. Each L1-L4 rating must cite specific evidence from eval-design.md fields.
  • Executive summary too long. It has one audience: busy stakeholders. Three risks, three actions, one page. Details go in the full report.
  • Ignoring cross-skill contradictions. When two skills disagree about the same dimension, that's the most important signal in the report — it reveals a structural issue in assumptions or methodology.
  • All recommendations marked P1. If everything is medium priority, nothing is. P0 = blocks ship. P1 = important this sprint. P2 = backlog.

Next Skills

After 04-eval-report:

  • Recommendations point to specific workflows (03-align-human, 01-eval-design, 07-redteam, etc.) based on identified weaknesses. This skill is the "end of loop" analysis — after addressing recommendations, run this skill again to track progress.

© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/eval_pipeline/04-eval-report of agentscope-ai/OpenJudge.

Open the folder on GitHubat commit d1e0642

Compare with similar skills

Eval Report next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Report compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Report this skillagentscope-ai/OpenJudge870—~2.4kAutomated safety check: PassApache-2.0
News Aggregator Skillcclank/news-aggregator-skill1.3k—~2.1kAutomated safety check: PassNone
AI Daily Newsgeekjourneyx/ai-daily-skill235—~2.3kAutomated safety check: PassNone
Vss Search ArchiveNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~3.6kAutomated safety check: PassApache-2.0
AnalyzeriBigQiang/feedgrab614—~1kAutomated safety check: PassMIT
Reportmicrosoft/data-formulator18k—~1.5kAutomated safety check: PassMIT

Similar skills

  • News Aggregator Skill

    cclank/news-aggregator-skill

    Comprehensive news aggregator that fetches, filters, and deeply analyzes real-time content from 44+ sources including Hacker News, Lobsters, Dev.to, GitHub, arXiv, Hugging Face Papers, AIHOT, TLDR…

    1.3k GitHub stars~2.1k tokensUpdated 4 mo ago
    Writing & ContentAuto-check passed
  • AI Daily News

    geekjourneyx/ai-daily-skill

    Fetches AI news from smol.ai RSS and generates structured markdown with intelligent summarization and categorization.

    235 GitHub stars~2.3k tokensUpdated today
    Writing & ContentAuto-check passed
  • Vss Search Archive

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when a user wants to search archived VSS video or ingest or delete a source for search.

    1.9k GitHub stars~3.6k tokensUpdated today
    Writing & ContentAuto-check passed
  • Analyzer

    iBigQiang/feedgrab

    Content Analyzer — any content (URL, text, transcript) into structured analysis report with actionable insights.

    614 GitHub stars~1k tokensUpdated 1 mo ago
    Writing & ContentAuto-check passed
  • Report

    microsoft/data-formulator

    Official

    Turn an exploration (threads, findings, charts) into a single Markdown report — note, blog post, executive summary, KPI dashboard, slide brief, or multi-section analytical report, with embedded…

    18k GitHub stars~1.5k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Tldr

    earlyaidopters/claudeclaw

    Summarize the current conversation into a TLDR note and save it to your notes folder.

    173 GitHub stars~953 tokensUpdated 5 mo ago
    Writing & ContentAuto-check passed

More from agentscope-ai/OpenJudge

All 19 skills in this repo
  • Align Human

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

    870 GitHub stars~3.1k tokensUpdated 28 days ago
    Auto-check passed
  • Prompt Regression

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

    870 GitHub stars~2.8k tokensUpdated 28 days ago
    Auto-check passed
  • RAG Eval

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.

    870 GitHub stars~2.4k tokensUpdated 28 days ago
    Auto-check passed
  • Claude Authenticity

    agentscope-ai/OpenJudge

    Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.

    870 GitHub starsUsed in 1 repo~5k tokens
    Auto-check passed
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    870 GitHub stars~2.8k tokensUpdated 28 days ago
    Auto-check: warnings
  • Find Skills Combo

    agentscope-ai/OpenJudge

    Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.

    870 GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check: warnings

Questions about Eval Report

What does Eval Report do?

A skill your agent uses when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive…. Eval Report is an agent skill from agentscope-ai/OpenJudge. Use when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive summary.

When should I use Eval Report?

Eval Report fits situations like: the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment; cross-skill signals; prioritized actions; an executive summary.

How do I install Eval Report in Claude Code?

Run `npx skills add agentscope-ai/OpenJudge --skill eval-report -a claude-code`. Or copy the skill folder (skills/eval_pipeline/04-eval-report in agentscope-ai/OpenJudge) into .claude/skills/eval-report in your project. Claude Code loads it when a task matches its description.

How do I install Eval Report in Codex?

Run `npx skills add agentscope-ai/OpenJudge --skill eval-report -a codex`. Or copy the skill folder (skills/eval_pipeline/04-eval-report in agentscope-ai/OpenJudge) into .agents/skills/eval-report in your project. Codex loads it when a task matches its description.

Can I use Eval Report in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill eval-report -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-report, .gemini/skills/eval-report, .github/skills/eval-report and .opencode/skills/eval-report in your project.

What does Eval Report need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Report is instructions for the agent only. Our summary lists: Python 3.

Does Eval Report access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Report safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Report use?

Eval Report is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval Report use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Report?

Skills that share tags, products or a category with Eval Report: News Aggregator Skill (cclank/news-aggregator-skill, 1.3k stars), AI Daily News (geekjourneyx/ai-daily-skill, 235 stars), Vss Search Archive (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars) and Analyzer (iBigQiang/feedgrab, 614 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Report?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 870 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.

Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.