Agent skill

Agent Evaluation Reporting

by sickn33 in sickn33/agentic-awesome-skills

A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

MITAuto-check passedAgent Workflows

Install Agent Evaluation Reporting

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill agent-evaluation-reporting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills agent-evaluation-reporting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-evaluation-reporting .claude/skills/agent-evaluation-reporting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-evaluation-reporting
GitHub stars
47k
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
936 words
Files
1
Skills in repo
1,394
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

  • Works in 6 steps: Freeze the comparison contract → Build a mutually exclusive outcome ledger → Lock each metric to a denominator → …
  • Summarizing agent evaluations where autonomous
  • SKILL.md covers Overview, When to Use This Skill, How It Works and Example, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Evaluation Reporting is an agent skill from sickn33/agentic-awesome-skills. Use when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Agent evaluation and testing. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Summarizing agent evaluations where autonomous
  • Invalid outcomes must remain distinct and comparable

Example prompts

  • “/agent-evaluation-reporting”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Freeze the comparison contract
  2. Build a mutually exclusive outcome ledger
  3. Lock each metric to a denominator
  4. Keep latency and cost populations honest
  5. Quantify uncertainty and comparability
  6. Map evidence to predeclared decision gates

What it can do on your machine

Read from SKILL.md and the folder at commit 1e53ce2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Evaluation Reporting loads about 2.1k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 936 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit 1e53ce2, republished under its MIT licence (© sickn33). 936 words, ~2,066 tokens.

Download SKILL.mdSave it as .claude/skills/agent-evaluation-reporting/SKILL.md (or your agent's skills folder).
name
agent-evaluation-reporting
description
Use when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.
category
agent-evaluation
risk
none
source
self
source_type
self
date_added
2026-08-18
author
Whxuan0701
tags
agent-evaluation, metrics, reporting, reliability, benchmarking
tools
claude, cursor, gemini, codex

Agent Evaluation Reporting

Overview

Turn raw agent evaluation runs into a decision-ready report without hiding failures or overstating capability. Keep outcome populations, denominators, latency populations, and experiment conditions explicit so readers can reproduce every headline number.

When to Use This Skill

  • Use when reporting benchmark, regression, pilot, or production evaluation runs for an AI agent.
  • Use when autonomous and human-assisted completions appear in the same result set.
  • Use when failures, timeouts, infrastructure-invalid runs, retries, or partial results affect the denominator.
  • Use when comparing two agents, prompts, harnesses, or releases and deciding whether the comparison is valid.

How It Works

Step 1: Freeze the comparison contract

Record the task set and sampling, model and provider, prompt or policy version, tool and harness versions, evaluator rubric, timeout and retry policy, token or cost budget, environment, and human-intervention policy. Assign the configuration a stable label or digest.

If a material condition differs between runs, mark the comparison as non-equivalent. Report a directional observation only; do not claim that the changed agent caused the difference.

Step 2: Build a mutually exclusive outcome ledger

Classify every scheduled attempt exactly once:

OutcomeMeaning
autonomous_successThe agent satisfied the evaluator without human intervention.
assisted_successThe task succeeded only after a human intervened.
failureThe run reached a terminal, evaluable failure.
timeoutThe run exhausted its declared time or step budget.
invalidThe agent never received a valid evaluation because the harness, environment, or input failed.

Preserve attempt ID, task ID or seed, retry index, parent attempt ID, configuration label, outcome, intervention count, duration, cost, evaluator evidence, and invalid reason when available. Never silently drop invalid or retried runs.

Also build a unique-task rollup. For each task, retain its first-attempt outcome and derive one eventual outcome after the predeclared retry policy finishes. An execution attempt may contribute once to attempt-level metrics, but a task may contribute only once to task-level completion metrics. If retry lineage or the retry policy is missing, do not report eventual task completion.

Step 3: Lock each metric to a denominator

Let N_all be all execution attempts, including retries, and N_eval = N_all - N_invalid be evaluable attempts. Let T_all be unique scheduled tasks and T_eval be tasks with a valid task-level outcome under the fixed retry policy. Report counts beside every rate.

text
autonomous attempt success = N_autonomous / N_eval
assisted attempt success   = N_assisted / N_eval
attempt non-completion     = (N_failure + N_timeout) / N_eval
invalid-attempt rate       = N_invalid / N_all
first-attempt completion   = T_first_attempt_completed / T_all
eventual task completion   = T_eventual_completed / T_eval
operational task delivery  = T_eventual_completed / T_all

Label attempt-level and unique-task metrics explicitly; never call an attempt-level rate workflow completion. Report the retry rate and attempts per task so policy-dependent gains remain visible. Check that evaluable attempt outcomes sum to N_eval, all attempt outcomes sum to N_all, and the task rollup sums to T_all.

If N_eval == 0, report every attempt capability rate as unavailable rather than dividing by zero, and mark any gate that depends on those rates inconclusive. Apply the same rule to any metric whose denominator is zero, including task-level rates when T_all == 0 or T_eval == 0.

Step 4: Keep latency and cost populations honest

Report autonomous-completion latency, assisted end-to-end latency, and failure time-to-terminal separately. A success-only P50 is not an overall P50, and subgroup medians cannot be averaged or weighted to reconstruct a combined median.

Calculate an all-run percentile only from per-run observations and state how timeouts are handled. If durations are right-censored, report the censoring policy or use an appropriate survival estimate. Apply the same population labels to token and cost metrics.

Show full SKILL.md (390 more words)Show less
Step 5: Quantify uncertainty and comparability

For stochastic evaluations, show sample size and an interval or repeated-run distribution beside headline rates. For comparisons, report the absolute delta and verify that both sides share the frozen contract from Step 1. If data is missing, conditions differ, or intervals are too wide, use inconclusive rather than choosing a winner.

Step 6: Map evidence to predeclared decision gates

Define readiness gates before reading the result, such as minimum autonomous success, maximum timeout rate, zero critical safety violations, and latency or cost bounds. Return pass, fail, or inconclusive for each gate.

Do not infer production readiness from a success rate alone. When no thresholds or risk requirements were supplied, state that readiness is not determined and list the missing gates.

Example

For 120 unique tasks with one attempt each, including 12 infrastructure-invalid runs, 48 autonomous successes, 24 assisted successes, 20 failures, and 16 timeouts:

text
Evaluable attempts:       108 / 120
Autonomous success:        48 / 108 = 44.4%
Assisted success:          24 / 108 = 22.2%
Attempt non-completion:     36 / 108 = 33.3%
First-attempt completion:   72 / 120 = 60.0%
Eventual task completion:   72 / 108 = 66.7% (no retries)
Operational task delivery: 72 / 120 = 60.0%
Infrastructure-invalid:    12 / 120 = 10.0%
Overall latency P50:       unavailable from subgroup aggregates
Readiness:                 inconclusive until gates are declared

Best Practices

  • Report counts, formulas, denominator labels, and exclusions together.
  • Separate autonomous capability from human-assisted workflow completion.
  • Preserve timeout and invalid-run rates even when publishing a valid-run score.
  • Pair aggregate metrics with failure categories and representative evidence.
  • Re-run both candidates under one frozen contract before making a causal improvement claim.

Limitations

  • This skill structures and interprets supplied evaluation evidence; it does not validate the evaluator or recreate missing run records.
  • Small or biased task sets can produce precise-looking but unrepresentative metrics.
  • Statistical significance does not establish production safety, user value, or acceptable cost.
  • Readiness remains inconclusive when acceptance thresholds, severity policy, or required evidence are absent.

Security & Safety Notes

  • Redact credentials, private prompts, personal data, and sensitive tool output from reports while retaining stable evidence references.
  • Treat critical safety violations as separate release gates rather than averaging them into a general quality score.

Common Pitfalls

  • Problem: Assisted completions are presented as autonomous success. Solution: Publish separate autonomous, assisted, and workflow-completion rates.
  • Problem: Timeouts or invalid runs disappear from the denominator. Solution: Reconcile the full outcome ledger against N_all before calculating metrics.
  • Problem: A faster success-only P50 is presented as a faster system. Solution: Label the population and report all-run time-to-terminal only from per-run data.
  • Problem: A release verdict is improvised after seeing results. Solution: Apply predeclared gates or return inconclusive.
  • @agent-evaluation - Design behavioral tests, benchmarks, and reliability evaluations.
  • @run-deep-swe - Execute reproducible DeepSWE benchmark runs before reporting their results.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/agent-evaluation-reporting of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit 1e53ce2

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Agent Evaluation Reporting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Evaluation Reporting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Evaluation Reporting this skillsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k62 repos~2.3kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers296k3 repos~1.7kAutomated safety check: PassMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Skill Release Gaterohitg00/ai-engineering-from-scratch65k—~1kAutomated safety check: PassMIT
CodeGraph Agent Evalcolbymchenry/codegraph73k—~950Automated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 62 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    296k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    65k GitHub stars~1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    73k GitHub stars~950 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,394 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Categories

Questions about Agent Evaluation Reporting

What does Agent Evaluation Reporting do?

A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable. Agent Evaluation Reporting is an agent skill from sickn33/agentic-awesome-skills. Use when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

When should I use Agent Evaluation Reporting?

Agent Evaluation Reporting fits situations like: summarizing agent evaluations where autonomous; invalid outcomes must remain distinct and comparable.

How do I install Agent Evaluation Reporting in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill agent-evaluation-reporting -a claude-code`. Or copy the skill folder (skills/agent-evaluation-reporting in sickn33/agentic-awesome-skills) into .claude/skills/agent-evaluation-reporting in your project. Claude Code loads it when a task matches its description.

How do I install Agent Evaluation Reporting in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill agent-evaluation-reporting -a codex`. Or copy the skill folder (skills/agent-evaluation-reporting in sickn33/agentic-awesome-skills) into .agents/skills/agent-evaluation-reporting in your project. Codex loads it when a task matches its description.

Can I use Agent Evaluation Reporting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill agent-evaluation-reporting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-evaluation-reporting, .gemini/skills/agent-evaluation-reporting, .github/skills/agent-evaluation-reporting and .opencode/skills/agent-evaluation-reporting in your project.

What does Agent Evaluation Reporting need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Evaluation Reporting is instructions for the agent only.

Does Agent Evaluation Reporting access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Evaluation Reporting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Evaluation Reporting use?

Agent Evaluation Reporting is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Evaluation Reporting use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Evaluation Reporting?

Skills that share tags, products or a category with Agent Evaluation Reporting: MCP Server Builder (anthropics/skills, 180k stars), Diagnosing Superpowers Sessions (obra/superpowers, 296k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars) and Skill Release Gate (rohitg00/ai-engineering-from-scratch, 65k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Evaluation Reporting?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,304 GitHub stars. The repository holds 1,394 skills in this directory. The repository was last updated on October 6, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.