Agent skill

AI Agent Evaluation Benchmarking

by sickn33 in sickn33/agentic-awesome-skills

Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks.

MITAuto-check passedAgent Workflows

Install AI Agent Evaluation Benchmarking

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill ai-agent-evaluation-benchmarking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills ai-agent-evaluation-benchmarking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-agent-evaluation-benchmarking .claude/skills/ai-agent-evaluation-benchmarking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-agent-evaluation-benchmarking
GitHub stars
47k
Used in
1 other repo
Token cost
~1.3k tokens
SKILL.md length
499 words
Files
1
Skills in repo
1,493
Repo updated
First seen
Licence
MIT

At a glance

Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks.

  • Works in 3 steps: Define the parameters, thresholds, and… → Select appropriate boundary enforcement… → Export standardized artifacts (CSV…
  • Tasks that involve Agent evaluation and testing
  • SKILL.md covers Overview, When to Use This Skill, How It Works and Field Reference, plus 9 more sections
  • Calls claude and gemini

What it does

AI Agent Evaluation Benchmarking is an agent skill from sickn33/agentic-awesome-skills. Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Agent evaluation and testing. It works with SQL. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Agent evaluation and testing

Example prompts

  • “/ai-agent-evaluation-benchmarking”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Define the parameters, thresholds, and identity bindings required for the target operational register.
  2. Select appropriate boundary enforcement values from validated enum select sets.
  3. Export standardized artifacts (CSV table, SQL DDL, JSON Schema) to integrate into validation CI pipelines.

What it can do on your machine

Read from SKILL.md and the folder at commit 680176d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude
    • gemini

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Agent Evaluation Benchmarking loads about 1.3k tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 499 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit 680176d, republished under its MIT licence (© sickn33). 499 words, ~1,347 tokens.

Download SKILL.mdSave it as .claude/skills/ai-agent-evaluation-benchmarking/SKILL.md (or your agent's skills folder).
name
ai-agent-evaluation-benchmarking
description
Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks.
category
engineering
risk
safe
source
self
source_type
self
date_added
2026-10-01
author
Ranjeet2063
tags
ai, agents, evaluation, benchmarks, llm, automation, quality
source_repo
Ranjeet2063/agentic-awesome-skills

AI Agent Capability Evaluation & Benchmarking

What it is: Standardizes multi-metric capability benchmarking, token cost efficiency, and regression monitoring across autonomous coding agents.

Overview

Provides a standardized, auditable framework and data model for AI Agent Capability Evaluation & Benchmarking operations across distributed engineering and decentralized application systems.

When to Use This Skill

  • When formalizing architectural contracts, security invariants, or operational limits for AI Agent Capability Evaluation & Benchmarking.
  • When cross-functional review is required between protocol developers, smart contract auditors, and AI engineering agents.
  • When generating reproducible CSV, SQL DDL, JSON Schema, and Notion property registers for tracking compliance.

How It Works

  1. Define the parameters, thresholds, and identity bindings required for the target operational register.
  2. Select appropriate boundary enforcement values from validated enum select sets.
  3. Export standardized artifacts (CSV table, SQL DDL, JSON Schema) to integrate into validation CI pipelines.

Field Reference

#Field NameTypeSQL TypeJSON Schema TypeNotion Property TypeExample Value
1Benchmark Run IDidSERIAL PRIMARY KEYintegerTextBENCH-001
2Evaluated Agent ModelselectVARCHAR(64)stringSelectClaude 3.7 Sonnet
3Benchmark Suite DomainselectVARCHAR(64)stringSelectSWE-bench Verified
4Tasks Evaluated CountnumberINTEGERnumberNumber100
5Pass Rate PercentagenumberNUMERIC(5,2)numberNumber78.40
6Tool Hallucination Rate %numberNUMERIC(5,2)numberNumber0.60
7Average Tokens Per TasknumberINTEGERnumberNumber42500
8Cost Per Solved Task USDcurrencyNUMERIC(8,4)numberNumber0.3420
9Regression VerdictselectVARCHAR(32)stringSelectSuperior
10Evaluation LeadtextVARCHAR(64)stringTextRanjeet2063
11Benchmark Execution DatedateDATEstring, format: dateDate2026-10-01

Select Options

Evaluated Agent Model

Claude 3.7 Sonnet | Claude 3.5 Sonnet | GPT-4o | Gemini 2.0 Flash | DeepSeek V3

Benchmark Suite Domain

SWE-bench Verified | WebArena | AgentBench | HumanEval-Rust | Web3AuditBench

Regression Verdict

Superior | Parity Baseline | Regression Failure

Relations

  • Audit Reference -> links to the formal review documentation or test repository.
  • Target Architecture -> links to the deployed contract or autonomous agent runtime component.
Show full SKILL.md (204 more words)Show less

Examples

Prompt

How do I configure and track AI Agent Capability Evaluation & Benchmarking for our production environment?

Recommended Next Step

Generate the unified field schema, SQL DDL migration, and JSON validation schema to register into your system catalog.

Workflow: Define criteria -> Run automated verification -> Record baseline -> Monitor invariants.

Best Practices

  • Enforce strict typing on numerical bounds and currency amounts; avoid unstructured free-text fields for critical states.
  • Re-run validation test suites on every state-altering commit or parameter change.
  • Keep example data synthetic and isolated from production cryptographic keys or private endpoints.

Limitations

  • Provides architectural specifications, data models, and verification schemas; does not execute direct transaction signing without authorized external tooling.
  • Requires network connectivity and valid RPC credentials when querying on-chain states.

Security & Safety Notes

  • All parameters declare risk: safe. No unauthorized state modification or privileged credential access is performed.
  • Use synthetic dummy keys and mock addresses in test suites and local verification scripts.

Common Pitfalls

  • Problem: Mismatched decimal precision between contract runtime and database register. Solution: Always verify decimals using the explicit field mapping in this reference.
  • Problem: Missing authorization checks prior to state update. Solution: Cross-validate against the Security Audit register before deployment.
  • @ai-agent-tool-routing - covers tool schema registration and retry policy.
  • @ai-prompt-regression-testing - covers prompt regression baselines and drift.
  • @ai-code-generation-guardrails - covers static guardrails for generated code.

Reusable Prompt

I want to establish a verified AI Agent Capability Evaluation & Benchmarking register for our production protocol.
Guide me through the required field parameters and output the corresponding SQL DDL and JSON Schema.

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ai-agent-evaluation-benchmarking of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit 680176d

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

AI Agent Evaluation Benchmarking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Agent Evaluation Benchmarking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Agent Evaluation Benchmarking this skillsickn33/agentic-awesome-skills47k1 repos~1.3kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers297k3 repos~1.7kAutomated safety check: PassMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Skill Release Gaterohitg00/ai-engineering-from-scratch66k—~1kAutomated safety check: PassMIT
CodeGraph Agent Evalcolbymchenry/codegraph74k—~950Automated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    297k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    66k GitHub stars~1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    74k GitHub stars~950 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,493 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Works with

Categories

Questions about AI Agent Evaluation Benchmarking

What does AI Agent Evaluation Benchmarking do?

Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks. AI Agent Evaluation Benchmarking is an agent skill from sickn33/agentic-awesome-skills. Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks.

When should I use AI Agent Evaluation Benchmarking?

AI Agent Evaluation Benchmarking fits situations like: tasks that involve Agent evaluation and testing.

How do I install AI Agent Evaluation Benchmarking in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill ai-agent-evaluation-benchmarking -a claude-code`. Or copy the skill folder (skills/ai-agent-evaluation-benchmarking in sickn33/agentic-awesome-skills) into .claude/skills/ai-agent-evaluation-benchmarking in your project. Claude Code loads it when a task matches its description.

How do I install AI Agent Evaluation Benchmarking in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill ai-agent-evaluation-benchmarking -a codex`. Or copy the skill folder (skills/ai-agent-evaluation-benchmarking in sickn33/agentic-awesome-skills) into .agents/skills/ai-agent-evaluation-benchmarking in your project. Codex loads it when a task matches its description.

Can I use AI Agent Evaluation Benchmarking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill ai-agent-evaluation-benchmarking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-agent-evaluation-benchmarking, .gemini/skills/ai-agent-evaluation-benchmarking, .github/skills/ai-agent-evaluation-benchmarking and .opencode/skills/ai-agent-evaluation-benchmarking in your project.

What does AI Agent Evaluation Benchmarking need to run?

Going by SKILL.md and its folder, AI Agent Evaluation Benchmarking needs the command-line tools its instructions call (claude and gemini).

Does AI Agent Evaluation Benchmarking access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AI Agent Evaluation Benchmarking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Agent Evaluation Benchmarking use?

AI Agent Evaluation Benchmarking is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Agent Evaluation Benchmarking use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AI Agent Evaluation Benchmarking?

Skills that share tags, products or a category with AI Agent Evaluation Benchmarking: MCP Server Builder (anthropics/skills, 180k stars), Diagnosing Superpowers Sessions (obra/superpowers, 297k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars) and Skill Release Gate (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Agent Evaluation Benchmarking?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,379 GitHub stars. The repository holds 1,493 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.