Topic · Agent Workflows

Best agent evaluation and testing skills for Claude Code, Codex and other agents.

Skills that test and benchmark how well agents and skills perform.
skills
131
official
23

Agent evaluation and testing skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Agent evaluation and testing skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

anthropics/skills180k62 repos~2.3kAutomated safety check: PassApache-2.02 days ago
2

Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

obra/superpowers296k3 repos~1.7kAutomated safety check: PassMITtoday
3

Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

alchaincyf/darwin-skill6.2k1 repo~4.7kAutomated safety check: PassMIT19 days ago
4

Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

HKUDS/OpenHarness16k1 repo~2.1kAutomated safety check: NotesMIT4 mo ago
5

Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

rohitg00/ai-engineering-from-scratch65k—~1kAutomated safety check: PassMITtoday
6

Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

colbymchenry/codegraph73k—~950Automated safety check: PassMITtoday
7

Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

dotnet/maui23k—~3.4kAutomated safety check: PassMITtoday
8

Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks.

aipoch/open-science5.4k—~1.7kAutomated safety check: PassApache-2.0today
9

Evaluates the design quality of an agent skill against official specifications and patterns from existing examples, scoring it and suggesting improvements.

shareAI-lab/Kode-CLI5.2k4 repos~7.5kAutomated safety check: PassApache-2.01 mo ago
10

Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan.

Galaxy-Dawn/claude-scholar5.7k1 repo~3kAutomated safety check: PassMIT14 days ago
11

Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit.

databricks-solutions/ai-dev-kit1.9k—~1.9kAutomated safety check: PassUnknown1 mo ago
12

Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop.

aiming-lab/AutoResearchClaw15k—~1.8kAutomated safety check: PassMIT1 mo ago
13

Evaluate skill quality, find the weakest dimension, and apply directed improvements.

Evol-ai/SkillCompass2161 repo~3.1kAutomated safety check: PassMIT5 mo ago
14
14.API CallerOfficial

Call any REST API dynamically. An agent skill from NVIDIA/SkillEvaluator.

NVIDIA/SkillEvaluator5441 repo~1.1kAutomated safety check: PassApache-2.0today
15

Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

windmill-labs/windmill18k—~969Automated safety check: NotesUnknowntoday
16

Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

langchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMITyesterday
17

Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

greyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.0today
18

Designs and verifies a deterministic grader that measures whether a GitHub Agentic Workflow run reached its real-world or repository outcome.

github/gh-aw5.3k—~6.8kAutomated safety check: PassMITtoday
19

End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine.

kayba-ai/agentic-context-engine2.6k—~1.4kAutomated safety check: PassApache-2.013 days ago
20

Coordinate multi-agent development pipelines using shared JSON task files.

win4r/team-tasks447—~2.9kAutomated safety check: PassNo licence8 mo ago
21

Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

bgauryy/octocode946—~2.1kAutomated safety check: PassMIT4 days ago
22

Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels.

code-yeongyu/senpi470—~2.7kAutomated safety check: NotesMITtoday
23

驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

OpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.01 mo ago
24

Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.

greyhaven-ai/autocontext1.3k—~892Automated safety check: PassApache-2.0today
25

Runs existing Harbor evaluation jobs against a local LobeHub build or LobeHub Cloud, with preflight checks, resume support, and failure triage.

lobehub/lobehub83k—~1kAutomated safety check: NotesUnknowntoday
26

Operates or analyzes a LoopX-managed benchmark experiment: launching runs, maintaining the experiment board, qualifying integrity, and writing case insights.

loopx-project/loopx6.2k—~3.6kAutomated safety check: PassApache-2.0today
27

AWS Bedrock AgentCore comprehensive expert for deploying and managing AI agents at scale.

zxkane/aws-skills3671 repo~2.5kAutomated safety check: PassMIT3 mo ago
28

Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports.

tech-leads-club/agent-skills7k—~3.9kAutomated safety check: PassCC-BY-4.017 days ago
29

Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches.

QwenLM/qwen-code28k—~1.1kAutomated safety check: PassApache-2.0today
30
30.Task ListOfficial

Required for 4+ step requests; add tasks at start and update status after each step.

NVIDIA/SkillEvaluator544—~460Automated safety check: PassApache-2.0today
31

A skill your agent uses when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI.

microsoft/work-iq1k—~2kAutomated safety check: NotesUnknowntoday
32

Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

affaan-m/ECC274k1 repo~623Automated safety check: PassMIT2 days ago
33

Walks you through drafting, testing, evaluating and tuning a skill for OpenCode, from an intake interview to description optimization.

antongulin/opencode-skill-creator172—~8.1kAutomated safety check: PassApache-2.06 days ago
34

Agent Script DSL for deterministic Agentforce agents. An agent skill from Jaganpro/sf-skills.

Jaganpro/sf-skills424—~3.8kAutomated safety check: PassMIT5 mo ago
35

Runs and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it.

edonadei/caliper206—~1.9kAutomated safety check: NotesMIT2 days ago
36

Ultimate Claude Code skill creator and architect. An agent skill from AgriciDaniel/skill-forge.

AgriciDaniel/skill-forge177—~1.9kAutomated safety check: NotesMIT6 mo ago
37

Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote.

Q00/ouroboros6.2k—~2.2kAutomated safety check: PassMITyesterday
38

Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

ory/lumen305—~497Automated safety check: PassUnknown1 mo ago
39

Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end.

microsoft/GitHub-Copilot-for-Azure2551 repo~6.7kAutomated safety check: PassMITtoday
40

A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

NVIDIA/SkillEvaluator544—~2.1kAutomated safety check: PassApache-2.0today
41

Run the weak-agent adversarial test harness against docx-cli.

kklimuk/docx-cli215—~6.1kAutomated safety check: NotesMIT11 days ago
42

Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

amd/gaia1.6k—~2.3kAutomated safety check: PassMITtoday
43

Runs caliper's smoke evals against the real agent CLIs after a harness or MCP change, with a dry-run plan, failure triage and a report to attach to the PR.

edonadei/caliper206—~664Automated safety check: PassMIT2 days ago
44

Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments.

AgentToolkit/altk-evolve122—~1.7kAutomated safety check: PassApache-2.0today
45

Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.

microsoft/waza1.4k—~2kAutomated safety check: PassMITyesterday
46
46.Expo Skill EvalOfficial

Evaluate Expo skills in this repo end-to-end - trigger accuracy, generated code quality, and runtime screenshots on iOS simulator and Android emulator via Expo Go (web optional).

expo/skills2.7k—~12kAutomated safety check: PassMITtoday
47

Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

malloydata/publisher116—~4.3kAutomated safety check: PassMITtoday
48

Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans.

NateBJones-Projects/OB14.7k—~1.8kAutomated safety check: PassUnknownyesterday

Questions, answered from the data.

What is the best agent evaluation and testing skill?

MCP Server Builder (official) from anthropics/skills ranks first of the 131 agent evaluation and testing skills listed here, with the highest score: its repository has 180k GitHub stars, 62 other GitHub owners carry a copy, its SKILL.md loads about 2.3k tokens and it passes the automated safety check with no findings. Next come Diagnosing Superpowers Sessions and Darwin Skill Optimizer.

Which agent evaluation and testing skills are official?

23 of the 131 agent evaluation and testing skills are official, published by the vendor's own GitHub organization: MCP Server Builder, Copilot Session Failure Analysis, API Caller, Agent Eval Engineering, Operational Value Designer and 18 more.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.