Topic · Agent Workflows
Best agent evaluation and testing skills for Claude Code, Codex and other agents.
- skills
- 131
- official
- 23
Agent evaluation and testing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation. | anthropics/ | 180k | 62 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 2 | Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers. | obra/ | 296k | 3 repos | ~1.7k | Automated safety check: Pass | MIT | today |
| 3 | Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints. | alchaincyf/ | 6.2k | 1 repo | ~4.7k | Automated safety check: Pass | MIT | 19 days ago |
| 4 | Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution. | HKUDS/ | 16k | 1 repo | ~2.1k | Automated safety check: Notes | MIT | 4 mo ago |
| 5 | Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability. | rohitg00/ | 65k | — | ~1k | Automated safety check: Pass | MIT | today |
| 6 | Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version. | colbymchenry/ | 73k | — | ~950 | Automated safety check: Pass | MIT | today |
| 7 | Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals. | dotnet/ | 23k | — | ~3.4k | Automated safety check: Pass | MIT | today |
| 8 | Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks. | aipoch/ | 5.4k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Evaluates the design quality of an agent skill against official specifications and patterns from existing examples, scoring it and suggesting improvements. | shareAI-lab/ | 5.2k | 4 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 10 | Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan. | Galaxy-Dawn/ | 5.7k | 1 repo | ~3k | Automated safety check: Pass | MIT | 14 days ago |
| 11 | 11.Skill Test Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit. | databricks-solutions/ | 1.9k | — | ~1.9k | Automated safety check: Pass | Unknown | 1 mo ago |
| 12 | Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop. | aiming-lab/ | 15k | — | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 13 | Evaluate skill quality, find the weakest dimension, and apply directed improvements. | Evol-ai/ | 216 | 1 repo | ~3.1k | Automated safety check: Pass | MIT | 5 mo ago |
| 14 | Call any REST API dynamically. An agent skill from NVIDIA/SkillEvaluator. | NVIDIA/ | 544 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 15 | Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons. | windmill-labs/ | 18k | — | ~969 | Automated safety check: Notes | Unknown | today |
| 16 | Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier. | langchain-ai/ | 1.3k | — | ~4k | Automated safety check: Pass | MIT | yesterday |
| 17 | Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI. | greyhaven-ai/ | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | today |
| 18 | Designs and verifies a deterministic grader that measures whether a GitHub Agentic Workflow run reached its real-world or repository outcome. | github/ | 5.3k | — | ~6.8k | Automated safety check: Pass | MIT | today |
| 19 | End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine. | kayba-ai/ | 2.6k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 20 | 20.Team Tasks Coordinate multi-agent development pipelines using shared JSON task files. | win4r/ | 447 | — | ~2.9k | Automated safety check: Pass | No licence | 8 mo ago |
| 21 | Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report. | bgauryy/ | 946 | — | ~2.1k | Automated safety check: Pass | MIT | 4 days ago |
| 22 | Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels. | code-yeongyu/ | 470 | — | ~2.7k | Automated safety check: Notes | MIT | today |
| 23 | 23.One Eval 驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。 | OpenDCAI/ | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 24 | 24.Autocontext Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met. | greyhaven-ai/ | 1.3k | — | ~892 | Automated safety check: Pass | Apache-2.0 | today |
| 25 | Runs existing Harbor evaluation jobs against a local LobeHub build or LobeHub Cloud, with preflight checks, resume support, and failure triage. | lobehub/ | 83k | — | ~1k | Automated safety check: Notes | Unknown | today |
| 26 | Operates or analyzes a LoopX-managed benchmark experiment: launching runs, maintaining the experiment board, qualifying integrity, and writing case insights. | loopx-project/ | 6.2k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | today |
| 27 | AWS Bedrock AgentCore comprehensive expert for deploying and managing AI agents at scale. | zxkane/ | 367 | 1 repo | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 28 | 28.Harness Eval Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports. | tech-leads-club/ | 7k | — | ~3.9k | Automated safety check: Pass | CC-BY-4.0 | 17 days ago |
| 29 | Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches. | QwenLM/ | 28k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 30 | Required for 4+ step requests; add tasks at start and update status after each step. | NVIDIA/ | 544 | — | ~460 | Automated safety check: Pass | Apache-2.0 | today |
| 31 | A skill your agent uses when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. | microsoft/ | 1k | — | ~2k | Automated safety check: Notes | Unknown | today |
| 32 | Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces. | affaan-m/ | 274k | 1 repo | ~623 | Automated safety check: Pass | MIT | 2 days ago |
| 33 | Walks you through drafting, testing, evaluating and tuning a skill for OpenCode, from an intake interview to description optimization. | antongulin/ | 172 | — | ~8.1k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 34 | Agent Script DSL for deterministic Agentforce agents. An agent skill from Jaganpro/sf-skills. | Jaganpro/ | 424 | — | ~3.8k | Automated safety check: Pass | MIT | 5 mo ago |
| 35 | Runs and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it. | edonadei/ | 206 | — | ~1.9k | Automated safety check: Notes | MIT | 2 days ago |
| 36 | 36.Skill Forge Ultimate Claude Code skill creator and architect. An agent skill from AgriciDaniel/skill-forge. | AgriciDaniel/ | 177 | — | ~1.9k | Automated safety check: Notes | MIT | 6 mo ago |
| 37 | Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote. | Q00/ | 6.2k | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 38 | Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch. | ory/ | 305 | — | ~497 | Automated safety check: Pass | Unknown | 1 mo ago |
| 39 | Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. | microsoft/ | 255 | 1 repo | ~6.7k | Automated safety check: Pass | MIT | today |
| 40 | A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. | NVIDIA/ | 544 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 41 | Run the weak-agent adversarial test harness against docx-cli. | kklimuk/ | 215 | — | ~6.1k | Automated safety check: Notes | MIT | 11 days ago |
| 42 | Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs. | amd/ | 1.6k | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 43 | Runs caliper's smoke evals against the real agent CLIs after a harness or MCP change, with a dry-run plan, failure triage and a report to attach to the PR. | edonadei/ | 206 | — | ~664 | Automated safety check: Pass | MIT | 2 days ago |
| 44 | Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments. | AgentToolkit/ | 122 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 45 | Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter. | microsoft/ | 1.4k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 46 | Evaluate Expo skills in this repo end-to-end - trigger accuracy, generated code quality, and runtime screenshots on iOS simulator and Android emulator via Expo Go (web optional). | expo/ | 2.7k | — | ~12k | Automated safety check: Pass | MIT | today |
| 47 | 47.Eval Answer Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer. | malloydata/ | 116 | — | ~4.3k | Automated safety check: Pass | MIT | today |
| 48 | Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans. | NateBJones-Projects/ | 4.7k | — | ~1.8k | Automated safety check: Pass | Unknown | yesterday |
Questions, answered from the data.
What is the best agent evaluation and testing skill?
MCP Server Builder (official) from anthropics/skills ranks first of the 131 agent evaluation and testing skills listed here, with the highest score: its repository has 180k GitHub stars, 62 other GitHub owners carry a copy, its SKILL.md loads about 2.3k tokens and it passes the automated safety check with no findings. Next come Diagnosing Superpowers Sessions and Darwin Skill Optimizer.
Which agent evaluation and testing skills are official?
23 of the 131 agent evaluation and testing skills are official, published by the vendor's own GitHub organization: MCP Server Builder, Copilot Session Failure Analysis, API Caller, Agent Eval Engineering, Operational Value Designer and 18 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in Agent Workflows
- MCP servers2,017
- Subagents1,015
- Agent instruction files737
- Agent memory507
- Multi-agent orchestration504
- Skill authoring502
- Planning495
- Brainstorming467
- Hooks and plugins363
- Autonomous loops325
- Context engineering297
- Session handoff262
- Requirements gathering234
- Task breakdown232
- Skill management205
- Human-in-the-loop approvals203
- Codebase knowledge for agents186
- Verification before completion169