Topic · Agent Workflows
Best agent evaluation and testing skills, page 2
Agent evaluation and testing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks. | amd/ | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 50 | 50.Benchflow Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow. | benchflow-ai/ | 353 | — | ~1.9k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 51 | Collects evidence from a failed or confusing skill-driven task and triggers skvm jit-optimize to propose concrete fixes to that skill's files. | SJTU-IPADS/ | 560 | — | ~2.9k | Automated safety check: Warn | MIT | 27 days ago |
| 52 | Runs Darwin Mode on an agent harness: mutates one policy file per generation in a sandbox, scores each variant against your tests, and archives only variants that measurably improve. | ruvnet/ | 97k | — | ~693 | Automated safety check: Pass | MIT | yesterday |
| 53 | Operational notes for generating, evaluating and debugging strategies in the autocontext grid_ctf scenario, with tier rules and parameter ranges that worked or failed. | greyhaven-ai/ | 1.3k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 54 | Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low. | microsoft/ | 1.4k | — | ~1.3k | Automated safety check: Pass | MIT | yesterday |
| 55 | Evaluate mathematical expressions and unit conversions. An agent skill from NVIDIA/SkillEvaluator. | NVIDIA/ | 544 | — | ~442 | Automated safety check: Pass | Apache-2.0 | today |
| 56 | Audits recent agent conversation history and turns repeated failures and successes into testable harness improvement proposals that later audits can check. | KonghaYao/ | 223 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | today |
| 57 | Run and maintain Con's terminal-agent benchmark against a live app session. | nowledge-co/ | 625 | — | ~822 | Automated safety check: Pass | MIT | yesterday |
| 58 | 58.Email Evals Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents. | tokencanopy/ | 192 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 59 | A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps… | microsoft/ | 138 | — | ~5.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 60 | Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship. | edonadei/ | 206 | — | ~2k | Automated safety check: Notes | MIT | 2 days ago |
| 61 | Create new skills, modify and improve existing skills, and measure skill performance. | ZS520L/ | 102 | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | 4 mo ago |
| 62 | Static safety audit of a SKILL.md that scores five dimensions and acts as a gate: skills below the pass line do not ship, whatever else they score. | openJiuwen-ai/ | 441 | — | ~3.5k | Automated safety check: Warn | Apache-2.0 | 7 days ago |
| 63 | 63.Skill Judge Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples. | shareAI-lab/ | 314 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 21 days ago |
| 64 | Analyze text content and produce statistics including word count, line count, character count, most frequent words, and readability metrics. | NVIDIA/ | 544 | — | ~453 | Automated safety check: Pass | Apache-2.0 | today |
| 65 | [omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. | rlaope/ | 3.2k | — | ~2.1k | Automated safety check: Pass | MIT | today |
| 66 | 66.Tester Test and validate AgentUse agents without real side effects. | agentuse/ | 205 | — | ~1.8k | Automated safety check: Notes | Unknown | yesterday |
| 67 | Evaluate catalog candidates against either the standard five agent-native criteria or the narrow operator-surface track, and check URL Onboarding. | haoruilee/ | 476 | — | ~2.8k | Automated safety check: Pass | CC0-1.0 | 2 days ago |
| 68 | A skill your agent uses when changing Deep Researcher Agent continuous integration, pre-commit, or contributor governance — editing .github/workflows/ (ci, ui, skills-eval, request-nvskills-ci)… | NVIDIA-AI-Blueprints/ | 883 | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 69 | Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately. | microsoft/ | 138 | — | ~22k | Automated safety check: Warn | MIT | 3 mo ago |
| 70 | Framework for measuring and tracking agent response quality over time. | vibeeval/ | 531 | — | ~2.9k | Automated safety check: Pass | MIT | 2 mo ago |
| 71 | Runbook for the five-arm Codex harness comparison on Terminal-Bench 4.0, pointing out four differences from SWE-Marathon that silently produce wrong scores. | loopx-project/ | 6.2k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 72 | Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles. | Orchestra-Research/ | 13k | 1 repo | ~3.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 73 | Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification. | bgauryy/ | 946 | — | ~1.6k | Automated safety check: Pass | MIT | 4 days ago |
| 74 | Evaluate agent and skill behavior or routing. An agent skill from TheGoat395/Codex-Skills. | TheGoat395/ | 126 | — | ~1.5k | Automated safety check: Pass | MIT | 27 days ago |
| 75 | Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js). | livekit-examples/ | 264 | 1 repo | ~1.9k | Automated safety check: Pass | MIT | 2 days ago |
| 76 | Run evaluation pipelines on Claude Code skills to test triggering accuracy, workflow correctness, and output quality. | AgriciDaniel/ | 177 | — | ~1.7k | Automated safety check: Pass | MIT | 6 mo ago |
| 77 | Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &… | microsoft/ | 138 | — | ~10k | Automated safety check: Pass | MIT | 3 mo ago |
| 78 | Pulls pattern, trend and benchmark data from the claude-view MCP server to present a behavioral analysis of your Claude Code usage, grouped by category. | tombelieber/ | 110 | — | ~3.1k | Automated safety check: Pass | MIT | 7 days ago |
| 79 | Runs Chat Customizations Evaluations analysis for the active customization file and summarizes the findings from the Problems panel. | microsoft/ | 141 | — | ~135 | Automated safety check: Pass | Unknown | 5 days ago |
| 80 | Points agents to the right sources when changing the Smithers workspace graph, generated docs, benchmark and eval evidence, or durable flows in the Smithers repository. | smithersai/ | 429 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 81 | Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result. | Prism-Shadow/ | 2.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 82 | Run and improve the c0 agent evaluation loop using the local and remote Cloudflare MCP harness, versioned history notes, and current Cloudflare product updates. | Consensys/ | 105 | — | ~1.2k | Automated safety check: Pass | LGPL-3.0 | 1 mo ago |
| 83 | 83.Bkit Evals Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results. | ww-w-ai/ | 600 | — | ~1k | Automated safety check: Notes | Apache-2.0 | 10 days ago |
| 84 | Analyzes a finished pull request that involved an agent to find what slowed or helped it, and recommends specific changes to instruction files, skills and docs. | dotnet/ | 23k | — | ~2.5k | Automated safety check: Pass | MIT | yesterday |
| 85 | Analyzes Copilot Studio evaluation results using Practical Guidance on Agent Evaluation's 10-step playbook (Steps 6, 7, and 9) plus Microsoft's triage diagnostics. | microsoft/ | 138 | — | ~9.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 86 | Improves a single skill through a test, fix and retest loop, keeping or reverting each change based on how the static and category scores move. | Donchitos/ | 26k | — | ~1.6k | Automated safety check: Notes | MIT | 8 days ago |
| 87 | Agentforce agent testing with dual-track workflow and 100-point scoring. | Jaganpro/ | 424 | — | ~2.2k | Automated safety check: Pass | MIT | 5 mo ago |
| 88 | 88.Write Skill Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it. | druxt/ | 114 | — | ~926 | Automated safety check: Pass | MIT | yesterday |
| 89 | Investigates what went wrong in a superpowers session by reading its transcript, reports findings with path and line citations, and can draft a GitHub issue or redacted bundle. | jnMetaCode/ | 8.3k | — | ~858 | Automated safety check: Pass | MIT | 3 days ago |
| 90 | Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates. | Arenukvern/ | 385 | — | ~2.4k | Automated safety check: Pass | MIT | 4 days ago |
| 91 | A skill your agent uses when evaluating, designing, or pressure-testing the business model of an AI agent product. | harperreed/ | 334 | — | ~3k | Automated safety check: Pass | No licence | 4 days ago |
| 92 | Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. | microsoft/ | 138 | — | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 93 | Agent skill for benchmark-suite - invoke with $agent-benchmark-suite | ruvnet/ | 74k | 2 repos | ~4.9k | Automated safety check: Pass | MIT | today |
| 94 | Agent skill for test-long-runner - invoke with $agent-test-long-runner | ruvnet/ | 74k | 2 repos | ~426 | Automated safety check: Pass | MIT | today |
| 95 | Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering… | fabricioctelles/ | 105 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 96 | 96.Skill Eval Measure whether a skill helps by comparing runs with and without it. | boshu2/ | 447 | 1 repo | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
Explore related skills
Category
More topics in Agent Workflows
- MCP servers2,017
- Subagents1,015
- Agent instruction files737
- Agent memory507
- Multi-agent orchestration504
- Skill authoring502
- Planning495
- Brainstorming467
- Hooks and plugins363
- Autonomous loops325
- Context engineering297
- Session handoff262
- Requirements gathering234
- Task breakdown232
- Skill management205
- Human-in-the-loop approvals203
- Codebase knowledge for agents186
- Verification before completion169