Topic · Agent Workflows

Best agent evaluation and testing skills, page 2

Skills #49–96 of 131, ranked by score.

Agent evaluation and testing skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Agent evaluation and testing skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

amd/gaia1.6k—~1.8kAutomated safety check: PassMITyesterday
50

Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

benchflow-ai/benchflow353—~1.9kAutomated safety check: NotesApache-2.02 days ago
51

Collects evidence from a failed or confusing skill-driven task and triggers skvm jit-optimize to propose concrete fixes to that skill's files.

SJTU-IPADS/SkVM560—~2.9kAutomated safety check: WarnMIT27 days ago
52

Runs Darwin Mode on an agent harness: mutates one policy file per generation in a sandbox, scores each variant against your tests, and archives only variants that measurably improve.

ruvnet/RuView97k—~693Automated safety check: PassMITyesterday
53

Operational notes for generating, evaluating and debugging strategies in the autocontext grid_ctf scenario, with tier rules and parameter ranges that worked or failed.

greyhaven-ai/autocontext1.3k—~1.3kAutomated safety check: PassApache-2.0yesterday
54
54.Waza InteractiveOfficial

Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

microsoft/waza1.4k—~1.3kAutomated safety check: PassMITyesterday
55
55.CalculatorOfficial

Evaluate mathematical expressions and unit conversions. An agent skill from NVIDIA/SkillEvaluator.

NVIDIA/SkillEvaluator544—~442Automated safety check: PassApache-2.0today
56

Audits recent agent conversation history and turns repeated failures and successes into testable harness improvement proposals that later audits can check.

KonghaYao/peri223—~3.5kAutomated safety check: PassApache-2.0today
57

Run and maintain Con's terminal-agent benchmark against a live app session.

nowledge-co/con-terminal625—~822Automated safety check: PassMITyesterday
58

Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.

tokencanopy/e2a192—~2.1kAutomated safety check: PassApache-2.0yesterday
59

A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps…

microsoft/eval-guide138—~5.9kAutomated safety check: PassMIT3 mo ago
60

Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship.

edonadei/caliper206—~2kAutomated safety check: NotesMIT2 days ago
61

Create new skills, modify and improve existing skills, and measure skill performance.

ZS520L/HanakoPro102—~7.4kAutomated safety check: PassApache-2.04 mo ago
62

Static safety audit of a SKILL.md that scores five dimensions and acts as a gate: skills below the pass line do not ship, whatever else they score.

openJiuwen-ai/agent-core441—~3.5kAutomated safety check: WarnApache-2.07 days ago
63

Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.

shareAI-lab/lab-skills314—~1.9kAutomated safety check: PassApache-2.021 days ago
64
64.Text AnalyzerOfficial

Analyze text content and produce statistics including word count, line count, character count, most frequent words, and readability metrics.

NVIDIA/SkillEvaluator544—~453Automated safety check: PassApache-2.0today
65

[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics.

rlaope/oh-my-hermes3.2k—~2.1kAutomated safety check: PassMITtoday
66

Test and validate AgentUse agents without real side effects.

agentuse/agentuse205—~1.8kAutomated safety check: NotesUnknownyesterday
67

Evaluate catalog candidates against either the standard five agent-native criteria or the narrow operator-surface track, and check URL Onboarding.

haoruilee/awesome-agent-native-services476—~2.8kAutomated safety check: PassCC0-1.02 days ago
68

A skill your agent uses when changing Deep Researcher Agent continuous integration, pre-commit, or contributor governance — editing .github/workflows/ (ci, ui, skills-eval, request-nvskills-ci)…

NVIDIA-AI-Blueprints/deep-researcher-agent883—~1.5kAutomated safety check: NotesApache-2.0yesterday
69
69.Eval GuideOfficial

Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

microsoft/eval-guide138—~22kAutomated safety check: WarnMIT3 mo ago
70

Framework for measuring and tracking agent response quality over time.

vibeeval/vibecosystem531—~2.9kAutomated safety check: PassMIT2 mo ago
71

Runbook for the five-arm Codex harness comparison on Terminal-Bench 4.0, pointing out four differences from SWE-Marathon that silently produce wrong scores.

loopx-project/loopx6.2k—~1.6kAutomated safety check: PassApache-2.0today
72

Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles.

Orchestra-Research/AI-Research-SKILLs13k1 repo~3.6kAutomated safety check: PassMIT3 mo ago
73

Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

bgauryy/octocode946—~1.6kAutomated safety check: PassMIT4 days ago
74

Evaluate agent and skill behavior or routing. An agent skill from TheGoat395/Codex-Skills.

TheGoat395/Codex-Skills126—~1.5kAutomated safety check: PassMIT27 days ago
75

Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js).

livekit-examples/agent-starter-python2641 repo~1.9kAutomated safety check: PassMIT2 days ago
76

Run evaluation pipelines on Claude Code skills to test triggering accuracy, workflow correctness, and output quality.

AgriciDaniel/skill-forge177—~1.7kAutomated safety check: PassMIT6 mo ago
77
77.Eval FaqOfficial

Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &…

microsoft/eval-guide138—~10kAutomated safety check: PassMIT3 mo ago
78

Pulls pattern, trend and benchmark data from the claude-view MCP server to present a behavioral analysis of your Claude Code usage, grouped by category.

tombelieber/claude-view110—~3.1kAutomated safety check: PassMIT7 days ago
79

Runs Chat Customizations Evaluations analysis for the active customization file and summarizes the findings from the Problems panel.

microsoft/vscode-chat-customizations-evaluation141—~135Automated safety check: PassUnknown5 days ago
80

Points agents to the right sources when changing the Smithers workspace graph, generated docs, benchmark and eval evidence, or durable flows in the Smithers repository.

smithersai/smithers429—~2.3kAutomated safety check: PassMITtoday
81

Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.

Prism-Shadow/penguin-harness2.5k—~2.5kAutomated safety check: PassApache-2.0yesterday
82

Run and improve the c0 agent evaluation loop using the local and remote Cloudflare MCP harness, versioned history notes, and current Cloudflare product updates.

Consensys/c0105—~1.2kAutomated safety check: PassLGPL-3.01 mo ago
83

Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.

ww-w-ai/bkit-claude-code600—~1kAutomated safety check: NotesApache-2.010 days ago
84
84.Learn From PROfficial

Analyzes a finished pull request that involved an agent to find what slowed or helped it, and recommends specific changes to instruction files, skills and docs.

dotnet/maui23k—~2.5kAutomated safety check: PassMITyesterday
85

Analyzes Copilot Studio evaluation results using Practical Guidance on Agent Evaluation's 10-step playbook (Steps 6, 7, and 9) plus Microsoft's triage diagnostics.

microsoft/eval-guide138—~9.9kAutomated safety check: PassMIT3 mo ago
86

Improves a single skill through a test, fix and retest loop, keeping or reverting each change based on how the static and category scores move.

Donchitos/Claude-Code-Game-Studios26k—~1.6kAutomated safety check: NotesMIT8 days ago
87

Agentforce agent testing with dual-track workflow and 100-point scoring.

Jaganpro/sf-skills424—~2.2kAutomated safety check: PassMIT5 mo ago
88

Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it.

druxt/druxt.js114—~926Automated safety check: PassMITyesterday
89

Investigates what went wrong in a superpowers session by reading its transcript, reports findings with path and line citations, and can draft a GitHub issue or redacted bundle.

jnMetaCode/superpowers-zh8.3k—~858Automated safety check: PassMIT3 days ago
90

Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.

Arenukvern/mcp_flutter385—~2.4kAutomated safety check: PassMIT4 days ago
91

A skill your agent uses when evaluating, designing, or pressure-testing the business model of an AI agent product.

harperreed/dotfiles334—~3kAutomated safety check: PassNo licence4 days ago
92

Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.

microsoft/eval-guide138—~2.3kAutomated safety check: PassMIT3 mo ago
93

Agent skill for benchmark-suite - invoke with $agent-benchmark-suite

ruvnet/ruflo74k2 repos~4.9kAutomated safety check: PassMITtoday
94

Agent skill for test-long-runner - invoke with $agent-test-long-runner

ruvnet/ruflo74k2 repos~426Automated safety check: PassMITtoday
95

Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering…

fabricioctelles/skills105—~3.8kAutomated safety check: PassApache-2.03 days ago
96

Measure whether a skill helps by comparing runs with and without it.

boshu2/agentops4471 repo~2.8kAutomated safety check: PassApache-2.0yesterday