Search
Agent evaluation and testing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks. | amd/ | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 50 | 50.Benchflow Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow. | benchflow-ai/ | 356 | — | ~1.9k | Automated safety check: Notes | Apache-2.0 | 4 days ago |
| 51 | Collects evidence from a failed or confusing skill-driven task and triggers skvm jit-optimize to propose concrete fixes to that skill's files. | SJTU-IPADS/ | 561 | — | ~2.9k | Automated safety check: Warn | MIT | 1 mo ago |
| 52 | Runs Darwin Mode on an agent harness: mutates one policy file per generation in a sandbox, scores each variant against your tests, and archives only variants that measurably improve. | ruvnet/ | 97k | — | ~693 | Automated safety check: Pass | MIT | today |
| 53 | Operational notes for generating, evaluating and debugging strategies in the autocontext grid_ctf scenario, with tier rules and parameter ranges that worked or failed. | greyhaven-ai/ | 1.3k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 54 | Operates or analyzes a LoopX-managed benchmark experiment: launching runs, maintaining the experiment board, qualifying integrity, and writing case insights. | loopx-project/ | 6.2k | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | today |
| 55 | Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low. | microsoft/ | 1.4k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 56 | Evaluate mathematical expressions and unit conversions. An agent skill from NVIDIA/SkillEvaluator. | NVIDIA/ | 554 | — | ~442 | Automated safety check: Pass | Apache-2.0 | today |
| 57 | Audits recent agent conversation history and turns repeated failures and successes into testable harness improvement proposals that later audits can check. | KonghaYao/ | 226 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | today |
| 58 | Run and maintain Con's terminal-agent benchmark against a live app session. | nowledge-co/ | 626 | — | ~822 | Automated safety check: Pass | MIT | today |
| 59 | 59.Email Evals Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents. | tokencanopy/ | 193 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 60 | A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps… | microsoft/ | 138 | — | ~5.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 61 | Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship. | edonadei/ | 209 | — | ~2.1k | Automated safety check: Notes | MIT | today |
| 62 | Create new skills, modify and improve existing skills, and measure skill performance. | ZS520L/ | 103 | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | 4 mo ago |
| 63 | Static safety audit of a SKILL.md that scores five dimensions and acts as a gate: skills below the pass line do not ship, whatever else they score. | openJiuwen-ai/ | 446 | — | ~3.5k | Automated safety check: Warn | Apache-2.0 | today |
| 64 | 64.Skill Judge Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples. | shareAI-lab/ | 315 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 24 days ago |
| 65 | Analyze text content and produce statistics including word count, line count, character count, most frequent words, and readability metrics. | NVIDIA/ | 554 | — | ~453 | Automated safety check: Pass | Apache-2.0 | today |
| 66 | [omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. | rlaope/ | 3.2k | — | ~2.1k | Automated safety check: Pass | MIT | today |
| 67 | 67.Tester Test and validate AgentUse agents without real side effects. | agentuse/ | 205 | — | ~1.8k | Automated safety check: Notes | Unknown | 3 days ago |
| 68 | Evaluate catalog candidates against either the standard five agent-native criteria or the narrow operator-surface track, and check URL Onboarding. | haoruilee/ | 477 | — | ~2.8k | Automated safety check: Pass | CC0-1.0 | yesterday |
| 69 | A skill your agent uses when changing Deep Researcher Agent continuous integration, pre-commit, or contributor governance — editing .github/workflows/ (ci, ui, skills-eval, request-nvskills-ci)… | NVIDIA-AI-Blueprints/ | 886 | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 70 | Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately. | microsoft/ | 138 | — | ~22k | Automated safety check: Warn | MIT | 3 mo ago |
| 71 | Framework for measuring and tracking agent response quality over time. | vibeeval/ | 532 | — | ~2.9k | Automated safety check: Pass | MIT | 2 mo ago |
| 72 | Runbook for the five-arm Codex harness comparison on Terminal-Bench 4.0, pointing out four differences from SWE-Marathon that silently produce wrong scores. | loopx-project/ | 6.2k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 73 | Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification. | bgauryy/ | 949 | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 74 | Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result. | Prism-Shadow/ | 2.5k | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 75 | Evaluate agent and skill behavior or routing. An agent skill from TheGoat395/Codex-Skills. | TheGoat395/ | 126 | — | ~1.5k | Automated safety check: Pass | MIT | 29 days ago |
| 76 | Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js). | livekit-examples/ | 264 | 1 repo | ~1.9k | Automated safety check: Pass | MIT | today |
| 77 | Run evaluation pipelines on Claude Code skills to test triggering accuracy, workflow correctness, and output quality. | AgriciDaniel/ | 179 | — | ~1.7k | Automated safety check: Pass | MIT | 6 mo ago |
| 78 | Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &… | microsoft/ | 138 | — | ~10k | Automated safety check: Pass | MIT | 3 mo ago |
| 79 | Pulls pattern, trend and benchmark data from the claude-view MCP server to present a behavioral analysis of your Claude Code usage, grouped by category. | tombelieber/ | 111 | — | ~3.1k | Automated safety check: Pass | MIT | 10 days ago |
| 80 | Runs Chat Customizations Evaluations analysis for the active customization file and summarizes the findings from the Problems panel. | microsoft/ | 142 | — | ~135 | Automated safety check: Pass | Unknown | 7 days ago |
| 81 | Creates, edits and benchmarks skills; supersedes the official skill-creator plugin, so when both are listed, use this one. | daymade/ | 1.4k | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | today |
| 82 | Points agents to the right sources when changing the Smithers workspace graph, generated docs, benchmark and eval evidence, or durable flows in the Smithers repository. | smithersai/ | 430 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 83 | Run and improve the c0 agent evaluation loop using the local and remote Cloudflare MCP harness, versioned history notes, and current Cloudflare product updates. | Consensys/ | 105 | — | ~1.2k | Automated safety check: Pass | LGPL-3.0 | 1 mo ago |
| 84 | 84.Bkit Evals Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results. | ww-w-ai/ | 601 | — | ~1k | Automated safety check: Notes | Apache-2.0 | 13 days ago |
| 85 | Analyzes a finished pull request that involved an agent to find what slowed or helped it, and recommends specific changes to instruction files, skills and docs. | dotnet/ | 23k | — | ~2.5k | Automated safety check: Pass | MIT | today |
| 86 | Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles. | Orchestra-Research/ | 13k | — | ~3.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 87 | Analyzes Copilot Studio evaluation results using Practical Guidance on Agent Evaluation's 10-step playbook (Steps 6, 7, and 9) plus Microsoft's triage diagnostics. | microsoft/ | 138 | — | ~9.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 88 | Improves a single skill through a test, fix and retest loop, keeping or reverting each change based on how the static and category scores move. | Donchitos/ | 26k | — | ~1.7k | Automated safety check: Notes | MIT | 2 days ago |
| 89 | Agentforce agent testing with dual-track workflow and 100-point scoring. | Jaganpro/ | 424 | — | ~2.2k | Automated safety check: Pass | MIT | 5 mo ago |
| 90 | 90.Write Skill Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it. | druxt/ | 114 | — | ~926 | Automated safety check: Pass | MIT | today |
| 91 | Investigates what went wrong in a superpowers session by reading its transcript, reports findings with path and line citations, and can draft a GitHub issue or redacted bundle. | jnMetaCode/ | 8.3k | — | ~858 | Automated safety check: Pass | MIT | 2 days ago |
| 92 | Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates. | Arenukvern/ | 387 | — | ~2.4k | Automated safety check: Pass | MIT | 7 days ago |
| 93 | A skill your agent uses when evaluating, designing, or pressure-testing the business model of an AI agent product. | harperreed/ | 334 | — | ~3k | Automated safety check: Pass | No licence | 7 days ago |
| 94 | Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. | microsoft/ | 138 | — | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 95 | Agent skill for benchmark-suite - invoke with $agent-benchmark-suite | ruvnet/ | 74k | 2 repos | ~4.9k | Automated safety check: Pass | MIT | today |
| 96 | Agent skill for test-long-runner - invoke with $agent-test-long-runner | ruvnet/ | 74k | 2 repos | ~426 | Automated safety check: Pass | MIT | today |