Topic · Agent Workflows
Best agent evaluation and testing skills, page 3
Agent evaluation and testing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Design reproducible evaluations for AI agents with representative task sets, explicit rubrics, appropriate graders, baselines, regression gates, and failure analysis. | seb1n/ | 206 | — | ~1.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 98 | 98.Skill Eval Measure whether a skill helps by comparing runs with and without it. | boshu2/ | 447 | 1 repo | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 99 | Evaluate agent behavior with versioned cases and explicit verifiers. | sickn33/ | 47k | 1 repo | ~2k | Automated safety check: Pass | MIT | today |
| 100 | A skill your agent uses when my-pi-agent tests, typecheck, lint, Pi CLI startup, OpenCode/MCP, Feishu channel, prompt rendering, memory/state, web-console/desktop build, or UI behavior fails… | skuramatata/ | 114 | — | ~777 | Automated safety check: Pass | No licence | 2 mo ago |
| 101 | 101.Agent Evaluation Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world… | davila7/ | 32k | 2 repos | ~519 | Automated safety check: Pass | MIT | today |
| 102 | Launch Devolutions Gateway from source for local coding-agent tests with loopback listeners by default; require explicit opt-in before binding to non-localhost addresses. | Devolutions/ | 162 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 103 | A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable. | sickn33/ | 47k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | today |
| 104 | 104.Agent Evaluation Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on re... | benjaminasterA/ | 368 | 1 repo | ~522 | Automated safety check: Pass | MIT | today |
| 105 | 105.Evaluate Evaluate an AI workflow against representative cases and its permitted actions. | suboss87/ | 954 | — | ~344 | Automated safety check: Pass | MIT | today |
| 106 | A skill your agent uses when my-pi-agent tests, typecheck, lint, /code workflow, verifier probes, Feishu channel handling, MCP bootstrap, skill install, or task resume behavior fails unexpectedly. | skuramatata/ | 114 | — | ~829 | Automated safety check: Pass | No licence | 2 mo ago |
| 107 | 107.Skill Creator Create, improve, and evaluate agent skills (SKILL.md plus reference files). | himself65/ | 3.4k | — | ~3.8k | Automated safety check: Pass | MIT | 3 days ago |
| 108 | Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks. | sickn33/ | 47k | 1 repo | ~1.3k | Automated safety check: Pass | MIT | today |
| 109 | 109.Run Deep Swe Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent. | sickn33/ | 47k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | today |
| 110 | 110.Sf Testing Apex test execution, coverage analysis, and test-fix loops with 120-point scoring. | Jaganpro/ | 424 | — | ~1.1k | Automated safety check: Pass | MIT | 5 mo ago |
| 111 | Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics. | affaan-m/ | 275k | — | ~1.5k | Automated safety check: Pass | MIT | 3 days ago |
| 112 | 112.Agent Designer A skill your agent uses when the user asks to design a multi-agent system, pick an orchestration pattern (supervisor/swarm/pipeline), generate tool schemas for agents, or evaluate agent execution… | alirezarezvani/ | 28k | — | ~1.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 113 | 113.Sf AI Agentforce Agentforce Builder metadata path for Builder-managed topics/actions, Prompt Builder templates, GenAiFunction/GenAiPlugin, Models API, and custom Lightning types. | Jaganpro/ | 424 | — | ~2.6k | Automated safety check: Pass | MIT | 5 mo ago |
| 114 | 法律 Skill 分层质量评测工具。消费 skill-lint 的通用质量结论,再用三份测试材料、通用六维度、场景微调和律师 taste 评估法律产出,并定位最小修复单元。本技能应在审查、回归验证或发布验收法律 Skill 时使用。不要用于代替通用 Skill lint、正式法律意见或跨场景排名。 | cat-xierluo/ | 713 | — | ~2.2k | Automated safety check: Pass | CC-BY-NC-4.0 | today |
| 115 | Apex test execution, coverage analysis, and test-fix loops with 120-point scoring. | forcedotcom/ | 1.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 116 | Generate human and AI agent test scripts from user journey specifications. | paralleldrive/ | 384 | — | ~971 | Automated safety check: Pass | MIT | 3 mo ago |
| 117 | 117.Toolkit Toolkit management: create and evaluate skills and agents, manage routing tables, generate Claude.md. | notque/ | 438 | — | ~3.1k | Automated safety check: Notes | MIT | 5 days ago |
| 118 | 118.Agent Evaluation Evaluate and improve Claude Code commands, skills, and agents. | NeoLabHQ/ | 1.7k | — | ~14k | Automated safety check: Pass | GPL-3.0 | 1 mo ago |
| 119 | 119.Ax Agent Rlm This skill helps an LLM generate correct AxAgent RLM/runtime code using @ax-llm/ax. | dosco/ | 107 | — | ~8.2k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 120 | 120.Map Skill Eval Evaluate a $map- skill's trigger accuracy and cost. An agent skill from azalio/map-framework. | azalio/ | 156 | — | ~2.7k | Automated safety check: Pass | MIT | today |
| 121 | 121.Map Skill Eval Evaluate a /map- skill's trigger accuracy and cost. An agent skill from azalio/map-framework. | azalio/ | 156 | — | ~2.7k | Automated safety check: Pass | MIT | today |
| 122 | 122.Phack Polyglot Run the instrumented specification search in the user's own statistical language — Stata (reghdfe, ivreghdfe, rdrobust, did2s), R (fixest, rdrobust, did2s), Python (statsmodels, linearmodels) or… | brycewang-stanford/ | 4.5k | — | ~2.1k | Automated safety check: Pass | Unknown | 2 days ago |
| 123 | Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. | vercel/ | 301 | — | ~3.6k | Automated safety check: Pass | Unknown | today |
| 124 | 124.Testing Testing: TDD, E2E, preferred patterns, test-value audits, verification, agent testing. | notque/ | 438 | — | ~4.7k | Automated safety check: Notes | MIT | 5 days ago |
| 125 | 125.Testing Strategy A skill your agent uses when planning or reviewing tests for a code change, choosing between unit, focused regression, integration, contract, end-to-end, performance, security, property-based, or… | AnastasiyaW/ | 154 | — | ~2k | Automated safety check: Pass | MIT | today |
| 126 | 126.Ak Test Set up testing and debug common issues in Agent Kernel projects. | yaalalabs/ | 191 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | today |
| 127 | 127.Eval Plan and run conversational AI agent evaluations with test generation and analysis. | mikeyobrien/ | 372 | — | ~9.7k | Automated safety check: Pass | MIT | 7 days ago |
| 128 | 128.Aeon Skill Evals Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation. | BankrBot/ | 1.2k | — | ~660 | Automated safety check: Pass | No licence | 2 days ago |
| 129 | MLflow 3 GenAI agent evaluation. An agent skill from databricks/databricks-agent-skills. | databricks/ | 345 | — | ~2.7k | Automated safety check: Pass | Unknown | today |
| 130 | 130.Bare Eval Run isolated eval and grading calls using CC 2.1.81 --bare mode. | yonatangross/ | 289 | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 131 | 131.Rllm Use rLLM for language-agent evaluation, dataset/task management, RL/SFT post-training, CLI setup, and gateway-backed rollout tracing. | majiayu000/ | 666 | 1 repo | ~977 | Automated safety check: Pass | Apache-2.0 | today |
| 132 | Routes any legal task to the right LLM, like OpenRouter but for legal work and grounded in benchmarks instead of brand loyalty. | lawve-ai/ | 836 | — | ~1.5k | Automated safety check: Pass | AGPL-3.0-or-later | 5 days ago |
| 133 | 133.Nw Agent Testing 5-layer testing approach for agent validation including adversarial testing, security validation, and prompt injection resistance | nWave-ai/ | 617 | — | ~882 | Automated safety check: Pass | MIT | 21 days ago |
Explore related skills
Category
More topics in Agent Workflows
- MCP servers2,051
- Subagents1,011
- Agent instruction files739
- Multi-agent orchestration526
- Agent memory505
- Skill authoring500
- Planning495
- Brainstorming463
- Hooks and plugins356
- Autonomous loops324
- Context engineering301
- Session handoff259
- Requirements gathering230
- Task breakdown227
- Human-in-the-loop approvals202
- Skill management195
- Codebase knowledge for agents190
- Verification before completion170