Search
Agent evaluation and testing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Agent skill for test-long-runner - invoke with $agent-test-long-runner | ruvnet/ | 74k | 2 repos | ~426 | Automated safety check: Pass | MIT | yesterday |
| 98 | Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering… | fabricioctelles/ | 106 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 99 | 99.Skill Eval Measure whether a skill helps by comparing runs with and without it. | boshu2/ | 448 | 1 repo | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 100 | 100.Agent Evaluation Evaluate agent behavior with versioned cases and explicit verifiers. | sickn33/ | 47k | 1 repo | ~2k | Automated safety check: Pass | MIT | today |
| 101 | A skill your agent uses when my-pi-agent tests, typecheck, lint, Pi CLI startup, OpenCode/MCP, Feishu channel, prompt rendering, memory/state, web-console/desktop build, or UI behavior fails… | skuramatata/ | 114 | — | ~777 | Automated safety check: Pass | No licence | 3 mo ago |
| 102 | 102.Agent Evaluation Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world… | davila7/ | 33k | 2 repos | ~519 | Automated safety check: Pass | MIT | today |
| 103 | 103.Agent Evaluation Design reproducible evaluations for AI agents with representative task sets, explicit rubrics, appropriate graders, baselines, regression gates, and failure analysis. | seb1n/ | 206 | — | ~1.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 104 | Launch Devolutions Gateway from source for local coding-agent tests with loopback listeners by default; require explicit opt-in before binding to non-localhost addresses. | Devolutions/ | 162 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 105 | A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable. | sickn33/ | 47k | 1 repo | ~2.1k | Automated safety check: Pass | MIT | 2 days ago |
| 106 | 106.Evaluate Evaluate an AI workflow against representative cases and its permitted actions. | suboss87/ | 957 | — | ~344 | Automated safety check: Pass | MIT | yesterday |
| 107 | A skill your agent uses when my-pi-agent tests, typecheck, lint, /code workflow, verifier probes, Feishu channel handling, MCP bootstrap, skill install, or task resume behavior fails unexpectedly. | skuramatata/ | 114 | — | ~829 | Automated safety check: Pass | No licence | 3 mo ago |
| 108 | 108.Sf Testing Apex test execution, coverage analysis, and test-fix loops with 120-point scoring. | Jaganpro/ | 424 | — | ~1.1k | Automated safety check: Pass | MIT | 5 mo ago |
| 109 | Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks. | sickn33/ | 47k | 1 repo | ~1.3k | Automated safety check: Pass | MIT | 2 days ago |
| 110 | 110.Run Deep Swe Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent. | sickn33/ | 47k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | 2 days ago |
| 111 | Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics. | affaan-m/ | 276k | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 112 | 112.Agent Designer A skill your agent uses when the user asks to design a multi-agent system, pick an orchestration pattern (supervisor/swarm/pipeline), generate tool schemas for agents, or evaluate agent execution… | alirezarezvani/ | 28k | — | ~1.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 113 | 113.Sf AI Agentforce Agentforce Builder metadata path for Builder-managed topics/actions, Prompt Builder templates, GenAiFunction/GenAiPlugin, Models API, and custom Lightning types. | Jaganpro/ | 424 | — | ~2.6k | Automated safety check: Pass | MIT | 5 mo ago |
| 114 | 114.Agent Evaluation Evaluate and improve Claude Code commands, skills, and agents. | NeoLabHQ/ | 1.8k | — | ~14k | Automated safety check: Pass | GPL-3.0 | 1 mo ago |
| 115 | 法律 Skill 分层质量评测工具。消费 skill-lint 的通用质量结论,再用三份测试材料、通用六维度、场景微调和律师 taste 评估法律产出,并定位最小修复单元。本技能应在审查、回归验证或发布验收法律 Skill 时使用。不要用于代替通用 Skill lint、正式法律意见或跨场景排名。 | cat-xierluo/ | 720 | — | ~2.2k | Automated safety check: Pass | CC-BY-NC-4.0 | yesterday |
| 116 | Apex test execution, coverage analysis, and test-fix loops with 120-point scoring. | forcedotcom/ | 1.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 117 | 117.Agent Evaluation Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on re... | benjaminasterA/ | 375 | 1 repo | ~522 | Automated safety check: Pass | MIT | yesterday |
| 118 | 118.Skill Creator Create, improve, and evaluate agent skills (SKILL.md plus reference files). | himself65/ | 3.4k | — | ~3.8k | Automated safety check: Pass | MIT | 6 days ago |
| 119 | Generate human and AI agent test scripts from user journey specifications. | paralleldrive/ | 384 | — | ~971 | Automated safety check: Pass | MIT | 4 mo ago |
| 120 | 120.Toolkit Toolkit management: create and evaluate skills and agents, manage routing tables, generate Claude.md. | notque/ | 441 | — | ~3.1k | Automated safety check: Notes | MIT | yesterday |
| 121 | 121.Ax Agent Rlm This skill helps an LLM generate correct AxAgent RLM/runtime code using @ax-llm/ax. | dosco/ | 107 | — | ~8.2k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 122 | 122.Map Skill Eval Evaluate a $map- skill's trigger accuracy and cost. An agent skill from azalio/map-framework. | azalio/ | 156 | — | ~2.7k | Automated safety check: Pass | MIT | 3 days ago |
| 123 | 123.Map Skill Eval Evaluate a /map- skill's trigger accuracy and cost. An agent skill from azalio/map-framework. | azalio/ | 156 | — | ~2.7k | Automated safety check: Pass | MIT | 3 days ago |
| 124 | 124.Phack Polyglot Run the instrumented specification search in the user's own statistical language — Stata (reghdfe, ivreghdfe, rdrobust, did2s), R (fixest, rdrobust, did2s), Python (statsmodels, linearmodels) or… | brycewang-stanford/ | 4.6k | — | ~2.1k | Automated safety check: Pass | Unknown | 5 days ago |
| 125 | Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. | vercel/ | 301 | — | ~3.6k | Automated safety check: Pass | Unknown | yesterday |
| 126 | 126.Skill Creator Create, edit, and evaluate agent skills iteratively. An agent skill from HezaoHezao/poirot. | HezaoHezao/ | 249 | — | ~838 | Automated safety check: Pass | MIT | 2 mo ago |
| 127 | 127.Testing Testing: TDD, E2E, preferred patterns, test-value audits, verification, agent testing. | notque/ | 441 | — | ~4.7k | Automated safety check: Notes | MIT | yesterday |
| 128 | 128.Testing Strategy A skill your agent uses when planning or reviewing tests for a code change, choosing between unit, focused regression, integration, contract, end-to-end, performance, security, property-based, or… | AnastasiyaW/ | 154 | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 129 | 129.Ak Test Set up testing and debug common issues in Agent Kernel projects. | yaalalabs/ | 192 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 130 | 130.Eval Plan and run conversational AI agent evaluations with test generation and analysis. | mikeyobrien/ | 373 | — | ~9.7k | Automated safety check: Pass | MIT | 10 days ago |
| 131 | 131.Aeon Skill Evals Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation. | BankrBot/ | 1.2k | — | ~660 | Automated safety check: Pass | No licence | yesterday |
| 132 | MLflow 3 GenAI agent evaluation. An agent skill from databricks/databricks-agent-skills. | databricks/ | 345 | — | ~2.7k | Automated safety check: Pass | Unknown | yesterday |
| 133 | 133.Bare Eval Run isolated eval and grading calls using CC 2.1.81 --bare mode. | yonatangross/ | 292 | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 134 | Routes any legal task to the right LLM, like OpenRouter but for legal work and grounded in benchmarks instead of brand loyalty. | lawve-ai/ | 847 | — | ~1.5k | Automated safety check: Pass | AGPL-3.0-or-later | 8 days ago |