Topic · Agent Workflows

Best agent evaluation and testing skills, page 3

Skills #97–133 of 133, ranked by score.

Agent evaluation and testing skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Agent evaluation and testing skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
97

Design reproducible evaluations for AI agents with representative task sets, explicit rubrics, appropriate graders, baselines, regression gates, and failure analysis.

seb1n/awesome-ai-agent-skills206—~1.4kAutomated safety check: PassMIT1 mo ago
98

Measure whether a skill helps by comparing runs with and without it.

boshu2/agentops4471 repo~2.8kAutomated safety check: PassApache-2.0today
99

Evaluate agent behavior with versioned cases and explicit verifiers.

sickn33/agentic-awesome-skills47k1 repo~2kAutomated safety check: PassMITtoday
100

A skill your agent uses when my-pi-agent tests, typecheck, lint, Pi CLI startup, OpenCode/MCP, Feishu channel, prompt rendering, memory/state, web-console/desktop build, or UI behavior fails…

skuramatata/my-pi-agent114—~777Automated safety check: PassNo licence2 mo ago
101

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world…

davila7/claude-code-templates32k2 repos~519Automated safety check: PassMITtoday
102

Launch Devolutions Gateway from source for local coding-agent tests with loopback listeners by default; require explicit opt-in before binding to non-localhost addresses.

Devolutions/devolutions-gateway162—~1.8kAutomated safety check: PassApache-2.0today
103

A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMITtoday
104

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on re...

benjaminasterA/antigravity-awesome-skills3681 repo~522Automated safety check: PassMITtoday
105

Evaluate an AI workflow against representative cases and its permitted actions.

suboss87/FDEOps954—~344Automated safety check: PassMITtoday
106

A skill your agent uses when my-pi-agent tests, typecheck, lint, /code workflow, verifier probes, Feishu channel handling, MCP bootstrap, skill install, or task resume behavior fails unexpectedly.

skuramatata/my-pi-agent114—~829Automated safety check: PassNo licence2 mo ago
107

Create, improve, and evaluate agent skills (SKILL.md plus reference files).

himself65/finance-skills3.4k—~3.8kAutomated safety check: PassMIT3 days ago
108

Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks.

sickn33/agentic-awesome-skills47k1 repo~1.3kAutomated safety check: PassMITtoday
109

Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.

sickn33/agentic-awesome-skills47k1 repo~1.2kAutomated safety check: PassMITtoday
110

Apex test execution, coverage analysis, and test-fix loops with 120-point scoring.

Jaganpro/sf-skills424—~1.1kAutomated safety check: PassMIT5 mo ago
111

Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

affaan-m/ECC275k—~1.5kAutomated safety check: PassMIT3 days ago
112

A skill your agent uses when the user asks to design a multi-agent system, pick an orchestration pattern (supervisor/swarm/pipeline), generate tool schemas for agents, or evaluate agent execution…

alirezarezvani/claude-skills28k—~1.1kAutomated safety check: PassMIT1 mo ago
113

Agentforce Builder metadata path for Builder-managed topics/actions, Prompt Builder templates, GenAiFunction/GenAiPlugin, Models API, and custom Lightning types.

Jaganpro/sf-skills424—~2.6kAutomated safety check: PassMIT5 mo ago
114

法律 Skill 分层质量评测工具。消费 skill-lint 的通用质量结论,再用三份测试材料、通用六维度、场景微调和律师 taste 评估法律产出,并定位最小修复单元。本技能应在审查、回归验证或发布验收法律 Skill 时使用。不要用于代替通用 Skill lint、正式法律意见或跨场景排名。

cat-xierluo/legal-skills713—~2.2kAutomated safety check: PassCC-BY-NC-4.0today
115

Apex test execution, coverage analysis, and test-fix loops with 120-point scoring.

forcedotcom/sf-skills1.1k—~1.9kAutomated safety check: PassApache-2.0today
116

Generate human and AI agent test scripts from user journey specifications.

paralleldrive/aidd384—~971Automated safety check: PassMIT3 mo ago
117

Toolkit management: create and evaluate skills and agents, manage routing tables, generate Claude.md.

notque/vexjoy-agent438—~3.1kAutomated safety check: NotesMIT5 days ago
118

Evaluate and improve Claude Code commands, skills, and agents.

NeoLabHQ/context-engineering-kit1.7k—~14kAutomated safety check: PassGPL-3.01 mo ago
119

This skill helps an LLM generate correct AxAgent RLM/runtime code using @ax-llm/ax.

dosco/aithy107—~8.2kAutomated safety check: PassApache-2.01 mo ago
120

Evaluate a $map- skill's trigger accuracy and cost. An agent skill from azalio/map-framework.

azalio/map-framework156—~2.7kAutomated safety check: PassMITtoday
121

Evaluate a /map- skill's trigger accuracy and cost. An agent skill from azalio/map-framework.

azalio/map-framework156—~2.7kAutomated safety check: PassMITtoday
122

Run the instrumented specification search in the user's own statistical language — Stata (reghdfe, ivreghdfe, rdrobust, did2s), R (fixest, rdrobust, did2s), Python (statsmodels, linearmodels) or…

brycewang-stanford/Auto-Empirical-Research-Skills4.5k—~2.1kAutomated safety check: PassUnknown2 days ago
123
123.Benchmark AgentsOfficial

Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

vercel/vercel-plugin301—~3.6kAutomated safety check: PassUnknowntoday
124

Testing: TDD, E2E, preferred patterns, test-value audits, verification, agent testing.

notque/vexjoy-agent438—~4.7kAutomated safety check: NotesMIT5 days ago
125

A skill your agent uses when planning or reviewing tests for a code change, choosing between unit, focused regression, integration, contract, end-to-end, performance, security, property-based, or…

AnastasiyaW/codex-claude-code-config154—~2kAutomated safety check: PassMITtoday
126

Set up testing and debug common issues in Agent Kernel projects.

yaalalabs/agent-kernel191—~3.9kAutomated safety check: PassApache-2.0today
127
127.Eval

Plan and run conversational AI agent evaluations with test generation and analysis.

mikeyobrien/rho372—~9.7kAutomated safety check: PassMIT7 days ago
128

Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation.

BankrBot/skills1.2k—~660Automated safety check: PassNo licence2 days ago
129

MLflow 3 GenAI agent evaluation. An agent skill from databricks/databricks-agent-skills.

databricks/databricks-agent-skills345—~2.7kAutomated safety check: PassUnknowntoday
130

Run isolated eval and grading calls using CC 2.1.81 --bare mode.

yonatangross/orchestkit289—~2.2kAutomated safety check: PassMITtoday
131
131.Rllm

Use rLLM for language-agent evaluation, dataset/task management, RL/SFT post-training, CLI setup, and gateway-backed rollout tracing.

majiayu000/claude-skill-registry6661 repo~977Automated safety check: PassApache-2.0today
132

Routes any legal task to the right LLM, like OpenRouter but for legal work and grounded in benchmarks instead of brand loyalty.

lawve-ai/awesome-legal-skills836—~1.5kAutomated safety check: PassAGPL-3.0-or-later5 days ago
133

5-layer testing approach for agent validation including adversarial testing, security validation, and prompt injection resistance

nWave-ai/nWave617—~882Automated safety check: PassMIT21 days ago