Search

Agent evaluation and testing

134 skills found, page 3.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
97

Agent skill for test-long-runner - invoke with $agent-test-long-runner

ruvnet/ruflo74k2 repos~426Automated safety check: PassMITyesterday
98

Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering…

fabricioctelles/skills106—~3.8kAutomated safety check: PassApache-2.0today
99

Measure whether a skill helps by comparing runs with and without it.

boshu2/agentops4481 repo~2.8kAutomated safety check: PassApache-2.0today
100

Evaluate agent behavior with versioned cases and explicit verifiers.

sickn33/agentic-awesome-skills47k1 repo~2kAutomated safety check: PassMITtoday
101

A skill your agent uses when my-pi-agent tests, typecheck, lint, Pi CLI startup, OpenCode/MCP, Feishu channel, prompt rendering, memory/state, web-console/desktop build, or UI behavior fails…

skuramatata/my-pi-agent114—~777Automated safety check: PassNo licence3 mo ago
102

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world…

davila7/claude-code-templates33k2 repos~519Automated safety check: PassMITtoday
103

Design reproducible evaluations for AI agents with representative task sets, explicit rubrics, appropriate graders, baselines, regression gates, and failure analysis.

seb1n/awesome-ai-agent-skills206—~1.4kAutomated safety check: PassMIT2 mo ago
104

Launch Devolutions Gateway from source for local coding-agent tests with loopback listeners by default; require explicit opt-in before binding to non-localhost addresses.

Devolutions/devolutions-gateway162—~1.8kAutomated safety check: PassApache-2.0yesterday
105

A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

sickn33/agentic-awesome-skills47k1 repo~2.1kAutomated safety check: PassMIT2 days ago
106

Evaluate an AI workflow against representative cases and its permitted actions.

suboss87/FDEOps957—~344Automated safety check: PassMITyesterday
107

A skill your agent uses when my-pi-agent tests, typecheck, lint, /code workflow, verifier probes, Feishu channel handling, MCP bootstrap, skill install, or task resume behavior fails unexpectedly.

skuramatata/my-pi-agent114—~829Automated safety check: PassNo licence3 mo ago
108

Apex test execution, coverage analysis, and test-fix loops with 120-point scoring.

Jaganpro/sf-skills424—~1.1kAutomated safety check: PassMIT5 mo ago
109

Autonomous AI agent benchmark evaluation register: task completion rates, planning accuracy, tool invocation precision, and cost benchmarks.

sickn33/agentic-awesome-skills47k1 repo~1.3kAutomated safety check: PassMIT2 days ago
110

Run reproducible DeepSWE coding-agent benchmark evaluations through OpenRouter and mini-swe-agent.

sickn33/agentic-awesome-skills47k1 repo~1.2kAutomated safety check: PassMIT2 days ago
111

Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

affaan-m/ECC276k—~1.5kAutomated safety check: PassMITyesterday
112

A skill your agent uses when the user asks to design a multi-agent system, pick an orchestration pattern (supervisor/swarm/pipeline), generate tool schemas for agents, or evaluate agent execution…

alirezarezvani/claude-skills28k—~1.1kAutomated safety check: PassMIT1 mo ago
113

Agentforce Builder metadata path for Builder-managed topics/actions, Prompt Builder templates, GenAiFunction/GenAiPlugin, Models API, and custom Lightning types.

Jaganpro/sf-skills424—~2.6kAutomated safety check: PassMIT5 mo ago
114

Evaluate and improve Claude Code commands, skills, and agents.

NeoLabHQ/context-engineering-kit1.8k—~14kAutomated safety check: PassGPL-3.01 mo ago
115

法律 Skill 分层质量评测工具。消费 skill-lint 的通用质量结论,再用三份测试材料、通用六维度、场景微调和律师 taste 评估法律产出,并定位最小修复单元。本技能应在审查、回归验证或发布验收法律 Skill 时使用。不要用于代替通用 Skill lint、正式法律意见或跨场景排名。

cat-xierluo/legal-skills720—~2.2kAutomated safety check: PassCC-BY-NC-4.0yesterday
116

Apex test execution, coverage analysis, and test-fix loops with 120-point scoring.

forcedotcom/sf-skills1.1k—~1.9kAutomated safety check: PassApache-2.0yesterday
117

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on re...

benjaminasterA/antigravity-awesome-skills3751 repo~522Automated safety check: PassMITyesterday
118

Create, improve, and evaluate agent skills (SKILL.md plus reference files).

himself65/finance-skills3.4k—~3.8kAutomated safety check: PassMIT6 days ago
119

Generate human and AI agent test scripts from user journey specifications.

paralleldrive/aidd384—~971Automated safety check: PassMIT4 mo ago
120

Toolkit management: create and evaluate skills and agents, manage routing tables, generate Claude.md.

notque/vexjoy-agent441—~3.1kAutomated safety check: NotesMITyesterday
121

This skill helps an LLM generate correct AxAgent RLM/runtime code using @ax-llm/ax.

dosco/aithy107—~8.2kAutomated safety check: PassApache-2.01 mo ago
122

Evaluate a $map- skill's trigger accuracy and cost. An agent skill from azalio/map-framework.

azalio/map-framework156—~2.7kAutomated safety check: PassMIT3 days ago
123

Evaluate a /map- skill's trigger accuracy and cost. An agent skill from azalio/map-framework.

azalio/map-framework156—~2.7kAutomated safety check: PassMIT3 days ago
124

Run the instrumented specification search in the user's own statistical language — Stata (reghdfe, ivreghdfe, rdrobust, did2s), R (fixest, rdrobust, did2s), Python (statsmodels, linearmodels) or…

brycewang-stanford/Auto-Empirical-Research-Skills4.6k—~2.1kAutomated safety check: PassUnknown5 days ago
125
125.Benchmark AgentsOfficial

Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

vercel/vercel-plugin301—~3.6kAutomated safety check: PassUnknownyesterday
126

Create, edit, and evaluate agent skills iteratively. An agent skill from HezaoHezao/poirot.

HezaoHezao/poirot249—~838Automated safety check: PassMIT2 mo ago
127

Testing: TDD, E2E, preferred patterns, test-value audits, verification, agent testing.

notque/vexjoy-agent441—~4.7kAutomated safety check: NotesMITyesterday
128

A skill your agent uses when planning or reviewing tests for a code change, choosing between unit, focused regression, integration, contract, end-to-end, performance, security, property-based, or…

AnastasiyaW/codex-claude-code-config154—~2kAutomated safety check: PassMITyesterday
129

Set up testing and debug common issues in Agent Kernel projects.

yaalalabs/agent-kernel192—~3.9kAutomated safety check: PassApache-2.0yesterday
130
130.Eval

Plan and run conversational AI agent evaluations with test generation and analysis.

mikeyobrien/rho373—~9.7kAutomated safety check: PassMIT10 days ago
131

Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation.

BankrBot/skills1.2k—~660Automated safety check: PassNo licenceyesterday
132

MLflow 3 GenAI agent evaluation. An agent skill from databricks/databricks-agent-skills.

databricks/databricks-agent-skills345—~2.7kAutomated safety check: PassUnknownyesterday
133

Run isolated eval and grading calls using CC 2.1.81 --bare mode.

yonatangross/orchestkit292—~2.2kAutomated safety check: PassMITyesterday
134

Routes any legal task to the right LLM, like OpenRouter but for legal work and grounded in benchmarks instead of brand loyalty.

lawve-ai/awesome-legal-skills847—~1.5kAutomated safety check: PassAGPL-3.0-or-later8 days ago