Search

Agent evaluation and testing

133 skills found, page 2.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

amd/gaia1.6k—~1.8kAutomated safety check: PassMITyesterday
50

Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

benchflow-ai/benchflow356—~1.9kAutomated safety check: NotesApache-2.04 days ago
51

Collects evidence from a failed or confusing skill-driven task and triggers skvm jit-optimize to propose concrete fixes to that skill's files.

SJTU-IPADS/SkVM561—~2.9kAutomated safety check: WarnMIT1 mo ago
52

Runs Darwin Mode on an agent harness: mutates one policy file per generation in a sandbox, scores each variant against your tests, and archives only variants that measurably improve.

ruvnet/RuView97k—~693Automated safety check: PassMITtoday
53

Operational notes for generating, evaluating and debugging strategies in the autocontext grid_ctf scenario, with tier rules and parameter ranges that worked or failed.

greyhaven-ai/autocontext1.3k—~1.3kAutomated safety check: PassApache-2.03 days ago
54

Operates or analyzes a LoopX-managed benchmark experiment: launching runs, maintaining the experiment board, qualifying integrity, and writing case insights.

loopx-project/loopx6.2k—~4.3kAutomated safety check: PassApache-2.0today
55
55.Waza InteractiveOfficial

Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

microsoft/waza1.4k—~1.3kAutomated safety check: PassMITtoday
56
56.CalculatorOfficial

Evaluate mathematical expressions and unit conversions. An agent skill from NVIDIA/SkillEvaluator.

NVIDIA/SkillEvaluator554—~442Automated safety check: PassApache-2.0today
57

Audits recent agent conversation history and turns repeated failures and successes into testable harness improvement proposals that later audits can check.

KonghaYao/peri226—~3.5kAutomated safety check: PassApache-2.0today
58

Run and maintain Con's terminal-agent benchmark against a live app session.

nowledge-co/con-terminal626—~822Automated safety check: PassMITtoday
59

Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.

tokencanopy/e2a193—~2.1kAutomated safety check: PassApache-2.04 days ago
60

A skill your agent uses when the user's Copilot Studio agent evaluations have come back and they need to interpret scores, diagnose root causes of underperforming test cases, find remediation steps…

microsoft/eval-guide138—~5.9kAutomated safety check: PassMIT3 mo ago
61

Interviews you about what a skill should do, writes an eval spec from your answers, then loops through run, diagnose and improve until the skill is ready to ship.

edonadei/caliper209—~2.1kAutomated safety check: NotesMITtoday
62

Create new skills, modify and improve existing skills, and measure skill performance.

ZS520L/HanakoPro103—~7.4kAutomated safety check: PassApache-2.04 mo ago
63

Static safety audit of a SKILL.md that scores five dimensions and acts as a gate: skills below the pass line do not ship, whatever else they score.

openJiuwen-ai/agent-core446—~3.5kAutomated safety check: WarnApache-2.0today
64

Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.

shareAI-lab/lab-skills315—~1.9kAutomated safety check: PassApache-2.024 days ago
65
65.Text AnalyzerOfficial

Analyze text content and produce statistics including word count, line count, character count, most frequent words, and readability metrics.

NVIDIA/SkillEvaluator554—~453Automated safety check: PassApache-2.0today
66

[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics.

rlaope/oh-my-hermes3.2k—~2.1kAutomated safety check: PassMITtoday
67

Test and validate AgentUse agents without real side effects.

agentuse/agentuse205—~1.8kAutomated safety check: NotesUnknown3 days ago
68

Evaluate catalog candidates against either the standard five agent-native criteria or the narrow operator-surface track, and check URL Onboarding.

haoruilee/awesome-agent-native-services477—~2.8kAutomated safety check: PassCC0-1.0yesterday
69

A skill your agent uses when changing Deep Researcher Agent continuous integration, pre-commit, or contributor governance — editing .github/workflows/ (ci, ui, skills-eval, request-nvskills-ci)…

NVIDIA-AI-Blueprints/deep-researcher-agent886—~1.5kAutomated safety check: NotesApache-2.0yesterday
70
70.Eval GuideOfficial

Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

microsoft/eval-guide138—~22kAutomated safety check: WarnMIT3 mo ago
71

Framework for measuring and tracking agent response quality over time.

vibeeval/vibecosystem532—~2.9kAutomated safety check: PassMIT2 mo ago
72

Runbook for the five-arm Codex harness comparison on Terminal-Bench 4.0, pointing out four differences from SWE-Marathon that silently produce wrong scores.

loopx-project/loopx6.2k—~1.6kAutomated safety check: PassApache-2.0today
73

Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

bgauryy/octocode949—~1.6kAutomated safety check: PassMITyesterday
74

Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.

Prism-Shadow/penguin-harness2.5k—~3.4kAutomated safety check: PassApache-2.0today
75

Evaluate agent and skill behavior or routing. An agent skill from TheGoat395/Codex-Skills.

TheGoat395/Codex-Skills126—~1.5kAutomated safety check: PassMIT29 days ago
76

Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js).

livekit-examples/agent-starter-python2641 repo~1.9kAutomated safety check: PassMITtoday
77

Run evaluation pipelines on Claude Code skills to test triggering accuracy, workflow correctness, and output quality.

AgriciDaniel/skill-forge179—~1.7kAutomated safety check: PassMIT6 mo ago
78
78.Eval FaqOfficial

Answers AI agent evaluation methodology questions with practical, opinionated guidance grounded primarily in Microsoft's agent evaluation ecosystem (MS Learn, Eval Scenario Library, Triage &…

microsoft/eval-guide138—~10kAutomated safety check: PassMIT3 mo ago
79

Pulls pattern, trend and benchmark data from the claude-view MCP server to present a behavioral analysis of your Claude Code usage, grouped by category.

tombelieber/claude-view111—~3.1kAutomated safety check: PassMIT10 days ago
80

Runs Chat Customizations Evaluations analysis for the active customization file and summarizes the findings from the Problems panel.

microsoft/vscode-chat-customizations-evaluation142—~135Automated safety check: PassUnknown7 days ago
81

Creates, edits and benchmarks skills; supersedes the official skill-creator plugin, so when both are listed, use this one.

daymade/claude-code-skills1.4k—~4.3kAutomated safety check: PassApache-2.0today
82

Points agents to the right sources when changing the Smithers workspace graph, generated docs, benchmark and eval evidence, or durable flows in the Smithers repository.

smithersai/smithers430—~2.3kAutomated safety check: PassMITtoday
83

Run and improve the c0 agent evaluation loop using the local and remote Cloudflare MCP harness, versioned history notes, and current Cloudflare product updates.

Consensys/c0105—~1.2kAutomated safety check: PassLGPL-3.01 mo ago
84

Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.

ww-w-ai/bkit-claude-code601—~1kAutomated safety check: NotesApache-2.013 days ago
85
85.Learn From PROfficial

Analyzes a finished pull request that involved an agent to find what slowed or helped it, and recommends specific changes to instruction files, skills and docs.

dotnet/maui23k—~2.5kAutomated safety check: PassMITtoday
86

Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles.

Orchestra-Research/AI-Research-SKILLs13k—~3.6kAutomated safety check: PassMIT3 mo ago
87

Analyzes Copilot Studio evaluation results using Practical Guidance on Agent Evaluation's 10-step playbook (Steps 6, 7, and 9) plus Microsoft's triage diagnostics.

microsoft/eval-guide138—~9.9kAutomated safety check: PassMIT3 mo ago
88

Improves a single skill through a test, fix and retest loop, keeping or reverting each change based on how the static and category scores move.

Donchitos/Claude-Code-Game-Studios26k—~1.7kAutomated safety check: NotesMIT2 days ago
89

Agentforce agent testing with dual-track workflow and 100-point scoring.

Jaganpro/sf-skills424—~2.2kAutomated safety check: PassMIT5 mo ago
90

Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it.

druxt/druxt.js114—~926Automated safety check: PassMITtoday
91

Investigates what went wrong in a superpowers session by reading its transcript, reports findings with path and line citations, and can draft a GitHub issue or redacted bundle.

jnMetaCode/superpowers-zh8.3k—~858Automated safety check: PassMIT2 days ago
92

Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.

Arenukvern/mcp_flutter387—~2.4kAutomated safety check: PassMIT7 days ago
93

A skill your agent uses when evaluating, designing, or pressure-testing the business model of an AI agent product.

harperreed/dotfiles334—~3kAutomated safety check: PassNo licence7 days ago
94

Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.

microsoft/eval-guide138—~2.3kAutomated safety check: PassMIT3 mo ago
95

Agent skill for benchmark-suite - invoke with $agent-benchmark-suite

ruvnet/ruflo74k2 repos~4.9kAutomated safety check: PassMITtoday
96

Agent skill for test-long-runner - invoke with $agent-test-long-runner

ruvnet/ruflo74k2 repos~426Automated safety check: PassMITtoday