Search
Agent evaluation and testing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation. | anthropics/ | 180k | 62 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 2 | Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers. | obra/ | 296k | 3 repos | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 3 | Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints. | alchaincyf/ | 6.2k | 1 repo | ~4.7k | Automated safety check: Pass | MIT | 19 days ago |
| 4 | Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution. | HKUDS/ | 16k | 1 repo | ~2.1k | Automated safety check: Notes | MIT | 4 mo ago |
| 5 | Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability. | rohitg00/ | 65k | — | ~1k | Automated safety check: Pass | MIT | yesterday |
| 6 | Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version. | colbymchenry/ | 73k | — | ~950 | Automated safety check: Pass | MIT | yesterday |
| 7 | Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals. | dotnet/ | 23k | — | ~3.4k | Automated safety check: Pass | MIT | yesterday |
| 8 | Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks. | aipoch/ | 5.4k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 9 | Evaluates the design quality of an agent skill against official specifications and patterns from existing examples, scoring it and suggesting improvements. | shareAI-lab/ | 5.2k | 4 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 10 | Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan. | Galaxy-Dawn/ | 5.7k | 1 repo | ~3k | Automated safety check: Pass | MIT | 14 days ago |
| 11 | 11.Skill Test Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit. | databricks-solutions/ | 1.9k | — | ~1.9k | Automated safety check: Pass | Unknown | 1 mo ago |
| 12 | Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop. | aiming-lab/ | 15k | — | ~1.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 13 | Evaluate skill quality, find the weakest dimension, and apply directed improvements. | Evol-ai/ | 216 | 1 repo | ~3.1k | Automated safety check: Pass | MIT | 5 mo ago |
| 14 | Call any REST API dynamically. An agent skill from NVIDIA/SkillEvaluator. | NVIDIA/ | 544 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 15 | Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons. | windmill-labs/ | 18k | — | ~969 | Automated safety check: Notes | Unknown | yesterday |
| 16 | Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier. | langchain-ai/ | 1.3k | — | ~4k | Automated safety check: Pass | MIT | 2 days ago |
| 17 | Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI. | greyhaven-ai/ | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 18 | Designs and verifies a deterministic grader that measures whether a GitHub Agentic Workflow run reached its real-world or repository outcome. | github/ | 5.3k | — | ~6.8k | Automated safety check: Pass | MIT | yesterday |
| 19 | End-to-end agent evaluation and improvement pipeline. An agent skill from kayba-ai/agentic-context-engine. | kayba-ai/ | 2.6k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 20 | 20.Team Tasks Coordinate multi-agent development pipelines using shared JSON task files. | win4r/ | 447 | — | ~2.9k | Automated safety check: Pass | No licence | 8 mo ago |
| 21 | Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report. | bgauryy/ | 946 | — | ~2.1k | Automated safety check: Pass | MIT | 4 days ago |
| 22 | Checks changes to the senpi coding agent by driving the real CLI from source in an isolated sandbox, over RPC, terminal UI, mock model and CLI smoke channels. | code-yeongyu/ | 470 | — | ~2.7k | Automated safety check: Notes | MIT | yesterday |
| 23 | 23.One Eval 驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。 | OpenDCAI/ | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 24 | 24.Autocontext Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met. | greyhaven-ai/ | 1.3k | — | ~892 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 25 | Runs existing Harbor evaluation jobs against a local LobeHub build or LobeHub Cloud, with preflight checks, resume support, and failure triage. | lobehub/ | 83k | — | ~1k | Automated safety check: Notes | Unknown | yesterday |
| 26 | Operates or analyzes a LoopX-managed benchmark experiment: launching runs, maintaining the experiment board, qualifying integrity, and writing case insights. | loopx-project/ | 6.2k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 27 | AWS Bedrock AgentCore comprehensive expert for deploying and managing AI agents at scale. | zxkane/ | 367 | 1 repo | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 28 | 28.Harness Eval Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports. | tech-leads-club/ | 7k | — | ~3.9k | Automated safety check: Pass | CC-BY-4.0 | 17 days ago |
| 29 | Runs a reference agent (Codex or Claude Code) and Qwen Code on the same scenario, captures HTTP and terminal traces, and compares them until behavior matches. | QwenLM/ | 28k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 30 | Required for 4+ step requests; add tasks at start and update status after each step. | NVIDIA/ | 544 | — | ~460 | Automated safety check: Pass | Apache-2.0 | today |
| 31 | A skill your agent uses when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. | microsoft/ | 1k | — | ~2k | Automated safety check: Notes | Unknown | today |
| 32 | Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces. | affaan-m/ | 274k | 1 repo | ~623 | Automated safety check: Pass | MIT | 2 days ago |
| 33 | Walks you through drafting, testing, evaluating and tuning a skill for OpenCode, from an intake interview to description optimization. | antongulin/ | 172 | — | ~8.1k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 34 | Agent Script DSL for deterministic Agentforce agents. An agent skill from Jaganpro/sf-skills. | Jaganpro/ | 424 | — | ~3.8k | Automated safety check: Pass | MIT | 5 mo ago |
| 35 | Runs and interprets a skill's Caliper eval: how often it succeeds over repeated attempts, whether it triggers at all, and whether it beats the agent without it. | edonadei/ | 206 | — | ~1.9k | Automated safety check: Notes | MIT | 2 days ago |
| 36 | 36.Skill Forge Ultimate Claude Code skill creator and architect. An agent skill from AgriciDaniel/skill-forge. | AgriciDaniel/ | 177 | — | ~1.9k | Automated safety check: Notes | MIT | 6 mo ago |
| 37 | Scores an agent's finished work with a three-stage pipeline: free mechanical checks, an advisory semantic review, and an optional multi-model consensus vote. | Q00/ | 6.2k | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 38 | Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch. | ory/ | 305 | — | ~497 | Automated safety check: Pass | Unknown | 1 mo ago |
| 39 | Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end. | microsoft/ | 255 | 1 repo | ~6.7k | Automated safety check: Pass | MIT | yesterday |
| 40 | A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. | NVIDIA/ | 544 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 41 | Run the weak-agent adversarial test harness against docx-cli. | kklimuk/ | 215 | — | ~6.1k | Automated safety check: Notes | MIT | 12 days ago |
| 42 | Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs. | amd/ | 1.6k | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 43 | Runs caliper's smoke evals against the real agent CLIs after a harness or MCP change, with a dry-run plan, failure triage and a report to attach to the PR. | edonadei/ | 206 | — | ~664 | Automated safety check: Pass | MIT | 2 days ago |
| 44 | Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments. | AgentToolkit/ | 122 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 45 | Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter. | microsoft/ | 1.4k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 46 | Evaluate Expo skills in this repo end-to-end - trigger accuracy, generated code quality, and runtime screenshots on iOS simulator and Android emulator via Expo Go (web optional). | expo/ | 2.7k | — | ~12k | Automated safety check: Pass | MIT | yesterday |
| 47 | 47.Eval Answer Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer. | malloydata/ | 116 | — | ~4.3k | Automated safety check: Pass | MIT | today |
| 48 | Designs, evaluates and improves the harness around an AI agent: tool permissions, approval gates, state, memory, evals and observability, with phased plans. | NateBJones-Projects/ | 4.7k | — | ~1.8k | Automated safety check: Pass | Unknown | yesterday |