Search
Python · Agent evaluation and testing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation. | anthropics/ | 180k | 64 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 2 | Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution. | HKUDS/ | 16k | 1 repo | ~2.1k | Automated safety check: Notes | MIT | 4 mo ago |
| 3 | 3.One Eval 驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。 | OpenDCAI/ | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 4 | Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs. | amd/ | 1.6k | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 5 | Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments. | AgentToolkit/ | 122 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Run and maintain Con's terminal-agent benchmark against a live app session. | nowledge-co/ | 625 | — | ~822 | Automated safety check: Pass | MIT | today |
| 7 | Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles. | Orchestra-Research/ | 13k | 1 repo | ~3.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 8 | Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js). | livekit-examples/ | 264 | 1 repo | ~1.9k | Automated safety check: Pass | MIT | 2 days ago |
| 9 | Run the instrumented specification search in the user's own statistical language — Stata (reghdfe, ivreghdfe, rdrobust, did2s), R (fixest, rdrobust, did2s), Python (statsmodels, linearmodels) or… | brycewang-stanford/ | 4.5k | — | ~2.1k | Automated safety check: Pass | Unknown | 2 days ago |