Search

Python · Agent evaluation and testing

9 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

anthropics/skills180k64 repos~2.3kAutomated safety check: PassApache-2.02 days ago
2

Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

HKUDS/OpenHarness16k1 repo~2.1kAutomated safety check: NotesMIT4 mo ago
3

驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

OpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.01 mo ago
4

Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

amd/gaia1.6k—~2.3kAutomated safety check: PassMITtoday
5

Contrasts successful and failed agent runs of the same task and derives guidelines backed by evidence from transcripts, tool calls and outcome judgments.

AgentToolkit/altk-evolve122—~1.7kAutomated safety check: PassApache-2.0today
6

Run and maintain Con's terminal-agent benchmark against a live app session.

nowledge-co/con-terminal625—~822Automated safety check: PassMITtoday
7

Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles.

Orchestra-Research/AI-Research-SKILLs13k1 repo~3.6kAutomated safety check: PassMIT3 mo ago
8

Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js).

livekit-examples/agent-starter-python2641 repo~1.9kAutomated safety check: PassMIT2 days ago
9

Run the instrumented specification search in the user's own statistical language — Stata (reghdfe, ivreghdfe, rdrobust, did2s), R (fixest, rdrobust, did2s), Python (statsmodels, linearmodels) or…

brycewang-stanford/Auto-Empirical-Research-Skills4.5k—~2.1kAutomated safety check: PassUnknown2 days ago