Search

Python · LLM evaluation

24 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

Orchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT3 mo ago
2

Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

FailproofAI/failproofai5.3k—~6kAutomated safety check: PassUnknownyesterday
3

Create a skill or bring an existing one up to the same standard (validate + asm eval fix loop); run evals, tune triggering.

luongnv89/asm954—~5.3kAutomated safety check: PassMIT2 days ago
4

Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

ai-evals-course/evals-skills1.5k—~2.2kAutomated safety check: PassApache-2.014 days ago
5

Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

microsoft/skills3.1k—~2.8kAutomated safety check: PassMITtoday
6

Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager.

wso2/agent-manager105—~710Automated safety check: PassApache-2.0today
7

Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.

timothywarner-org/claude-code224—~696Automated safety check: NotesMIT2 mo ago
8
8.Logfire EvalsOfficial

Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire.

pydantic/skills140—~3.6kAutomated safety check: PassMIT7 days ago
9

Create a new built-in classification evaluator for Phoenix evals.

Arize-ai/phoenix12k—~2.3kAutomated safety check: PassApache-2.0today
10

Bump the next release-please version for a Phoenix Python package (arize-phoenix, arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) by opening a PR with a Release-As commit footer.

Arize-ai/phoenix12k—~708Automated safety check: PassApache-2.0today
11

Audit recent changes to Phoenix's user-facing surfaces (Python clients, TypeScript clients, CLI, REST/GraphQL APIs) and patch the three external-facing agent skills — phoenix-tracing, phoenix-cli…

Arize-ai/phoenix12k—~5.1kAutomated safety check: PassUnknowntoday
12

A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.

tikalk/adlc-team-skills141—~1.4kAutomated safety check: PassMITtoday
13

Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm…

comet-ml/opik-mcp220—~3kAutomated safety check: NotesApache-2.0today
14
14.Phoenix EvalsOfficial

Build and run evaluators for AI/LLM applications using Phoenix.

github/awesome-copilot40k2 repos~1.1kAutomated safety check: PassApache-2.0today
15
15.Model EvaluationOfficial

Generates python code that evaluates SageMaker models. An agent skill from awslabs/agent-plugins.

awslabs/agent-plugins915—~1.3kAutomated safety check: PassApache-2.0today
16

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1582 repos~3.1kAutomated safety check: PassMITyesterday
17
17.Eval Driven DevOfficial

Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~4.4kAutomated safety check: WarnMITtoday
18

Configures and runs LLM evaluation using Promptfoo framework.

daymade/claude-code-skills1.4k—~3kAutomated safety check: PassMITtoday
19

Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

ClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMITtoday
20

Compile an Anthropic-style skill — a directory with a SKILL.md and optional references/ — into a deterministic, runnable workflow via the rote CLI.

ccplugins/awesome-claude-code-plugins970—~1.3kAutomated safety check: NotesApache-2.01 mo ago
21

A skill your agent uses when integrating NVIDIA NeMo Fabric into a consumer application, service, evaluation harness, or platform through the typed Python SDK — translating the consumer's own…

NVIDIA/skills3.5k—~5.5kAutomated safety check: PassApache-2.0today
22

Best practices for building AI agents with Google's Agent Development Kit (ADK) in Python, covering agent design, tools, sessions, memory, artifacts, evaluation, and deployment.

Mindrally/skills269—~2.5kAutomated safety check: PassApache-2.0today
23

Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit…

datadog-labs/agent-skills177—~25kAutomated safety check: PassMITtoday
24

Build type-safe AI agents and graph-based workflows with PydanticAI and PydanticGraph.

magnus919/agent-skills116—~4kAutomated safety check: PassMITtoday