Search
Python · LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints. | Orchestra-Research/ | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | 3 mo ago |
| 2 | Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs. | FailproofAI/ | 5.3k | — | ~6k | Automated safety check: Pass | Unknown | yesterday |
| 3 | Create a skill or bring an existing one up to the same standard (validate + asm eval fix loop); run evals, tune triggering. | luongnv89/ | 954 | — | ~5.3k | Automated safety check: Pass | MIT | 2 days ago |
| 4 | Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data. | ai-evals-course/ | 1.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 5 | Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes. | microsoft/ | 3.1k | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 6 | Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager. | wso2/ | 105 | — | ~710 | Automated safety check: Pass | Apache-2.0 | today |
| 7 | Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships. | timothywarner-org/ | 224 | — | ~696 | Automated safety check: Notes | MIT | 2 mo ago |
| 8 | Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire. | pydantic/ | 140 | — | ~3.6k | Automated safety check: Pass | MIT | 7 days ago |
| 9 | Create a new built-in classification evaluator for Phoenix evals. | Arize-ai/ | 12k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 10 | Bump the next release-please version for a Phoenix Python package (arize-phoenix, arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) by opening a PR with a Release-As commit footer. | Arize-ai/ | 12k | — | ~708 | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Audit recent changes to Phoenix's user-facing surfaces (Python clients, TypeScript clients, CLI, REST/GraphQL APIs) and patch the three external-facing agent skills — phoenix-tracing, phoenix-cli… | Arize-ai/ | 12k | — | ~5.1k | Automated safety check: Pass | Unknown | today |
| 12 | A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness. | tikalk/ | 141 | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 13 | Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm… | comet-ml/ | 220 | — | ~3k | Automated safety check: Notes | Apache-2.0 | today |
| 14 | Build and run evaluators for AI/LLM applications using Phoenix. | github/ | 40k | 2 repos | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 15 | Generates python code that evaluates SageMaker models. An agent skill from awslabs/agent-plugins. | awslabs/ | 915 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 16 | lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 158 | 2 repos | ~3.1k | Automated safety check: Pass | MIT | yesterday |
| 17 | Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot. | github/ | 40k | 1 repo | ~4.4k | Automated safety check: Warn | MIT | today |
| 18 | Configures and runs LLM evaluation using Promptfoo framework. | daymade/ | 1.4k | — | ~3k | Automated safety check: Pass | MIT | today |
| 19 | Eval-driven skill tuning. An agent skill from ClawBio/ClawBio. | ClawBio/ | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 20 | 20.Compile Compile an Anthropic-style skill — a directory with a SKILL.md and optional references/ — into a deterministic, runnable workflow via the rote CLI. | ccplugins/ | 970 | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | 1 mo ago |
| 21 | A skill your agent uses when integrating NVIDIA NeMo Fabric into a consumer application, service, evaluation harness, or platform through the typed Python SDK — translating the consumer's own… | NVIDIA/ | 3.5k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 22 | 22.Google Adk Best practices for building AI agents with Google's Agent Development Kit (ADK) in Python, covering agent design, tools, sessions, memory, artifacts, evaluation, and deployment. | Mindrally/ | 269 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | today |
| 23 | Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit… | datadog-labs/ | 177 | — | ~25k | Automated safety check: Pass | MIT | today |
| 24 | 24.Pydanticai Build type-safe AI agents and graph-based workflows with PydanticAI and PydanticGraph. | magnus919/ | 116 | — | ~4k | Automated safety check: Pass | MIT | today |