Topic · AI & LLM Engineering
Best LLM evaluation skills, page 5
LLM evaluation skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 193 | 193.Create Skill Author a new Dex skill that actually fires and passes the quality bar. | davekilleen/ | 493 | — | ~2.9k | Automated safety check: Pass | Unknown | 6 days ago |
| 194 | This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise… | aiskillstore/ | 430 | 4 repos | ~4.2k | Automated safety check: Pass | No licence | yesterday |
| 195 | 195.Eval Harness Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k… | affaan-m/ | 275k | — | ~2.2k | Automated safety check: Pass | MIT | 3 days ago |
| 196 | 196.Evals Init A skill your agent uses when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline. | tikalk/ | 141 | — | ~934 | Automated safety check: Pass | MIT | 2 days ago |
| 197 | 197.Eval Evaluate and rank agent results by metric or LLM judge for an AgentHub session. | alirezarezvani/ | 28k | 1 repo | ~618 | Automated safety check: Pass | MIT | 1 mo ago |
| 198 | Entry point for Pydantic Logfire — an observability, monitoring, and evals platform. | pydantic/ | 140 | — | ~1.5k | Automated safety check: Pass | MIT | 7 days ago |
| 199 | 199.Evals Build a regression + eval harness for AI-written code and AI features. | Houseofmvps/ | 123 | — | ~1.1k | Automated safety check: Notes | MIT | 3 mo ago |
| 200 | 200.Respond To Eval Turn student course evaluations (free-text + numeric) into an actionable teaching-improvement plan — the teaching analogue of /respond-to-referees. | pedrohcgs/ | 1.6k | — | ~2.6k | Automated safety check: Notes | MIT | 10 days ago |
| 201 | Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 3 days ago |
| 202 | 202.Agenthub Multi-agent collaboration plugin that spawns N parallel subagents competing on the same task via git worktree isolation. | alirezarezvani/ | 28k | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 203 | Teaches how to write and run evals on the products/posthogai/evalharness/ harness — sandboxed agent suites that execute the real coding agent in a Docker or Modal sandbox against a seeded Hedgebox… | PostHog/ | 40k | — | ~4k | Automated safety check: Notes | Unknown | today |
| 204 | Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. | google/ | 21k | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | today |
| 205 | 205.Skill Creator Create, improve, and evaluate agent skills (SKILL.md plus reference files). | himself65/ | 3.4k | — | ~3.8k | Automated safety check: Pass | MIT | 3 days ago |
| 206 | 206.Evals Specify A skill your agent uses when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria. | tikalk/ | 141 | — | ~915 | Automated safety check: Pass | MIT | 2 days ago |
| 207 | Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics. | affaan-m/ | 275k | — | ~1.5k | Automated safety check: Pass | MIT | 3 days ago |
| 208 | 用于医疗应用部署的患者安全评估工具。针对CDSS准确性、PHI暴露、临床工作流完整性和集成合规性的自动化测试套件。在安全故障时阻止部署。 | affaan-m/ | 275k | — | ~1.4k | Automated safety check: Pass | MIT | 3 days ago |
| 209 | 209.Metric Design A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining… | agentscope-ai/ | 868 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 210 | 210.Aidd Riteway AI Teaches agents how to write correct riteway ai prompt evals (.sudo files) for multi-step flows that involve tool calls. | paralleldrive/ | 384 | — | ~1.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 211 | Run Azure SDK QA bot evaluations on curated datasets locally, including a single test case. | Azure/ | 134 | — | ~1.1k | Automated safety check: Notes | MIT | today |
| 212 | lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 139 | 2 repos | ~3.1k | Automated safety check: Pass | MIT | today |
| 213 | Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified. | PostHog/ | 40k | — | ~6.7k | Automated safety check: Pass | Unknown | today |
| 214 | Create offline evaluation tests for Output SDK workflows using @outputai/evals. | growthxai/ | 440 | — | ~3.8k | Automated safety check: Notes | Apache-2.0 | today |
| 215 | 215.Agent Evals Build automated evaluation suites for AI agents using golden datasets, rubrics, and regression gates. | sickn33/ | 47k | 2 repos | ~3.1k | Automated safety check: Warn | MIT | yesterday |
| 216 | 216.Benchmark Runner Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs. | RConsortium/ | 118 | — | ~5.3k | Automated safety check: Pass | MIT | 4 days ago |
| 217 | Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot. | github/ | 40k | 1 repo | ~4.4k | Automated safety check: Warn | MIT | today |
| 218 | Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded… | PostHog/ | 40k | — | ~1.5k | Automated safety check: Pass | Unknown | today |
| 219 | INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith. | langchain-ai/ | 1.3k | — | ~8.7k | Automated safety check: Notes | MIT | 2 days ago |
| 220 | 220.ML Machine learning and LLM engineering judgment, distilled from a stronger model - invoke when DECIDING whether/how to use ML or an LLM for a task (prompt vs RAG vs fine-tune vs classical); working… | telagod/ | 243 | 1 repo | ~566 | Automated safety check: Pass | MIT | 2 mo ago |
| 221 | 221.Autoresearch Autonomously optimize any Claude Code skill by running it repeatedly, scoring outputs against binary evals, mutating the prompt, and keeping improvements. | pedronauck/ | 634 | — | ~3.7k | Automated safety check: Pass | No licence | 23 days ago |
| 222 | 222.Testing Boss Author or review software tests and LLM/agent evals; choose test placement and mocks, diagnose flaky CI, or repair brittle suites. | pedronauck/ | 634 | 1 repo | ~636 | Automated safety check: Pass | No licence | 23 days ago |
| 223 | 223.Evals Validate A skill your agent uses when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy). | tikalk/ | 141 | — | ~850 | Automated safety check: Pass | MIT | 2 days ago |
| 224 | Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment). | PostHog/ | 40k | — | ~5.7k | Automated safety check: Pass | Unknown | today |
| 225 | Set up an LLM-judge evaluation that extracts canonical use cases for a PostHog feature at scale and streams the results to a Slack channel as a live feed. | PostHog/ | 40k | — | ~7.6k | Automated safety check: Pass | Unknown | today |
| 226 | A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection. | Aperivue/ | 329 | — | ~2.4k | Automated safety check: Pass | MIT | 3 days ago |
| 227 | 227.Eval Creator [Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them. | pskoett/ | 311 | — | ~2.6k | Automated safety check: Pass | No licence | 3 days ago |
| 228 | 228.Pre Flight Check [Beta] Session-start scan that surfaces relevant learnings, recent errors, and eval status before work begins. | pskoett/ | 311 | — | ~1.4k | Automated safety check: Pass | No licence | 3 days ago |
| 229 | 229.Readme Showcase Build or refresh a product README showcase using a seeded Bag of Words workspace, polished in-product screenshots, and repository-ready visual assets. | bagofwords1/ | 458 | — | ~1.8k | Automated safety check: Pass | Unknown | today |
| 230 | 230.Yao Meta Skill Create, refactor, evaluate, and package agent skills from workflows, prompts, transcripts, docs, or notes. | aiskillstore/ | 430 | 2 repos | ~806 | Automated safety check: Pass | MIT | yesterday |
| 231 | Audit an existing eval suite for trustworthiness. An agent skill from growthxai/output. | growthxai/ | 440 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | today |
| 232 | Design effective LLM judge .prompt files for evaluators. An agent skill from growthxai/output. | growthxai/ | 440 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | today |
| 233 | Configures and runs LLM evaluation using Promptfoo framework. | daymade/ | 1.4k | — | ~3k | Automated safety check: Pass | MIT | today |
| 234 | Eval-driven skill tuning. An agent skill from ClawBio/ClawBio. | ClawBio/ | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 235 | Patterns for continuous autonomous agent loops with quality gates, evals, and recovery controls. | majiayu000/ | 666 | 6 repos | ~268 | Automated safety check: Pass | MIT | today |
| 236 | Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy… | open-thoughts/ | 301 | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 237 | 237.Testing Skill validation framework PLUS daily test-suite health and regression intelligence. | inbrainfun/ | 142 | 1 repo | ~2k | Automated safety check: Pass | Unknown | 2 mo ago |
| 238 | 238.Benchmark A skill your agent uses when choosing between engines or models for an agent pipeline and the answer must come from measurement on your own data, not from marketing pages: speech recognition for… | serejaris/ | 229 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 239 | 239.LLM Eval Harness Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression. | daymade/ | 1.4k | — | ~4.7k | Automated safety check: Pass | MIT | today |
| 240 | 240.Prompt Optimize A skill your agent uses when the user has a prompt that feeds a system they can already score, and wants that prompt automatically improved to raise the score against their own evaluation command. | gaasher/ | 174 | — | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- LLM inference and serving372
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- GPU and accelerator computing176
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23