Topic · AI & LLM Engineering
Best LLM evaluation skills, page 3
LLM evaluation skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Finds top models for a task from official Hugging Face benchmark leaderboards, filters them to what fits your hardware, and returns a comparison table with scores. | huggingface/ | 11k | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 98 | Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link. | comet-ml/ | 219 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 99 | This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs… | pifferologo/ | 129 | 1 repo | ~6.8k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 100 | Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 101 | 101.Analyze Eval Investigate a single failing eval from the convex-evals system. | get-convex/ | 129 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 102 | 102.Skill Creator Create new skills, modify and improve existing skills, and measure skill performance. | ZS520L/ | 102 | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | 4 mo ago |
| 103 | 103.SEO Audit Audit a website and deliver a one-page, plain-language SEO report anyone can act on, centered on a single do-this-week action. | Jwuthri/ | 1.5k | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 104 | This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated… | guanyang/ | 975 | 2 repos | ~4.2k | Automated safety check: Pass | MIT | yesterday |
| 105 | 105.Add Evaluator Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager. | wso2/ | 105 | — | ~710 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 106 | 106.Eval Debate 测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。 | YIKUAIBANZI/ | 121 | — | ~740 | Automated safety check: Pass | MIT | 6 mo ago |
| 107 | 107.Skill Creator A skill your agent uses when the user wants to work on a Claude Code skill file (SKILL.md): writing one from scratch, testing whether an existing one works well, running evals or benchmarks… | avibebuilder/ | 120 | — | ~8.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 108 | 108.Agent Eval Cases Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic. | agentailor/ | 132 | — | ~5.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 109 | Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов). | AnastasiyaW/ | 154 | — | ~3.8k | Automated safety check: Pass | MIT | today |
| 110 | Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately. | microsoft/ | 138 | — | ~22k | Automated safety check: Warn | MIT | 3 mo ago |
| 111 | 111.Analyze Run Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations. | get-convex/ | 129 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 112 | 112.Test Afm Binary Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU… | scouzi1966/ | 345 | — | ~3.8k | Automated safety check: Pass | MIT | 3 days ago |
| 113 | 113.Project Evals Best practices for creating expectations and grader files to evaluate guidance quality. | GoogleChrome/ | 1.1k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 114 | 114.LLM Judge AI quality judge that scores agent responses 0-10 across helpfulness, accuracy, completeness, and clarity. | Atmosphere/ | 3.8k | — | ~333 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 115 | Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships. | timothywarner-org/ | 224 | — | ~696 | Automated safety check: Notes | MIT | 2 mo ago |
| 116 | 116.LLM Evaluation Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing. | davila7/ | 32k | 13 repos | ~3.5k | Automated safety check: Pass | MIT | today |
| 117 | Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification. | bgauryy/ | 949 | — | ~1.6k | Automated safety check: Pass | MIT | 5 days ago |
| 118 | Triage failing @novu/agent-evals scenarios to decide whether a failure is real or flaky, and whether to fix the playbook/prompt or the test (grader, tape, scenario, or judge). | novuhq/ | 40k | — | ~1.5k | Automated safety check: Pass | Unknown | today |
| 119 | 119.Woo AI Smoke Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring. | woocommerce/ | 358 | 1 repo | ~7.4k | Automated safety check: Notes | GPL-2.0 | today |
| 120 | 120.Skill Forge Eval Run evaluation pipelines on Claude Code skills to test triggering accuracy, workflow correctness, and output quality. | AgriciDaniel/ | 177 | — | ~1.7k | Automated safety check: Pass | MIT | 6 mo ago |
| 121 | 121.Loop Architect Design well-structured agent loops with best-practice coaching and cross-model review gates before you run them. | fabricioctelles/ | 106 | — | ~2.1k | Automated safety check: Notes | MIT | 4 days ago |
| 122 | 122.Eval Harness Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed. | Archive228/ | 756 | — | ~876 | Automated safety check: Pass | MIT | 2 mo ago |
| 123 | Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI. | diegosouzapw/ | 74k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 124 | 124.Eval Loop Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint. | malloydata/ | 116 | — | ~7.8k | Automated safety check: Pass | MIT | today |
| 125 | The step-by-step procedure for running this repo's Promptfoo evals with a generator and a judge over provider APIs. | maplibre/ | 151 | — | ~3.3k | Automated safety check: Pass | Unknown | 4 days ago |
| 126 | Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire. | pydantic/ | 140 | — | ~3.6k | Automated safety check: Pass | MIT | 7 days ago |
| 127 | 127.Nexus Testing Verify a Nexus change — pick a Jest lane, write a test that can actually fail, run the in-app Obsidian CLI loop, drive the eval harness, or fix a shipped-docs drift failure. | ProfSynapse/ | 156 | — | ~994 | Automated safety check: Pass | MIT | today |
| 128 | 128.Write A Spec Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer. | different-ai/ | 24k | — | ~3.3k | Automated safety check: Pass | Unknown | today |
| 129 | Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow. | githits-com/ | 114 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 130 | 130.Rde Eval Run a targeted local React Doctor Evals loop against an uncommitted rule change. | millionco/ | 15k | — | ~510 | Automated safety check: Pass | Unknown | today |
| 131 | 131.AI Evals Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies. | RefoundAI/ | 1.4k | — | ~1.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 132 | Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files. | elastic/ | 21k | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 133 | Trigger an on-demand @kbn/evals Buildkite run by describing what you want in plain English. | elastic/ | 21k | — | ~2.6k | Automated safety check: Notes | Unknown | today |
| 134 | 134.Eval Creator CI [Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). | pskoett/ | 311 | — | ~2.2k | Automated safety check: Pass | No licence | 3 days ago |
| 135 | 135.Paperclip Evals Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification. | paperclipai/ | 99k | — | ~839 | Automated safety check: Pass | MIT | today |
| 136 | 136.Bkit Evals Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results. | ww-w-ai/ | 601 | — | ~1k | Automated safety check: Notes | Apache-2.0 | 11 days ago |
| 137 | 137.Eval Plan Builds an eval plan for an AI feature - error analysis on real traces, failure-mode ranking, code checks and binary LLM judges validated against human labels, CI tests, production monitoring, and a… | menkesu/ | 429 | — | ~4.5k | Automated safety check: Pass | Unknown | 2 days ago |
| 138 | You MUST load this skill before the skill-creator skill AND before making ANY change to, or conducting a review of ANY Agent Skill. | sammcj/ | 162 | — | ~9.8k | Automated safety check: Pass | Apache-2.0 | today |
| 139 | Audit supervised fine-tuning datasets against the behavior and task they are meant to teach. | tokenbender/ | 367 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 140 | 140.Agent Eval A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual… | ericrisco/ | 167 | — | ~3.2k | Automated safety check: Pass | MIT | today |
| 141 | Set up compliance exports, drift detection, evaluations, scoring, and learning analytics | ucsandman/ | 310 | — | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 142 | Patterns and techniques for evaluating and improving AI agent outputs. | github/ | 40k | 4 repos | ~1.5k | Automated safety check: Pass | MIT | today |
| 143 | Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate. | ai-evals-course/ | 1.5k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 144 | Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats. | glebis/ | 389 | — | ~2.9k | Automated safety check: Pass | MIT | 12 days ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- LLM inference and serving372
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- GPU and accelerator computing176
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23