Topic · AI & LLM Engineering

Best LLM evaluation skills for Claude Code, Codex and other agents.

Skills that build evals and benchmarks to measure the quality of model and agent output.
skills
303
official
41

LLM evaluation skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM evaluation skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1
1.Skill CreatorOfficial

Create new skills, modify and improve existing skills, and measure skill performance.

Azure/azqr79489 repos~8.2kAutomated safety check: PassApache-2.02 days ago
2

Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself.

JuliusBrussee/caveman110k1 repo~975Automated safety check: PassApache-2.0today
3

Turns a Caveman report-only optimization observation into one minimal code change and a paired baseline evaluation, after the operator picks which to pursue.

JuliusBrussee/caveman110k1 repo~1.2kAutomated safety check: PassApache-2.0today
4

Provides context about the CoStrict evals system structure in this monorepo.

zgsm-ai/costrict4.4k1 repo~1.9kAutomated safety check: PassApache-2.07 days ago
5

Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

Orchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT3 mo ago
6

Create, improve, or evaluate an existing skill from workflows, prompts, SOPs, scripts.

yaojingang/yao-meta-skill2.7k—~768Automated safety check: PassMIT1 mo ago
7

Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

microsoft/skills3.1k6 repos~2.8kAutomated safety check: PassMITyesterday
8

Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.

chatboxai/chatbox42k—~758Automated safety check: PassGPL-3.013 days ago
9

Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

FailproofAI/failproofai5.3k—~6kAutomated safety check: PassUnknownyesterday
10

A skill your agent uses when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use'…

tripleyak/SkillForge905—~2.3kAutomated safety check: NotesMIT2 mo ago
11

Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

Jeffallan/claude-skills12k1 repo~1.7kAutomated safety check: PassMIT4 days ago
12

A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.

DenisSergeevitch/agents-best-practices2.4k—~7.4kAutomated safety check: PassMIT2 days ago
13

Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

ksimback/looper710—~2.7kAutomated safety check: NotesMIT1 mo ago
14

Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

huggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.06 days ago
15

Research, create, improve, migrate, evaluate, package, install-check, govern, and safely publish qiaomu-flavored agent skills from workflows, prompts, transcripts, docs, SOPs, runbooks, scripts, or…

joeseesun/qiaomu-meta-skill383—~2.8kAutomated safety check: PassMIT2 mo ago
16

Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

windmill-labs/windmill18k—~969Automated safety check: NotesUnknowntoday
17

Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

langchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMITyesterday
18

Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

cloudnative-co/claude-code-starter-kit15110 repos~1.3kAutomated safety check: PassMITtoday
19

Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

GoogleCloudPlatform/vertex-ai-samples791—~2kAutomated safety check: PassApache-2.0today
20

Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

bgauryy/octocode946—~2.1kAutomated safety check: PassMIT4 days ago
21

Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each…

ankit/stylebot1.6k—~1.1kAutomated safety check: PassMITtoday
22

亚马逊卖家专用的 skill 创建器(中文)。当用户想把一个亚马逊运营/自媒体/日常工作流程变成可复用的 skill 时使用。触发场景包括但不限于:用户说"我想做一个 skill""把这个流程变成 skill""帮我写个自动化""优化我已有的 skill""给这个工作流做个自动化",即使用户没用"skill"这个词,只要在描述"以后每次都这样做"的重复性工作时也应触发。本 skill…

zach22-1999/amazon-skills2061 repo~3.9kAutomated safety check: PassApache-2.01 mo ago
23

Create new skills, modify and improve existing skills, and measure skill performance.

AgentTeam-TaichuAI/ScienceClaw670—~10kAutomated safety check: PassApache-2.05 mo ago
24

SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.

benchflow-ai/benchflow353—~4.5kAutomated safety check: PassApache-2.0yesterday
25

Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

ai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.012 days ago
26

Create a skill or bring an existing one up to the same standard (validate + asm eval fix loop); run evals, tune triggering.

luongnv89/asm953—~5.3kAutomated safety check: PassMITyesterday
27

Create, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor.

smixs/skill-conductor179—~6.6kAutomated safety check: PassMIT2 mo ago
28

A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…

sun461941-hub/ai-project-copilot100—~3kAutomated safety check: PassMIT1 mo ago
29
29.Commerce EvalsOfficial

Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.

anthropics/commerce-agents3.2k—~1.8kAutomated safety check: PassApache-2.05 days ago
30

Evaluate verified findings from merge-ready, Greptile, pull-request, CI, security, billing, and other code reviews, then promote durable review gaps into the version-controlled .greptile…

Jwuthri/Tracely-ai1.5k—~1.1kAutomated safety check: PassMITyesterday
31

Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes.

frontier-harness-eval/eval297—~8kAutomated safety check: PassNo licence29 days ago
32

Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.

greyhaven-ai/autocontext1.3k—~892Automated safety check: PassApache-2.0today
33

Sends one prompt to Claude, GPT through the Codex CLI and Gemini, then tabulates response time, token use and cost, with an optional judged quality score.

garrytan/gstack136k—~4kAutomated safety check: NotesMITtoday
34

Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

ai-evals-course/evals-skills1.5k—~2.5kAutomated safety check: PassApache-2.012 days ago
35

Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports.

tech-leads-club/agent-skills7k—~3.9kAutomated safety check: PassCC-BY-4.017 days ago
36

A skill your agent uses when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI.

microsoft/work-iq1k—~2kAutomated safety check: NotesUnknowntoday
37

Audit an agent's context layout against the four places: system prompt, tools, history, tail.

undefined-ui/second-brain-os999—~810Automated safety check: PassMIT8 days ago
38

Turn the current conversation's workflow into a reusable agent skill.

Undertone0809/rudder292—~3.6kAutomated safety check: PassApache-2.0today
39

Plan a documentation eval for supabase/evals, where a docs guide is the subject under test.

supabase/evals143—~2.9kAutomated safety check: PassApache-2.0today
40

Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

ai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.012 days ago
41

Converts test suites from external eval frameworks into the Margin Eval suite format.

Margin-Lab/evals161—~3.7kAutomated safety check: PassAGPL-3.02 mo ago
42

Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~2.1kAutomated safety check: NotesApache-2.0today
43
43.Fde

Keeps the engagement record for client work. An agent skill from suboss87/FDEOps.

suboss87/FDEOps953—~2.9kAutomated safety check: PassMITyesterday
44
44.Review SkillOfficial

Review a proposed Agent Skill for structural validity and content quality before publishing.

mongodb/agent-skills190—~1.5kAutomated safety check: NotesApache-2.0yesterday
45

Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.

benchflow-ai/benchflow353—~4kAutomated safety check: PassApache-2.0yesterday
46

Sets up and runs side-by-side model evaluations with frozen questions, isolated tools and traceable results, using your own question bank or benchmark subsets.

liucongg/liucong-skills248—~752Automated safety check: PassApache-2.029 days ago
47

Build DSPy 3.2.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget.

intertwine/dspy-agent-skills278—~1.7kAutomated safety check: PassMIT1 mo ago
48

Prepare the environment and run the LLM-driven agent evals (e2e/agent-evals/) against a chosen sim-use binary.

lycorp-jp/sim-use1.4k—~1.4kAutomated safety check: PassApache-2.0yesterday

Questions, answered from the data.

What is the best LLM evaluation skill?

Skill Creator (official) from Azure/azqr ranks first of the 303 LLM evaluation skills listed here, with the highest score: its repository has 794 GitHub stars, 89 other GitHub owners carry a copy, its SKILL.md loads about 8.2k tokens and it passes the automated safety check with no findings. Next come Caveman Experiment Manager and Caveman Optimization Evaluator.

Which LLM evaluation skills are official?

41 of the 303 LLM evaluation skills are official, published by the vendor's own GitHub organization: Skill Creator, Azure AI Projects Python SDK, Hugging Face Local Model Evals, Agent Eval Engineering, Commerce Evals and 36 more.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.