Topic · AI & LLM Engineering
Best LLM evaluation skills for Claude Code, Codex and other agents.
- skills
- 303
- official
- 41
LLM evaluation skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Create new skills, modify and improve existing skills, and measure skill performance. | Azure/ | 794 | 89 repos | ~8.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 2 | Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself. | JuliusBrussee/ | 110k | 1 repo | ~975 | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Turns a Caveman report-only optimization observation into one minimal code change and a paired baseline evaluation, after the operator picks which to pursue. | JuliusBrussee/ | 110k | 1 repo | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | Provides context about the CoStrict evals system structure in this monorepo. | zgsm-ai/ | 4.4k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 5 | Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints. | Orchestra-Research/ | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Create, improve, or evaluate an existing skill from workflows, prompts, SOPs, scripts. | yaojingang/ | 2.7k | — | ~768 | Automated safety check: Pass | MIT | 1 mo ago |
| 7 | Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes. | microsoft/ | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | yesterday |
| 8 | Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures. | chatboxai/ | 42k | — | ~758 | Automated safety check: Pass | GPL-3.0 | 13 days ago |
| 9 | Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs. | FailproofAI/ | 5.3k | — | ~6k | Automated safety check: Pass | Unknown | yesterday |
| 10 | 10.Skillforge A skill your agent uses when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use'… | tripleyak/ | 905 | — | ~2.3k | Automated safety check: Notes | MIT | 2 mo ago |
| 11 | Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment. | Jeffallan/ | 12k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | 4 days ago |
| 12 | A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain. | DenisSergeevitch/ | 2.4k | — | ~7.4k | Automated safety check: Pass | MIT | 2 days ago |
| 13 | 13.Looper Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council. | ksimback/ | 710 | — | ~2.7k | Automated safety check: Notes | MIT | 1 mo ago |
| 14 | Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends. | huggingface/ | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 15 | Research, create, improve, migrate, evaluate, package, install-check, govern, and safely publish qiaomu-flavored agent skills from workflows, prompts, transcripts, docs, SOPs, runbooks, scripts, or… | joeseesun/ | 383 | — | ~2.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 16 | Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons. | windmill-labs/ | 18k | — | ~969 | Automated safety check: Notes | Unknown | today |
| 17 | Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier. | langchain-ai/ | 1.3k | — | ~4k | Automated safety check: Pass | MIT | yesterday |
| 18 | 18.Eval Harness Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles. | cloudnative-co/ | 151 | 10 repos | ~1.3k | Automated safety check: Pass | MIT | today |
| 19 | Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK. | GoogleCloudPlatform/ | 791 | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 20 | Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report. | bgauryy/ | 946 | — | ~2.1k | Automated safety check: Pass | MIT | 4 days ago |
| 21 | 21.Eval Chat Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each… | ankit/ | 1.6k | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 22 | 亚马逊卖家专用的 skill 创建器(中文)。当用户想把一个亚马逊运营/自媒体/日常工作流程变成可复用的 skill 时使用。触发场景包括但不限于:用户说"我想做一个 skill""把这个流程变成 skill""帮我写个自动化""优化我已有的 skill""给这个工作流做个自动化",即使用户没用"skill"这个词,只要在描述"以后每次都这样做"的重复性工作时也应触发。本 skill… | zach22-1999/ | 206 | 1 repo | ~3.9k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 23 | Create new skills, modify and improve existing skills, and measure skill performance. | AgentTeam-TaichuAI/ | 670 | — | ~10k | Automated safety check: Pass | Apache-2.0 | 5 mo ago |
| 24 | 24.Task Creator SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric. | benchflow-ai/ | 353 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 25 | Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data. | ai-evals-course/ | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 26 | Create a skill or bring an existing one up to the same standard (validate + asm eval fix loop); run evals, tune triggering. | luongnv89/ | 953 | — | ~5.3k | Automated safety check: Pass | MIT | yesterday |
| 27 | Create, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor. | smixs/ | 179 | — | ~6.6k | Automated safety check: Pass | MIT | 2 mo ago |
| 28 | A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient… | sun461941-hub/ | 100 | — | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 29 | Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures. | anthropics/ | 3.2k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 30 | Evaluate verified findings from merge-ready, Greptile, pull-request, CI, security, billing, and other code reviews, then promote durable review gaps into the version-controlled .greptile… | Jwuthri/ | 1.5k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 31 | Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes. | frontier-harness-eval/ | 297 | — | ~8k | Automated safety check: Pass | No licence | 29 days ago |
| 32 | 32.Autocontext Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met. | greyhaven-ai/ | 1.3k | — | ~892 | Automated safety check: Pass | Apache-2.0 | today |
| 33 | Sends one prompt to Claude, GPT through the Codex CLI and Gemini, then tabulates response time, token use and cost, with an optional judged quality score. | garrytan/ | 136k | — | ~4k | Automated safety check: Notes | MIT | today |
| 34 | Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes. | ai-evals-course/ | 1.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 35 | 35.Harness Eval Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports. | tech-leads-club/ | 7k | — | ~3.9k | Automated safety check: Pass | CC-BY-4.0 | 17 days ago |
| 36 | A skill your agent uses when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. | microsoft/ | 1k | — | ~2k | Automated safety check: Notes | Unknown | today |
| 37 | Audit an agent's context layout against the four places: system prompt, tools, history, tail. | undefined-ui/ | 999 | — | ~810 | Automated safety check: Pass | MIT | 8 days ago |
| 38 | Turn the current conversation's workflow into a reusable agent skill. | Undertone0809/ | 292 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | today |
| 39 | Plan a documentation eval for supabase/evals, where a docs guide is the subject under test. | supabase/ | 143 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 40 | 40.Evaluate RAG Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately. | ai-evals-course/ | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 41 | Converts test suites from external eval frameworks into the Margin Eval suite format. | Margin-Lab/ | 161 | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | 2 mo ago |
| 42 | Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and… | NVIDIA-AI-Blueprints/ | 1.9k | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | today |
| 43 | 43.Fde Keeps the engagement record for client work. An agent skill from suboss87/FDEOps. | suboss87/ | 953 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 44 | Review a proposed Agent Skill for structural validity and content quality before publishing. | mongodb/ | 190 | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 45 | Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes. | benchflow-ai/ | 353 | — | ~4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 46 | Sets up and runs side-by-side model evaluations with frozen questions, isolated tools and traceable results, using your own question bank or benchmark subsets. | liucongg/ | 248 | — | ~752 | Automated safety check: Pass | Apache-2.0 | 29 days ago |
| 47 | Build DSPy 3.2.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget. | intertwine/ | 278 | — | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 48 | 48.Run Evals Prepare the environment and run the LLM-driven agent evals (e2e/agent-evals/) against a chosen sim-use binary. | lycorp-jp/ | 1.4k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
Questions, answered from the data.
What is the best LLM evaluation skill?
Skill Creator (official) from Azure/azqr ranks first of the 303 LLM evaluation skills listed here, with the highest score: its repository has 794 GitHub stars, 89 other GitHub owners carry a copy, its SKILL.md loads about 8.2k tokens and it passes the automated safety check with no findings. Next come Caveman Experiment Manager and Caveman Optimization Evaluator.
Which LLM evaluation skills are official?
41 of the 303 LLM evaluation skills are official, published by the vendor's own GitHub organization: Skill Creator, Azure AI Projects Python SDK, Hugging Face Local Model Evals, Agent Eval Engineering, Commerce Evals and 36 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- LLM inference and serving364
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Computer vision206
- Model hubs and datasets180
- GPU and accelerator computing171
- Diffusion and image models167
- Natural language processing141
- Reinforcement learning67
- AI interpretability23