Topic · AI & LLM Engineering

Best LLM evaluation skills, page 3

Skills #97–144 of 308, ranked by score.

LLM evaluation skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM evaluation skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
97

Finds top models for a task from official Hugging Face benchmark leaderboards, filters them to what fits your hardware, and returns a comparison table with scores.

huggingface/skills11k2 repos~1.5kAutomated safety check: PassApache-2.07 days ago
98

Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.

comet-ml/opik-mcp219—~2.5kAutomated safety check: NotesApache-2.0yesterday
99

This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs…

pifferologo/cloud-agents-cli1291 repo~6.8kAutomated safety check: PassApache-2.01 mo ago
100

Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
101

Investigate a single failing eval from the convex-evals system.

get-convex/convex-evals129—~1.1kAutomated safety check: PassApache-2.0today
102

Create new skills, modify and improve existing skills, and measure skill performance.

ZS520L/HanakoPro102—~7.4kAutomated safety check: PassApache-2.04 mo ago
103

Audit a website and deliver a one-page, plain-language SEO report anyone can act on, centered on a single do-this-week action.

Jwuthri/Tracely-ai1.5k—~1.5kAutomated safety check: PassMITyesterday
104

This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…

guanyang/open-agent-hub9752 repos~4.2kAutomated safety check: PassMITyesterday
105

Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager.

wso2/agent-manager105—~710Automated safety check: PassApache-2.0yesterday
106

测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。

YIKUAIBANZI/forge-skill121—~740Automated safety check: PassMIT6 mo ago
107

A skill your agent uses when the user wants to work on a Claude Code skill file (SKILL.md): writing one from scratch, testing whether an existing one works well, running evals or benchmarks…

avibebuilder/claude-prime120—~8.2kAutomated safety check: PassMIT4 mo ago
108

Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.

agentailor/fullstack-langgraph-nextjs-agent132—~5.3kAutomated safety check: PassMIT1 mo ago
109

Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

AnastasiyaW/codex-claude-code-config154—~3.8kAutomated safety check: PassMITtoday
110
110.Eval GuideOfficial

Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

microsoft/eval-guide138—~22kAutomated safety check: WarnMIT3 mo ago
111

Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.

get-convex/convex-evals129—~2.1kAutomated safety check: PassApache-2.0today
112

Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…

scouzi1966/maclocal-api345—~3.8kAutomated safety check: PassMIT3 days ago
113

Best practices for creating expectations and grader files to evaluate guidance quality.

GoogleChrome/modern-web-guidance-src1.1k—~2.3kAutomated safety check: PassApache-2.0today
114

AI quality judge that scores agent responses 0-10 across helpfulness, accuracy, completeness, and clarity.

Atmosphere/atmosphere3.8k—~333Automated safety check: PassApache-2.02 days ago
115

Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.

timothywarner-org/claude-code224—~696Automated safety check: NotesMIT2 mo ago
116

Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

davila7/claude-code-templates32k13 repos~3.5kAutomated safety check: PassMITtoday
117

Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

bgauryy/octocode949—~1.6kAutomated safety check: PassMIT5 days ago
118

Triage failing @novu/agent-evals scenarios to decide whether a failure is real or flaky, and whether to fix the playbook/prompt or the test (grader, tape, scenario, or judge).

novuhq/novu40k—~1.5kAutomated safety check: PassUnknowntoday
119

Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

woocommerce/woocommerce-ios3581 repo~7.4kAutomated safety check: NotesGPL-2.0today
120

Run evaluation pipelines on Claude Code skills to test triggering accuracy, workflow correctness, and output quality.

AgriciDaniel/skill-forge177—~1.7kAutomated safety check: PassMIT6 mo ago
121

Design well-structured agent loops with best-practice coaching and cross-model review gates before you run them.

fabricioctelles/skills106—~2.1kAutomated safety check: NotesMIT4 days ago
122

Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.

Archive228/loopkit756—~876Automated safety check: PassMIT2 mo ago
123

Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

diegosouzapw/OmniRoute74k—~1.3kAutomated safety check: PassMITtoday
124

Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

malloydata/publisher116—~7.8kAutomated safety check: PassMITtoday
125

The step-by-step procedure for running this repo's Promptfoo evals with a generator and a judge over provider APIs.

maplibre/maplibre-agent-skills151—~3.3kAutomated safety check: PassUnknown4 days ago
126
126.Logfire EvalsOfficial

Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire.

pydantic/skills140—~3.6kAutomated safety check: PassMIT7 days ago
127

Verify a Nexus change — pick a Jest lane, write a test that can actually fail, run the in-app Obsidian CLI loop, drive the eval harness, or fix a shipped-docs drift failure.

ProfSynapse/nexus156—~994Automated safety check: PassMITtoday
128

Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer.

different-ai/openwork24k—~3.3kAutomated safety check: PassUnknowntoday
129

Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.

githits-com/githits-cli114—~3.1kAutomated safety check: PassApache-2.0yesterday
130

Run a targeted local React Doctor Evals loop against an uncommitted rule change.

millionco/react-doctor15k—~510Automated safety check: PassUnknowntoday
131

Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

RefoundAI/lenny-skills1.4k—~1.7kAutomated safety check: PassMIT2 mo ago
132
132.Evals Create SuiteOfficial

Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files.

elastic/kibana21k—~1.7kAutomated safety check: PassUnknowntoday
133

Trigger an on-demand @kbn/evals Buildkite run by describing what you want in plain English.

elastic/kibana21k—~2.6kAutomated safety check: NotesUnknowntoday
134

[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).

pskoett/pskoett-ai-skills311—~2.2kAutomated safety check: PassNo licence3 days ago
135

Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

paperclipai/paperclip99k—~839Automated safety check: PassMITtoday
136

Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.

ww-w-ai/bkit-claude-code601—~1kAutomated safety check: NotesApache-2.011 days ago
137

Builds an eval plan for an AI feature - error analysis on real traces, failure-mode ranking, code checks and binary LLM judges validated against human labels, CI tests, production monitoring, and a…

menkesu/awesome-pm-skills429—~4.5kAutomated safety check: PassUnknown2 days ago
138

You MUST load this skill before the skill-creator skill AND before making ANY change to, or conducting a review of ANY Agent Skill.

sammcj/agentic-coding162—~9.8kAutomated safety check: PassApache-2.0today
139

Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.

tokenbender/agent-guides367—~2.7kAutomated safety check: PassApache-2.02 mo ago
140

A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

ericrisco/rsc-harness167—~3.2kAutomated safety check: PassMITtoday
141

Set up compliance exports, drift detection, evaluations, scoring, and learning analytics

ucsandman/DashClaw310—~1.8kAutomated safety check: PassMIT2 days ago
142
142.Agentic EvalOfficial

Patterns and techniques for evaluating and improving AI agent outputs.

github/awesome-copilot40k4 repos~1.5kAutomated safety check: PassMITtoday
143

Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate.

ai-evals-course/evals-skills1.5k—~3.7kAutomated safety check: PassApache-2.014 days ago
144

Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.

glebis/claude-skills389—~2.9kAutomated safety check: PassMIT12 days ago