Topic · AI & LLM Engineering

Best LLM evaluation skills, page 6

Skills #241–288 of 308, ranked by score.

LLM evaluation skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM evaluation skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
241

Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic.

agentsope/SkillAlchemy459—~6.3kAutomated safety check: PassMIT1 mo ago
242

Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.

agentsope/SkillAlchemy459—~6.4kAutomated safety check: PassMIT1 mo ago
243

Agent observability, evals, feedback, and experiments. An agent skill from BuilderIO/agent-native.

BuilderIO/agent-native7.1k—~7kAutomated safety check: PassNo licencetoday
244

This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".

borghei/Claude-Skills881—~1.9kAutomated safety check: PassMITyesterday
245

Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits.

growthxai/output440—~2.5kAutomated safety check: NotesApache-2.0today
246

Compile an Anthropic-style skill — a directory with a SKILL.md and optional references/ — into a deterministic, runnable workflow via the rote CLI.

ccplugins/awesome-claude-code-plugins968—~1.3kAutomated safety check: NotesApache-2.01 mo ago
247

[omh] Missed route or run lessons to record: classify and review self-improvement store routes as an auxiliary review lane before durable writes, then record workflow attempts as metadata-only…

rlaope/oh-my-hermes3.2k—~2kAutomated safety check: PassMITtoday
248

Audit a white paper or long-form technical document against a research-grounded best-practices checklist.

glebis/claude-skills389—~752Automated safety check: PassMIT12 days ago
249

6 production-ready AI engineering workflows: prompt evaluation (8-dimension scoring), context budget planning, RAG pipeline design, agent security audit (65-point checklist), eval harness building…

majiayu000/claude-skill-registry6663 repos~1.7kAutomated safety check: PassMITtoday
250

Run and score the agent p-hacking benchmark. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills.

brycewang-stanford/Auto-Empirical-Research-Skills4.5k—~1.6kAutomated safety check: PassUnknown3 days ago
251

Entry point for the p-hacking skills suite. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills.

brycewang-stanford/Auto-Empirical-Research-Skills4.5k—~1.7kAutomated safety check: PassUnknown3 days ago
252

Report content-quality trends over time from logged evaluations: weekly score trend charts, a content-type leaderboard, per-dimension performance breakdown, statistically flagged regression alerts…

indranilbanerjee/digital-marketing-pro8551 repo~2.5kAutomated safety check: PassMIT4 days ago
253

Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it.

davepoon/buildwithclaude3.6k—~1.3kAutomated safety check: PassMIT2 days ago
254
254.Benchmark AgentsOfficial

Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

vercel/vercel-plugin301—~3.6kAutomated safety check: PassUnknownyesterday
255

Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates.

JasonColapietro/suede-creator-skills127—~3.3kAutomated safety check: PassMITyesterday
256
256.Run EvalsOfficial

Run evaluation tests for prompt quality. An agent skill from Azure/azure-sdk-tools.

Azure/azure-sdk-tools134—~977Automated safety check: PassMITtoday
257

Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.

Dynatrace/dynatrace-for-ai161—~4.5kAutomated safety check: PassApache-2.07 days ago
258

A skill your agent uses when integrating NVIDIA NeMo Fabric into a consumer application, service, evaluation harness, or platform through the typed Python SDK — translating the consumer's own…

NVIDIA/skills3.5k—~5.5kAutomated safety check: PassApache-2.0yesterday
259

A skill your agent uses when designing, auditing, refactoring, or explaining an agentic harness for any domain, especially when work must continue from a measured gap to verified completion.

AnastasiyaW/codex-claude-code-config154—~5.4kAutomated safety check: PassMITtoday
260
260.Agents OptimizeOfficial

A skill your agent uses when measuring or improving agent quality and performance — set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization.

aws/agent-toolkit-for-aws2.8k—~914Automated safety check: NotesApache-2.0today
261
261.RAG EvalOfficial

Filesystem RAG benchmarks: corpus/, train.json, evaluaterag.py (RAGAS quality).

NVIDIA/skills3.5k—~2.3kAutomated safety check: NotesApache-2.0yesterday
262

Benchmark AI models across 60+ academic evaluation suites and metrics

wentorai/research-plugins2981 repo~2kAutomated safety check: PassMIT3 mo ago
263

Best practices for building AI agents with Google's Agent Development Kit (ADK) in Python, covering agent design, tools, sessions, memory, artifacts, evaluation, and deployment.

Mindrally/skills268—~2.5kAutomated safety check: PassApache-2.01 mo ago
264

Iterate on RAG systems with structured evals instead of eyeballing.

glebis/claude-skills389—~1.5kAutomated safety check: PassMIT12 days ago
265

Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision).

softspark/ai-toolkit179—~1.1kAutomated safety check: NotesApache-2.0yesterday
266

Generate an evals file for a drafted skill and measure whether its trigger description fires on the right requests: five to eight phrases that should trigger it, five that should not, three golden…

mohitagw15856/pm-claude-skills1.4k—~1.2kAutomated safety check: PassMITyesterday
267

Write a new SecSkills skill end to end — choosing the plugin bucket and skill tier, writing a description that triggers correctly without stealing traffic from siblings, the required sections…

trilwu/secskills156—~3.1kAutomated safety check: PassMIT1 mo ago
268

Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit…

datadog-labs/agent-skills177—~25kAutomated safety check: PassMITtoday
269

Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory…

jeremylongshore/tons-of-skills-marketplace2.8k—~3.7kAutomated safety check: PassMITtoday
270

Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation.

BankrBot/skills1.2k—~660Automated safety check: PassNo licence3 days ago
271

A skill your agent uses when creating a new Claude skill from scratch, editing or improving an existing skill, or measuring skill performance with evals and benchmarks.

curiositech/some_claude_skills243—~7.2kAutomated safety check: PassApache-2.01 mo ago
272

A skill your agent uses when one prompt must give the same right answer across reruns, models, and pasted-in hostile input: forcing a fixed schema, picking the few-shot set, ordering the prompt…

ericrisco/rsc-harness167—~2.4kAutomated safety check: PassMITtoday
273

Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets.

majiayu000/claude-skill-registry6661 repo~3.1kAutomated safety check: PassMITtoday
274
274.Mastra

A skill your agent uses when working with Mastra - the TypeScript AI framework for building agents, workflows, tools, and AI-powered applications.

majiayu000/claude-skill-registry6661 repo~3.2kAutomated safety check: PassMITtoday
275

Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource.

stella/stella258—~2.7kAutomated safety check: PassApache-2.0today
276

TRIGGER for .flow files, UiPath Flow / Maestro Flow / Maestro Automate build/edit requests, and adding or listing IXP model/document-extraction nodes for a Flow.

UiPath/skills167—~6.5kAutomated safety check: NotesMITtoday
277

AI 评测工程专家 Owner — 当任务涉及模型/Prompt 评测、模型选择、黄金集、评分量表、LLM-as-judge、Judge 校准、重复采样、方差、质量-成本-延迟权衡、提示词回归或模型升级回归时使用;要求把概率性结果转化为可复现、可比较且防污染的评测证据。

devcodex-labs/devcodex439—~434Automated safety check: PassAGPL-3.021 days ago
278

适用于 Claude Code 会话的正规评测框架(Evaluation Framework),实现了评测驱动开发(Eval-Driven Development, EDD)原则

xu-xiang/everything-claude-code-zh2k—~974Automated safety check: PassMIT7 mo ago
279

为 Claude Code 会话提供的正式评测框架,实现了评测驱动开发(EDD)原则. An agent skill from xu-xiang/everything-claude-code-zh.

xu-xiang/everything-claude-code-zh2k—~904Automated safety check: PassMIT7 mo ago
280

Author and validate Vally evals for Agent Skills under .github/skills.

Azure/azure-sdk-tools134—~946Automated safety check: PassMITtoday
281

Author and validate hermetic single-tool Vally evals under evals/tools.

Azure/azure-sdk-tools134—~893Automated safety check: PassMITtoday
282

Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows.

Azure/azure-sdk-tools134—~972Automated safety check: PassMITtoday
283

Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes.

yonatangross/orchestkit289—~3.6kAutomated safety check: NotesMITtoday
284

A skill your agent uses when coordinating continuous improvement loops (team-levelup + change + evals feedback + cleanup) targeting team-ai-directives — includes build-to-delete pruning and…

tikalk/adlc-team-skills141—~1.5kAutomated safety check: PassMIT2 days ago
285

A skill your agent uses when 王建硕 wants to evaluate whether a change to VoiceDrop's 挖矿 system prompt is actually better than the live version — runs the local eval harness (golden fixtures ×…

jianshuo/claude-skills130—~475Automated safety check: PassMIT1 mo ago
286

A skill your agent uses when building durable AI agents or agentic workflows with Inngest and AgentKit, including model calls, tool calls, multi-agent networks, human approval, realtime progress…

Asymmetric-al/core381—~2.6kAutomated safety check: PassAGPL-3.0today
287

A skill your agent uses when analyzing an existing TypeScript or JavaScript codebase to decide where and how to introduce Inngest.

Asymmetric-al/core381—~3.1kAutomated safety check: PassAGPL-3.0today
288

Version tracking for Agent Skills bundles and their associated files across sessions, surfaces, and platforms.

LeoYeAI/openclaw-master-skills2.2k—~4.8kAutomated safety check: PassMIT2 mo ago