Topic · AI & LLM Engineering

Best LLM evaluation skills, page 5

Skills #193–240 of 308, ranked by score.

LLM evaluation skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM evaluation skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
193

Author a new Dex skill that actually fires and passes the quality bar.

davekilleen/Dex493—~2.9kAutomated safety check: PassUnknown6 days ago
194

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise…

aiskillstore/marketplace4304 repos~4.2kAutomated safety check: PassNo licenceyesterday
195

Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

affaan-m/ECC275k—~2.2kAutomated safety check: PassMIT3 days ago
196

A skill your agent uses when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline.

tikalk/adlc-team-skills141—~934Automated safety check: PassMIT2 days ago
197
197.Eval

Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

alirezarezvani/claude-skills28k1 repo~618Automated safety check: PassMIT1 mo ago
198
198.Logfire SetupOfficial

Entry point for Pydantic Logfire — an observability, monitoring, and evals platform.

pydantic/skills140—~1.5kAutomated safety check: PassMIT7 days ago
199
199.Evals

Build a regression + eval harness for AI-written code and AI features.

Houseofmvps/ultraship123—~1.1kAutomated safety check: NotesMIT3 mo ago
200

Turn student course evaluations (free-text + numeric) into an actionable teaching-improvement plan — the teaching analogue of /respond-to-referees.

pedrohcgs/claude-code-my-workflow1.6k—~2.6kAutomated safety check: NotesMIT10 days ago
201

Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines.

wshobson/agents40k—~2kAutomated safety check: PassMIT3 days ago
202

Multi-agent collaboration plugin that spawns N parallel subagents competing on the same task via git worktree isolation.

alirezarezvani/claude-skills28k—~2kAutomated safety check: PassMIT1 mo ago
203
203.Writing EvalsOfficial

Teaches how to write and run evals on the products/posthogai/evalharness/ harness — sandboxed agent suites that execute the real coding agent in a Docker or Modal sandbox against a seeded Hedgebox…

PostHog/posthog40k—~4kAutomated safety check: NotesUnknowntoday
204

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology.

google/skills21k—~7.4kAutomated safety check: PassApache-2.0today
205

Create, improve, and evaluate agent skills (SKILL.md plus reference files).

himself65/finance-skills3.4k—~3.8kAutomated safety check: PassMIT3 days ago
206

A skill your agent uses when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria.

tikalk/adlc-team-skills141—~915Automated safety check: PassMIT2 days ago
207

Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

affaan-m/ECC275k—~1.5kAutomated safety check: PassMIT3 days ago
208

用于医疗应用部署的患者安全评估工具。针对CDSS准确性、PHI暴露、临床工作流完整性和集成合规性的自动化测试套件。在安全故障时阻止部署。

affaan-m/ECC275k—~1.4kAutomated safety check: PassMIT3 days ago
209

A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining…

agentscope-ai/OpenJudge868—~5.1kAutomated safety check: PassApache-2.027 days ago
210

Teaches agents how to write correct riteway ai prompt evals (.sudo files) for multi-step flows that involve tool calls.

paralleldrive/aidd384—~1.6kAutomated safety check: PassMIT3 mo ago
211

Run Azure SDK QA bot evaluations on curated datasets locally, including a single test case.

Azure/azure-sdk-tools134—~1.1kAutomated safety check: NotesMITtoday
212

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1392 repos~3.1kAutomated safety check: PassMITtoday
213

Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified.

PostHog/posthog40k—~6.7kAutomated safety check: PassUnknowntoday
214

Create offline evaluation tests for Output SDK workflows using @outputai/evals.

growthxai/output440—~3.8kAutomated safety check: NotesApache-2.0today
215

Build automated evaluation suites for AI agents using golden datasets, rubrics, and regression gates.

sickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: WarnMITyesterday
216

Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs.

RConsortium/pharma-skills118—~5.3kAutomated safety check: PassMIT4 days ago
217
217.Eval Driven DevOfficial

Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~4.4kAutomated safety check: WarnMITtoday
218

Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded…

PostHog/posthog40k—~1.5kAutomated safety check: PassUnknowntoday
219

INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.

langchain-ai/langchain-skills1.3k—~8.7kAutomated safety check: NotesMIT2 days ago
220
220.ML

Machine learning and LLM engineering judgment, distilled from a stronger model - invoke when DECIDING whether/how to use ML or an LLM for a task (prompt vs RAG vs fine-tune vs classical); working…

telagod/code-abyss2431 repo~566Automated safety check: PassMIT2 mo ago
221

Autonomously optimize any Claude Code skill by running it repeatedly, scoring outputs against binary evals, mutating the prompt, and keeping improvements.

pedronauck/skills634—~3.7kAutomated safety check: PassNo licence23 days ago
222

Author or review software tests and LLM/agent evals; choose test placement and mocks, diagnose flaky CI, or repair brittle suites.

pedronauck/skills6341 repo~636Automated safety check: PassNo licence23 days ago
223

A skill your agent uses when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy).

tikalk/adlc-team-skills141—~850Automated safety check: PassMIT2 days ago
224

Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment).

PostHog/posthog40k—~5.7kAutomated safety check: PassUnknowntoday
225
225.Feature Usage FeedOfficial

Set up an LLM-judge evaluation that extracts canonical use cases for a PostHog feature at scale and streams the results to a Slack channel as a live feed.

PostHog/posthog40k—~7.6kAutomated safety check: PassUnknowntoday
226

A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.

Aperivue/medsci-skills329—~2.4kAutomated safety check: PassMIT3 days ago
227

[Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them.

pskoett/pskoett-ai-skills311—~2.6kAutomated safety check: PassNo licence3 days ago
228

[Beta] Session-start scan that surfaces relevant learnings, recent errors, and eval status before work begins.

pskoett/pskoett-ai-skills311—~1.4kAutomated safety check: PassNo licence3 days ago
229

Build or refresh a product README showcase using a seeded Bag of Words workspace, polished in-product screenshots, and repository-ready visual assets.

bagofwords1/bagofwords458—~1.8kAutomated safety check: PassUnknowntoday
230

Create, refactor, evaluate, and package agent skills from workflows, prompts, transcripts, docs, or notes.

aiskillstore/marketplace4302 repos~806Automated safety check: PassMITyesterday
231

Audit an existing eval suite for trustworthiness. An agent skill from growthxai/output.

growthxai/output440—~2.5kAutomated safety check: NotesApache-2.0today
232

Design effective LLM judge .prompt files for evaluators. An agent skill from growthxai/output.

growthxai/output440—~3.2kAutomated safety check: PassApache-2.0today
233

Configures and runs LLM evaluation using Promptfoo framework.

daymade/claude-code-skills1.4k—~3kAutomated safety check: PassMITtoday
234

Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

ClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMITyesterday
235

Patterns for continuous autonomous agent loops with quality gates, evals, and recovery controls.

majiayu000/claude-skill-registry6666 repos~268Automated safety check: PassMITtoday
236

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

open-thoughts/OpenThoughts-Agent301—~3.7kAutomated safety check: PassApache-2.010 days ago
237

Skill validation framework PLUS daily test-suite health and regression intelligence.

inbrainfun/inbrain1421 repo~2kAutomated safety check: PassUnknown2 mo ago
238

A skill your agent uses when choosing between engines or models for an agent pipeline and the answer must come from measurement on your own data, not from marketing pages: speech recognition for…

serejaris/personal-corp-os229—~1.8kAutomated safety check: PassMITyesterday
239

Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.

daymade/claude-code-skills1.4k—~4.7kAutomated safety check: PassMITtoday
240

A skill your agent uses when the user has a prompt that feeds a system they can already score, and wants that prompt automatically improved to raise the score against their own evaluation command.

gaasher/Agent-Loop-Skills174—~2.1kAutomated safety check: PassMIT3 mo ago