Search

LLM evaluation

306 skills found, page 5.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
193

6 production-ready AI engineering workflows: prompt evaluation (8-dimension scoring), context budget planning, RAG pipeline design, agent security audit (65-point checklist), eval harness building…

sickn33/agentic-awesome-skills47k2 repos~1.9kAutomated safety check: PassMIT2 days ago
194

Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

affaan-m/ECC276k—~2.2kAutomated safety check: PassMITyesterday
195
195.Model EvaluationOfficial

Generates python code that evaluates SageMaker models. An agent skill from awslabs/agent-plugins.

awslabs/agent-plugins916—~1.3kAutomated safety check: PassApache-2.0yesterday
196

A skill your agent uses when you want guided bottom-up error analysis, structured failure taxonomy discovery, or comprehensive trace coding before documenting eval criteria.

tikalk/adlc-team-skills141—~915Automated safety check: PassMIT2 days ago
197
197.Logfire SetupOfficial

Entry point for Pydantic Logfire — an observability, monitoring, and evals platform.

pydantic/skills140—~1.5kAutomated safety check: PassMIT10 days ago
198
198.Evals

Build a regression + eval harness for AI-written code and AI features.

Houseofmvps/ultraship123—~1.1kAutomated safety check: NotesMIT3 mo ago
199

Turn student course evaluations (free-text + numeric) into an actionable teaching-improvement plan — the teaching analogue of /respond-to-referees.

pedrohcgs/claude-code-my-workflow1.7k—~2.6kAutomated safety check: NotesMIT13 days ago
200

Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines.

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
201

Create offline evaluation tests for Output SDK workflows using @outputai/evals.

growthxai/output442—~3.8kAutomated safety check: NotesApache-2.0yesterday
202

Multi-agent collaboration plugin that spawns N parallel subagents competing on the same task via git worktree isolation.

alirezarezvani/claude-skills28k—~2kAutomated safety check: PassMIT1 mo ago
203

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise…

aiskillstore/marketplace4333 repos~4.2kAutomated safety check: PassNo licenceyesterday
204
204.Writing EvalsOfficial

Teaches how to write and run evals on the products/posthogai/evalharness/ harness — sandboxed agent suites that execute the real coding agent in a Docker or Modal sandbox against a seeded Hedgebox…

PostHog/posthog40k—~4kAutomated safety check: NotesUnknownyesterday
205

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology.

google/skills21k—~7.4kAutomated safety check: PassApache-2.02 days ago
206

Author a new Dex skill that actually fires and passes the quality bar.

davekilleen/Dex494—~2.9kAutomated safety check: PassMIT2 days ago
207

lm-eval-harness: benchmark LLMs (MMLU, GSM8K, etc.). An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1712 repos~3.1kAutomated safety check: PassMIT3 days ago
208

A skill your agent uses when a goldset with graders is ready to run — executes the evaluation pyramid and validates evaluator quality (SLA compliance, TPR/TNR, statistical accuracy).

tikalk/adlc-team-skills141—~850Automated safety check: PassMIT2 days ago
209

Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

affaan-m/ECC276k—~1.5kAutomated safety check: PassMITyesterday
210

用于医疗应用部署的患者安全评估工具。针对CDSS准确性、PHI暴露、临床工作流完整性和集成合规性的自动化测试套件。在安全故障时阻止部署。

affaan-m/ECC276k—~1.4kAutomated safety check: PassMITyesterday
211

A skill your agent uses when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining…

agentscope-ai/OpenJudge871—~5.1kAutomated safety check: PassApache-2.01 mo ago
212

Teaches agents how to write correct riteway ai prompt evals (.sudo files) for multi-step flows that involve tool calls.

paralleldrive/aidd384—~1.6kAutomated safety check: PassMIT4 mo ago
213

Run Azure SDK QA bot evaluations on curated datasets locally, including a single test case.

Azure/azure-sdk-tools134—~1.1kAutomated safety check: NotesMITyesterday
214

Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified.

PostHog/posthog40k—~6.7kAutomated safety check: PassUnknownyesterday
215

Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

woocommerce/woocommerce-ios358—~7.4kAutomated safety check: NotesGPL-2.02 days ago
216

Build automated evaluation suites for AI agents using golden datasets, rubrics, and regression gates.

sickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: WarnMIT2 days ago
217

Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs.

RConsortium/pharma-skills120—~5.3kAutomated safety check: PassMIT7 days ago
218

Evaluate LLM systems using automated metrics, LLM-as-judge, and benchmarks.

ancoleman/ai-design-components525—~4.7kAutomated safety check: PassMIT10 mo ago
219
219.Eval Driven DevOfficial

Improve AI application with evaluation-driven development. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~4.4kAutomated safety check: WarnMIT2 days ago
220

Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded…

PostHog/posthog40k—~1.5kAutomated safety check: PassUnknownyesterday
221

INVOKE THIS SKILL when building, testing, or deploying Managed Deep Agents in LangSmith.

langchain-ai/langchain-skills1.3k—~8.7kAutomated safety check: NotesMIT2 days ago
222

Autonomously optimize any Claude Code skill by running it repeatedly, scoring outputs against binary evals, mutating the prompt, and keeping improvements.

pedronauck/skills634—~3.7kAutomated safety check: PassNo licence27 days ago
223

Author or review software tests and LLM/agent evals; choose test placement and mocks, diagnose flaky CI, or repair brittle suites.

pedronauck/skills6341 repo~636Automated safety check: PassNo licence27 days ago
224

Create, improve, and evaluate agent skills (SKILL.md plus reference files).

himself65/finance-skills3.4k—~3.8kAutomated safety check: PassMIT6 days ago
225

Investigate AI observability evaluations — hog (deterministic code-based), llmjudge (LLM-prompt-based), and sentiment (user-message sentiment).

PostHog/posthog40k—~5.7kAutomated safety check: PassUnknownyesterday
226
226.Feature Usage FeedOfficial

Set up an LLM-judge evaluation that extracts canonical use cases for a PostHog feature at scale and streams the results to a Slack channel as a live feed.

PostHog/posthog40k—~7.6kAutomated safety check: PassUnknownyesterday
227

A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.

Aperivue/medsci-skills333—~2.4kAutomated safety check: PassMIT6 days ago
228

[Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them.

pskoett/pskoett-ai-skills315—~2.6kAutomated safety check: PassNo licence6 days ago
229

[Beta] Session-start scan that surfaces relevant learnings, recent errors, and eval status before work begins.

pskoett/pskoett-ai-skills315—~1.4kAutomated safety check: PassNo licence6 days ago
230
230.Agents CLI AquaOfficial

Work with the Ambient Quality Agent (AQuA) added to this agents-cli project: augment the agents-cli agent with AQuA or attach the agent to an AQuA deployed elsewhere, read the quality insights it…

google/adk-recipes10k—~5.5kAutomated safety check: PassApache-2.0yesterday
231

Build or refresh a product README showcase using a seeded Bag of Words workspace, polished in-product screenshots, and repository-ready visual assets.

bagofwords1/bagofwords459—~1.8kAutomated safety check: PassUnknownyesterday
232

Create, refactor, evaluate, and package agent skills from workflows, prompts, transcripts, docs, or notes.

aiskillstore/marketplace4332 repos~806Automated safety check: PassMITyesterday
233

Audit an existing eval suite for trustworthiness. An agent skill from growthxai/output.

growthxai/output442—~2.5kAutomated safety check: NotesApache-2.0yesterday
234

Design effective LLM judge .prompt files for evaluators. An agent skill from growthxai/output.

growthxai/output442—~3.2kAutomated safety check: PassApache-2.0yesterday
235
235.Eval

Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

alirezarezvani/claude-skills28k—~618Automated safety check: PassMIT1 mo ago
236

Configures and runs LLM evaluation using Promptfoo framework.

daymade/claude-code-skills1.4k—~3kAutomated safety check: PassMITyesterday
237

Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

ClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMIT2 days ago
238

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

open-thoughts/OpenThoughts-Agent301—~3.7kAutomated safety check: PassApache-2.013 days ago
239

Skill validation framework PLUS daily test-suite health and regression intelligence.

inbrainfun/inbrain1421 repo~2kAutomated safety check: PassUnknown2 mo ago
240

A skill your agent uses when choosing between engines or models for an agent pipeline and the answer must come from measurement on your own data, not from marketing pages: speech recognition for…

serejaris/personal-corp-os229—~1.8kAutomated safety check: PassMIT4 days ago