Search

LLM evaluation

306 skills found, page 6.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
241

Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression.

daymade/claude-code-skills1.4k—~4.7kAutomated safety check: PassMITyesterday
242

A skill your agent uses when the user has a prompt that feeds a system they can already score, and wants that prompt automatically improved to raise the score against their own evaluation command.

gaasher/Agent-Loop-Skills174—~2.1kAutomated safety check: PassMIT3 mo ago
243

This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".

borghei/Claude-Skills891—~1.9kAutomated safety check: PassMIT4 days ago
244

Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic.

agentsope/SkillAlchemy436—~6.3kAutomated safety check: PassMIT2 days ago
245

Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.

agentsope/SkillAlchemy436—~6.4kAutomated safety check: PassMIT2 days ago
246

Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits.

growthxai/output442—~2.5kAutomated safety check: NotesApache-2.0yesterday
247

Agent observability, evals, feedback, and experiments. An agent skill from BuilderIO/agent-native.

BuilderIO/agent-native7.1k—~7.3kAutomated safety check: PassNo licenceyesterday
248

Compile an Anthropic-style skill — a directory with a SKILL.md and optional references/ — into a deterministic, runnable workflow via the rote CLI.

ccplugins/awesome-claude-code-plugins970—~1.3kAutomated safety check: NotesApache-2.02 mo ago
249

[omh] Missed route or run lessons to record: classify and review self-improvement store routes as an auxiliary review lane before durable writes, then record workflow attempts as metadata-only…

rlaope/oh-my-hermes3.2k—~2kAutomated safety check: PassMITyesterday
250

Audit a white paper or long-form technical document against a research-grounded best-practices checklist.

glebis/claude-skills391—~752Automated safety check: PassMIT3 days ago
251

Run and score the agent p-hacking benchmark. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills.

brycewang-stanford/Auto-Empirical-Research-Skills4.6k—~1.6kAutomated safety check: PassUnknown6 days ago
252

Entry point for the p-hacking skills suite. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills.

brycewang-stanford/Auto-Empirical-Research-Skills4.6k—~1.7kAutomated safety check: PassUnknown6 days ago
253

Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it.

davepoon/buildwithclaude3.6k—~1.3kAutomated safety check: PassMIT2 days ago
254
254.Benchmark AgentsOfficial

Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.

vercel/vercel-plugin301—~3.6kAutomated safety check: PassUnknownyesterday
255

Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates.

JasonColapietro/suede-creator-skills127—~3.3kAutomated safety check: PassMITyesterday
256
256.Run EvalsOfficial

Run evaluation tests for prompt quality. An agent skill from Azure/azure-sdk-tools.

Azure/azure-sdk-tools134—~977Automated safety check: PassMITyesterday
257

Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup.

Dynatrace/dynatrace-for-ai163—~4.5kAutomated safety check: PassApache-2.010 days ago
258

A skill your agent uses when integrating NVIDIA NeMo Fabric into a consumer application, service, evaluation harness, or platform through the typed Python SDK — translating the consumer's own…

NVIDIA/skills3.6k—~5.8kAutomated safety check: PassApache-2.02 days ago
259

A skill your agent uses when designing, auditing, refactoring, or explaining an agentic harness for any domain, especially when work must continue from a measured gap to verified completion.

AnastasiyaW/codex-claude-code-config154—~5.4kAutomated safety check: PassMIT2 days ago
260
260.Agents OptimizeOfficial

A skill your agent uses when measuring or improving agent quality and performance — set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization.

aws/agent-toolkit-for-aws2.8k—~914Automated safety check: NotesApache-2.0yesterday
261
261.ML

Machine learning and LLM engineering judgment, distilled from a stronger model - invoke when DECIDING whether/how to use ML or an LLM for a task (prompt vs RAG vs fine-tune vs classical); working…

telagod/code-abyss244—~566Automated safety check: PassMIT2 mo ago
262

Benchmark AI models across 60+ academic evaluation suites and metrics

wentorai/research-plugins2981 repo~2kAutomated safety check: PassMIT3 mo ago
263
263.RAG EvalOfficial

Filesystem RAG benchmarks: corpus/, train.json, evaluaterag.py (RAGAS quality).

NVIDIA/skills3.6k—~2.3kAutomated safety check: NotesApache-2.02 days ago
264

Best practices for building AI agents with Google's Agent Development Kit (ADK) in Python, covering agent design, tools, sessions, memory, artifacts, evaluation, and deployment.

Mindrally/skills271—~2.5kAutomated safety check: PassApache-2.02 days ago
265

Iterate on RAG systems with structured evals instead of eyeballing.

glebis/claude-skills391—~1.5kAutomated safety check: PassMIT3 days ago
266

Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision).

softspark/ai-toolkit179—~1.1kAutomated safety check: NotesApache-2.03 days ago
267

Generate an evals file for a drafted skill and measure whether its trigger description fires on the right requests: five to eight phrases that should trigger it, five that should not, three golden…

mohitagw15856/pm-claude-skills1.4k—~1.2kAutomated safety check: PassMIT2 days ago
268

Write a new SecSkills skill end to end — choosing the plugin bucket and skill tier, writing a description that triggers correctly without stealing traffic from siblings, the required sections…

trilwu/secskills157—~3.1kAutomated safety check: PassMIT1 mo ago
269

Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit…

datadog-labs/agent-skills177—~25kAutomated safety check: PassMIT2 days ago
270

Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory…

jeremylongshore/tons-of-skills-marketplace2.8k—~3.7kAutomated safety check: PassMITyesterday
271

Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation.

BankrBot/skills1.2k—~660Automated safety check: PassNo licenceyesterday
272

A skill your agent uses when creating a new Claude skill from scratch, editing or improving an existing skill, or measuring skill performance with evals and benchmarks.

curiositech/some_claude_skills244—~7.2kAutomated safety check: PassApache-2.01 mo ago
273

A skill your agent uses when one prompt must give the same right answer across reruns, models, and pasted-in hostile input: forcing a fixed schema, picking the few-shot set, ordering the prompt…

ericrisco/rsc-harness180—~2.4kAutomated safety check: PassMIT2 days ago
274

Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource.

stella/stella258—~2.7kAutomated safety check: PassApache-2.0yesterday
275

Report content quality trends from logged evals, with regression alerts.

indranilbanerjee/digital-marketing-pro8621 repo~2.4kAutomated safety check: PassMITyesterday
276

TRIGGER for .flow files, UiPath Flow / Maestro Flow / Maestro Automate build/edit requests, and adding or listing IXP model/document-extraction nodes for a Flow.

UiPath/skills167—~6.5kAutomated safety check: NotesMITyesterday
277

AI 评测工程专家 Owner — 当任务涉及模型/Prompt 评测、模型选择、黄金集、评分量表、LLM-as-judge、Judge 校准、重复采样、方差、质量-成本-延迟权衡、提示词回归或模型升级回归时使用;要求把概率性结果转化为可复现、可比较且防污染的评测证据。

devcodex-labs/devcodex439—~434Automated safety check: PassAGPL-3.024 days ago
278

适用于 Claude Code 会话的正规评测框架(Evaluation Framework),实现了评测驱动开发(Eval-Driven Development, EDD)原则

xu-xiang/everything-claude-code-zh2k—~974Automated safety check: PassMIT7 mo ago
279

为 Claude Code 会话提供的正式评测框架,实现了评测驱动开发(EDD)原则. An agent skill from xu-xiang/everything-claude-code-zh.

xu-xiang/everything-claude-code-zh2k—~904Automated safety check: PassMIT7 mo ago
280

Author and validate Vally evals for Agent Skills under .github/skills.

Azure/azure-sdk-tools134—~946Automated safety check: PassMITyesterday
281

Author and validate hermetic single-tool Vally evals under evals/tools.

Azure/azure-sdk-tools134—~893Automated safety check: PassMITyesterday
282

Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows.

Azure/azure-sdk-tools134—~972Automated safety check: PassMITyesterday
283

Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes.

yonatangross/orchestkit292—~3.6kAutomated safety check: NotesMITyesterday
284

A skill your agent uses when coordinating continuous improvement loops (team-levelup + change + evals feedback + cleanup) targeting team-ai-directives — includes build-to-delete pruning and…

tikalk/adlc-team-skills141—~1.5kAutomated safety check: PassMIT2 days ago
285

A skill your agent uses when 王建硕 wants to evaluate whether a change to VoiceDrop's 挖矿 system prompt is actually better than the live version — runs the local eval harness (golden fixtures ×…

jianshuo/claude-skills131—~475Automated safety check: PassMIT1 mo ago
286

Version tracking for Agent Skills bundles and their associated files across sessions, surfaces, and platforms.

LeoYeAI/openclaw-master-skills2.2k—~4.8kAutomated safety check: PassMIT2 mo ago
287

A skill your agent uses when building durable AI agents or agentic workflows with Inngest and AgentKit, including model calls, tool calls, multi-agent networks, human approval, realtime progress…

Asymmetric-al/core381—~2.6kAutomated safety check: PassAGPL-3.0yesterday
288

A skill your agent uses when analyzing an existing TypeScript or JavaScript codebase to decide where and how to introduce Inngest.

Asymmetric-al/core381—~3.1kAutomated safety check: PassAGPL-3.0yesterday