Search
LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 241 | 241.LLM Eval Harness Tests/benchmarks a third-party LLM endpoint (OpenAI- or Anthropic-compatible): availability, fidelity, speed, concurrency, protocol compliance, quality regression. | daymade/ | 1.4k | — | ~4.7k | Automated safety check: Pass | MIT | yesterday |
| 242 | 242.Prompt Optimize A skill your agent uses when the user has a prompt that feeds a system they can already score, and wants that prompt automatically improved to raise the score against their own evaluation command. | gaasher/ | 174 | — | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 243 | This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality". | borghei/ | 891 | — | ~1.9k | Automated safety check: Pass | MIT | 4 days ago |
| 244 | Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic. | agentsope/ | 436 | — | ~6.3k | Automated safety check: Pass | MIT | 2 days ago |
| 245 | Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy. | agentsope/ | 436 | — | ~6.4k | Automated safety check: Pass | MIT | 2 days ago |
| 246 | Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. | growthxai/ | 442 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 247 | 247.Observability Agent observability, evals, feedback, and experiments. An agent skill from BuilderIO/agent-native. | BuilderIO/ | 7.1k | — | ~7.3k | Automated safety check: Pass | No licence | yesterday |
| 248 | 248.Compile Compile an Anthropic-style skill — a directory with a SKILL.md and optional references/ — into a deterministic, runnable workflow via the rote CLI. | ccplugins/ | 970 | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | 2 mo ago |
| 249 | [omh] Missed route or run lessons to record: classify and review self-improvement store routes as an auxiliary review lane before durable writes, then record workflow attempts as metadata-only… | rlaope/ | 3.2k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 250 | 250.Whitepaper Audit Audit a white paper or long-form technical document against a research-grounded best-practices checklist. | glebis/ | 391 | — | ~752 | Automated safety check: Pass | MIT | 3 days ago |
| 251 | Run and score the agent p-hacking benchmark. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills. | brycewang-stanford/ | 4.6k | — | ~1.6k | Automated safety check: Pass | Unknown | 6 days ago |
| 252 | 252.Phack Router Entry point for the p-hacking skills suite. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills. | brycewang-stanford/ | 4.6k | — | ~1.7k | Automated safety check: Pass | Unknown | 6 days ago |
| 253 | 253.Onboard Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it. | davepoon/ | 3.6k | — | ~1.3k | Automated safety check: Pass | MIT | 2 days ago |
| 254 | Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. | vercel/ | 301 | — | ~3.6k | Automated safety check: Pass | Unknown | yesterday |
| 255 | 255.Suede AI Eval Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates. | JasonColapietro/ | 127 | — | ~3.3k | Automated safety check: Pass | MIT | yesterday |
| 256 | Run evaluation tests for prompt quality. An agent skill from Azure/azure-sdk-tools. | Azure/ | 134 | — | ~977 | Automated safety check: Pass | MIT | yesterday |
| 257 | 257.Dt Obs Genai Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup. | Dynatrace/ | 163 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 258 | A skill your agent uses when integrating NVIDIA NeMo Fabric into a consumer application, service, evaluation harness, or platform through the typed Python SDK — translating the consumer's own… | NVIDIA/ | 3.6k | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 259 | A skill your agent uses when designing, auditing, refactoring, or explaining an agentic harness for any domain, especially when work must continue from a measured gap to verified completion. | AnastasiyaW/ | 154 | — | ~5.4k | Automated safety check: Pass | MIT | 2 days ago |
| 260 | A skill your agent uses when measuring or improving agent quality and performance — set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization. | aws/ | 2.8k | — | ~914 | Automated safety check: Notes | Apache-2.0 | yesterday |
| 261 | 261.ML Machine learning and LLM engineering judgment, distilled from a stronger model - invoke when DECIDING whether/how to use ML or an LLM for a task (prompt vs RAG vs fine-tune vs classical); working… | telagod/ | 244 | — | ~566 | Automated safety check: Pass | MIT | 2 mo ago |
| 262 | Benchmark AI models across 60+ academic evaluation suites and metrics | wentorai/ | 298 | 1 repo | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 263 | Filesystem RAG benchmarks: corpus/, train.json, evaluaterag.py (RAGAS quality). | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 264 | 264.Google Adk Best practices for building AI agents with Google's Agent Development Kit (ADK) in Python, covering agent design, tools, sessions, memory, artifacts, evaluation, and deployment. | Mindrally/ | 271 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 265 | 265.RAG Eval Iterate on RAG systems with structured evals instead of eyeballing. | glebis/ | 391 | — | ~1.5k | Automated safety check: Pass | MIT | 3 days ago |
| 266 | 266.Evaluate Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision). | softspark/ | 179 | — | ~1.1k | Automated safety check: Notes | Apache-2.0 | 3 days ago |
| 267 | 267.Promoter Test Generate an evals file for a drafted skill and measure whether its trigger description fires on the right requests: five to eight phrases that should trigger it, five that should not, three golden… | mohitagw15856/ | 1.4k | — | ~1.2k | Automated safety check: Pass | MIT | 2 days ago |
| 268 | Write a new SecSkills skill end to end — choosing the plugin bucket and skill tier, writing a description that triggers correctly without stealing traffic from siblings, the required sections… | trilwu/ | 157 | — | ~3.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 269 | Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit… | datadog-labs/ | 177 | — | ~25k | Automated safety check: Pass | MIT | 2 days ago |
| 270 | Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory… | jeremylongshore/ | 2.8k | — | ~3.7k | Automated safety check: Pass | MIT | yesterday |
| 271 | 271.Aeon Skill Evals Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation. | BankrBot/ | 1.2k | — | ~660 | Automated safety check: Pass | No licence | yesterday |
| 272 | 272.Skill Creator A skill your agent uses when creating a new Claude skill from scratch, editing or improving an existing skill, or measuring skill performance with evals and benchmarks. | curiositech/ | 244 | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 273 | A skill your agent uses when one prompt must give the same right answer across reruns, models, and pasted-in hostile input: forcing a fixed schema, picking the few-shot set, ordering the prompt… | ericrisco/ | 180 | — | ~2.4k | Automated safety check: Pass | MIT | 2 days ago |
| 274 | 274.Conventions MCP Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource. | stella/ | 258 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 275 | 275.Quality Report Report content quality trends from logged evals, with regression alerts. | indranilbanerjee/ | 862 | 1 repo | ~2.4k | Automated safety check: Pass | MIT | yesterday |
| 276 | TRIGGER for .flow files, UiPath Flow / Maestro Flow / Maestro Automate build/edit requests, and adding or listing IXP model/document-extraction nodes for a Flow. | UiPath/ | 167 | — | ~6.5k | Automated safety check: Notes | MIT | yesterday |
| 277 | AI 评测工程专家 Owner — 当任务涉及模型/Prompt 评测、模型选择、黄金集、评分量表、LLM-as-judge、Judge 校准、重复采样、方差、质量-成本-延迟权衡、提示词回归或模型升级回归时使用;要求把概率性结果转化为可复现、可比较且防污染的评测证据。 | devcodex-labs/ | 439 | — | ~434 | Automated safety check: Pass | AGPL-3.0 | 24 days ago |
| 278 | 278.Eval Harness 适用于 Claude Code 会话的正规评测框架(Evaluation Framework),实现了评测驱动开发(Eval-Driven Development, EDD)原则 | xu-xiang/ | 2k | — | ~974 | Automated safety check: Pass | MIT | 7 mo ago |
| 279 | 279.Eval Harness 为 Claude Code 会话提供的正式评测框架,实现了评测驱动开发(EDD)原则. An agent skill from xu-xiang/everything-claude-code-zh. | xu-xiang/ | 2k | — | ~904 | Automated safety check: Pass | MIT | 7 mo ago |
| 280 | Author and validate Vally evals for Agent Skills under .github/skills. | Azure/ | 134 | — | ~946 | Automated safety check: Pass | MIT | yesterday |
| 281 | Author and validate hermetic single-tool Vally evals under evals/tools. | Azure/ | 134 | — | ~893 | Automated safety check: Pass | MIT | yesterday |
| 282 | Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows. | Azure/ | 134 | — | ~972 | Automated safety check: Pass | MIT | yesterday |
| 283 | 283.Error Analysis Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes. | yonatangross/ | 292 | — | ~3.6k | Automated safety check: Notes | MIT | yesterday |
| 284 | 284.Factory Learn A skill your agent uses when coordinating continuous improvement loops (team-levelup + change + evals feedback + cleanup) targeting team-ai-directives — includes build-to-delete pruning and… | tikalk/ | 141 | — | ~1.5k | Automated safety check: Pass | MIT | 2 days ago |
| 285 | A skill your agent uses when 王建硕 wants to evaluate whether a change to VoiceDrop's 挖矿 system prompt is actually better than the live version — runs the local eval harness (golden fixtures ×… | jianshuo/ | 131 | — | ~475 | Automated safety check: Pass | MIT | 1 mo ago |
| 286 | 286.Skill Provenance Version tracking for Agent Skills bundles and their associated files across sessions, surfaces, and platforms. | LeoYeAI/ | 2.2k | — | ~4.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 287 | 287.Inngest Agents A skill your agent uses when building durable AI agents or agentic workflows with Inngest and AgentKit, including model calls, tool calls, multi-agent networks, human approval, realtime progress… | Asymmetric-al/ | 381 | — | ~2.6k | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 288 | A skill your agent uses when analyzing an existing TypeScript or JavaScript codebase to decide where and how to introduce Inngest. | Asymmetric-al/ | 381 | — | ~3.1k | Automated safety check: Pass | AGPL-3.0 | yesterday |