Topic · AI & LLM Engineering
Best LLM evaluation skills, page 6
LLM evaluation skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 241 | Build and govern a 50-200 example domain-specific held-out benchmark sampled from real traffic. | agentsope/ | 459 | — | ~6.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 242 | Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy. | agentsope/ | 459 | — | ~6.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 243 | 243.Observability Agent observability, evals, feedback, and experiments. An agent skill from BuilderIO/agent-native. | BuilderIO/ | 7.1k | — | ~7k | Automated safety check: Pass | No licence | today |
| 244 | This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality". | borghei/ | 881 | — | ~1.9k | Automated safety check: Pass | MIT | yesterday |
| 245 | Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. | growthxai/ | 440 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | today |
| 246 | 246.Compile Compile an Anthropic-style skill — a directory with a SKILL.md and optional references/ — into a deterministic, runnable workflow via the rote CLI. | ccplugins/ | 968 | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | 1 mo ago |
| 247 | [omh] Missed route or run lessons to record: classify and review self-improvement store routes as an auxiliary review lane before durable writes, then record workflow attempts as metadata-only… | rlaope/ | 3.2k | — | ~2k | Automated safety check: Pass | MIT | today |
| 248 | 248.Whitepaper Audit Audit a white paper or long-form technical document against a research-grounded best-practices checklist. | glebis/ | 389 | — | ~752 | Automated safety check: Pass | MIT | 12 days ago |
| 249 | 6 production-ready AI engineering workflows: prompt evaluation (8-dimension scoring), context budget planning, RAG pipeline design, agent security audit (65-point checklist), eval harness building… | majiayu000/ | 666 | 3 repos | ~1.7k | Automated safety check: Pass | MIT | today |
| 250 | Run and score the agent p-hacking benchmark. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills. | brycewang-stanford/ | 4.5k | — | ~1.6k | Automated safety check: Pass | Unknown | 3 days ago |
| 251 | 251.Phack Router Entry point for the p-hacking skills suite. An agent skill from brycewang-stanford/Auto-Empirical-Research-Skills. | brycewang-stanford/ | 4.5k | — | ~1.7k | Automated safety check: Pass | Unknown | 3 days ago |
| 252 | 252.Quality Report Report content-quality trends over time from logged evaluations: weekly score trend charts, a content-type leaderboard, per-dimension performance breakdown, statistically flagged regression alerts… | indranilbanerjee/ | 855 | 1 repo | ~2.5k | Automated safety check: Pass | MIT | 4 days ago |
| 253 | 253.Onboard Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it. | davepoon/ | 3.6k | — | ~1.3k | Automated safety check: Pass | MIT | 2 days ago |
| 254 | Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. | vercel/ | 301 | — | ~3.6k | Automated safety check: Pass | Unknown | yesterday |
| 255 | 255.Suede AI Eval Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates. | JasonColapietro/ | 127 | — | ~3.3k | Automated safety check: Pass | MIT | yesterday |
| 256 | Run evaluation tests for prompt quality. An agent skill from Azure/azure-sdk-tools. | Azure/ | 134 | — | ~977 | Automated safety check: Pass | MIT | today |
| 257 | 257.Dt Obs Genai Analyze & debug GenAI/LLM apps: token cost & caching by prompt, model & provider; latency/errors; agent & tool loops/failures; conversations; guardrails; evaluations; OpenTelemetry/dt-evals setup. | Dynatrace/ | 161 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 258 | A skill your agent uses when integrating NVIDIA NeMo Fabric into a consumer application, service, evaluation harness, or platform through the typed Python SDK — translating the consumer's own… | NVIDIA/ | 3.5k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 259 | A skill your agent uses when designing, auditing, refactoring, or explaining an agentic harness for any domain, especially when work must continue from a measured gap to verified completion. | AnastasiyaW/ | 154 | — | ~5.4k | Automated safety check: Pass | MIT | today |
| 260 | A skill your agent uses when measuring or improving agent quality and performance — set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization. | aws/ | 2.8k | — | ~914 | Automated safety check: Notes | Apache-2.0 | today |
| 261 | Filesystem RAG benchmarks: corpus/, train.json, evaluaterag.py (RAGAS quality). | NVIDIA/ | 3.5k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 262 | Benchmark AI models across 60+ academic evaluation suites and metrics | wentorai/ | 298 | 1 repo | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 263 | 263.Google Adk Best practices for building AI agents with Google's Agent Development Kit (ADK) in Python, covering agent design, tools, sessions, memory, artifacts, evaluation, and deployment. | Mindrally/ | 268 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 264 | 264.RAG Eval Iterate on RAG systems with structured evals instead of eyeballing. | glebis/ | 389 | — | ~1.5k | Automated safety check: Pass | MIT | 12 days ago |
| 265 | 265.Evaluate Evaluates RAG retrieval and LLM-as-judge metrics (faithfulness, relevancy, context precision). | softspark/ | 179 | — | ~1.1k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 266 | 266.Promoter Test Generate an evals file for a drafted skill and measure whether its trigger description fires on the right requests: five to eight phrases that should trigger it, five that should not, three golden… | mohitagw15856/ | 1.4k | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 267 | Write a new SecSkills skill end to end — choosing the plugin bucket and skill tier, writing a description that triggers correctly without stealing traffic from siblings, the required sections… | trilwu/ | 156 | — | ~3.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 268 | Bootstrap evaluators from production traces — by default propose online LLM-judge evaluators and, after you confirm, create them in Datadog as disabled drafts (never auto-enabled); on request emit… | datadog-labs/ | 177 | — | ~25k | Automated safety check: Pass | MIT | today |
| 269 | Build reproducible evaluation pipelines for LangChain 1.0 chains and LangGraph 1.0 agents — golden datasets, LangSmith evaluate(), ragas RAG metrics, deepeval LLM-as-judge, agent trajectory… | jeremylongshore/ | 2.8k | — | ~3.7k | Automated safety check: Pass | MIT | today |
| 270 | 270.Aeon Skill Evals Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation. | BankrBot/ | 1.2k | — | ~660 | Automated safety check: Pass | No licence | 3 days ago |
| 271 | 271.Skill Creator A skill your agent uses when creating a new Claude skill from scratch, editing or improving an existing skill, or measuring skill performance with evals and benchmarks. | curiositech/ | 243 | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 272 | A skill your agent uses when one prompt must give the same right answer across reruns, models, and pasted-in hostile input: forcing a fixed schema, picking the few-shot set, ordering the prompt… | ericrisco/ | 167 | — | ~2.4k | Automated safety check: Pass | MIT | today |
| 273 | 273.Agent Harness Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets. | majiayu000/ | 666 | 1 repo | ~3.1k | Automated safety check: Pass | MIT | today |
| 274 | 274.Mastra A skill your agent uses when working with Mastra - the TypeScript AI framework for building agents, workflows, tools, and AI-powered applications. | majiayu000/ | 666 | 1 repo | ~3.2k | Automated safety check: Pass | MIT | today |
| 275 | 275.Conventions MCP Apply when adding or changing an MCP tool, a capability the CLI generates, a tool input or output schema, a tool description, an error envelope, or an agent-facing reference resource. | stella/ | 258 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 276 | TRIGGER for .flow files, UiPath Flow / Maestro Flow / Maestro Automate build/edit requests, and adding or listing IXP model/document-extraction nodes for a Flow. | UiPath/ | 167 | — | ~6.5k | Automated safety check: Notes | MIT | today |
| 277 | AI 评测工程专家 Owner — 当任务涉及模型/Prompt 评测、模型选择、黄金集、评分量表、LLM-as-judge、Judge 校准、重复采样、方差、质量-成本-延迟权衡、提示词回归或模型升级回归时使用;要求把概率性结果转化为可复现、可比较且防污染的评测证据。 | devcodex-labs/ | 439 | — | ~434 | Automated safety check: Pass | AGPL-3.0 | 21 days ago |
| 278 | 278.Eval Harness 适用于 Claude Code 会话的正规评测框架(Evaluation Framework),实现了评测驱动开发(Eval-Driven Development, EDD)原则 | xu-xiang/ | 2k | — | ~974 | Automated safety check: Pass | MIT | 7 mo ago |
| 279 | 279.Eval Harness 为 Claude Code 会话提供的正式评测框架,实现了评测驱动开发(EDD)原则. An agent skill from xu-xiang/everything-claude-code-zh. | xu-xiang/ | 2k | — | ~904 | Automated safety check: Pass | MIT | 7 mo ago |
| 280 | Author and validate Vally evals for Agent Skills under .github/skills. | Azure/ | 134 | — | ~946 | Automated safety check: Pass | MIT | today |
| 281 | Author and validate hermetic single-tool Vally evals under evals/tools. | Azure/ | 134 | — | ~893 | Automated safety check: Pass | MIT | today |
| 282 | Author and validate multi-tool, multi-turn, mock, and live Vally scenarios under evals/workflows. | Azure/ | 134 | — | ~972 | Automated safety check: Pass | MIT | today |
| 283 | 283.Error Analysis Evals-first error analysis for LLM apps: clusters real Langfuse or JSONL traces into a human-confirmed failure taxonomy with counts, then recommends binary pass/fail evals for recurring named modes. | yonatangross/ | 289 | — | ~3.6k | Automated safety check: Notes | MIT | today |
| 284 | 284.Factory Learn A skill your agent uses when coordinating continuous improvement loops (team-levelup + change + evals feedback + cleanup) targeting team-ai-directives — includes build-to-delete pruning and… | tikalk/ | 141 | — | ~1.5k | Automated safety check: Pass | MIT | 2 days ago |
| 285 | A skill your agent uses when 王建硕 wants to evaluate whether a change to VoiceDrop's 挖矿 system prompt is actually better than the live version — runs the local eval harness (golden fixtures ×… | jianshuo/ | 130 | — | ~475 | Automated safety check: Pass | MIT | 1 mo ago |
| 286 | 286.Inngest Agents A skill your agent uses when building durable AI agents or agentic workflows with Inngest and AgentKit, including model calls, tool calls, multi-agent networks, human approval, realtime progress… | Asymmetric-al/ | 381 | — | ~2.6k | Automated safety check: Pass | AGPL-3.0 | today |
| 287 | A skill your agent uses when analyzing an existing TypeScript or JavaScript codebase to decide where and how to introduce Inngest. | Asymmetric-al/ | 381 | — | ~3.1k | Automated safety check: Pass | AGPL-3.0 | today |
| 288 | 288.Skill Provenance Version tracking for Agent Skills bundles and their associated files across sessions, surfaces, and platforms. | LeoYeAI/ | 2.2k | — | ~4.8k | Automated safety check: Pass | MIT | 2 mo ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- LLM inference and serving372
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- GPU and accelerator computing176
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23