Search
LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 98 | 98.Analyze Eval Investigate a single failing eval from the convex-evals system. | get-convex/ | 130 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 99 | Best practices for creating expectations and grader files to evaluate guidance quality. | GoogleChrome/ | 1.1k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 100 | 100.Skill Creator Create new skills, modify and improve existing skills, and measure skill performance. | ZS520L/ | 103 | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | 4 mo ago |
| 101 | 101.Add Evaluator Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager. | wso2/ | 108 | — | ~710 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 102 | This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated… | guanyang/ | 977 | 2 repos | ~4.2k | Automated safety check: Pass | MIT | today |
| 103 | 103.Eval Debate 测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。 | YIKUAIBANZI/ | 122 | — | ~740 | Automated safety check: Pass | MIT | 6 mo ago |
| 104 | 104.Skill Creator A skill your agent uses when the user wants to work on a Claude Code skill file (SKILL.md): writing one from scratch, testing whether an existing one works well, running evals or benchmarks… | avibebuilder/ | 120 | — | ~8.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 105 | 105.Agent Eval Cases Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic. | agentailor/ | 132 | — | ~5.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 106 | Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов). | AnastasiyaW/ | 154 | — | ~3.8k | Automated safety check: Pass | MIT | today |
| 107 | Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately. | microsoft/ | 138 | — | ~22k | Automated safety check: Warn | MIT | 3 mo ago |
| 108 | 108.Analyze Run Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations. | get-convex/ | 130 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 109 | 109.Test Afm Binary Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU… | scouzi1966/ | 346 | — | ~3.8k | Automated safety check: Pass | MIT | yesterday |
| 110 | 110.LLM Judge AI quality judge that scores agent responses 0-10 across helpfulness, accuracy, completeness, and clarity. | Atmosphere/ | 3.8k | — | ~333 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 111 | Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships. | timothywarner-org/ | 224 | — | ~696 | Automated safety check: Notes | MIT | 2 mo ago |
| 112 | 112.Benchmark Memory Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring | Abilityai/ | 109 | — | ~3.1k | Automated safety check: Notes | MIT | 3 days ago |
| 113 | Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification. | bgauryy/ | 949 | — | ~1.6k | Automated safety check: Pass | MIT | 2 days ago |
| 114 | Triage failing @novu/agent-evals scenarios to decide whether a failure is real or flaky, and whether to fix the playbook/prompt or the test (grader, tape, scenario, or judge). | novuhq/ | 40k | — | ~1.5k | Automated safety check: Pass | Unknown | yesterday |
| 115 | 115.LLM Evaluation Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing. | davila7/ | 33k | 12 repos | ~3.5k | Automated safety check: Pass | MIT | today |
| 116 | 116.Skill Forge Eval Run evaluation pipelines on Claude Code skills to test triggering accuracy, workflow correctness, and output quality. | AgriciDaniel/ | 179 | — | ~1.7k | Automated safety check: Pass | MIT | 6 mo ago |
| 117 | Patterns for continuous autonomous agent loops with quality gates, evals, and recovery controls. | affaan-m/ | 277k | 5 repos | ~298 | Automated safety check: Pass | MIT | today |
| 118 | 118.Loop Architect Design well-structured agent loops with best-practice coaching and cross-model review gates before you run them. | fabricioctelles/ | 106 | — | ~2.1k | Automated safety check: Notes | MIT | today |
| 119 | 119.Eval Harness Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed. | Archive228/ | 755 | — | ~876 | Automated safety check: Pass | MIT | 2 mo ago |
| 120 | 120.Eval Loop Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint. | malloydata/ | 116 | — | ~7.8k | Automated safety check: Pass | MIT | today |
| 121 | The step-by-step procedure for running this repo's Promptfoo evals with a generator and a judge over provider APIs. | maplibre/ | 151 | — | ~3.3k | Automated safety check: Pass | Unknown | yesterday |
| 122 | Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire. | pydantic/ | 140 | — | ~3.6k | Automated safety check: Pass | MIT | 10 days ago |
| 123 | 123.Nexus Testing Verify a Nexus change — pick a Jest lane, write a test that can actually fail, run the in-app Obsidian CLI loop, drive the eval harness, or fix a shipped-docs drift failure. | ProfSynapse/ | 157 | — | ~994 | Automated safety check: Pass | MIT | yesterday |
| 124 | 124.Write A Spec Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer. | different-ai/ | 24k | — | ~3.3k | Automated safety check: Pass | Unknown | today |
| 125 | Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow. | githits-com/ | 115 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 126 | 126.Rde Eval Run a targeted local React Doctor Evals loop against an uncommitted rule change. | millionco/ | 15k | — | ~510 | Automated safety check: Pass | Unknown | today |
| 127 | 127.AI Evals Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies. | RefoundAI/ | 1.4k | — | ~1.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 128 | Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files. | elastic/ | 21k | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 129 | Trigger an on-demand @kbn/evals Buildkite run by describing what you want in plain English. | elastic/ | 21k | — | ~2.6k | Automated safety check: Notes | Unknown | today |
| 130 | 130.Eval Creator CI [Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). | pskoett/ | 315 | — | ~2.2k | Automated safety check: Pass | No licence | 6 days ago |
| 131 | 131.Paperclip Evals Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification. | paperclipai/ | 100k | — | ~839 | Automated safety check: Pass | MIT | today |
| 132 | 132.Bkit Evals Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results. | ww-w-ai/ | 601 | — | ~1k | Automated safety check: Notes | Apache-2.0 | 14 days ago |
| 133 | 133.Eval Plan Builds an eval plan for an AI feature - error analysis on real traces, failure-mode ranking, code checks and binary LLM judges validated against human labels, CI tests, production monitoring, and a… | menkesu/ | 434 | — | ~4.5k | Automated safety check: Pass | Unknown | 5 days ago |
| 134 | You MUST load this skill before the skill-creator skill AND before making ANY change to, or conducting a review of ANY Agent Skill. | sammcj/ | 162 | — | ~9.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 135 | Audit supervised fine-tuning datasets against the behavior and task they are meant to teach. | tokenbender/ | 367 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 136 | 136.Agent Eval A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual… | ericrisco/ | 190 | — | ~3.2k | Automated safety check: Pass | MIT | yesterday |
| 137 | 137.Evals Analyze A skill your agent uses when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog. | tikalk/ | 141 | — | ~1.2k | Automated safety check: Pass | MIT | 2 days ago |
| 138 | Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment. | Jeffallan/ | 12k | — | ~1.7k | Automated safety check: Pass | MIT | 8 days ago |
| 139 | 139.Prompt Engineer Designs, tests and refines LLM prompts: zero-shot, few-shot and chain-of-thought patterns, system prompts, structured output schemas and evaluation test suites. | Jeffallan/ | 12k | — | ~1.5k | Automated safety check: Pass | MIT | 8 days ago |
| 140 | 140.RAG Architect Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step. | Jeffallan/ | 12k | — | ~2k | Automated safety check: Pass | MIT | 8 days ago |
| 141 | Set up compliance exports, drift detection, evaluations, scoring, and learning analytics | ucsandman/ | 311 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 142 | Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate. | ai-evals-course/ | 1.5k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 143 | Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval. | henryalouf/ | 157 | — | ~1.6k | Automated safety check: Pass | MIT | 4 mo ago |
| 144 | Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats. | glebis/ | 391 | — | ~2.9k | Automated safety check: Pass | MIT | 3 days ago |