Topic · AI & LLM Engineering
Best LLM evaluation skills, page 4
LLM evaluation skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval. | henryalouf/ | 157 | — | ~1.6k | Automated safety check: Pass | MIT | 4 mo ago |
| 146 | Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats. | glebis/ | 390 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 147 | Create a new built-in classification evaluator for Phoenix evals. | Arize-ai/ | 12k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 148 | Write, extend, and debug PXI Playwright E2E tests for Phoenix. | Arize-ai/ | 12k | — | ~2.6k | Automated safety check: Pass | Unknown | today |
| 149 | Bump the next release-please version for a Phoenix Python package (arize-phoenix, arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) by opening a PR with a Release-As commit footer. | Arize-ai/ | 12k | — | ~708 | Automated safety check: Pass | Apache-2.0 | today |
| 150 | Audit recent changes to Phoenix's user-facing surfaces (Python clients, TypeScript clients, CLI, REST/GraphQL APIs) and patch the three external-facing agent skills — phoenix-tracing, phoenix-cli… | Arize-ai/ | 12k | — | ~5.1k | Automated safety check: Pass | Unknown | today |
| 151 | Maintain the bundled TypeScript package docs that ship inside Phoenix npm packages. | Arize-ai/ | 12k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 152 | 152.Bkit Explore Browse installed bkit skills, agents, and evals via lib/discovery/explorer.js (filesystem scan, no subprocess). | ww-w-ai/ | 601 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 153 | Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report. | NVIDIA/ | 3.5k | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | today |
| 154 | Patterns and techniques for evaluating and improving AI agent outputs. | github/ | 40k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | today |
| 155 | Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository. | dotnet/ | 5.6k | 1 repo | ~6.1k | Automated safety check: Pass | MIT | today |
| 156 | 156.Watch Monitor arXiv for new papers relevant to the GLIDE project (prediction-powered inference, active statistical inference, LLM evaluation debiasing, proxy annotation bias correction). | EmertonData/ | 119 | — | ~3.5k | Automated safety check: Pass | Unknown | today |
| 157 | 157.Evals Clarify A skill your agent uses when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json). | tikalk/ | 141 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 158 | 158.Write Skill Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it. | druxt/ | 114 | — | ~926 | Automated safety check: Pass | MIT | today |
| 159 | Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates. | Arenukvern/ | 386 | — | ~2.4k | Automated safety check: Pass | MIT | 6 days ago |
| 160 | Find out what is going wrong in LLM or agent traffic by reading sampled Phoenix traces, spans, or sessions, writing free-form notes (open coding), then grouping the notes into a few narrow… | Arize-ai/ | 12k | — | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 161 | 161.Eval Harness Claude Codeセッションの正式な評価フレームワークで、評価駆動開発(EDD)の原則を実装します. An agent skill from affaan-m/ECC. | affaan-m/ | 276k | 2 repos | ~884 | Automated safety check: Pass | MIT | 4 days ago |
| 162 | 162.Eval Harness 평가 주도 개발(EDD) 원칙을 구현하는 Claude Code 세션용 공식 평가 프레임워크. An agent skill from affaan-m/ECC. | affaan-m/ | 276k | 2 repos | ~1.2k | Automated safety check: Pass | MIT | 4 days ago |
| 163 | Patient safety evaluation harness for healthcare application deployments. | affaan-m/ | 276k | 1 repo | ~2k | Automated safety check: Pass | MIT | 4 days ago |
| 164 | 164.Eval Loop This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target. | jacob-dietle/ | 111 | — | ~5.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 165 | 165.Evals Start Entry point for evals. An agent skill from ai-evals-course/evals-skills. | ai-evals-course/ | 1.5k | — | ~412 | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 166 | 166.Agent Eval Tests A skill your agent uses when writing, editing, or reviewing evalite-scored agent evals in packages/core/compute/assistant-evals/src/evals. | dxos/ | 525 | — | ~4.1k | Automated safety check: Pass | Unknown | today |
| 167 | Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. | microsoft/ | 138 | — | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 168 | 168.Evals Implement A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness. | tikalk/ | 141 | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 169 | 169.Opik Online Eval Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm… | comet-ml/ | 220 | — | ~3k | Automated safety check: Notes | Apache-2.0 | today |
| 170 | eve framework guidance for durable AI agents and agent-powered applications. | vercel/ | 301 | 5 repos | ~1.2k | Automated safety check: Pass | Unknown | yesterday |
| 171 | 171.Autoresearch Autonomously optimize a Claude Code skill or agent system by running it repeatedly, scoring outputs against evals, mutating owned artifacts (prompt, references, scripts, agent definitions), and… | byungjunjang/ | 120 | — | ~6k | Automated safety check: Warn | No licence | 2 mo ago |
| 172 | Scaffold a synthetic gold-standard annotation project for evaluating OpenMed NER and de-identification models — label schema, annotation guidelines, BRAT or Label Studio config, and disjoint… | maziyarpanahi/ | 5.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 173 | Evaluate an OpenMed de-identification or clinical NER model against the leakage-first release gates G1a through G8, which gate releases on residual PHI leakage rather than on F1. | maziyarpanahi/ | 5.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 174 | 174.Run Evals do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals. | Devin-AXIS/ | 6.8k | — | ~851 | Automated safety check: Pass | Unknown | yesterday |
| 175 | Build custom LLM evaluation pipelines using the OpenJudge framework. | agentscope-ai/ | 870 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 28 days ago |
| 176 | Debug LLM applications using the Phoenix CLI. An agent skill from github/awesome-copilot. | github/ | 40k | 1 repo | ~4k | Automated safety check: Pass | Apache-2.0 | today |
| 177 | Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware. | sickn33/ | 47k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 178 | Curated upstream guidance for Huggingface Community Evals; use when the workflow matches the user goal. | sickn33/ | 47k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | today |
| 179 | UiPath API Workflow assistant — author, run, validate, package, publish, deploy, and troubleshoot JSON workflows for uip api-workflow. | UiPath/ | 168 | — | ~7.5k | Automated safety check: Notes | MIT | today |
| 180 | Andrej Karpathy 视角顾问。以 Karpathy 的心智模型和工程师式判断,为 AI 产品经理场景做分析。 | SpaceZephyr/ | 160 | — | ~1.2k | Automated safety check: Pass | No licence | 3 mo ago |
| 181 | 181.Eval Harness Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi | affaan-m/ | 276k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | 4 days ago |
| 182 | Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing. | sickn33/ | 47k | 2 repos | ~3.1k | Automated safety check: Pass | MIT | today |
| 183 | Evaluate retrieval and citation behavior for RAG pipelines from deterministic JSONL fixtures. | davepoon/ | 3.6k | — | ~813 | Automated safety check: Pass | MIT | today |
| 184 | 184.Datasets Understand what a Phoenix dataset is and reason well about its examples, outputs, splits, and how it feeds evaluators and experiments. | Arize-ai/ | 12k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 185 | 185.Evals Init A skill your agent uses when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline. | tikalk/ | 141 | — | ~934 | Automated safety check: Pass | MIT | today |
| 186 | Guardrailed DELETE of auto-registered eval sandboxjobs rows that DID score but FAILED the harvest gate — partial evals (valid-complete <90% or non-benign infra-error 10%). | open-thoughts/ | 301 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 187 | 187.Eval Harness 克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则 | affaan-m/ | 276k | 3 repos | ~916 | Automated safety check: Pass | MIT | 4 days ago |
| 188 | Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset. | NVIDIA-AI-Blueprints/ | 1.9k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | today |
| 189 | 189.Tui Validate Validates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation. | mikeyobrien/ | 3.2k | — | ~3k | Automated safety check: Pass | MIT | 4 days ago |
| 190 | Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and… | github/ | 40k | 1 repo | ~8.1k | Automated safety check: Notes | MIT | today |
| 191 | 191.Spec Optimize Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first. | leo-kuang-ai/ | 107 | — | ~13k | Automated safety check: Pass | MIT | today |
| 192 | AI agent and LLM system engineering reference covering single-agent dev (ReAct, tool calling, plan-execute), multi-agent coordination (swarm, role decomposition, file locking), LLM security (prompt… | telagod/ | 243 | — | ~691 | Automated safety check: Pass | MIT | 2 mo ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents546
- Deep learning408
- Embeddings373
- LLM inference and serving372
- Prompt engineering348
- Retrieval-augmented generation348
- Fine-tuning308
- Speech recognition and synthesis305
- Structured output and tool calling268
- Model routing and gateways257
- LLM cost and token optimization254
- LLM API integration250
- LLM observability242
- LLM guardrails217
- Computer vision195
- Model hubs and datasets179
- GPU and accelerator computing172
- Diffusion and image models165
- Natural language processing134
- Reinforcement learning66
- AI interpretability23