Topic · AI & LLM Engineering

Best LLM evaluation skills, page 4

Skills #145–192 of 308, ranked by score.

LLM evaluation skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM evaluation skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
145

Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval.

henryalouf/ruflow157—~1.6kAutomated safety check: PassMIT4 mo ago
146

Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.

glebis/claude-skills390—~2.9kAutomated safety check: PassMITyesterday
147

Create a new built-in classification evaluator for Phoenix evals.

Arize-ai/phoenix12k—~2.3kAutomated safety check: PassApache-2.0today
148

Write, extend, and debug PXI Playwright E2E tests for Phoenix.

Arize-ai/phoenix12k—~2.6kAutomated safety check: PassUnknowntoday
149

Bump the next release-please version for a Phoenix Python package (arize-phoenix, arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) by opening a PR with a Release-As commit footer.

Arize-ai/phoenix12k—~708Automated safety check: PassApache-2.0today
150

Audit recent changes to Phoenix's user-facing surfaces (Python clients, TypeScript clients, CLI, REST/GraphQL APIs) and patch the three external-facing agent skills — phoenix-tracing, phoenix-cli…

Arize-ai/phoenix12k—~5.1kAutomated safety check: PassUnknowntoday
151

Maintain the bundled TypeScript package docs that ship inside Phoenix npm packages.

Arize-ai/phoenix12k—~2.2kAutomated safety check: PassApache-2.0today
152

Browse installed bkit skills, agents, and evals via lib/discovery/explorer.js (filesystem scan, no subprocess).

ww-w-ai/bkit-claude-code601—~1.1kAutomated safety check: PassApache-2.012 days ago
153

Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report.

NVIDIA/skills3.5k—~1.3kAutomated safety check: NotesApache-2.0today
154
154.Agentic EvalOfficial

Patterns and techniques for evaluating and improving AI agent outputs.

github/awesome-copilot40k3 repos~1.5kAutomated safety check: PassMITtoday
155
155.Create Skill TestOfficial

Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository.

dotnet/skills5.6k1 repo~6.1kAutomated safety check: PassMITtoday
156
156.Watch

Monitor arXiv for new papers relevant to the GLIDE project (prediction-powered inference, active statistical inference, LLM evaluation debiasing, proxy annotation bias correction).

EmertonData/glide119—~3.5kAutomated safety check: PassUnknowntoday
157

A skill your agent uses when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json).

tikalk/adlc-team-skills141—~1.1kAutomated safety check: PassMITtoday
158

Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it.

druxt/druxt.js114—~926Automated safety check: PassMITtoday
159

Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.

Arenukvern/mcp_flutter386—~2.4kAutomated safety check: PassMIT6 days ago
160

Find out what is going wrong in LLM or agent traffic by reading sampled Phoenix traces, spans, or sessions, writing free-form notes (open coding), then grouping the notes into a few narrow…

Arize-ai/phoenix12k—~6.4kAutomated safety check: PassApache-2.0today
161

Claude Codeセッションの正式な評価フレームワークで、評価駆動開発(EDD)の原則を実装します. An agent skill from affaan-m/ECC.

affaan-m/ECC276k2 repos~884Automated safety check: PassMIT4 days ago
162

평가 주도 개발(EDD) 원칙을 구현하는 Claude Code 세션용 공식 평가 프레임워크. An agent skill from affaan-m/ECC.

affaan-m/ECC276k2 repos~1.2kAutomated safety check: PassMIT4 days ago
163

Patient safety evaluation harness for healthcare application deployments.

affaan-m/ECC276k1 repo~2kAutomated safety check: PassMIT4 days ago
164

This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target.

jacob-dietle/context-os111—~5.2kAutomated safety check: PassMIT1 mo ago
165

Entry point for evals. An agent skill from ai-evals-course/evals-skills.

ai-evals-course/evals-skills1.5k—~412Automated safety check: PassApache-2.014 days ago
166

A skill your agent uses when writing, editing, or reviewing evalite-scored agent evals in packages/core/compute/assistant-evals/src/evals.

dxos/dxos525—~4.1kAutomated safety check: PassUnknowntoday
167
167.Eval Suite PlannerOfficial

Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description.

microsoft/eval-guide138—~2.3kAutomated safety check: PassMIT3 mo ago
168

A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness.

tikalk/adlc-team-skills141—~1.4kAutomated safety check: PassMITtoday
169

Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm…

comet-ml/opik-mcp220—~3kAutomated safety check: NotesApache-2.0today
170
170.EveOfficial

eve framework guidance for durable AI agents and agent-powered applications.

vercel/vercel-plugin3015 repos~1.2kAutomated safety check: PassUnknownyesterday
171

Autonomously optimize a Claude Code skill or agent system by running it repeatedly, scoring outputs against evals, mutating owned artifacts (prompt, references, scripts, agent definitions), and…

byungjunjang/jangpm-meta-skills120—~6kAutomated safety check: WarnNo licence2 mo ago
172

Scaffold a synthetic gold-standard annotation project for evaluating OpenMed NER and de-identification models — label schema, annotation guidelines, BRAT or Label Studio config, and disjoint…

maziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0today
173

Evaluate an OpenMed de-identification or clinical NER model against the leakage-first release gates G1a through G8, which gate releases on residual PHI leakage rather than on F1.

maziyarpanahi/openmed5.5k—~2kAutomated safety check: PassApache-2.0today
174

do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals.

Devin-AXIS/iPolloWork6.8k—~851Automated safety check: PassUnknownyesterday
175

Build custom LLM evaluation pipelines using the OpenJudge framework.

agentscope-ai/OpenJudge870—~1.3kAutomated safety check: PassApache-2.028 days ago
176
176.Phoenix CLIOfficial

Debug LLM applications using the Phoenix CLI. An agent skill from github/awesome-copilot.

github/awesome-copilot40k1 repo~4kAutomated safety check: PassApache-2.0today
177

Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware.

sickn33/agentic-awesome-skills47k1 repo~1.9kAutomated safety check: PassApache-2.0today
178

Curated upstream guidance for Huggingface Community Evals; use when the workflow matches the user goal.

sickn33/agentic-awesome-skills47k1 repo~1.7kAutomated safety check: PassMITtoday
179

UiPath API Workflow assistant — author, run, validate, package, publish, deploy, and troubleshoot JSON workflows for uip api-workflow.

UiPath/skills168—~7.5kAutomated safety check: NotesMITtoday
180

Andrej Karpathy 视角顾问。以 Karpathy 的心智模型和工程师式判断,为 AI 产品经理场景做分析。

SpaceZephyr/career.skill160—~1.2kAutomated safety check: PassNo licence3 mo ago
181

Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi

affaan-m/ECC276k1 repo~1.7kAutomated safety check: PassMIT4 days ago
182

Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.

sickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: PassMITtoday
183

Evaluate retrieval and citation behavior for RAG pipelines from deterministic JSONL fixtures.

davepoon/buildwithclaude3.6k—~813Automated safety check: PassMITtoday
184

Understand what a Phoenix dataset is and reason well about its examples, outputs, splits, and how it feeds evaluators and experiments.

Arize-ai/phoenix12k—~1.6kAutomated safety check: PassUnknowntoday
185

A skill your agent uses when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline.

tikalk/adlc-team-skills141—~934Automated safety check: PassMITtoday
186

Guardrailed DELETE of auto-registered eval sandboxjobs rows that DID score but FAILED the harvest gate — partial evals (valid-complete <90% or non-benign infra-error 10%).

open-thoughts/OpenThoughts-Agent301—~2.8kAutomated safety check: PassApache-2.010 days ago
187

克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则

affaan-m/ECC276k3 repos~916Automated safety check: PassMIT4 days ago
188

Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset.

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~2.5kAutomated safety check: PassApache-2.0today
189

Validates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation.

mikeyobrien/ralph-orchestrator3.2k—~3kAutomated safety check: PassMIT4 days ago
190
190.Arize EvaluatorOfficial

Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

github/awesome-copilot40k1 repo~8.1kAutomated safety check: NotesMITtoday
191

Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.

leo-kuang-ai/spec-first107—~13kAutomated safety check: PassMITtoday
192

AI agent and LLM system engineering reference covering single-agent dev (ReAct, tool calling, plan-execute), multi-agent coordination (swarm, role decomposition, file locking), LLM security (prompt…

telagod/code-abyss243—~691Automated safety check: PassMIT2 mo ago