Search
LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | Create a new built-in classification evaluator for Phoenix evals. | Arize-ai/ | 12k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 146 | Write, extend, and debug PXI Playwright E2E tests for Phoenix. | Arize-ai/ | 12k | — | ~2.6k | Automated safety check: Pass | Unknown | yesterday |
| 147 | Bump the next release-please version for a Phoenix Python package (arize-phoenix, arize-phoenix-client, arize-phoenix-evals, arize-phoenix-otel) by opening a PR with a Release-As commit footer. | Arize-ai/ | 12k | — | ~708 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 148 | Audit recent changes to Phoenix's user-facing surfaces (Python clients, TypeScript clients, CLI, REST/GraphQL APIs) and patch the three external-facing agent skills — phoenix-tracing, phoenix-cli… | Arize-ai/ | 12k | — | ~5.1k | Automated safety check: Pass | Unknown | yesterday |
| 149 | Maintain the bundled TypeScript package docs that ship inside Phoenix npm packages. | Arize-ai/ | 12k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 150 | Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository. | dotnet/ | 5.6k | 1 repo | ~6.1k | Automated safety check: Pass | MIT | yesterday |
| 151 | 151.Bkit Explore Browse installed bkit skills, agents, and evals via lib/discovery/explorer.js (filesystem scan, no subprocess). | ww-w-ai/ | 601 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 152 | Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 153 | Patterns and techniques for evaluating and improving AI agent outputs. | github/ | 40k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | 2 days ago |
| 154 | 154.Watch Monitor arXiv for new papers relevant to the GLIDE project (prediction-powered inference, active statistical inference, LLM evaluation debiasing, proxy annotation bias correction). | EmertonData/ | 119 | — | ~3.5k | Automated safety check: Pass | Unknown | 2 days ago |
| 155 | 155.Write Skill Creates or changes a druxt.js contributor skill in .agents/skills, with its evals and the tests that gate it. | druxt/ | 114 | — | ~926 | Automated safety check: Pass | MIT | today |
| 156 | 156.Evals Clarify A skill your agent uses when draft eval criteria need refining, clustering, and acceptance into the published goldset with an isolated holdout split (goldset.md + goldset.json). | tikalk/ | 141 | — | ~1.1k | Automated safety check: Pass | MIT | 2 days ago |
| 157 | Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates. | Arenukvern/ | 387 | — | ~2.4k | Automated safety check: Pass | MIT | 8 days ago |
| 158 | Find out what is going wrong in LLM or agent traffic by reading sampled Phoenix traces, spans, or sessions, writing free-form notes (open coding), then grouping the notes into a few narrow… | Arize-ai/ | 12k | — | ~6.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 159 | 159.Eval Harness Claude Codeセッションの正式な評価フレームワークで、評価駆動開発(EDD)の原則を実装します. An agent skill from affaan-m/ECC. | affaan-m/ | 277k | 2 repos | ~884 | Automated safety check: Pass | MIT | today |
| 160 | 160.Eval Harness 평가 주도 개발(EDD) 원칙을 구현하는 Claude Code 세션용 공식 평가 프레임워크. An agent skill from affaan-m/ECC. | affaan-m/ | 277k | 2 repos | ~1.2k | Automated safety check: Pass | MIT | today |
| 161 | Patient safety evaluation harness for healthcare application deployments. | affaan-m/ | 277k | 1 repo | ~2k | Automated safety check: Pass | MIT | today |
| 162 | 162.Eval Loop This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target. | jacob-dietle/ | 111 | — | ~5.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 163 | 163.Evals Start Entry point for evals. An agent skill from ai-evals-course/evals-skills. | ai-evals-course/ | 1.5k | — | ~412 | Automated safety check: Pass | Apache-2.0 | 17 days ago |
| 164 | 164.Agent Eval Tests A skill your agent uses when writing, editing, or reviewing evalite-scored agent evals in packages/core/compute/assistant-evals/src/evals. | dxos/ | 526 | — | ~4.1k | Automated safety check: Pass | Unknown | today |
| 165 | Plan standalone — populates the Eval Suite Planning & Logging Template from an Agent Vision or plain-English agent description. | microsoft/ | 138 | — | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 166 | 166.Evals Implement A skill your agent uses when a published goldset is ready to become executable graders and configs — Python graders and metrics with auto-run unit tests verifying grader correctness. | tikalk/ | 141 | — | ~1.4k | Automated safety check: Pass | MIT | 2 days ago |
| 167 | 167.Opik Online Eval Take a judge live on production traffic — create an Opik online evaluation rule (LLM-as-judge or Python metric) on a project with sampling, filters, variable mapping, and a cost cap, then confirm… | comet-ml/ | 220 | — | ~3k | Automated safety check: Notes | Apache-2.0 | 3 days ago |
| 168 | eve framework guidance for durable AI agents and agent-powered applications. | vercel/ | 301 | 5 repos | ~1.2k | Automated safety check: Pass | Unknown | yesterday |
| 169 | 169.Autoresearch Autonomously optimize a Claude Code skill or agent system by running it repeatedly, scoring outputs against evals, mutating owned artifacts (prompt, references, scripts, agent definitions), and… | byungjunjang/ | 120 | — | ~6k | Automated safety check: Warn | No licence | 2 mo ago |
| 170 | Scaffold a synthetic gold-standard annotation project for evaluating OpenMed NER and de-identification models — label schema, annotation guidelines, BRAT or Label Studio config, and disjoint… | maziyarpanahi/ | 5.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 171 | Evaluate an OpenMed de-identification or clinical NER model against the leakage-first release gates G1a through G8, which gate releases on residual PHI leakage rather than on F1. | maziyarpanahi/ | 5.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 172 | Build custom LLM evaluation pipelines using the OpenJudge framework. | agentscope-ai/ | 871 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 173 | 173.Run Evals do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals. | Devin-AXIS/ | 6.8k | — | ~851 | Automated safety check: Pass | Unknown | yesterday |
| 174 | Debug LLM applications using the Phoenix CLI. An agent skill from github/awesome-copilot. | github/ | 40k | 1 repo | ~4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 175 | Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware. | sickn33/ | 47k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 176 | Curated upstream guidance for Huggingface Community Evals; use when the workflow matches the user goal. | sickn33/ | 47k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 177 | UiPath API Workflow assistant — author, run, validate, package, publish, deploy, and troubleshoot JSON workflows for uip api-workflow. | UiPath/ | 167 | — | ~7.5k | Automated safety check: Notes | MIT | today |
| 178 | Andrej Karpathy 视角顾问。以 Karpathy 的心智模型和工程师式判断,为 AI 产品经理场景做分析。 | SpaceZephyr/ | 160 | — | ~1.2k | Automated safety check: Pass | No licence | 3 mo ago |
| 179 | 179.Eval Harness Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi | affaan-m/ | 277k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | today |
| 180 | Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing. | sickn33/ | 47k | 2 repos | ~3.1k | Automated safety check: Pass | MIT | yesterday |
| 181 | Evaluate retrieval and citation behavior for RAG pipelines from deterministic JSONL fixtures. | davepoon/ | 3.6k | — | ~813 | Automated safety check: Pass | MIT | yesterday |
| 182 | 182.Datasets Understand what a Phoenix dataset is and reason well about its examples, outputs, splits, and how it feeds evaluators and experiments. | Arize-ai/ | 12k | — | ~1.6k | Automated safety check: Pass | Unknown | yesterday |
| 183 | 183.Evals Init A skill your agent uses when standing up evals/{system}/ for the first time — scaffolds the EDD directory structure, picks PromptFoo or DeepEval by tech stack, and generates a security baseline. | tikalk/ | 141 | — | ~934 | Automated safety check: Pass | MIT | 2 days ago |
| 184 | Guardrailed DELETE of auto-registered eval sandboxjobs rows that DID score but FAILED the harvest gate — partial evals (valid-complete <90% or non-benign infra-error 10%). | open-thoughts/ | 301 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 185 | 185.Eval Harness 克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则 | affaan-m/ | 277k | 3 repos | ~916 | Automated safety check: Pass | MIT | today |
| 186 | Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset. | NVIDIA-AI-Blueprints/ | 1.9k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 187 | 187.Tui Validate Validates Terminal User Interface (TUI) output using freeze for screenshot capture and LLM-as-judge for semantic validation. | mikeyobrien/ | 3.2k | — | ~3k | Automated safety check: Pass | MIT | 6 days ago |
| 188 | Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and… | github/ | 40k | 1 repo | ~8.1k | Automated safety check: Notes | MIT | 2 days ago |
| 189 | 189.Spec Optimize Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first. | leo-kuang-ai/ | 107 | — | ~13k | Automated safety check: Pass | MIT | 3 days ago |
| 190 | AI agent and LLM system engineering reference covering single-agent dev (ReAct, tool calling, plan-execute), multi-agent coordination (swarm, role decomposition, file locking), LLM security (prompt… | telagod/ | 244 | — | ~691 | Automated safety check: Pass | MIT | 2 mo ago |
| 191 | Build and run evaluators for AI/LLM applications using Phoenix. | github/ | 40k | 2 repos | ~1.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 192 | 192.LLM Ops LLM Operations -- RAG, embeddings, vector databases, fine-tuning, prompt engineering avancado, custos de LLM, evals de qualidade e arquiteturas de IA para producao. | davila7/ | 33k | 3 repos | ~2k | Automated safety check: Pass | MIT | yesterday |