Search

LLM evaluation

306 skills found, page 3.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
97

Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT3 mo ago
98

Investigate a single failing eval from the convex-evals system.

get-convex/convex-evals130—~1.1kAutomated safety check: PassApache-2.0yesterday
99

Best practices for creating expectations and grader files to evaluate guidance quality.

GoogleChrome/modern-web-guidance-src1.1k—~2.3kAutomated safety check: PassApache-2.0today
100

Create new skills, modify and improve existing skills, and measure skill performance.

ZS520L/HanakoPro103—~7.4kAutomated safety check: PassApache-2.04 mo ago
101

Add a new evaluator to the amp-evaluation Python library. An agent skill from wso2/agent-manager.

wso2/agent-manager108—~710Automated safety check: PassApache-2.02 days ago
102

This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…

guanyang/open-agent-hub9772 repos~4.2kAutomated safety check: PassMITtoday
103

测试 use-self 替身会议的辩论质量。给定 persona + 3 个决策场景,运行完整三阶段辩论并按 5 个维度评分,输出质量报告。

YIKUAIBANZI/forge-skill122—~740Automated safety check: PassMIT6 mo ago
104

A skill your agent uses when the user wants to work on a Claude Code skill file (SKILL.md): writing one from scratch, testing whether an existing one works well, running evals or benchmarks…

avibebuilder/claude-prime120—~8.2kAutomated safety check: PassMIT4 mo ago
105

Decide which AI agent behaviors are worth an eval case, then write those cases — harness-, framework-, and language-agnostic.

agentailor/fullstack-langgraph-nextjs-agent132—~5.3kAutomated safety check: PassMIT1 mo ago
106

Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

AnastasiyaW/codex-claude-code-config154—~3.8kAutomated safety check: PassMITtoday
107
107.Eval GuideOfficial

Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

microsoft/eval-guide138—~22kAutomated safety check: WarnMIT3 mo ago
108

Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations.

get-convex/convex-evals130—~2.1kAutomated safety check: PassApache-2.0yesterday
109

Test a pre-built afm binary at any path — runs pre-flight safety checks, then any combination of unit tests, assertions, smart analysis, promptfoo evals, batch validation, OpenAI compat, GPU…

scouzi1966/maclocal-api346—~3.8kAutomated safety check: PassMITyesterday
110

AI quality judge that scores agent responses 0-10 across helpfulness, accuracy, completeness, and clarity.

Atmosphere/atmosphere3.8k—~333Automated safety check: PassApache-2.02 days ago
111

Score a Python generative-AI app's outputs on groundedness, relevance, coherence, and safety before it ships.

timothywarner-org/claude-code224—~696Automated safety check: NotesMIT2 mo ago
112

Systematic benchmarking framework for Local Brain Search memory system with LLM-as-judge scoring

Abilityai/cornelius109—~3.1kAutomated safety check: NotesMIT3 days ago
113

Runs a measurable keep-or-discard improvement loop against a runnable sensor, from framing a goal and KPI through baseline, judging and held-out verification.

bgauryy/octocode949—~1.6kAutomated safety check: PassMIT2 days ago
114

Triage failing @novu/agent-evals scenarios to decide whether a failure is real or flaky, and whether to fix the playbook/prompt or the test (grader, tape, scenario, or judge).

novuhq/novu40k—~1.5kAutomated safety check: PassUnknownyesterday
115

Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

davila7/claude-code-templates33k12 repos~3.5kAutomated safety check: PassMITtoday
116

Run evaluation pipelines on Claude Code skills to test triggering accuracy, workflow correctness, and output quality.

AgriciDaniel/skill-forge179—~1.7kAutomated safety check: PassMIT6 mo ago
117

Patterns for continuous autonomous agent loops with quality gates, evals, and recovery controls.

affaan-m/ECC277k5 repos~298Automated safety check: PassMITtoday
118

Design well-structured agent loops with best-practice coaching and cross-model review gates before you run them.

fabricioctelles/skills106—~2.1kAutomated safety check: NotesMITtoday
119

Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.

Archive228/loopkit755—~876Automated safety check: PassMIT2 mo ago
120

Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

malloydata/publisher116—~7.8kAutomated safety check: PassMITtoday
121

The step-by-step procedure for running this repo's Promptfoo evals with a generator and a judge over provider APIs.

maplibre/maplibre-agent-skills151—~3.3kAutomated safety check: PassUnknownyesterday
122
122.Logfire EvalsOfficial

Run offline Python (pydanticevals) or Node.js (logfire/evals) evaluations and review them in Logfire.

pydantic/skills140—~3.6kAutomated safety check: PassMIT10 days ago
123

Verify a Nexus change — pick a Jest lane, write a test that can actually fail, run the in-app Obsidian CLI loop, drive the eval harness, or fix a shipped-docs drift failure.

ProfSynapse/nexus157—~994Automated safety check: PassMITyesterday
124

Write or extend an E2E journey spec in evals/specs that proves a PR's change to a human reviewer.

different-ai/openwork24k—~3.3kAutomated safety check: PassUnknowntoday
125

Inspect, query, compare, or explicitly export GitHits agent-eval history in Braintrust using the repository's verified workflow.

githits-com/githits-cli115—~3.1kAutomated safety check: PassApache-2.02 days ago
126

Run a targeted local React Doctor Evals loop against an uncommitted rule change.

millionco/react-doctor15k—~510Automated safety check: PassUnknowntoday
127

Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

RefoundAI/lenny-skills1.4k—~1.7kAutomated safety check: PassMIT2 mo ago
128
128.Evals Create SuiteOfficial

Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files.

elastic/kibana21k—~1.7kAutomated safety check: PassUnknowntoday
129

Trigger an on-demand @kbn/evals Buildkite run by describing what you want in plain English.

elastic/kibana21k—~2.6kAutomated safety check: NotesUnknowntoday
130

[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).

pskoett/pskoett-ai-skills315—~2.2kAutomated safety check: PassNo licence6 days ago
131

Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

paperclipai/paperclip100k—~839Automated safety check: PassMITtoday
132

Run skill evals via evals/runner.js — wrapper validates skill names, captures stdout/stderr, persists JSON results.

ww-w-ai/bkit-claude-code601—~1kAutomated safety check: NotesApache-2.014 days ago
133

Builds an eval plan for an AI feature - error analysis on real traces, failure-mode ranking, code checks and binary LLM judges validated against human labels, CI tests, production monitoring, and a…

menkesu/awesome-pm-skills434—~4.5kAutomated safety check: PassUnknown5 days ago
134

You MUST load this skill before the skill-creator skill AND before making ANY change to, or conducting a review of ANY Agent Skill.

sammcj/agentic-coding162—~9.8kAutomated safety check: PassApache-2.02 days ago
135

Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.

tokenbender/agent-guides367—~2.7kAutomated safety check: PassApache-2.02 mo ago
136

A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

ericrisco/rsc-harness190—~3.2kAutomated safety check: PassMITyesterday
137

A skill your agent uses when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog.

tikalk/adlc-team-skills141—~1.2kAutomated safety check: PassMIT2 days ago
138

Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

Jeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT8 days ago
139

Designs, tests and refines LLM prompts: zero-shot, few-shot and chain-of-thought patterns, system prompts, structured output schemas and evaluation test suites.

Jeffallan/claude-skills12k—~1.5kAutomated safety check: PassMIT8 days ago
140

Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

Jeffallan/claude-skills12k—~2kAutomated safety check: PassMIT8 days ago
141

Set up compliance exports, drift detection, evaluations, scoring, and learning analytics

ucsandman/DashClaw311—~1.8kAutomated safety check: PassMITyesterday
142

Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate.

ai-evals-course/evals-skills1.5k—~3.7kAutomated safety check: PassApache-2.016 days ago
143

Run local evaluations for Hugging Face Hub models with inspect-ai or lighteval.

henryalouf/ruflow157—~1.6kAutomated safety check: PassMIT4 mo ago
144

Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.

glebis/claude-skills391—~2.9kAutomated safety check: PassMIT3 days ago