Topic · AI & LLM Engineering

Best LLM evaluation skills, page 2

Skills #49–96 of 308, ranked by score.

LLM evaluation skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

LLM evaluation skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Build DSPy 3.2.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget.

intertwine/dspy-agent-skills278—~1.7kAutomated safety check: PassMIT1 mo ago
50

Prepare the environment and run the LLM-driven agent evals (e2e/agent-evals/) against a chosen sim-use binary.

lycorp-jp/sim-use1.4k—~1.4kAutomated safety check: PassApache-2.0today
51

Audit an agent's context layout against the four places: system prompt, tools, history, tail.

undefined-ui/second-brain-os1k—~802Automated safety check: PassMITtoday
52
52.Eval

Evaluate and score agent behavior against a golden reference.

agentevals-dev/agentevals162—~904Automated safety check: PassApache-2.0today
53

End-to-end GECX/CXAS/CES conversational agent lifecycle -- build agents from requirements (PRD-to-agent), create and run evals (goldens, simulations, tool tests, callback tests), debug failures, and…

GoogleCloudPlatform/cxas-scrapi107—~2.4kAutomated safety check: PassApache-2.0today
54

A skill your agent uses for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or…

theowenyoung/home1151 repo~1.4kAutomated safety check: PassApache-2.011 days ago
55

Propose an improved version of a prompt registered in a self-hosted AgentX (AgentX-trace-eval) instance, using real low-rated evaluation results as evidence, then publish it as a new version once…

AgentX-ai/AgentX-Trace-Eval106—~2kAutomated safety check: PassUnknowntoday
56

Runs the `autoctx` CLI to improve an approach to a task over several generations, score or refine a single output and inspect what a run produced.

greyhaven-ai/autocontext1.3k—~964Automated safety check: PassApache-2.0yesterday
57

Create, refine, and benchmark agent skills. An agent skill from feiskyer/claude-code-settings.

feiskyer/claude-code-settings1.7k—~7.6kAutomated safety check: PassApache-2.010 days ago
58

Designs, tests and refines LLM prompts: zero-shot, few-shot and chain-of-thought patterns, system prompts, structured output schemas and evaluation test suites.

Jeffallan/claude-skills12k1 repo~1.5kAutomated safety check: PassMIT4 days ago
59

Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…

evaleval/every_eval_ever133—~2.5kAutomated safety check: PassMITtoday
60

Install and run a verifiers environment — smoke testing during development and full benchmark evals.

PrimeIntellect-ai/prime-envs130—~4.6kAutomated safety check: PassApache-2.0today
61

Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

ory/lumen305—~497Automated safety check: PassUnknown1 mo ago
62

Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

ai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.013 days ago
63

Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing…

ProfSynapse/nexus156—~1kAutomated safety check: PassMITtoday
64

Create new skills, modify and improve existing skills, and measure skill performance.

deepklarity/harness-kit100—~3.2kAutomated safety check: NotesMIT2 mo ago
65

Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.

Orchestra-Research/AI-Research-SKILLs13k5 repos~2.9kAutomated safety check: PassMIT3 mo ago
66

Build DSPy evaluation harnesses with rich-feedback metrics that are essential for GEPA optimization.

intertwine/dspy-agent-skills278—~1.5kAutomated safety check: PassMIT1 mo ago
67

Scaffold a first eval suite for an agent: mine real failures into cases, write behavioural checks over traces, and generate the runner.

undefined-ui/second-brain-os1k—~786Automated safety check: PassMITtoday
68

Write and review content for the OpenSEO website (web/) — blog posts, guides, feature pages, FAQs.

Jwuthri/Tracely-ai1.5k—~922Automated safety check: PassMITtoday
69

Design, implement, validate, and calibrate a new eval for the convex-evals suite.

get-convex/convex-evals129—~4.6kAutomated safety check: PassApache-2.0today
70

Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.

starkyru/learn-ai107—~1.9kAutomated safety check: PassMIT2 mo ago
71

Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step.

Jeffallan/claude-skills12k1 repo~2kAutomated safety check: PassMIT4 days ago
72

Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.

microsoft/waza1.4k—~2kAutomated safety check: PassMITyesterday
73

Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

ai-evals-course/evals-skills1.5k—~2.2kAutomated safety check: PassApache-2.013 days ago
74

Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

amd/gaia1.6k—~1.8kAutomated safety check: PassMITtoday
75

Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/).

Arize-ai/phoenix12k—~7.1kAutomated safety check: PassApache-2.0today
76

Run a quality benchmark of the /translate skill by selecting stratified test keys, capturing ground truth, translating, judging with sub-agents, and compiling a regression report.

shapeshift/web206—~1.6kAutomated safety check: PassMITtoday
77

Scaffolds evaluation suites for the Axiom AI SDK: eval files, scorers, flag schemas and axiom.config.ts, generated from plain descriptions of an AI capability.

openclaw/clawhub9.5k—~4.1kAutomated safety check: WarnMITtoday
78

Find the decisions in a pipeline that do not need the expensive model and propose the gate for each: a rule, a classic classifier, or a small model, with fail-closed routing.

undefined-ui/second-brain-os1k—~854Automated safety check: PassMITtoday
79

Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

benchflow-ai/benchflow353—~1.9kAutomated safety check: NotesApache-2.02 days ago
80

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution.

Orchestra-Research/AI-Research-SKILLs13k3 repos~3.1kAutomated safety check: PassMIT3 mo ago
81

Benchmark Claude Code skill performance with variance analysis, tracking pass rate, execution time, and token usage across iterations.

AgriciDaniel/skill-forge177—~1.4kAutomated safety check: PassMIT6 mo ago
82

Run a skill's evals and report results. An agent skill from redhat-cop/vault-config-operator.

redhat-cop/vault-config-operator167—~2.1kAutomated safety check: PassApache-2.0yesterday
83

Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

get-convex/convex-evals129—~1.5kAutomated safety check: NotesApache-2.0today
84

Evaluate any output file against a structured evals.yaml assertions file and produce a score report with per-assertion pass/fail results.

digipulse-engineering/GAAI-framework163—~1.7kAutomated safety check: PassUnknown8 days ago
85

Create new skills, improve existing skills, and measure skill performance.

SpectrAI-Initiative/InnoClaw396—~8.4kAutomated safety check: PassApache-2.01 mo ago
86
86.Waza InteractiveOfficial

Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

microsoft/waza1.4k—~1.3kAutomated safety check: PassMITyesterday
87

Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.

ai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.013 days ago
88

Create, evaluate, improve, and benchmark content skills using the local Skill Lab workflow.

ceilf6/FrontAgent120—~747Automated safety check: PassMIT8 days ago
89

Inspect and debug live streaming agent sessions to understand what the agent did.

agentevals-dev/agentevals162—~534Automated safety check: PassApache-2.0today
90

测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。

YIKUAIBANZI/forge-skill121—~531Automated safety check: PassMIT6 mo ago
91

Build custom LLM evaluation pipelines using the OpenJudge framework.

agentscope-ai/OpenJudge8681 repo~1.3kAutomated safety check: PassApache-2.027 days ago
92

Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

open-thoughts/OpenThoughts-Agent301—~3.1kAutomated safety check: PassApache-2.09 days ago
93

Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.

Leeroo-AI/superml195—~3.8kAutomated safety check: PassApache-2.06 mo ago
94

為 Twinkle Eval 新增一個評測 benchmark(IFEval、BFCL、RAGAS、Text2SQL、Vision MCQ 之類)。涵蓋 CLAUDE.md §6 的完整強制流程:先建 Milestone 與 6 個 Issue、準備 example dataset、實作 Extractor + Scorer 並註冊 PRESETS、與參考框架做分數與速度對比、撰寫…

ai-twinkle/Eval117—~1.8kAutomated safety check: PassMIT22 days ago
95

Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV).

yaalalabs/agent-kernel191—~3.4kAutomated safety check: PassApache-2.0today
96

Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.

tokencanopy/e2a192—~2.1kAutomated safety check: PassApache-2.02 days ago