Search

LLM evaluation

306 skills found, page 2.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

A skill your agent uses for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or…

theowenyoung/home1151 repo~1.4kAutomated safety check: PassApache-2.014 days ago
50

Propose an improved version of a prompt registered in a self-hosted AgentX (AgentX-trace-eval) instance, using real low-rated evaluation results as evidence, then publish it as a new version once…

AgentX-ai/AgentX-Trace-Eval106—~2kAutomated safety check: PassUnknown3 days ago
51

Runs the `autoctx` CLI to improve an approach to a task over several generations, score or refine a single output and inspect what a run produced.

greyhaven-ai/autocontext1.3k—~964Automated safety check: PassApache-2.04 days ago
52

Create, refine, and benchmark agent skills. An agent skill from feiskyer/claude-code-settings.

feiskyer/claude-code-settings1.7k—~7.6kAutomated safety check: PassApache-2.014 days ago
53

Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars…

evaleval/every_eval_ever135—~2.5kAutomated safety check: PassMIT3 days ago
54

Install and run a verifiers environment — smoke testing during development and full benchmark evals.

PrimeIntellect-ai/prime-envs131—~4.6kAutomated safety check: PassApache-2.0yesterday
55

Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch.

ory/lumen307—~497Automated safety check: PassUnknown2 mo ago
56

Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

ai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.016 days ago
57

Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing…

ProfSynapse/nexus157—~1kAutomated safety check: PassMITyesterday
58

A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

NVIDIA/SkillEvaluator554—~2.1kAutomated safety check: PassApache-2.0yesterday
59

Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~2.1kAutomated safety check: NotesApache-2.0yesterday
60

Create new skills, modify and improve existing skills, and measure skill performance.

deepklarity/harness-kit100—~3.2kAutomated safety check: NotesMIT2 mo ago
61

Build DSPy evaluation harnesses with rich-feedback metrics that are essential for GEPA optimization.

intertwine/dspy-agent-skills278—~1.5kAutomated safety check: PassMIT1 mo ago
62

Find the decisions in a pipeline that do not need the expensive model and propose the gate for each: a rule, a classic classifier, or a small model, with fail-closed routing.

undefined-ui/second-brain-os1k—~854Automated safety check: PassMITyesterday
63

Design, implement, validate, and calibrate a new eval for the convex-evals suite.

get-convex/convex-evals130—~4.6kAutomated safety check: PassApache-2.0yesterday
64

Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover.

starkyru/learn-ai107—~1.9kAutomated safety check: PassMIT2 mo ago
65

Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.9kAutomated safety check: PassMIT3 mo ago
66

Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.

microsoft/waza1.4k—~2kAutomated safety check: PassMITyesterday
67

Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

ai-evals-course/evals-skills1.5k—~2.2kAutomated safety check: PassApache-2.016 days ago
68

Ship and spec AI features, LLM products, agents, copilots, and generative UX — including when to use a model vs.

andreaskelm/pm-brain234—~1.8kAutomated safety check: PassUnknown2 days ago
69

Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/).

Arize-ai/phoenix12k—~7.1kAutomated safety check: PassApache-2.0yesterday
70

Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

amd/gaia1.6k—~1.8kAutomated safety check: PassMITtoday
71

Run a quality benchmark of the /translate skill by selecting stratified test keys, capturing ground truth, translating, judging with sub-agents, and compiling a regression report.

shapeshift/web206—~1.6kAutomated safety check: PassMITyesterday
72

Scaffolds evaluation suites for the Axiom AI SDK: eval files, scorers, flag schemas and axiom.config.ts, generated from plain descriptions of an AI capability.

openclaw/clawhub9.5k—~4.1kAutomated safety check: WarnMIT2 days ago
73

Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets.

borghei/Claude-Skills891—~3.1kAutomated safety check: PassMIT4 days ago
74

Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent.

undefined-ui/second-brain-os1k—~754Automated safety check: PassMITyesterday
75

Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow.

benchflow-ai/benchflow356—~1.9kAutomated safety check: NotesApache-2.05 days ago
76

Benchmark Claude Code skill performance with variance analysis, tracking pass rate, execution time, and token usage across iterations.

AgriciDaniel/skill-forge179—~1.4kAutomated safety check: PassMIT6 mo ago
77

Run a skill's evals and report results. An agent skill from redhat-cop/vault-config-operator.

redhat-cop/vault-config-operator167—~2.1kAutomated safety check: PassApache-2.0yesterday
78

Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs.

get-convex/convex-evals130—~1.5kAutomated safety check: NotesApache-2.0yesterday
79

Evaluate any output file against a structured evals.yaml assertions file and produce a score report with per-assertion pass/fail results.

digipulse-engineering/GAAI-framework163—~1.7kAutomated safety check: PassUnknown12 days ago
80

Create new skills, improve existing skills, and measure skill performance.

SpectrAI-Initiative/InnoClaw396—~8.4kAutomated safety check: PassApache-2.02 mo ago
81

Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

diegosouzapw/OmniRoute75k—~1.3kAutomated safety check: PassMITtoday
82
82.Waza InteractiveOfficial

Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

microsoft/waza1.4k—~1.3kAutomated safety check: PassMITyesterday
83

Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.

ai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.016 days ago
84

Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

microsoft/skills3.1k—~2.8kAutomated safety check: PassMITyesterday
85

Create, evaluate, improve, and benchmark content skills using the local Skill Lab workflow.

ceilf6/FrontAgent120—~747Automated safety check: PassMIT11 days ago
86

Inspect and debug live streaming agent sessions to understand what the agent did.

agentevals-dev/agentevals163—~534Automated safety check: PassApache-2.0yesterday
87

测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。

YIKUAIBANZI/forge-skill122—~531Automated safety check: PassMIT6 mo ago
88

Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

open-thoughts/OpenThoughts-Agent301—~3.1kAutomated safety check: PassApache-2.012 days ago
89

Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim.

Leeroo-AI/superml195—~3.8kAutomated safety check: PassApache-2.06 mo ago
90

為 Twinkle Eval 新增一個評測 benchmark(IFEval、BFCL、RAGAS、Text2SQL、Vision MCQ 之類)。涵蓋 CLAUDE.md §6 的完整強制流程:先建 Milestone 與 6 個 Issue、準備 example dataset、實作 Extractor + Scorer 並註冊 PRESETS、與參考框架做分數與速度對比、撰寫…

ai-twinkle/Eval117—~1.8kAutomated safety check: PassMIT26 days ago
91

Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV).

yaalalabs/agent-kernel192—~3.4kAutomated safety check: PassApache-2.02 days ago
92

Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents.

tokencanopy/e2a193—~2.1kAutomated safety check: PassApache-2.0yesterday
93

Finds top models for a task from official Hugging Face benchmark leaderboards, filters them to what fits your hardware, and returns a comparison table with scores.

huggingface/skills11k2 repos~1.5kAutomated safety check: PassApache-2.03 days ago
94

Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link.

comet-ml/opik-mcp220—~2.5kAutomated safety check: NotesApache-2.03 days ago
95

This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs…

pifferologo/cloud-agents-cli1291 repo~6.8kAutomated safety check: PassApache-2.01 mo ago
96

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution.

Orchestra-Research/AI-Research-SKILLs13k2 repos~3.1kAutomated safety check: PassMIT3 mo ago