Search
LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | 49.Openai Docs A skill your agent uses for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or… | theowenyoung/ | 115 | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 50 | Propose an improved version of a prompt registered in a self-hosted AgentX (AgentX-trace-eval) instance, using real low-rated evaluation results as evidence, then publish it as a new version once… | AgentX-ai/ | 106 | — | ~2k | Automated safety check: Pass | Unknown | 3 days ago |
| 51 | Runs the `autoctx` CLI to improve an approach to a task over several generations, score or refine a single output and inspect what a run produced. | greyhaven-ai/ | 1.3k | — | ~964 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 52 | Create, refine, and benchmark agent skills. An agent skill from feiskyer/claude-code-settings. | feiskyer/ | 1.7k | — | ~7.6k | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 53 | Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars… | evaleval/ | 135 | — | ~2.5k | Automated safety check: Pass | MIT | 3 days ago |
| 54 | 54.Evaluation Install and run a verifiers environment — smoke testing during development and full benchmark evals. | PrimeIntellect-ai/ | 131 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 55 | Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch. | ory/ | 307 | — | ~497 | Automated safety check: Pass | Unknown | 2 mo ago |
| 56 | Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries. | ai-evals-course/ | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 57 | Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing… | ProfSynapse/ | 157 | — | ~1k | Automated safety check: Pass | MIT | yesterday |
| 58 | A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. | NVIDIA/ | 554 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 59 | Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and… | NVIDIA-AI-Blueprints/ | 1.9k | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 60 | Create new skills, modify and improve existing skills, and measure skill performance. | deepklarity/ | 100 | — | ~3.2k | Automated safety check: Notes | MIT | 2 mo ago |
| 61 | Build DSPy evaluation harnesses with rich-feedback metrics that are essential for GEPA optimization. | intertwine/ | 278 | — | ~1.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 62 | 62.Gate Check Find the decisions in a pipeline that do not need the expensive model and propose the gate for each: a rule, a classic classifier, or a small model, with fail-closed routing. | undefined-ui/ | 1k | — | ~854 | Automated safety check: Pass | MIT | yesterday |
| 63 | 63.Add Eval Design, implement, validate, and calibrate a new eval for the convex-evals suite. | get-convex/ | 130 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 64 | Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover. | starkyru/ | 107 | — | ~1.9k | Automated safety check: Pass | MIT | 2 mo ago |
| 65 | Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics. | Orchestra-Research/ | 13k | 4 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 66 | Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter. | microsoft/ | 1.4k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 67 | Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data. | ai-evals-course/ | 1.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 68 | Ship and spec AI features, LLM products, agents, copilots, and generative UX — including when to use a model vs. | andreaskelm/ | 234 | — | ~1.8k | Automated safety check: Pass | Unknown | 2 days ago |
| 69 | Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/). | Arize-ai/ | 12k | — | ~7.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 70 | Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks. | amd/ | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 71 | Run a quality benchmark of the /translate skill by selecting stratified test keys, capturing ground truth, translating, judging with sub-agents, and compiling a regression report. | shapeshift/ | 206 | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 72 | Scaffolds evaluation suites for the Axiom AI SDK: eval files, scorers, flag schemas and axiom.config.ts, generated from plain descriptions of an AI capability. | openclaw/ | 9.5k | — | ~4.1k | Automated safety check: Warn | MIT | 2 days ago |
| 73 | Test and evaluation harness for AI agents — scenario suites, deterministic replay, regression diffing, cost and latency budgets. | borghei/ | 891 | — | ~3.1k | Automated safety check: Pass | MIT | 4 days ago |
| 74 | 74.Goal Test Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent. | undefined-ui/ | 1k | — | ~754 | Automated safety check: Pass | MIT | yesterday |
| 75 | 75.Benchflow Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow. | benchflow-ai/ | 356 | — | ~1.9k | Automated safety check: Notes | Apache-2.0 | 5 days ago |
| 76 | Benchmark Claude Code skill performance with variance analysis, tracking pass rate, execution time, and token usage across iterations. | AgriciDaniel/ | 179 | — | ~1.4k | Automated safety check: Pass | MIT | 6 mo ago |
| 77 | Run a skill's evals and report results. An agent skill from redhat-cop/vault-config-operator. | redhat-cop/ | 167 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 78 | 78.Add Model Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs. | get-convex/ | 130 | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 79 | 79.Eval Run Evaluate any output file against a structured evals.yaml assertions file and produce a score report with per-assertion pass/fail results. | digipulse-engineering/ | 163 | — | ~1.7k | Automated safety check: Pass | Unknown | 12 days ago |
| 80 | Create new skills, improve existing skills, and measure skill performance. | SpectrAI-Initiative/ | 396 | — | ~8.4k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 81 | Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI. | diegosouzapw/ | 75k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 82 | Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low. | microsoft/ | 1.4k | — | ~1.3k | Automated safety check: Pass | MIT | yesterday |
| 83 | Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format. | ai-evals-course/ | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 84 | Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes. | microsoft/ | 3.1k | — | ~2.8k | Automated safety check: Pass | MIT | yesterday |
| 85 | Create, evaluate, improve, and benchmark content skills using the local Skill Lab workflow. | ceilf6/ | 120 | — | ~747 | Automated safety check: Pass | MIT | 11 days ago |
| 86 | 86.Inspect Inspect and debug live streaming agent sessions to understand what the agent did. | agentevals-dev/ | 163 | — | ~534 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 87 | 测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。 | YIKUAIBANZI/ | 122 | — | ~531 | Automated safety check: Pass | MIT | 6 mo ago |
| 88 | Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100… | open-thoughts/ | 301 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 89 | Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim. | Leeroo-AI/ | 195 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 90 | 為 Twinkle Eval 新增一個評測 benchmark(IFEval、BFCL、RAGAS、Text2SQL、Vision MCQ 之類)。涵蓋 CLAUDE.md §6 的完整強制流程:先建 Milestone 與 6 個 Issue、準備 example dataset、實作 Extractor + Scorer 並註冊 PRESETS、與參考框架做分數與速度對比、撰寫… | ai-twinkle/ | 117 | — | ~1.8k | Automated safety check: Pass | MIT | 26 days ago |
| 91 | Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV). | yaalalabs/ | 192 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 92 | 92.Email Evals Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents. | tokencanopy/ | 193 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 93 | Finds top models for a task from official Hugging Face benchmark leaderboards, filters them to what fits your hardware, and returns a comparison table with scores. | huggingface/ | 11k | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 94 | Build an LLM evaluation and run it against the app, returning an Opik experiment with scores and its link. | comet-ml/ | 220 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 3 days ago |
| 95 | This skill should be used when the user wants to "run an evaluation", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs… | pifferologo/ | 129 | 1 repo | ~6.8k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 96 | Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. | Orchestra-Research/ | 13k | 2 repos | ~3.1k | Automated safety check: Pass | MIT | 3 mo ago |