Topic · AI & LLM Engineering
Best LLM evaluation skills, page 2
LLM evaluation skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Build DSPy 3.2.x programs through spec, program, metric and baseline; extend to optimization and export when requested and justified by task budget. | intertwine/ | 278 | — | ~1.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 50 | 50.Run Evals Prepare the environment and run the LLM-driven agent evals (e2e/agent-evals/) against a chosen sim-use binary. | lycorp-jp/ | 1.4k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 51 | Audit an agent's context layout against the four places: system prompt, tools, history, tail. | undefined-ui/ | 1k | — | ~802 | Automated safety check: Pass | MIT | today |
| 52 | 52.Eval Evaluate and score agent behavior against a golden reference. | agentevals-dev/ | 162 | — | ~904 | Automated safety check: Pass | Apache-2.0 | today |
| 53 | End-to-end GECX/CXAS/CES conversational agent lifecycle -- build agents from requirements (PRD-to-agent), create and run evals (goldens, simulations, tool tests, callback tests), debug failures, and… | GoogleCloudPlatform/ | 107 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 54 | 54.Openai Docs A skill your agent uses for Codex models/pricing, scheduled tasks, skills, settings, setup, troubleshooting, customization, automations, and self-knowledge—including 'you,' 'your,' 'this app,' or… | theowenyoung/ | 115 | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | 11 days ago |
| 55 | Propose an improved version of a prompt registered in a self-hosted AgentX (AgentX-trace-eval) instance, using real low-rated evaluation results as evidence, then publish it as a new version once… | AgentX-ai/ | 106 | — | ~2k | Automated safety check: Pass | Unknown | today |
| 56 | Runs the `autoctx` CLI to improve an approach to a task over several generations, score or refine a single output and inspect what a run produced. | greyhaven-ai/ | 1.3k | — | ~964 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 57 | Create, refine, and benchmark agent skills. An agent skill from feiskyer/claude-code-settings. | feiskyer/ | 1.7k | — | ~7.6k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 58 | Designs, tests and refines LLM prompts: zero-shot, few-shot and chain-of-thought patterns, system prompts, structured output schemas and evaluation test suites. | Jeffallan/ | 12k | 1 repo | ~1.5k | Automated safety check: Pass | MIT | 4 days ago |
| 59 | Convert an evaluation dataset or leaderboard into the Every Eval Ever (EEE) schema — aggregate .json logs (eval.schema.json) and optional instance samples.jsonl sidecars… | evaleval/ | 133 | — | ~2.5k | Automated safety check: Pass | MIT | today |
| 60 | 60.Evaluation Install and run a verifiers environment — smoke testing during development and full benchmark evals. | PrimeIntellect-ai/ | 130 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | today |
| 61 | Adds a new task to the bench-swe pipeline from a real GitHub bug-fix issue or pull request, then checks the generated task file and patch. | ory/ | 305 | — | ~497 | Automated safety check: Pass | Unknown | 1 mo ago |
| 62 | Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries. | ai-evals-course/ | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 63 | Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing… | ProfSynapse/ | 156 | — | ~1k | Automated safety check: Pass | MIT | today |
| 64 | Create new skills, modify and improve existing skills, and measure skill performance. | deepklarity/ | 100 | — | ~3.2k | Automated safety check: Notes | MIT | 2 mo ago |
| 65 | Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics. | Orchestra-Research/ | 13k | 5 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 66 | Build DSPy evaluation harnesses with rich-feedback metrics that are essential for GEPA optimization. | intertwine/ | 278 | — | ~1.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 67 | Scaffold a first eval suite for an agent: mine real failures into cases, write behavioural checks over traces, and generate the runner. | undefined-ui/ | 1k | — | ~786 | Automated safety check: Pass | MIT | today |
| 68 | Write and review content for the OpenSEO website (web/) — blog posts, guides, feature pages, FAQs. | Jwuthri/ | 1.5k | — | ~922 | Automated safety check: Pass | MIT | today |
| 69 | 69.Add Eval Design, implement, validate, and calibrate a new eval for the convex-evals suite. | get-convex/ | 129 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | today |
| 70 | Analyze a job description (pasted text OR a URL) and find the AI/ML/GenAI topics it requires that this learn-ai course does NOT yet cover. | starkyru/ | 107 | — | ~1.9k | Automated safety check: Pass | MIT | 2 mo ago |
| 71 | Designs retrieval-augmented generation systems: document chunking, embeddings, vector store setup, hybrid search, reranking and retrieval evaluation, with checks at each step. | Jeffallan/ | 12k | 1 repo | ~2k | Automated safety check: Pass | MIT | 4 days ago |
| 72 | Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter. | microsoft/ | 1.4k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 73 | Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data. | ai-evals-course/ | 1.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 74 | Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks. | amd/ | 1.6k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 75 | Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/). | Arize-ai/ | 12k | — | ~7.1k | Automated safety check: Pass | Apache-2.0 | today |
| 76 | Run a quality benchmark of the /translate skill by selecting stratified test keys, capturing ground truth, translating, judging with sub-agents, and compiling a regression report. | shapeshift/ | 206 | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 77 | Scaffolds evaluation suites for the Axiom AI SDK: eval files, scorers, flag schemas and axiom.config.ts, generated from plain descriptions of an AI capability. | openclaw/ | 9.5k | — | ~4.1k | Automated safety check: Warn | MIT | today |
| 78 | 78.Gate Check Find the decisions in a pipeline that do not need the expensive model and propose the gate for each: a rule, a classic classifier, or a small model, with fail-closed routing. | undefined-ui/ | 1k | — | ~854 | Automated safety check: Pass | MIT | today |
| 79 | 79.Benchflow Run agent benchmarks, create tasks, analyze results, and manage agents using BenchFlow. | benchflow-ai/ | 353 | — | ~1.9k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 80 | Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. | Orchestra-Research/ | 13k | 3 repos | ~3.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 81 | Benchmark Claude Code skill performance with variance analysis, tracking pass rate, execution time, and token usage across iterations. | AgriciDaniel/ | 177 | — | ~1.4k | Automated safety check: Pass | MIT | 6 mo ago |
| 82 | Run a skill's evals and report results. An agent skill from redhat-cop/vault-config-operator. | redhat-cop/ | 167 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 83 | 83.Add Model Add a new model to the convex-evals coding leaderboard, and optionally the decision benchmark, through a PR, then dispatch its baseline runs. | get-convex/ | 129 | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | today |
| 84 | 84.Eval Run Evaluate any output file against a structured evals.yaml assertions file and produce a score report with per-assertion pass/fail results. | digipulse-engineering/ | 163 | — | ~1.7k | Automated safety check: Pass | Unknown | 8 days ago |
| 85 | Create new skills, improve existing skills, and measure skill performance. | SpectrAI-Initiative/ | 396 | — | ~8.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 86 | Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low. | microsoft/ | 1.4k | — | ~1.3k | Automated safety check: Pass | MIT | yesterday |
| 87 | Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format. | ai-evals-course/ | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 88 | Create, evaluate, improve, and benchmark content skills using the local Skill Lab workflow. | ceilf6/ | 120 | — | ~747 | Automated safety check: Pass | MIT | 8 days ago |
| 89 | 89.Inspect Inspect and debug live streaming agent sessions to understand what the agent did. | agentevals-dev/ | 162 | — | ~534 | Automated safety check: Pass | Apache-2.0 | today |
| 90 | 测试 use-persona 的角色扮演一致性。给定 persona + 10 个对话场景,生成回复并按 5 个维度评分,输出一致性报告。 | YIKUAIBANZI/ | 121 | — | ~531 | Automated safety check: Pass | MIT | 6 mo ago |
| 91 | Build custom LLM evaluation pipelines using the OpenJudge framework. | agentscope-ai/ | 868 | 1 repo | ~1.3k | Automated safety check: Pass | Apache-2.0 | 27 days ago |
| 92 | Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100… | open-thoughts/ | 301 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 93 | Checks training code, configs and math against documented framework behavior before an expensive run, citing a knowledge base or official docs for every claim. | Leeroo-AI/ | 195 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 94 | 為 Twinkle Eval 新增一個評測 benchmark(IFEval、BFCL、RAGAS、Text2SQL、Vision MCQ 之類)。涵蓋 CLAUDE.md §6 的完整強制流程:先建 Milestone 與 6 個 Issue、準備 example dataset、實作 Extractor + Scorer 並註冊 PRESETS、與參考框架做分數與速度對比、撰寫… | ai-twinkle/ | 117 | — | ~1.8k | Automated safety check: Pass | MIT | 22 days ago |
| 95 | Step-by-step guide for adding a new built-in test evaluator provider to Agent Kernel (beyond DeepEval, Opik and JEV). | yaalalabs/ | 191 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 96 | 96.Email Evals Author and safely run deterministic email-agent evaluation suites with dedicated e2a test agents. | tokencanopy/ | 192 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- LLM inference and serving372
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- GPU and accelerator computing176
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23