Search
LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Create new skills, modify and improve existing skills, and measure skill performance. | Azure/ | 795 | 89 repos | ~8.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 2 | Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself. | JuliusBrussee/ | 110k | 1 repo | ~975 | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Turns a Caveman report-only optimization observation into one minimal code change and a paired baseline evaluation, after the operator picks which to pursue. | JuliusBrussee/ | 110k | 1 repo | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | Provides context about the CoStrict evals system structure in this monorepo. | zgsm-ai/ | 4.4k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 5 | Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints. | Orchestra-Research/ | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Create, improve, or evaluate an existing skill from workflows, prompts, SOPs, scripts. | yaojingang/ | 2.7k | — | ~768 | Automated safety check: Pass | MIT | 1 mo ago |
| 7 | Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes. | microsoft/ | 3.1k | 6 repos | ~2.8k | Automated safety check: Pass | MIT | 2 days ago |
| 8 | Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures. | chatboxai/ | 42k | — | ~758 | Automated safety check: Pass | GPL-3.0 | 14 days ago |
| 9 | Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs. | FailproofAI/ | 5.3k | — | ~6k | Automated safety check: Pass | Unknown | yesterday |
| 10 | 10.Skillforge A skill your agent uses when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use'… | tripleyak/ | 905 | — | ~2.3k | Automated safety check: Notes | MIT | 2 mo ago |
| 11 | Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment. | Jeffallan/ | 12k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | 5 days ago |
| 12 | A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain. | DenisSergeevitch/ | 2.4k | — | ~7.4k | Automated safety check: Pass | MIT | 4 days ago |
| 13 | 13.Looper Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council. | ksimback/ | 710 | — | ~2.7k | Automated safety check: Notes | MIT | 2 mo ago |
| 14 | Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends. | huggingface/ | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 15 | Research, create, improve, migrate, evaluate, package, install-check, govern, and safely publish qiaomu-flavored agent skills from workflows, prompts, transcripts, docs, SOPs, runbooks, scripts, or… | joeseesun/ | 383 | — | ~2.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 16 | Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons. | windmill-labs/ | 18k | — | ~969 | Automated safety check: Notes | Unknown | yesterday |
| 17 | Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier. | langchain-ai/ | 1.3k | — | ~4k | Automated safety check: Pass | MIT | 3 days ago |
| 18 | 18.Eval Harness Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles. | cloudnative-co/ | 152 | 10 repos | ~1.3k | Automated safety check: Pass | MIT | today |
| 19 | Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK. | GoogleCloudPlatform/ | 791 | — | ~2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 20 | 20.Goal Test Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent. | undefined-ui/ | 1k | 1 repo | ~754 | Automated safety check: Pass | MIT | yesterday |
| 21 | Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report. | bgauryy/ | 949 | — | ~2.1k | Automated safety check: Pass | MIT | 5 days ago |
| 22 | 22.Eval Chat Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each… | ankit/ | 1.6k | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 23 | 亚马逊卖家专用的 skill 创建器(中文)。当用户想把一个亚马逊运营/自媒体/日常工作流程变成可复用的 skill 时使用。触发场景包括但不限于:用户说"我想做一个 skill""把这个流程变成 skill""帮我写个自动化""优化我已有的 skill""给这个工作流做个自动化",即使用户没用"skill"这个词,只要在描述"以后每次都这样做"的重复性工作时也应触发。本 skill… | zach22-1999/ | 207 | 1 repo | ~3.9k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 24 | Create new skills, modify and improve existing skills, and measure skill performance. | AgentTeam-TaichuAI/ | 670 | — | ~10k | Automated safety check: Pass | Apache-2.0 | 5 mo ago |
| 25 | Write LLM evaluation spec files with datasets, tasks, and evaluators using the @kbn/evals Playwright fixture. | elastic/ | 21k | — | ~2.3k | Automated safety check: Pass | Unknown | today |
| 26 | 26.Task Creator SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric. | benchflow-ai/ | 353 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 27 | Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data. | ai-evals-course/ | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 28 | Create a skill or bring an existing one up to the same standard (validate + asm eval fix loop); run evals, tune triggering. | luongnv89/ | 953 | — | ~5.3k | Automated safety check: Pass | MIT | 2 days ago |
| 29 | Create, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor. | smixs/ | 179 | — | ~6.6k | Automated safety check: Pass | MIT | 2 mo ago |
| 30 | A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient… | sun461941-hub/ | 100 | — | ~3k | Automated safety check: Pass | MIT | 1 mo ago |
| 31 | Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures. | anthropics/ | 3.2k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 32 | Evaluate verified findings from merge-ready, Greptile, pull-request, CI, security, billing, and other code reviews, then promote durable review gaps into the version-controlled .greptile… | Jwuthri/ | 1.5k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 33 | Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes. | frontier-harness-eval/ | 297 | — | ~8k | Automated safety check: Pass | No licence | 1 mo ago |
| 34 | 34.Autocontext Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met. | greyhaven-ai/ | 1.3k | — | ~892 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 35 | A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. | NVIDIA/ | 548 | 1 repo | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 36 | Sends one prompt to Claude, GPT through the Codex CLI and Gemini, then tabulates response time, token use and cost, with an optional judged quality score. | garrytan/ | 136k | — | ~4k | Automated safety check: Notes | MIT | today |
| 37 | Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes. | ai-evals-course/ | 1.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 38 | 38.Harness Eval Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports. | tech-leads-club/ | 7k | — | ~3.9k | Automated safety check: Pass | CC-BY-4.0 | 18 days ago |
| 39 | A skill your agent uses when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. | microsoft/ | 1k | — | ~2k | Automated safety check: Notes | Unknown | yesterday |
| 40 | Turn the current conversation's workflow into a reusable agent skill. | Undertone0809/ | 292 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | today |
| 41 | Plan a documentation eval for supabase/evals, where a docs guide is the subject under test. | supabase/ | 143 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 42 | 42.Evaluate RAG Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately. | ai-evals-course/ | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 43 | Converts test suites from external eval frameworks into the Margin Eval suite format. | Margin-Lab/ | 161 | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | 2 mo ago |
| 44 | Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and… | NVIDIA-AI-Blueprints/ | 1.9k | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | today |
| 45 | 45.Fde Keeps the engagement record for client work. An agent skill from suboss87/FDEOps. | suboss87/ | 954 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 46 | Review a proposed Agent Skill for structural validity and content quality before publishing. | mongodb/ | 190 | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 47 | Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes. | benchflow-ai/ | 353 | — | ~4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 48 | Sets up and runs side-by-side model evaluations with frozen questions, isolated tools and traceable results, using your own question bank or benchmark subsets. | liucongg/ | 248 | — | ~752 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |