Search

LLM evaluation

308 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1
1.Skill CreatorOfficial

Create new skills, modify and improve existing skills, and measure skill performance.

Azure/azqr79589 repos~8.2kAutomated safety check: PassApache-2.03 days ago
2

Reads the state and results of Caveman Cloud experiments and reports one recommendation or a block, without changing an experiment's lifecycle itself.

JuliusBrussee/caveman110k1 repo~975Automated safety check: PassApache-2.0today
3

Turns a Caveman report-only optimization observation into one minimal code change and a paired baseline evaluation, after the operator picks which to pursue.

JuliusBrussee/caveman110k1 repo~1.2kAutomated safety check: PassApache-2.0today
4

Provides context about the CoStrict evals system structure in this monorepo.

zgsm-ai/costrict4.4k1 repo~1.9kAutomated safety check: PassApache-2.08 days ago
5

Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

Orchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT3 mo ago
6

Create, improve, or evaluate an existing skill from workflows, prompts, SOPs, scripts.

yaojingang/yao-meta-skill2.7k—~768Automated safety check: PassMIT1 mo ago
7

Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

microsoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT2 days ago
8

Runs and debugs evaluations of how Chatbox models answer questions about large attached files, using synthetic and real long-document fixtures.

chatboxai/chatbox42k—~758Automated safety check: PassGPL-3.014 days ago
9

Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs.

FailproofAI/failproofai5.3k—~6kAutomated safety check: PassUnknownyesterday
10

A skill your agent uses when creating, improving, finding, or auditing agent skills - the user says 'create a skill', 'do I have a skill for X', 'improve the X skill', 'which skill should I use'…

tripleyak/SkillForge905—~2.3kAutomated safety check: NotesMIT2 mo ago
11

Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

Jeffallan/claude-skills12k1 repo~1.7kAutomated safety check: PassMIT5 days ago
12

A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain.

DenisSergeevitch/agents-best-practices2.4k—~7.4kAutomated safety check: PassMIT4 days ago
13

Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

ksimback/looper710—~2.7kAutomated safety check: NotesMIT2 mo ago
14

Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

huggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.07 days ago
15

Research, create, improve, migrate, evaluate, package, install-check, govern, and safely publish qiaomu-flavored agent skills from workflows, prompts, transcripts, docs, SOPs, runbooks, scripts, or…

joeseesun/qiaomu-meta-skill383—~2.8kAutomated safety check: PassMIT2 mo ago
16

Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

windmill-labs/windmill18k—~969Automated safety check: NotesUnknownyesterday
17

Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

langchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT3 days ago
18

Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

cloudnative-co/claude-code-starter-kit15210 repos~1.3kAutomated safety check: PassMITtoday
19

Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

GoogleCloudPlatform/vertex-ai-samples791—~2kAutomated safety check: PassApache-2.0yesterday
20

Turn a vague task into a testable definition of done and generate an executable goal-test script for it, optionally with a bounded retry loop around a headless agent.

undefined-ui/second-brain-os1k1 repo~754Automated safety check: PassMITyesterday
21

Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

bgauryy/octocode949—~2.1kAutomated safety check: PassMIT5 days ago
22

Measure whether a change makes Chat's styling replies better, with yarn eval:chat — it runs styling requests on recorded real pages for two versions of Stylebot and measures on the page whether each…

ankit/stylebot1.6k—~1.1kAutomated safety check: PassMITtoday
23

亚马逊卖家专用的 skill 创建器(中文)。当用户想把一个亚马逊运营/自媒体/日常工作流程变成可复用的 skill 时使用。触发场景包括但不限于:用户说"我想做一个 skill""把这个流程变成 skill""帮我写个自动化""优化我已有的 skill""给这个工作流做个自动化",即使用户没用"skill"这个词,只要在描述"以后每次都这样做"的重复性工作时也应触发。本 skill…

zach22-1999/amazon-skills2071 repo~3.9kAutomated safety check: PassApache-2.01 mo ago
24

Create new skills, modify and improve existing skills, and measure skill performance.

AgentTeam-TaichuAI/ScienceClaw670—~10kAutomated safety check: PassApache-2.05 mo ago
25
25.Evals Write SpecOfficial

Write LLM evaluation spec files with datasets, tasks, and evaluators using the @kbn/evals Playwright fixture.

elastic/kibana21k—~2.3kAutomated safety check: PassUnknowntoday
26

SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.

benchflow-ai/benchflow353—~4.5kAutomated safety check: PassApache-2.03 days ago
27

Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

ai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.014 days ago
28

Create a skill or bring an existing one up to the same standard (validate + asm eval fix loop); run evals, tune triggering.

luongnv89/asm953—~5.3kAutomated safety check: PassMIT2 days ago
29

Create, edit, evaluate, and package agent skills. An agent skill from smixs/skill-conductor.

smixs/skill-conductor179—~6.6kAutomated safety check: PassMIT2 mo ago
30

A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…

sun461941-hub/ai-project-copilot100—~3kAutomated safety check: PassMIT1 mo ago
31
31.Commerce EvalsOfficial

Authoring and running behavioral evals for a shopping or merchant agent, covering the case shape, authoring rules, code graders and judges, the run pattern, and poisoned fixtures.

anthropics/commerce-agents3.2k—~1.8kAutomated safety check: PassApache-2.06 days ago
32

Evaluate verified findings from merge-ready, Greptile, pull-request, CI, security, billing, and other code reviews, then promote durable review gaps into the version-controlled .greptile…

Jwuthri/Tracely-ai1.5k—~1.1kAutomated safety check: PassMITyesterday
33

Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes.

frontier-harness-eval/eval297—~8kAutomated safety check: PassNo licence1 mo ago
34

Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.

greyhaven-ai/autocontext1.3k—~892Automated safety check: PassApache-2.02 days ago
35

A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

NVIDIA/SkillEvaluator5481 repo~2.1kAutomated safety check: PassApache-2.0yesterday
36

Sends one prompt to Claude, GPT through the Codex CLI and Gemini, then tabulates response time, token use and cost, with an optional judged quality score.

garrytan/gstack136k—~4kAutomated safety check: NotesMITtoday
37

Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

ai-evals-course/evals-skills1.5k—~2.5kAutomated safety check: PassApache-2.014 days ago
38

Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports.

tech-leads-club/agent-skills7k—~3.9kAutomated safety check: PassCC-BY-4.018 days ago
39

A skill your agent uses when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI.

microsoft/work-iq1k—~2kAutomated safety check: NotesUnknownyesterday
40

Turn the current conversation's workflow into a reusable agent skill.

Undertone0809/rudder292—~3.6kAutomated safety check: PassApache-2.0today
41

Plan a documentation eval for supabase/evals, where a docs guide is the subject under test.

supabase/evals143—~2.9kAutomated safety check: PassApache-2.0yesterday
42

Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

ai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.014 days ago
43

Converts test suites from external eval frameworks into the Margin Eval suite format.

Margin-Lab/evals161—~3.7kAutomated safety check: PassAGPL-3.02 mo ago
44

Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

NVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~2.1kAutomated safety check: NotesApache-2.0today
45
45.Fde

Keeps the engagement record for client work. An agent skill from suboss87/FDEOps.

suboss87/FDEOps954—~2.9kAutomated safety check: PassMITyesterday
46
46.Review SkillOfficial

Review a proposed Agent Skill for structural validity and content quality before publishing.

mongodb/agent-skills190—~1.5kAutomated safety check: NotesApache-2.0yesterday
47

Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes.

benchflow-ai/benchflow353—~4kAutomated safety check: PassApache-2.03 days ago
48

Sets up and runs side-by-side model evaluations with frozen questions, isolated tools and traceable results, using your own question bank or benchmark subsets.

liucongg/liucong-skills248—~752Automated safety check: PassApache-2.01 mo ago