Search

Education · LLM evaluation

19 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric.

benchflow-ai/benchflow356—~4.5kAutomated safety check: PassApache-2.05 days ago
2

A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

NVIDIA/SkillEvaluator554—~2.1kAutomated safety check: PassApache-2.0yesterday
3

This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…

guanyang/open-agent-hub9772 repos~4.2kAutomated safety check: PassMITtoday
4

Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository.

dotnet/skills5.6k1 repo~6.1kAutomated safety check: PassMITyesterday
5
5.Agentic EvalOfficial

Patterns and techniques for evaluating and improving AI agent outputs.

github/awesome-copilot40k3 repos~1.5kAutomated safety check: PassMIT2 days ago
6

This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target.

jacob-dietle/context-os111—~5.2kAutomated safety check: PassMIT1 mo ago
7

Turn student course evaluations (free-text + numeric) into an actionable teaching-improvement plan — the teaching analogue of /respond-to-referees.

pedrohcgs/claude-code-my-workflow1.7k—~2.6kAutomated safety check: NotesMIT13 days ago
8

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise…

aiskillstore/marketplace4333 repos~4.2kAutomated safety check: PassNo licenceyesterday
9

Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring.

woocommerce/woocommerce-ios358—~7.4kAutomated safety check: NotesGPL-2.0yesterday
10

A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.

Aperivue/medsci-skills333—~2.4kAutomated safety check: PassMIT5 days ago
11

Configures and runs LLM evaluation using Promptfoo framework.

daymade/claude-code-skills1.4k—~3kAutomated safety check: PassMITyesterday
12

Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

ClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMITyesterday
13

This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".

borghei/Claude-Skills891—~1.9kAutomated safety check: PassMIT4 days ago
14

Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy.

agentsope/SkillAlchemy436—~6.4kAutomated safety check: PassMIT2 days ago
15

Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates.

JasonColapietro/suede-creator-skills127—~3.3kAutomated safety check: PassMITyesterday
16

Design an evaluation plan for an LLM or AI feature before shipping it.

mohitagw15856/pm-claude-skills1.4k—~996Automated safety check: PassMITyesterday
17

Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output.

mohitagw15856/pm-claude-skills1.4k—~1.1kAutomated safety check: PassMITyesterday
18

Creates assessments with varied question types (MCQ, code-completion, debugging, projects) aligned to learning objectives with meaningful distractors based on common misconceptions.

aiskillstore/marketplace433—~4.3kAutomated safety check: PassNo licenceyesterday
19

A skill your agent uses when authoring a NEW rsc skill or editing an existing one — scoping it to one job, writing the description that decides whether it ever loads, splitting the body into…

ericrisco/rsc-harness180—~4.3kAutomated safety check: PassMITyesterday