Search
Education · LLM evaluation
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | SkillsBench task authoring — walk a contributor from idea to submission-ready task following CONTRIBUTING.md and the task-implementation rubric. | benchflow-ai/ | 356 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 2 | A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation. | NVIDIA/ | 554 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 3 | This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated… | guanyang/ | 977 | 2 repos | ~4.2k | Automated safety check: Pass | MIT | today |
| 4 | Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository. | dotnet/ | 5.6k | 1 repo | ~6.1k | Automated safety check: Pass | MIT | yesterday |
| 5 | Patterns and techniques for evaluating and improving AI agent outputs. | github/ | 40k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | 2 days ago |
| 6 | This skill should be used when a specific quality problem (UX, data, architecture, feature) needs systematic diagnosis and iterative fixing toward a defined target. | jacob-dietle/ | 111 | — | ~5.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 7 | Turn student course evaluations (free-text + numeric) into an actionable teaching-improvement plan — the teaching analogue of /respond-to-referees. | pedrohcgs/ | 1.7k | — | ~2.6k | Automated safety check: Notes | MIT | 13 days ago |
| 8 | This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise… | aiskillstore/ | 433 | 3 repos | ~4.2k | Automated safety check: Pass | No licence | yesterday |
| 9 | Evaluate WooAIAssistant against a structured scenario suite with hard invariants + LLM-as-judge rubric scoring. | woocommerce/ | 358 | — | ~7.4k | Automated safety check: Notes | GPL-2.0 | yesterday |
| 10 | A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection. | Aperivue/ | 333 | — | ~2.4k | Automated safety check: Pass | MIT | 5 days ago |
| 11 | Configures and runs LLM evaluation using Promptfoo framework. | daymade/ | 1.4k | — | ~3k | Automated safety check: Pass | MIT | yesterday |
| 12 | Eval-driven skill tuning. An agent skill from ClawBio/ClawBio. | ClawBio/ | 1.2k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 13 | This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality". | borghei/ | 891 | — | ~1.9k | Automated safety check: Pass | MIT | 4 days ago |
| 14 | Decomposed, multi-criteria metric design for LLM pipelines. An agent skill from agentsope/SkillAlchemy. | agentsope/ | 436 | — | ~6.4k | Automated safety check: Pass | MIT | 2 days ago |
| 15 | Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates. | JasonColapietro/ | 127 | — | ~3.3k | Automated safety check: Pass | MIT | yesterday |
| 16 | 16.AI Eval Plan Design an evaluation plan for an LLM or AI feature before shipping it. | mohitagw15856/ | 1.4k | — | ~996 | Automated safety check: Pass | MIT | yesterday |
| 17 | Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output. | mohitagw15856/ | 1.4k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 18 | Creates assessments with varied question types (MCQ, code-completion, debugging, projects) aligned to learning objectives with meaningful distractors based on common misconceptions. | aiskillstore/ | 433 | — | ~4.3k | Automated safety check: Pass | No licence | yesterday |
| 19 | 19.Author Skill A skill your agent uses when authoring a NEW rsc skill or editing an existing one — scoping it to one job, writing the description that decides whether it ever loads, splitting the body into… | ericrisco/ | 180 | — | ~4.3k | Automated safety check: Pass | MIT | yesterday |