Repository
ai-evals-course/evals-skills agent skills
- skills
- 9
- GitHub stars
- 1.5k
GitHub description: “Skills that guide AI coding agents to help you build product-specific AI evals.”
- Stars
- 1,479 (101 forks)
- Licence
- Apache-2.0
- Last push
- Sep 2026
- Created
- Jun 2026
Install all skills
npx skills add ai-evals-course/evals-skillsAdd --skill <name> for a single skill and -a <agent> to choose the agent (see the agent guides).
Skills in ai-evals-course/evals-skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data. | ai-evals-course/ | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 2 | Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes. | ai-evals-course/ | 1.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 3 | Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately. | ai-evals-course/ | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 4 | Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries. | ai-evals-course/ | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 5 | Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data. | ai-evals-course/ | 1.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 6 | Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format. | ai-evals-course/ | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 7 | Write code evaluators for known failure modes with objective rules. | ai-evals-course/ | 1.5k | — | ~385 | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 8 | Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate. | ai-evals-course/ | 1.5k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 9 | Entry point for evals. An agent skill from ai-evals-course/evals-skills. | ai-evals-course/ | 1.5k | — | ~412 | Automated safety check: Pass | Apache-2.0 | 16 days ago |
Questions, answered from the data.
What is the best skill in ai-evals-course/evals-skills?
LLM Trace Review Interface from ai-evals-course/evals-skills ranks first of the 9 skills in ai-evals-course/evals-skills listed here, with the highest score: its repository has 1.5k GitHub stars, its SKILL.md loads about 1.4k tokens and it passes the automated safety check with no findings. Next come LLM Eval Pipeline Audit and Evaluate RAG.
Are the skills in ai-evals-course/evals-skills official?
None yet. All 9 skills in ai-evals-course/evals-skills listed here come from community repositories; a skill counts as official when the product's own GitHub organization publishes it.
How do I install all skills from ai-evals-course/evals-skills?
Run npx skills add ai-evals-course/evals-skills in your project: the open-source skills CLI installs the repository's skills into your coding agent's skills folder. To install a single skill, open its page here for the exact command.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.