Search
By ai-evals-course
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data. | ai-evals-course/ | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 2 | Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes. | ai-evals-course/ | 1.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 3 | Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately. | ai-evals-course/ | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 4 | Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries. | ai-evals-course/ | 1.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 5 | Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data. | ai-evals-course/ | 1.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 6 | Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format. | ai-evals-course/ | 1.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 7 | Write code evaluators for known failure modes with objective rules. | ai-evals-course/ | 1.5k | — | ~385 | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 8 | Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate. | ai-evals-course/ | 1.5k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 16 days ago |
| 9 | Entry point for evals. An agent skill from ai-evals-course/evals-skills. | ai-evals-course/ | 1.5k | — | ~412 | Automated safety check: Pass | Apache-2.0 | 16 days ago |