GitHub organization

Agent skills by ai-evals-course

Every agent skill ai-evals-course publishes on GitHub, ranked by score, with the repositories they come from.
skills
9
repository
1

Repositories by ai-evals-course

Skills by ai-evals-course, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Skills by ai-evals-course, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

ai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.012 days ago
2

Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

ai-evals-course/evals-skills1.5k—~2.5kAutomated safety check: PassApache-2.012 days ago
3

Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

ai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.012 days ago
4

Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

ai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.012 days ago
5

Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

ai-evals-course/evals-skills1.5k—~2.2kAutomated safety check: PassApache-2.012 days ago
6

Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.

ai-evals-course/evals-skills1.5k—~1.9kAutomated safety check: PassApache-2.012 days ago
7

Write code evaluators for known failure modes with objective rules.

ai-evals-course/evals-skills1.5k—~385Automated safety check: PassApache-2.012 days ago
8

Guides an interactive error analysis of LLM outputs: studies the dataset, builds a review interface, picks diverse samples and organizes the failure modes you annotate.

ai-evals-course/evals-skills1.5k—~3.7kAutomated safety check: PassApache-2.012 days ago
9

Entry point for evals. An agent skill from ai-evals-course/evals-skills.

ai-evals-course/evals-skills1.5k—~412Automated safety check: PassApache-2.012 days ago

Questions, answered from the data.

What is the best skill by ai-evals-course?

LLM Trace Review Interface from ai-evals-course/evals-skills ranks first of the 9 skills by ai-evals-course listed here, with the highest score: its repository has 1.5k GitHub stars, its SKILL.md loads about 1.4k tokens and it passes the automated safety check with no findings. Next come LLM Eval Pipeline Audit and Evaluate RAG.

Are ai-evals-course's skills official?

None yet. All 9 skills by ai-evals-course listed here come from community repositories; a skill counts as official when the product's own GitHub organization publishes it.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.