Agent skill

Evals Start

by ai-evals-course in ai-evals-course/evals-skills

Entry point for evals. An agent skill from ai-evals-course/evals-skills.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Evals Start

skills CLI
$ npx skills add ai-evals-course/evals-skills --skill evals-start -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-evals-course/evals-skills evals-start --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-evals-course/evals-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evals-start .claude/skills/evals-start && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evals-start
GitHub stars
1.5k
Token cost
~412 tokens
SKILL.md length
200 words
Files
2
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Entry point for evals. An agent skill from ai-evals-course/evals-skills.

  • The user asks for help with evals
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Does not know where to begin
  • Asks for something no other skill in this plugin matches

What it does

Evals Start is an agent skill from ai-evals-course/evals-skills. Entry point for evals. Use when the user asks for help with evals, does not know where to begin, or asks for something no other skill in this plugin matches. Do NOT use when a more specific skill in this plugin already matches; load that skill directly.

Its SKILL.md is about 410 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Skills that guide AI coding agents to help you build product-specific AI evals. The licence is Apache-2.0.

When your agent uses it

  • The user asks for help with evals
  • Does not know where to begin
  • Asks for something no other skill in this plugin matches
  • A more specific skill in this plugin already matches

Example prompts

  • “/evals-start”

What it can do on your machine

Read from SKILL.md and the folder at commit 80d5f7b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evals Start loads about 412 tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 200 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~412

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-evals-course/evals-skills at commit 80d5f7b, republished under its Apache-2.0 licence (© ai-evals-course). 200 words, ~412 tokens.

Download SKILL.mdSave it as .claude/skills/evals-start/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
evals-start
description
Entry point for evals. Use when the user asks for help with evals, does not know where to begin, or asks for something no other skill in this plugin matches. Do NOT use when a more specific skill in this plugin already matches; load that skill directly.

Evals Start

This plugin splits eval work into targeted skills. Your job here is small: find the row below that matches the user's situation, tell the user which skill you are loading and why, then load that skill and follow its workflow from start to finish instead of improvising your own version of it.

SituationSkill to load
Has traces, wants to find failure modes, no established taxonomy yeterror-discovery
Has an existing eval pipeline and wants to know if it can be trustedeval-audit
Has a known failure mode that code can check (e.g., format, schema, regex, execution)write-code-eval
Has a known failure mode and wants an LLM judge for itwrite-judge-prompt
Has an LLM judge or evaluator and wants to check its qualityvalidate-evaluator
Has no traces to review yetgenerate-synthetic-data, then error-discovery
Wants a custom annotation interface for some other labeling taskbuild-review-interface
Wants to evaluate a RAG pipelineevaluate-rag

Most requests that mention error analysis with traces in hand mean error-discovery. New users with an existing pipeline usually need eval-audit first. This file holds only routing. When in doubt about which row fits, ask the user instead of guessing. The workflow lives in the targeted skill.

© ai-evals-course, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/evals-start of ai-evals-course/evals-skills.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 80d5f7b

Compare with similar skills

Evals Start next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evals Start compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evals Start this skillai-evals-course/evals-skills1.5k—~412Automated safety check: PassApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from ai-evals-course/evals-skills

All 9 skills in this repo
  • LLM Trace Review Interface

    ai-evals-course/evals-skills

    Builds a browser-based annotation page for reviewing LLM traces one at a time with pass/fail labels, notes and saved results, tailored to your data.

    1.5k GitHub stars~1.4k tokensUpdated 15 days ago
    Auto-check passed
  • LLM Eval Pipeline Audit

    ai-evals-course/evals-skills

    Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

    1.5k GitHub stars~2.5k tokensUpdated 15 days ago
    Auto-check passed
  • Evaluate RAG

    ai-evals-course/evals-skills

    Guides evaluation of a RAG system by diagnosing failures in traces, building a retrieval test set and scoring retrieval and generation separately.

    1.5k GitHub stars~1.9k tokensUpdated 15 days ago
    Auto-check passed
  • Synthetic Eval Data Generator

    ai-evals-course/evals-skills

    Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

    1.5k GitHub stars~1.4k tokensUpdated 15 days ago
    Auto-check passed
  • LLM Judge Validation

    ai-evals-course/evals-skills

    Checks an LLM judge against human labels using train, dev and test splits, TPR and TNR, and a bias correction applied to production data.

    1.5k GitHub stars~2.2k tokensUpdated 15 days ago
    Auto-check passed
  • LLM-as-Judge Prompt Writer

    ai-evals-course/evals-skills

    Designs a binary Pass/Fail LLM-as-Judge prompt for one subjective failure mode, built from a task statement, clear definitions, labeled examples and a structured output format.

    1.5k GitHub stars~1.9k tokensUpdated 15 days ago
    Auto-check passed

Questions about Evals Start

What does Evals Start do?

Entry point for evals. An agent skill from ai-evals-course/evals-skills. Evals Start is an agent skill from ai-evals-course/evals-skills. Entry point for evals.

When should I use Evals Start?

Evals Start fits situations like: the user asks for help with evals; does not know where to begin; asks for something no other skill in this plugin matches; A more specific skill in this plugin already matches.

How do I install Evals Start in Claude Code?

Run `npx skills add ai-evals-course/evals-skills --skill evals-start -a claude-code`. Or copy the skill folder (skills/evals-start in ai-evals-course/evals-skills) into .claude/skills/evals-start in your project. Claude Code loads it when a task matches its description.

How do I install Evals Start in Codex?

Run `npx skills add ai-evals-course/evals-skills --skill evals-start -a codex`. Or copy the skill folder (skills/evals-start in ai-evals-course/evals-skills) into .agents/skills/evals-start in your project. Codex loads it when a task matches its description.

Can I use Evals Start in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-evals-course/evals-skills --skill evals-start -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evals-start, .gemini/skills/evals-start, .github/skills/evals-start and .opencode/skills/evals-start in your project.

What does Evals Start need to run?

SKILL.md names no scripts, command-line tools or credentials: Evals Start is instructions for the agent only.

Does Evals Start access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evals Start safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evals Start use?

Evals Start is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evals Start use?

About 412 tokens (SKILL.md is roughly 1.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evals Start?

Skills that share tags, products or a category with Evals Start: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evals Start?

ai-evals-course (a GitHub organization) maintains it in ai-evals-course/evals-skills, which has 1,472 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on September 24, 2026.

Source: ai-evals-course/evals-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.