Evaluate a named AI Analyst configuration across a frozen suite.

MITAuto-check passed

Install Eval

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/eval .claude/skills/eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval
GitHub stars
304
Token cost
~2.4k tokens
SKILL.md length
1,134 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Evaluate a named AI Analyst configuration across a frozen suite.

  • Works in 8 steps: For a complete-analysis case, load the… → Select the exposure, purpose, named… → Use… → …
  • The user asks to run an eval suite
  • SKILL.md covers Before running, Run, Run a full-analysis… and Run the focused SQL…, plus 3 more sections
  • Calls python3 and python

What it does

Eval is an agent skill from ai-analyst-lab/ai-analyst. Evaluate a named AI Analyst configuration across a frozen suite. Use when the user asks to run an eval suite, compare a change, inspect system accuracy, or run working or heldout capability and regression cases.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • The user asks to run an eval suite
  • Compare a change
  • Inspect system accuracy
  • Heldout capability and regression cases

Example prompts

  • “/eval”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. For a complete-analysis case, load the question-only case from evals/cases/public/. Focused component and calculation suites live under…
  2. Select the exposure, purpose, named cases, and any slice before the run starts. Exposure is working or heldout. Purpose is capability or…
  3. Use helpers.evals.controller.EvaluationController to launch and record the trials.
  4. Give each trial only its public task, permitted system files, permitted data, and permitted tools.
  5. Lock every trial output before grading begins.
  6. Grade deterministic criteria first. Keep model-based grades separate.
  7. Preserve pass, fail, blocked, error, invalid, and unknown as different results.
  8. Report every case and slice before discussing the aggregate.

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval loads about 2.4k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 1,134 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 1,134 words, ~2,411 tokens.

Download SKILL.mdSave it as .claude/skills/eval/SKILL.md (or your agent's skills folder).
name
eval
description
Evaluate a named AI Analyst configuration across a frozen suite. Use when the user asks to run an eval suite, compare a change, inspect system accuracy, or run working or heldout capability and regression cases.

Evaluate the system

Before running

Name the exact system under test. Record its model, instructions, skills, agents, helpers, knowledge, workflow, tools, connector configuration, and data snapshot.

Use one of these modes honestly:

  • Working mode supports iteration. Its references may be visible to the evaluator, but never to the child trial before its output is locked.
  • Course heldout mode sends locked outputs to the course-controlled grader. The expected results do not live in the student clone.
  • A local visible answer file is development material. Do not call it a secret heldout evaluation.

Run

  1. For a complete-analysis case, load the question-only case from evals/cases/public/. Focused component and calculation suites live under evals/focused/public/.
  2. Select the exposure, purpose, named cases, and any slice before the run starts. Exposure is working or heldout. Purpose is capability or regression. Do not treat these as one dimension.
  3. Use helpers.evals.controller.EvaluationController to launch and record the trials.
  4. Give each trial only its public task, permitted system files, permitted data, and permitted tools.
  5. Lock every trial output before grading begins.
  6. Grade deterministic criteria first. Keep model-based grades separate.
  7. Preserve pass, fail, blocked, error, invalid, and unknown as different results.
  8. Report every case and slice before discussing the aggregate.

The local controller is available through python3 -m helpers.evals.cli run-suite. Use --model claude-opus-4-6, --exposure, optional --purpose, and repeated --case-id arguments when selecting a subset. General code access is not required for routing or contract cases. When local data analysis requires --allow-code, state that local process isolation is not the same as course-heldout answer isolation.

For a reviewed working suite with local references, lock the trial outputs first, then grade them with python3 -m helpers.evals.cli grade-suite. Pass the run ID, public manifest, and reviewed reference file. Never copy the reference file into the trial workspace.

Run a full-analysis development case

Full-analysis cases live under evals/cases/public/. They produce a complete analysis bundle rather than one scalar answer.

  1. Start a new run. Pass the public case directory you intend to run. For example:

    bash
    python3 -m helpers.evals.full_analysis start \
      --case evals/cases/public/novamart-monthly-operating-review-001/v1
  2. Read the generated RUN-INSTRUCTIONS.md, public case, and result schema. The start command creates the trial's draft/ directory but does not begin the analysis trace. The public case can define a scalar, a table, or several rows and columns.

  3. Before querying data, start exactly one analysis trace with the trial draft as its output directory. Keep that analysis ID for every query, finding, receipt, and trace artifact in the trial.

  4. Perform the analysis through the existing AI Analyst system. Save exactly the files requested by that case in its draft directory. Do not inspect any course reference repository before the first run is locked.

  5. Register the reported findings, build the trace, and confirm the trace HTML path. Do not start a second analysis trace.

  6. Lock the run:

    bash
    python3 -m helpers.evals.full_analysis lock --run-id <run-id>

    Locking verifies the Snowflake snapshot, validates the case-specific output contract, requires a complete trace, copies the output bundle, hashes every artifact, and snapshots the analysis record, query log, action log, provenance, receipt, and trace HTML. Never modify the locked submission or trace.

  7. After the course evaluation repository is released, run its grader against the locked run. The grader executes the submitted read-only SQL and the reviewed reference SQL against the same Snowflake snapshot, then compares their returned values. Different SQL can pass when it returns the same reviewed result. Read the resulting student-report.md and individual grade records.

  8. Diagnose failures using the saved trace. Start any changed system as a new development run:

    bash
    python3 -m helpers.evals.full_analysis start \
      --case evals/cases/public/novamart-monthly-operating-review-001/v1 \
      --baseline-run-id <baseline-run-id> \
      --intended-change "<one concrete system change>"
  9. After grading the candidate, compare the two immutable runs:

    bash
    python3 -m helpers.evals.full_analysis compare \
      --before-run-id <baseline-run-id> \
      --after-run-id <candidate-run-id>

The shared course answers make these development cases. A genuinely held-out set must remain outside the system and be graded through a course-controlled boundary. Trace checks diagnose why a result passed or failed. They do not override a failed output grade.

Show full SKILL.md (555 more words)Show less

Run the focused SQL development set

Focused SQL suites test calculations without requiring a complete brief, chart, and trace for every case. They are optional component tests. Never label their result system accuracy or combine them with complete-analysis accuracy.

  1. Run the public tasks in isolated workspaces. The controller copies the analytical system and each public task, but not the private references:

    bash
    python3 -m helpers.evals.cli run-suite \
      --manifest evals/focused/public/session6-sql-development.yaml \
      --project-root . \
      --runs-root working/evals/runs \
      --exposure working \
      --trials 1 \
      --parallelism 4 \
      --model claude-opus-4-6 \
      --timeout 600 \
      --allow-code
  2. Record the run ID printed by the command. Confirm all eight trial records are locked before introducing the references.

  3. After the sibling ai-analyst-course-evals repository is available, grade the locked run:

    bash
    python3 -m helpers.evals.cli grade-suite \
      --run-id <run-id> \
      --manifest evals/focused/public/session6-sql-development.yaml \
      --references ../ai-analyst-course-evals/focused-cases/session6-sql-development.yaml \
      --project-root . \
      --runs-root working/evals/runs
  4. Open working/evals/runs/<run-id>/report.html, then inspect the case results and slice summary in manifest.json.

These focused cases grade one reported calculation each. They do not establish that the system can produce an acceptable complete analysis. Report the focused SQL suite and complete-analysis suite as two different measurements.

Run the Session 6 complete-analysis set

The complete set runs one isolated, traceable analysis for each case. Every child process receives the analytical system and public case, but no reviewed answer or grader repository.

  1. Start the set in the foreground:

    bash
    python3 -m helpers.evals.full_suite \
      --suite evals/suites/session-6-complete-analysis.yaml \
      --model claude-opus-4-6 \
      --parallelism 4
  2. The command records requested and actual parallelism, creates a separate temporary workspace for every case, runs each analysis, locks its output and trace, and copies the immutable run into working/evals/runs/.

  3. Open the suite manifest under working/evals/suites/<suite-run-id>/manifest.json. Keep every locked, blocked, invalid, and errored case visible.

  4. Only after the analytical runs are locked, use the sibling course evaluator:

    bash
    ../ai-analyst-course-evals/.venv/bin/python -m course_evals grade-suite \
      --manifest working/evals/suites/<suite-run-id>/manifest.json \
      --parallelism 4
  5. Open working/evals/suites/<suite-run-id>/grades/suite-report.html. Read every case before the aggregate. Then inspect overall case accuracy, gate-level accuracy, and slices by domain, task type, complexity, data shape, and primary analytical risk.

An execution error remains in the denominator. Do not silently rerun or remove a failed case to improve the score. If a transient problem justifies another attempt, preserve the first run and record the reason for the new suite run.

Design before running

When a user is creating a new case, interview them for the intended user, decision, consequence if wrong, observable criteria, independent reference plan, grader per criterion, human-review boundary, task and risk slices, and lifecycle status. Preserve the user's decisions rather than silently choosing for them.

Keep a new case proposed until its reference and graders receive independent review. Validate its structure with python3 -m helpers.evals.cli validate-case. Structural validation does not verify the reference or promote the case.

When a user is assembling a proposed set from a candidate pool, require an explicit selection, at least one rejected candidate with a reason, and named missing coverage. Validate it with python3 -m helpers.evals.cli validate-suite. Do not replace the user's proposed set with a canonical set during comparison.

Compare a change

Hold the suite, data snapshot, model, evaluator, tools, and trial count fixed. Name one intended system change. If more than one material input changed, label the comparison confounded rather than attributing the score movement.

Use --intended-change context for a candidate context run. Do not expose expected values, reference queries, private grader prompts, or a heldout answer key in the report.

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/eval of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval this skillai-analyst-lab/ai-analyst304—~2.4kAutomated safety check: PassMIT
Evals Create Suiteelastic/kibana21k—~1.7kAutomated safety check: PassCustom licence
Eval-Driven Development Harnessaffaan-m/ECC275k—~1.5kAutomated safety check: PassMIT
Evalalirezarezvani/claude-skills28k1 repos~618Automated safety check: PassMIT
Eval Harnessaffaan-m/ECC275k—~2.2kAutomated safety check: PassMIT
OmniRoute CLI Evalsdiegosouzapw/OmniRoute74k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Evals Create Suite

    elastic/kibana

    Official

    Scaffold a new LLM evaluation suite package with Playwright config, evaluate fixture, and package files.

    21k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

    275k GitHub stars~1.5k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Eval

    alirezarezvani/claude-skills

    Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

    28k GitHub starsUsed in 1 repo~618 tokens
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

    275k GitHub stars~2.2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • OmniRoute CLI Evals

    diegosouzapw/OmniRoute

    Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

    74k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi

    275k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 7 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed

Questions about Eval

What does Eval do?

Evaluate a named AI Analyst configuration across a frozen suite. Eval is an agent skill from ai-analyst-lab/ai-analyst. Evaluate a named AI Analyst configuration across a frozen suite.

When should I use Eval?

Eval fits situations like: the user asks to run an eval suite; compare a change; inspect system accuracy; heldout capability and regression cases.

How do I install Eval in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill eval -a claude-code`. Or copy the skill folder (.claude/skills/eval in ai-analyst-lab/ai-analyst) into .claude/skills/eval in your project. Claude Code loads it when a task matches its description.

How do I install Eval in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill eval -a codex`. Or copy the skill folder (.claude/skills/eval in ai-analyst-lab/ai-analyst) into .agents/skills/eval in your project. Codex loads it when a task matches its description.

Can I use Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval, .gemini/skills/eval, .github/skills/eval and .opencode/skills/eval in your project.

What does Eval need to run?

Going by SKILL.md and its folder, Eval needs the command-line tools its instructions call (python3 and python). Our summary lists: Python 3.

Does Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval use?

Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Eval use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval?

Skills that share tags, products or a category with Eval: Evals Create Suite (elastic/kibana, 21k stars), Eval-Driven Development Harness (affaan-m/ECC, 275k stars), Eval (alirezarezvani/claude-skills, 28k stars) and Eval Harness (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.