Evaluate any output file against a structured evals.yaml assertions file and produce a score report with per-assertion pass/fail results.

Custom licenceAuto-check passedAI & LLM Engineering

Install Eval Run

skills CLI
$ npx skills add digipulse-engineering/GAAI-framework --skill eval-run -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install digipulse-engineering/GAAI-framework eval-run --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/digipulse-engineering/GAAI-framework.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.gaai/core/skills/cross/eval-run .claude/skills/eval-run && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-run
GitHub stars
163
Token cost
~1.7k tokens
SKILL.md length
633 words
Files
2 (incl. references)
Skills in repo
55
Repo updated
First seen
Licence
Custom licence

At a glance

Evaluate any output file against a structured evals.yaml assertions file and produce a score report with per-assertion pass/fail results.

  • Works in 4 steps: Load inputs → Run code assertions → Run llm-judge assertions → …
  • Tasks that involve LLM evaluation
  • SKILL.md covers Purpose / When to Activate, Process, Quality Checks and Outputs, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Eval Run is an agent skill from digipulse-engineering/GAAI-framework. Evaluate any output file against a structured evals.yaml assertions file and produce a score report with per-assertion pass/fail results. Activate when the Discovery Agent runs the Skill Optimize protocol to measure output quality or detect regressions after skill instruction changes.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/evals-format.md`). Compatibility notes: Works with any filesystem-based AI coding agent

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Turns AI coding tools into reliable software delivery systems. Drop a .gaai/ folder into any project — Discovery defines what to build, Delivery executes autonomously until…

When your agent uses it

  • Tasks that involve LLM evaluation

Example prompts

  • “/eval-run”

Requirements

  • Compatibility (from SKILL.md): Works with any filesystem-based AI coding agent

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Load inputs
  2. Run code assertions
  3. Run llm-judge assertions
  4. Compile score report

What it can do on your machine

Read from SKILL.md and the folder at commit a26ea7a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown and yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Works with any filesystem-based AI coding agent

    From compatibility in the SKILL.md frontmatter.

Context cost

Eval Run loads about 1.7k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 74 tokens; SKILL.md has 633 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 633 words (~1,738 tokens).

name
eval-run
compatibility
Works with any filesystem-based AI coding agent
license
ELv2
metadata.author
gaai-framework
metadata.version
1.0
metadata.category
cross
metadata.track
cross-cutting
metadata.id
SKILL-CRS-025
metadata.updated_at
2026-03-15
metadata.status
experimental
outputs
score report (YAML or structured Markdown) with per-assertion pass/fail, total score, and failed assertion details

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file (references) in .gaai/core/skills/cross/eval-run of digipulse-engineering/GAAI-framework.

  • SKILL.md
  • references/evals-format.md

Open the folder on GitHubat commit a26ea7a

Compare with similar skills

Eval Run next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Run compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Run this skilldigipulse-engineering/GAAI-framework163—~1.7kAutomated safety check: PassCustom licence
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Azure AI Projects Python SDKmicrosoft/skills3.1k6 repos~2.8kAutomated safety check: PassMIT
Fine-Tuning ExpertJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from digipulse-engineering/GAAI-framework

All 55 skills in this repo
  • Pattern Transfer

    digipulse-engineering/GAAI-framework

    Discover structurally similar patterns across domains, assess transfer viability via structural invariant checking, and propose domain adaptations with risk gates.

    163 GitHub stars~2.3k tokensUpdated 8 days ago
    Auto-check passed
  • Skill Optimize

    digipulse-engineering/GAAI-framework

    Run a structured evaluate-analyze-improve cycle on any GAAI skill to measure quality, detect regressions, and propose targeted improvements.

    163 GitHub stars~2k tokensUpdated 8 days ago
    Auto-check passed
  • Abort Safe Handler

    digipulse-engineering/GAAI-framework

    Orchestrator-level Stage 4 entry gate for /gaai:bootstrap. An agent skill from digipulse-engineering/GAAI-framework.

    163 GitHub stars~2.6k tokensUpdated 8 days ago
    Auto-check passed
  • Create Skill

    digipulse-engineering/GAAI-framework

    Guide creation of a new GAAI skill following the agentskills.io spec and GAAI best practices.

    163 GitHub stars~1.7k tokensUpdated 8 days ago
    Auto-check passed
  • Discovery High Level Plan

    digipulse-engineering/GAAI-framework

    Transform vague or high-level human intent into a governed Discovery action plan.

    163 GitHub stars~570 tokensUpdated 8 days ago
    Auto-check passed
  • Memory Compact

    digipulse-engineering/GAAI-framework

    Emergency single-pass memory compression when context window pressure is high mid-task.

    163 GitHub stars~874 tokensUpdated 8 days ago
    Auto-check passed

Questions about Eval Run

What does Eval Run do?

Evaluate any output file against a structured evals.yaml assertions file and produce a score report with per-assertion pass/fail results. Eval Run is an agent skill from digipulse-engineering/GAAI-framework.yaml assertions file and produce a score report with per-assertion pass/fail results.

When should I use Eval Run?

Eval Run fits situations like: tasks that involve LLM evaluation.

How do I install Eval Run in Claude Code?

Run `npx skills add digipulse-engineering/GAAI-framework --skill eval-run -a claude-code`. Or copy the skill folder (.gaai/core/skills/cross/eval-run in digipulse-engineering/GAAI-framework) into .claude/skills/eval-run in your project. Claude Code loads it when a task matches its description.

How do I install Eval Run in Codex?

Run `npx skills add digipulse-engineering/GAAI-framework --skill eval-run -a codex`. Or copy the skill folder (.gaai/core/skills/cross/eval-run in digipulse-engineering/GAAI-framework) into .agents/skills/eval-run in your project. Codex loads it when a task matches its description.

Can I use Eval Run in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add digipulse-engineering/GAAI-framework --skill eval-run -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-run, .gemini/skills/eval-run, .github/skills/eval-run and .opencode/skills/eval-run in your project.

What does Eval Run need to run?

SKILL.md names no scripts, command-line tools or credentials: Eval Run is instructions for the agent only. Compatibility (from SKILL.md): Works with any filesystem-based AI coding agent.

Does Eval Run access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Run safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Run use?

Eval Run has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Eval Run use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Eval Run?

Skills that share tags, products or a category with Eval Run: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Azure AI Projects Python SDK (microsoft/skills, 3.1k stars), Fine-Tuning Expert (Jeffallan/claude-skills, 12k stars) and Looper (ksimback/looper, 710 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Run?

digipulse-engineering (a GitHub organization) maintains it in digipulse-engineering/GAAI-framework, which has 163 GitHub stars. The repository holds 55 skills in this directory. The repository was last updated on September 29, 2026.

Source: digipulse-engineering/GAAI-framework on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.