Agent skill

Eval Creator

by pskoett in pskoett/pskoett-ai-skills

[Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them.

No licenceAuto-check passedAgent Workflows

Install Eval Creator

skills CLI
$ npx skills add pskoett/pskoett-ai-skills --skill eval-creator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pskoett/pskoett-ai-skills eval-creator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pskoett/pskoett-ai-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-creator .claude/skills/eval-creator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-creator
GitHub stars
311
Token cost
~2.6k tokens
SKILL.md length
887 words
Files
1
Skills in repo
24
Repo updated
First seen
Licence
None found

At a glance

[Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them.

  • Works in 6 steps: Check precondition — if not met, mark as… → Execute verification method → Compare result to expected → …
  • A learning is promoted and has a clear pass/fail condition
  • SKILL.md covers When to Use, Eval Directory Structure, Creating an Eval Case and Running Evals, plus 6 more sections
  • Calls python

What it does

Eval Creator is an agent skill from pskoett/pskoett-ai-skills. [Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them. Turns failures into test cases that prevent silent regression. This is the outer loop's regress-test step. Use when a learning is promoted and has a clear pass/fail condition, or on cadence to verify promoted rules still hold.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Test generation and LLM evaluation.

When your agent uses it

  • A learning is promoted and has a clear pass/fail condition
  • On cadence to verify promoted rules still hold

Example prompts

  • “/eval-creator”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Check precondition — if not met, mark as skip
  2. Execute verification method
  3. Compare result to expected
  4. Check evidence strength — a presence/structural method cannot satisfy a
  5. Update last-run and last-result in the eval case file
  6. Update EVAL_INDEX.md with the result

What it can do on your machine

Read from SKILL.md and the folder at commit 5a836dc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Creator loads about 2.6k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 887 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 887 words (~2,598 tokens).

“Turns promoted learnings into permanent eval cases. Runs regression checks to verify promoted rules hold. This is the outer loop's regress-test step.”

— opening of SKILL.md by pskoett
name
eval-creator

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/eval-creator of pskoett/pskoett-ai-skills.

Open the folder on GitHubat commit 5a836dc

Compare with similar skills

Eval Creator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Creator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Creator this skillpskoett/pskoett-ai-skills311—~2.6kAutomated safety check: PassNone
Eval Guidemicrosoft/eval-guide138—~22kAutomated safety check: WarnMIT
Synthetic Eval Data Generatorai-evals-course/evals-skills1.5k—~1.4kAutomated safety check: PassApache-2.0
Axiom Eval Writeropenclaw/clawhub9.5k—~4.1kAutomated safety check: WarnMIT
SDK AI Bot Run EvaluationAzure/azure-sdk-tools134—~1.1kAutomated safety check: NotesMIT
Testinginbrainfun/inbrain1421 repos~2kAutomated safety check: PassCustom licence

Similar skills

  • Eval Guide

    microsoft/eval-guide

    Official

    Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

    138 GitHub stars~22k tokensUpdated 3 mo ago
    Testing & QAAuto-check: warnings
  • Synthetic Eval Data Generator

    ai-evals-course/evals-skills

    Builds diverse synthetic test inputs for LLM pipeline evaluation by defining failure-focused dimensions, drafting tuples with you and turning them into realistic queries.

    1.5k GitHub stars~1.4k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Axiom Eval Writer

    openclaw/clawhub

    Scaffolds evaluation suites for the Axiom AI SDK: eval files, scorers, flag schemas and axiom.config.ts, generated from plain descriptions of an AI capability.

    9.5k GitHub stars~4.1k tokensUpdated today
    AI & LLM EngineeringAuto-check: warnings
  • SDK AI Bot Run Evaluation

    Azure/azure-sdk-tools

    Official

    Run Azure SDK QA bot evaluations on curated datasets locally, including a single test case.

    134 GitHub stars~1.1k tokensUpdated today
    Knowledge ManagementAuto-check: notes
  • Testing

    inbrainfun/inbrain

    Skill validation framework PLUS daily test-suite health and regression intelligence.

    142 GitHub starsUsed in 1 repo~2k tokens
    Testing & QAAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    791 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from pskoett/pskoett-ai-skills

All 24 skills in this repo
  • Self Improvement

    pskoett/pskoett-ai-skills

    Captures learnings, errors, corrections, and feature requests to enable continuous improvement.

    311 GitHub stars~5k tokensUpdated 2 days ago
    Auto-check passed
  • Self Improvement

    pskoett/pskoett-ai-skills

    Captures learnings, errors, corrections, and feature requests to enable continuous improvement.

    311 GitHub stars~5.4k tokensUpdated 2 days ago
    Auto-check passed
  • Self Healing

    pskoett/pskoett-ai-skills

    Active runtime recovery for coding agents: when something breaks mid-task, diagnose the root cause, write a fix, VERIFY by re-running the broken thing, then file a HEAL- entry to .learnings/HEALS.md…

    311 GitHub stars~5.3k tokensUpdated 2 days ago
    Auto-check: notes
  • Skill Tester

    pskoett/pskoett-ai-skills

    Validates all interactive skills in this repo against the Agent Skills spec, project conventions, and structural requirements.

    311 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Skill Tester CI

    pskoett/pskoett-ai-skills

    Validates all CI skills in this repo. An agent skill from pskoett/pskoett-ai-skills.

    311 GitHub stars~845 tokensUpdated 2 days ago
    Auto-check passed
  • Context Decay

    pskoett/pskoett-ai-skills

    Revalidate persistent agent knowledge by classifying why it can become stale and how quickly it changes, then retain, revise, externalize, or retire it.

    311 GitHub stars~3.3k tokensUpdated 2 days ago
    Auto-check passed

Questions about Eval Creator

What does Eval Creator do?

[Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them. Eval Creator is an agent skill from pskoett/pskoett-ai-skills. [Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them.

When should I use Eval Creator?

Eval Creator fits situations like: A learning is promoted and has a clear pass/fail condition; on cadence to verify promoted rules still hold.

How do I install Eval Creator in Claude Code?

Run `npx skills add pskoett/pskoett-ai-skills --skill eval-creator -a claude-code`. Or copy the skill folder (skills/eval-creator in pskoett/pskoett-ai-skills) into .claude/skills/eval-creator in your project. Claude Code loads it when a task matches its description.

How do I install Eval Creator in Codex?

Run `npx skills add pskoett/pskoett-ai-skills --skill eval-creator -a codex`. Or copy the skill folder (skills/eval-creator in pskoett/pskoett-ai-skills) into .agents/skills/eval-creator in your project. Codex loads it when a task matches its description.

Can I use Eval Creator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pskoett/pskoett-ai-skills --skill eval-creator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-creator, .gemini/skills/eval-creator, .github/skills/eval-creator and .opencode/skills/eval-creator in your project.

What does Eval Creator need to run?

Going by SKILL.md and its folder, Eval Creator needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Eval Creator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Eval Creator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Creator use?

No licence was found for Eval Creator or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Eval Creator use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Eval Creator?

Skills that share tags, products or a category with Eval Creator: Eval Guide (microsoft/eval-guide, 138 stars), Synthetic Eval Data Generator (ai-evals-course/evals-skills, 1.5k stars), Axiom Eval Writer (openclaw/clawhub, 9.5k stars) and SDK AI Bot Run Evaluation (Azure/azure-sdk-tools, 134 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Creator?

pskoett (a GitHub user) maintains it in pskoett/pskoett-ai-skills, which has 311 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 5, 2026.

Source: pskoett/pskoett-ai-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.