Agent skill

Eval Creator CI

by pskoett in pskoett/pskoett-ai-skills

[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).

No licenceAuto-check passedAI & LLM Engineering

Install Eval Creator CI

skills CLI
$ npx skills add pskoett/pskoett-ai-skills --skill eval-creator-ci -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pskoett/pskoett-ai-skills eval-creator-ci --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pskoett/pskoett-ai-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/eval-creator-ci .claude/skills/eval-creator-ci && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
eval-creator-ci
GitHub stars
314
Token cost
~2.2k tokens
SKILL.md length
794 words
Files
2 (incl. references)
Skills in repo
24
Repo updated
First seen
Licence
None found

At a glance

[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).

  • Works in 6 steps: Eval execution is read-only for code —… → Eval case creation writes to .evals/… → Headless — no interactive prompts, no… → …
  • : you want automated regression testing of promoted rules in CI/headless pipelines
  • SKILL.md covers Install, Purpose, Context Limitation (Important) and Prerequisites, plus 9 more sections
  • Calls gh and npx

What it does

Eval Creator CI is an agent skill from pskoett/pskoett-ai-skills. [Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). Runs all eval cases in .evals/ on a schedule or per-PR, reports pass/fail results, and can block merges on regressions. Also creates new eval cases from promoted patterns flagged by learning-aggregator-ci. Use when: you want automated regression testing of promoted rules in CI/headless pipelines. For interactive eval creation and runs, use eval-creator.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/workflow-example.md`).

It sits in AI & LLM Engineering, covering LLM evaluation and QA and bug reports. It works with GitHub.

When your agent uses it

  • : you want automated regression testing of promoted rules in CI/headless pipelines
  • Tasks that involve LLM evaluation
  • Tasks that involve QA and bug reports

Example prompts

  • “/eval-creator-ci”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Eval execution is read-only for code — eval cases read files and run check commands but do not modify source code
  2. Eval case creation writes to .evals/ only — when creating new evals from promotion candidates
  3. Headless — no interactive prompts, no approval gates
  4. Structured output — emit results as YAML under eval_creator_ci key
  5. Gate policy — can fail the check run on eval regressions (configurable)
  6. Single comment — post one consolidated results comment per run

What it can do on your machine

Read from SKILL.md and the folder at commit 5a836dc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Eval Creator CI loads about 2.2k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 112 tokens; SKILL.md has 794 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 794 words (~2,233 tokens).

name
eval-creator-ci

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file (references) in skills/eval-creator-ci of pskoett/pskoett-ai-skills.

  • SKILL.md
  • references/workflow-example.md

Open the folder on GitHubat commit 5a836dc

Compare with similar skills

Eval Creator CI next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Eval Creator CI compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Eval Creator CI this skillpskoett/pskoett-ai-skills314—~2.2kAutomated safety check: PassNone
RAG Observability Evalssickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: PassMIT
Benchmark RunnerRConsortium/pharma-skills119—~5.3kAutomated safety check: PassMIT
Onboarddavepoon/buildwithclaude3.6k—~1.3kAutomated safety check: PassMIT
Octocode Benchmark Runnerbgauryy/octocode949—~2.1kAutomated safety check: PassMIT
AI Project Copilotsun461941-hub/ai-project-copilot100—~3kAutomated safety check: PassMIT

Similar skills

  • RAG Observability Evals

    sickn33/agentic-awesome-skills

    Monitor and evaluate RAG systems with retrieval quality metrics, groundedness checks, hallucination detection, and continuous regression testing.

    47k GitHub starsUsed in 2 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Benchmark Runner

    RConsortium/pharma-skills

    Auto-discover all skills with evals in RConsortium/pharma-skills, benchmark each with vs.

    119 GitHub stars~5.3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Onboard

    davepoon/buildwithclaude

    Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it.

    3.6k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • AI Project Copilot

    sun461941-hub/ai-project-copilot

    A skill your agent uses to turn an AI idea or existing repository into a credible open-source product and to run evidence-first repository engineering across codebase discovery, context-efficient…

    100 GitHub stars~3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Frontierharness Eval

    frontier-harness-eval/eval

    Benchmark a third-party coding-agent harness against FrontierHarness Eval using Runta runtimes.

    301 GitHub stars~8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from pskoett/pskoett-ai-skills

All 24 skills in this repo
  • Self Improvement

    pskoett/pskoett-ai-skills

    Captures learnings, errors, corrections, and feature requests to enable continuous improvement.

    314 GitHub stars~5k tokensUpdated 4 days ago
    Auto-check passed
  • Self Improvement

    pskoett/pskoett-ai-skills

    Captures learnings, errors, corrections, and feature requests to enable continuous improvement.

    314 GitHub stars~5.4k tokensUpdated 4 days ago
    Auto-check passed
  • Self Healing

    pskoett/pskoett-ai-skills

    Active runtime recovery for coding agents: when something breaks mid-task, diagnose the root cause, write a fix, VERIFY by re-running the broken thing, then file a HEAL- entry to .learnings/HEALS.md…

    314 GitHub stars~5.3k tokensUpdated 4 days ago
    Auto-check: notes
  • Skill Tester

    pskoett/pskoett-ai-skills

    Validates all interactive skills in this repo against the Agent Skills spec, project conventions, and structural requirements.

    314 GitHub stars~1.3k tokensUpdated 4 days ago
    Auto-check passed
  • Skill Tester CI

    pskoett/pskoett-ai-skills

    Validates all CI skills in this repo. An agent skill from pskoett/pskoett-ai-skills.

    314 GitHub stars~845 tokensUpdated 4 days ago
    Auto-check passed
  • Context Decay

    pskoett/pskoett-ai-skills

    Revalidate persistent agent knowledge by classifying why it can become stale and how quickly it changes, then retain, revise, externalize, or retire it.

    314 GitHub stars~3.3k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Eval Creator CI

What does Eval Creator CI do?

[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). Eval Creator CI is an agent skill from pskoett/pskoett-ai-skills. [Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows).

When should I use Eval Creator CI?

Eval Creator CI fits situations like: : you want automated regression testing of promoted rules in CI/headless pipelines; tasks that involve LLM evaluation; tasks that involve QA and bug reports.

How do I install Eval Creator CI in Claude Code?

Run `npx skills add pskoett/pskoett-ai-skills --skill eval-creator-ci -a claude-code`. Or copy the skill folder (skills/eval-creator-ci in pskoett/pskoett-ai-skills) into .claude/skills/eval-creator-ci in your project. Claude Code loads it when a task matches its description.

How do I install Eval Creator CI in Codex?

Run `npx skills add pskoett/pskoett-ai-skills --skill eval-creator-ci -a codex`. Or copy the skill folder (skills/eval-creator-ci in pskoett/pskoett-ai-skills) into .agents/skills/eval-creator-ci in your project. Codex loads it when a task matches its description.

Can I use Eval Creator CI in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pskoett/pskoett-ai-skills --skill eval-creator-ci -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/eval-creator-ci, .gemini/skills/eval-creator-ci, .github/skills/eval-creator-ci and .opencode/skills/eval-creator-ci in your project.

What does Eval Creator CI need to run?

Going by SKILL.md and its folder, Eval Creator CI needs the command-line tools its instructions call (gh and npx). Our summary lists: Node.js.

Does Eval Creator CI access the network?

SKILL.md contains no URLs. Its commands use gh and npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Eval Creator CI safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Eval Creator CI use?

No licence was found for Eval Creator CI or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Eval Creator CI use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Eval Creator CI?

Skills that share tags, products or a category with Eval Creator CI: RAG Observability Evals (sickn33/agentic-awesome-skills, 47k stars), Benchmark Runner (RConsortium/pharma-skills, 119 stars), Onboard (davepoon/buildwithclaude, 3.6k stars) and Octocode Benchmark Runner (bgauryy/octocode, 949 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Eval Creator CI?

pskoett (a GitHub user) maintains it in pskoett/pskoett-ai-skills, which has 314 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 5, 2026.

Source: pskoett/pskoett-ai-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.