Agent skill

Aeon Skill Evals

by BankrBot in BankrBot/skills

Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation.

No licenceAuto-check passedTesting & QA

Install Aeon Skill Evals

skills CLI
$ npx skills add BankrBot/skills --skill aeon-skill-evals -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BankrBot/skills aeon-skill-evals --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BankrBot/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/aeon-skill-evals .claude/skills/aeon-skill-evals && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aeon-skill-evals
GitHub stars
1.2k
Token cost
~660 tokens
SKILL.md length
199 words
Files
2
Skills in repo
106
Repo updated
First seen
Licence
None found

At a glance

Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation.

  • Tasks that involve Agent evaluation and testing
  • SKILL.md covers Manifest format, Operations, Regression states and Bootstrap mode, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve LLM evaluation

What it does

Aeon Skill Evals is an agent skill from BankrBot/skills. Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation. Detects regressions by diffing vs prior runs (NEWFAIL / NEWPASS / CHRONIC / STABLEFAIL). Bootstrap mode generates a starter manifest from a skill's recent successful runs so manifests aren't written speculatively. Triggers: "evaluate this skill's output", "check skill X for regressions", "bootstrap evals for Y", "did this skill output pass quality gates".

Its SKILL.md is about 660 tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `catalog.json`).

It sits in Testing & QA, covering Agent evaluation and testing, LLM evaluation and Quality gates. The repository describes itself as: Bankr Skills equip builders with plug-and-play tools to build more powerful agents.

When your agent uses it

  • Tasks that involve Agent evaluation and testing
  • Tasks that involve LLM evaluation
  • Tasks that involve Quality gates

Example prompts

  • “s recent successful runs so manifests aren”
  • “evaluate this skill”
  • “check skill X for regressions”
  • “/aeon-skill-evals”

What it can do on your machine

Read from SKILL.md and the folder at commit dc47eed. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Aeon Skill Evals loads about 660 tokens when it runs. Until then it costs about 135 tokens; SKILL.md has 199 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~135
When it runs · the whole SKILL.md, loaded when a task matches
~660

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 199 words (~660 tokens).

“Quality net for installed skills. Each skill can declare an assertion manifest; outputs are checked against it; failing assertions surface regressions and route concrete fixes.”

— opening of SKILL.md by BankrBot
name
aeon-skill-evals

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file in aeon-skill-evals of BankrBot/skills.

  • SKILL.md
  • catalog.json

Open the folder on GitHubat commit dc47eed

Compare with similar skills

Aeon Skill Evals next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Aeon Skill Evals compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Aeon Skill Evals this skillBankrBot/skills1.2k—~660Automated safety check: PassNone
Skill Eval ImproveArenukvern/mcp_flutter387—~2.4kAutomated safety check: PassMIT
Eval Guidemicrosoft/eval-guide138—~22kAutomated safety check: WarnMIT
Windmill AI Evalswindmill-labs/windmill18k—~969Automated safety check: NotesCustom licence
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Octocode Benchmark Runnerbgauryy/octocode949—~2.1kAutomated safety check: PassMIT

Similar skills

  • Skill Eval Improve

    Arenukvern/mcp_flutter

    Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.

    387 GitHub stars~2.4k tokensUpdated 7 days ago
    Agent WorkflowsAuto-check passed
  • Eval Guide

    microsoft/eval-guide

    Official

    Eval enablement accelerator — help customers think through "what does good look like" for their AI agent, then generate a structured eval plan and test cases they can use immediately.

    138 GitHub stars~22k tokensUpdated 3 mo ago
    Testing & QAAuto-check: warnings
  • Windmill AI Evals

    windmill-labs/windmill

    Writes and runs black-box benchmark cases for Windmill's flow, app, script, CLI and global AI generation modes, including before-and-after comparisons.

    18k GitHub stars~969 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Runs blind pairwise comparisons of Octocode against a gh-based baseline over markdown research questions, scored by total characters through the model rather than self-report.

    949 GitHub stars~2.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Autocontext

    greyhaven-ai/autocontext

    Runs LLM-based rubric judging on agent output and loops revise-and-rejudge rounds until a quality threshold is met.

    1.3k GitHub stars~892 tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

More from BankrBot/skills

All 106 skills in this repo
  • OnchainKit App Builder

    BankrBot/skills

    Builds onchain apps with Coinbase's OnchainKit React components and TypeScript utilities: wallet connection, identity, token swaps, transactions and checkout flows.

    1.2k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Lists AgenticBets prediction markets on Base, shows odds, places UP or DOWN bets on token prices in USDC and claims winnings through the Bankr wallet API.

    1.2k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • AI2Human Task Router

    BankrBot/skills

    Creates AI2Human tasks for steps that need a real person, such as manual QA or local checks, and returns a task URL the agent can track.

    1.2k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Lets an agent inspect and operate AZZLE V2 tasks on Base through Bankr, from posting and claiming to funding, delivery, release and disputes, behind verified deployment pins.

    1.2k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • B20 Console

    BankrBot/skills

    Inspect B20 token contract addresses on Base through B20 Console.

    1.2k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Capacitr

    BankrBot/skills

    Paste a URL or free text and get matched Polymarket / Hyperliquid / Deribit markets with Quotient edge scores.

    1.2k GitHub stars~3k tokensUpdated today
    Auto-check passed

Questions about Aeon Skill Evals

What does Aeon Skill Evals do?

Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation. Aeon Skill Evals is an agent skill from BankrBot/skills. Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation.

When should I use Aeon Skill Evals?

Aeon Skill Evals fits situations like: tasks that involve Agent evaluation and testing; tasks that involve LLM evaluation; tasks that involve Quality gates.

How do I install Aeon Skill Evals in Claude Code?

Run `npx skills add BankrBot/skills --skill aeon-skill-evals -a claude-code`. Or copy the skill folder (aeon-skill-evals in BankrBot/skills) into .claude/skills/aeon-skill-evals in your project. Claude Code loads it when a task matches its description.

How do I install Aeon Skill Evals in Codex?

Run `npx skills add BankrBot/skills --skill aeon-skill-evals -a codex`. Or copy the skill folder (aeon-skill-evals in BankrBot/skills) into .agents/skills/aeon-skill-evals in your project. Codex loads it when a task matches its description.

Can I use Aeon Skill Evals in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BankrBot/skills --skill aeon-skill-evals -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aeon-skill-evals, .gemini/skills/aeon-skill-evals, .github/skills/aeon-skill-evals and .opencode/skills/aeon-skill-evals in your project.

What does Aeon Skill Evals need to run?

SKILL.md names no scripts, command-line tools or credentials: Aeon Skill Evals is instructions for the agent only.

Does Aeon Skill Evals access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Aeon Skill Evals safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Aeon Skill Evals use?

No licence was found for Aeon Skill Evals or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Aeon Skill Evals use?

About 660 tokens (SKILL.md is roughly 2.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Aeon Skill Evals?

Skills that share tags, products or a category with Aeon Skill Evals: Skill Eval Improve (Arenukvern/mcp_flutter, 387 stars), Eval Guide (microsoft/eval-guide, 138 stars), Windmill AI Evals (windmill-labs/windmill, 18k stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Aeon Skill Evals?

BankrBot (a GitHub organization) maintains it in BankrBot/skills, which has 1,202 GitHub stars. The repository holds 106 skills in this directory. The repository was last updated on October 10, 2026.

Source: BankrBot/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.