Aeon Skill Evals
BankrBot/skills
Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation.
Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.
$ npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Arenukvern/mcp_flutter skill-eval-improve --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Arenukvern/mcp_flutter.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/skill-eval-improve .claude/skills/skill-eval-improve && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skill-eval-improve" agent skill from https://github.com/Arenukvern/mcp_flutter/tree/main/.agents/skills/skill-eval-improve into .claude/skills/skill-eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-eval-improve", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Arenukvern/mcp_flutter/tree/main/.agents/skills/skill-eval-improveType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Arenukvern/mcp_flutter skill-eval-improve --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arenukvern/mcp_flutter.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/skill-eval-improve .agents/skills/skill-eval-improve && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skill-eval-improve" agent skill from https://github.com/Arenukvern/mcp_flutter/tree/main/.agents/skills/skill-eval-improve into .agents/skills/skill-eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-eval-improve", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Arenukvern/mcp_flutter skill-eval-improve --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arenukvern/mcp_flutter.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/skill-eval-improve .cursor/skills/skill-eval-improve && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skill-eval-improve" agent skill from https://github.com/Arenukvern/mcp_flutter/tree/main/.agents/skills/skill-eval-improve into .cursor/skills/skill-eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-eval-improve", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Arenukvern/mcp_flutter.git --path .agents/skills/skill-eval-improve--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Arenukvern/mcp_flutter skill-eval-improve --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arenukvern/mcp_flutter.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/skill-eval-improve .gemini/skills/skill-eval-improve && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skill-eval-improve" agent skill from https://github.com/Arenukvern/mcp_flutter/tree/main/.agents/skills/skill-eval-improve into .gemini/skills/skill-eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-eval-improve", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Arenukvern/mcp_flutter skill-eval-improveInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Arenukvern/mcp_flutter.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/skill-eval-improve .github/skills/skill-eval-improve && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skill-eval-improve" agent skill from https://github.com/Arenukvern/mcp_flutter/tree/main/.agents/skills/skill-eval-improve into .github/skills/skill-eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-eval-improve", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Arenukvern/mcp_flutter skill-eval-improve --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Arenukvern/mcp_flutter.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/skill-eval-improve .opencode/skills/skill-eval-improve && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skill-eval-improve" agent skill from https://github.com/Arenukvern/mcp_flutter/tree/main/.agents/skills/skill-eval-improve into .opencode/skills/skill-eval-improve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-eval-improve", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skill-eval-improveImproves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.
Skill Eval Improve is an agent skill from Arenukvern/mcp_flutter. Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates. Use when tuning skill quality, routing, or adopting Chrome/Microsoft T-named quality gates—not for bulk validate-only or SkillOpt automation.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including reference files (for example `evals/cases/duplicated-guidance-compression-trigger.yaml`, `evals/cases/improve-routing-trigger.yaml` and `evals/cases/tool-loop-drift-trigger.yaml`).
It sits in Agent Workflows, covering Agent evaluation and testing, Quality gates and LLM evaluation. It works with pnpm. The repository describes itself as: MCP Toolkit for Flutter AI Agent Driven Development (MCP/CLI + custom client side tools) - via closed feedback loop (visual & semantic snapshot) and high client side… The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 62f3ee1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pnpmnpxFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
microsoft.github.ioarxiv.orggithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skill Eval Improve loads about 2.4k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 827 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Arenukvern/mcp_flutter at commit 62f3ee1, republished under its MIT licence (© Arenukvern). 827 words, ~2,388 tokens.
.claude/skills/skill-eval-improve/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Improve skills measurably: baseline → measure → bounded edit → re-validate. Combine local tooling, Codex plugin-eval (when installed), and research-backed loops (SkillOpt).
SKILL.md, high token cost, weak outcomespnpm run validate only (skill-authoring-lifecycle for audit); do not start benchmark or SkillOpt loops.Cursor scope (optional): activate when editing under skills/** or scripts/validate-skills.mjs.
| Layer | Expert | Tool / method | Cost |
|---|---|---|---|
| 0 — Gate | Lint | pnpm run validate, skill-authoring-lifecycle | seconds |
| 0b — Rules | Routing/docs SSOT | pnpm run eval (T1 behavior-critical YAML cases) | seconds |
| 1 — Static | Structure | Codex plugin-eval analyze (if available) | seconds |
| 2 — Human | Behavior | 3–5 prompts with/without skill | minutes |
| 3 — Measured | Usage | plugin-eval benchmark + measurement-plan | minutes–hours |
| 4 — Evolve | Text optimization | SkillOpt-style bounded edits + held-out gate | hours |
| 5 — Navigate | Telemetry / dogfood | Current steward benchmark scenarios on compact traces | seconds |
Use the cheapest layer that answers the question. Do not skip layer 0.
pnpm run validate
pnpm run validate:json # CI / automationFix all error: lines. Treat warn: (missing sources.md, long SKILL.md) seriously.
| Tier | Skills | CI |
|---|---|---|
| T1 — Behavior-critical eval-gated | Routing/procedure skills where drift can change agent decisions, claims, delegation, governance, or evidence boundaries | pnpm run eval + validate |
| T2 — Structural validate-only | All others | pnpm run validate |
T1 behavior-critical currently includes harness-engineering-lifecycle, mcp-harness-repo-maintainer, mixture-of-experts, multi-agent-handoff, plugin-marketplace-setup, repo-quality-system-lifecycle, repository-governance-lifecycle, skill-authoring-lifecycle, skill-eval-improve, steward-continuity-boundary-lifecycle, and vision-alignment-foresight. Each requires evals/cases/*.yaml (≥2) + references/evals.md. Schema: eval-case-schema.md.
pnpm run eval
pnpm run eval -- --skill mcp-harness-repo-maintainer
pnpm run eval:jsonChrome eval design (failure modes, rubrics, objective vs judge): references/chrome-eval-design.md.
CI does not run LLM judges. Subjective quality stays in references/evals.md (layer 2+).
When Codex plugin-eval is installed (~/.codex/plugins/.../plugin-eval):
# Chat-first router
plugin-eval start skills/<name> --request "Evaluate this skill." --format markdown
# Static report
plugin-eval analyze skills/<name> --format markdown
# Token budget explanation
plugin-eval explain-budget skills/<name> --format markdown
# Starter benchmark config
plugin-eval init-benchmark skills/<name>
plugin-eval benchmark skills/<name> --dry-runHand off rewrite plans to plugin-eval’s improve-skill skill after analyze --brief-out.
Details: references/plugin-eval.md.
evals/cases/*.yaml (CI rules) and references/evals.md (behavior log).Split ~60% train (edit against) / 40% held-out (gate acceptance)—mirrors SkillOpt selection gate.
SkillOpt treats SKILL.md as trainable text with a frozen agent:
Rollout (tasks + current skill) → Reflect (failures vs successes)
→ Bounded edit (add/delete/replace under budget) → Held-out gate (keep only if better)Skill Steward manual adaptation (no GPU cluster required):
| Step | Action |
|---|---|
| 1 | Baseline: held-out pass rate without skill |
| 2 | With skill: same tasks, log pass rate |
| 3 | Reflect: list 1–3 concrete failure modes |
| 4 | Bounded edit: ≤10% line churn or one new section; no wholesale rewrite |
| 5 | Re-run held-out only; keep edit only if improved |
| 6 | Record outcome in references/evals.md + sources.md |
Paper: https://arxiv.org/abs/2605.23904 · Site: https://microsoft.github.io/SkillOpt/
Related: SkillLens (model-generated skills study).
| Resource | Use |
|---|---|
| SkillsBench | Inspiration for paired vanilla vs skill-augmented tasks |
| skillgrade | Regression testing skill quality (mgechev) |
| Claude authoring best practices | Eval-before-write workflow |
At 10,000x scale, NLP prompt evaluation fails because LLMs suffer Cognitive Overload navigating massive toolsets. Runtime dogfood should objectively assert their logical trajectory using deterministic traces.
return_to_goal_step, required artifacts, and negative checks for unrelated actions.steward benchmark --scenario <id> --json for runtime dogfood scenarios. Do not put product runtime scenarios under T1 behavior-critical skill evals.The current steward eval --name registered-eval path is legacy/experimental. Skill quality remains pnpm run eval; runtime dogfood belongs to steward benchmark, where durability_blocked is valid blocked evidence when contract inputs are modified or untracked, not proof of runtime behavior.
- [ ] sources.md cites plugin-eval + SkillOpt if used
- [ ] pnpm run validate
- [ ] T1 behavior-critical: `pnpm run eval` + cases updated
- [ ] plugin-eval analyze (optional)
- [ ] 3+ prompt evals documented in references/evals.md
- [ ] Bounded edit applied; held-out improved
- [ ] skill-authoring-lifecycle checklist
- [ ] PR mentions eval deltaname / description (routing)—must include what + whenreferences/sources.mdreferences/ (SKILL.md < 500 lines)references/sources.md rowspnpm run eval and claiming agent behavior is proven| Skill | Role |
|---|---|
skill-source-citations | Save research links |
skill-authoring-lifecycle | Scaffold |
skill-authoring-lifecycle | Pre-merge audit |
npx skills add arenukvern/skill_steward --skill skill-eval-improve© Arenukvern, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (references) in .agents/skills/skill-eval-improve of Arenukvern/mcp_flutter.
Open the folder on GitHubat commit 62f3ee1
Skill Eval Improve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skill Eval Improve this skillArenukvern/mcp_flutter | 386 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Aeon Skill EvalsBankrBot/skills | 1.2k | — | ~660 | Automated safety check: Pass | None | |
| Waza Skill Evaluatormicrosoft/waza | 1.4k | — | ~2k | Automated safety check: Pass | MIT | |
| Waza Interactivemicrosoft/waza | 1.4k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Continuous Agent Loopaffaan-m/ECC | 276k | 5 repos | ~298 | Automated safety check: Pass | MIT | |
| Skill Release Gaterohitg00/ai-engineering-from-scratch | 66k | — | ~1k | Automated safety check: Pass | MIT |
BankrBot/skills
Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation.
microsoft/waza
Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.
microsoft/waza
Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.
affaan-m/ECC
Patterns for continuous autonomous agent loops with quality gates, evals, and recovery controls.
rohitg00/ai-engineering-from-scratch
Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.
github/gh-aw
Designs and verifies a deterministic grader that measures whether a GitHub Agentic Workflow run reached its real-world or repository outcome.
Arenukvern/mcp_flutter
Design, implement, and integrate generalized validation harnesses across a producer-consumer boundary after a local harness contract exists.
Arenukvern/mcp_flutter
Run a Mixture of Experts (MoE) audit on any topic, plan, codebase, evidence archive, or process.
Arenukvern/mcp_flutter
Plan and document handoffs, parent lane contracts, and parallel batch contracts between specialized AI agents (foreman, workers, reviewers).
Arenukvern/mcp_flutter
Designs public or private Agent Skill and plugin marketplaces for Cursor, Claude Code, Codex, Zed, Open Plugin, and npx skills—manifest layout, install matrix, and Skill Steward vs product boundaries.
Arenukvern/mcp_flutter
Chooses ecosystem-native release and changelog tooling (Changesets, Melos, release-plz) plus binary distribution (GitHub Release tarballs, install.sh) when the product is an executable.
Arenukvern/mcp_flutter
Master orchestration for repository governance, North Star impact, sub-Star boundaries, and repair-first or evidence-first drift checks.
Works with
Categories
Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates. Skill Eval Improve is an agent skill from Arenukvern/mcp_flutter. Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.
Skill Eval Improve fits situations like: tuning skill quality; adopting Chrome/Microsoft T-named quality gates—not for bulk validate-only; skillOpt automation.
Run `npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a claude-code`. Or copy the skill folder (.agents/skills/skill-eval-improve in Arenukvern/mcp_flutter) into .claude/skills/skill-eval-improve in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a codex`. Or copy the skill folder (.agents/skills/skill-eval-improve in Arenukvern/mcp_flutter) into .agents/skills/skill-eval-improve in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-eval-improve, .gemini/skills/skill-eval-improve, .github/skills/skill-eval-improve and .opencode/skills/skill-eval-improve in your project.
Going by SKILL.md and its folder, Skill Eval Improve needs the command-line tools its instructions call (pnpm and npx). Our summary lists: Node.js.
SKILL.md names 3 domains. As links in the text: microsoft.github.io, arxiv.org and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Skill Eval Improve is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Skill Eval Improve: Aeon Skill Evals (BankrBot/skills, 1.2k stars), Waza Skill Evaluator (microsoft/waza, 1.4k stars), Waza Interactive (microsoft/waza, 1.4k stars) and Continuous Agent Loop (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Arenukvern (a GitHub user) maintains it in Arenukvern/mcp_flutter, which has 386 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 3, 2026.
Source: Arenukvern/mcp_flutter on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.