Agent skill

Skill Eval Improve

by Arenukvern in Arenukvern/mcp_flutter

Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.

MITAuto-check passedAgent Workflows

Install Skill Eval Improve

skills CLI
$ npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Arenukvern/mcp_flutter skill-eval-improve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Arenukvern/mcp_flutter.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/skill-eval-improve .claude/skills/skill-eval-improve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-eval-improve
GitHub stars
386
Token cost
~2.4k tokens
SKILL.md length
827 words
Files
12 (incl. references)
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.

  • Works in 4 steps: Write 3–5 representative user prompts… → Run agent without skill → record failures. → Run with skill → record improvements and… → …
  • Tuning skill quality
  • SKILL.md covers When to use, When not to use, Mixture of experts (evaluation… and Layer 0 — Skill Steward…, plus 13 more sections
  • Calls pnpm and npx

What it does

Skill Eval Improve is an agent skill from Arenukvern/mcp_flutter. Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates. Use when tuning skill quality, routing, or adopting Chrome/Microsoft T-named quality gates—not for bulk validate-only or SkillOpt automation.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including reference files (for example `evals/cases/duplicated-guidance-compression-trigger.yaml`, `evals/cases/improve-routing-trigger.yaml` and `evals/cases/tool-loop-drift-trigger.yaml`).

It sits in Agent Workflows, covering Agent evaluation and testing, Quality gates and LLM evaluation. It works with pnpm. The repository describes itself as: MCP Toolkit for Flutter AI Agent Driven Development (MCP/CLI + custom client side tools) - via closed feedback loop (visual & semantic snapshot) and high client side… The licence is MIT.

When your agent uses it

  • Tuning skill quality
  • Adopting Chrome/Microsoft T-named quality gates—not for bulk validate-only
  • SkillOpt automation

Example prompts

  • “Use the skill-eval-improve skill to improve Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with…”
  • “/skill-eval-improve”

Requirements

  • Node.js

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Write 3–5 representative user prompts (should trigger + should not trigger).
  2. Run agent without skill → record failures.
  3. Run with skill → record improvements and new failures.
  4. Mirror prompts in evals/cases/*.yaml (CI rules) and references/evals.md (behavior log).

What it can do on your machine

Read from SKILL.md and the folder at commit 62f3ee1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • microsoft.github.io
    • arxiv.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Eval Improve loads about 2.4k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 827 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Arenukvern/mcp_flutter at commit 62f3ee1, republished under its MIT licence (© Arenukvern). 827 words, ~2,388 tokens.

Download SKILL.mdSave it as .claude/skills/skill-eval-improve/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
skill-eval-improve
description
Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates. Use when tuning skill quality, routing, or adopting Chrome/Microsoft T-named quality gates—not for bulk validate-only or SkillOpt automation.
license
MIT
type
governance
metadata.author
skill-steward
metadata.version
1.1.0
metadata.category
marketplace
paths
skills/**/evals/**, scripts/eval-skill.mjs, scripts/eval-tiers.mjs

Skill eval & improve

Improve skills measurably: baseline → measure → bounded edit → re-validate. Combine local tooling, Codex plugin-eval (when installed), and research-backed loops (SkillOpt).

When to use

  • Skill triggers wrong or never loads (description routing)
  • Bloated SKILL.md, high token cost, weak outcomes
  • After adding a new procedure—need regression checks
  • Porting patterns from product MCP / plugin-eval research into Skill Steward skills

When not to use

  • Bulk repo validation — e.g. “validate every skill in this repo” → pnpm run validate only (skill-authoring-lifecycle for audit); do not start benchmark or SkillOpt loops.
  • Automated SkillOpt / cluster training — Skill Steward documents a manual bounded-edit loop; no overnight optimizer pipeline.
  • Creating a new skill — use skill-authoring-lifecycle first; eval-improve applies after a skill exists.

Cursor scope (optional): activate when editing under skills/** or scripts/validate-skills.mjs.

Mixture of experts (evaluation stack)

LayerExpertTool / methodCost
0 — GateLintpnpm run validate, skill-authoring-lifecycleseconds
0b — RulesRouting/docs SSOTpnpm run eval (T1 behavior-critical YAML cases)seconds
1 — StaticStructureCodex plugin-eval analyze (if available)seconds
2 — HumanBehavior3–5 prompts with/without skillminutes
3 — MeasuredUsageplugin-eval benchmark + measurement-planminutes–hours
4 — EvolveText optimizationSkillOpt-style bounded edits + held-out gatehours
5 — NavigateTelemetry / dogfoodCurrent steward benchmark scenarios on compact tracesseconds

Use the cheapest layer that answers the question. Do not skip layer 0.

Layer 0 — Skill Steward validator (always)

bash
pnpm run validate
pnpm run validate:json   # CI / automation

Fix all error: lines. Treat warn: (missing sources.md, long SKILL.md) seriously.

T-named skill quality gates (ADR 0011, ADR 0027)

TierSkillsCI
T1 — Behavior-critical eval-gatedRouting/procedure skills where drift can change agent decisions, claims, delegation, governance, or evidence boundariespnpm run eval + validate
T2 — Structural validate-onlyAll otherspnpm run validate

T1 behavior-critical currently includes harness-engineering-lifecycle, mcp-harness-repo-maintainer, mixture-of-experts, multi-agent-handoff, plugin-marketplace-setup, repo-quality-system-lifecycle, repository-governance-lifecycle, skill-authoring-lifecycle, skill-eval-improve, steward-continuity-boundary-lifecycle, and vision-alignment-foresight. Each requires evals/cases/*.yaml (≥2) + references/evals.md. Schema: eval-case-schema.md.

Layer 0b — Rule-based cases (T1 behavior-critical CI)

bash
pnpm run eval
pnpm run eval -- --skill mcp-harness-repo-maintainer
pnpm run eval:json

Chrome eval design (failure modes, rubrics, objective vs judge): references/chrome-eval-design.md.

CI does not run LLM judges. Subjective quality stays in references/evals.md (layer 2+).

Layer 1 — Codex plugin-eval (local)

When Codex plugin-eval is installed (~/.codex/plugins/.../plugin-eval):

bash
# Chat-first router
plugin-eval start skills/<name> --request "Evaluate this skill." --format markdown

# Static report
plugin-eval analyze skills/<name> --format markdown

# Token budget explanation
plugin-eval explain-budget skills/<name> --format markdown

# Starter benchmark config
plugin-eval init-benchmark skills/<name>
plugin-eval benchmark skills/<name> --dry-run

Hand off rewrite plans to plugin-eval’s improve-skill skill after analyze --brief-out.

Details: references/plugin-eval.md.

Layer 2 — Human prompt suite (required for material edits)

  1. Write 3–5 representative user prompts (should trigger + should not trigger).
  2. Run agent without skill → record failures.
  3. Run with skill → record improvements and new failures.
  4. Mirror prompts in evals/cases/*.yaml (CI rules) and references/evals.md (behavior log).

Split ~60% train (edit against) / 40% held-out (gate acceptance)—mirrors SkillOpt selection gate.

Layer 3 — SkillOpt-inspired improve loop (research)

SkillOpt treats SKILL.md as trainable text with a frozen agent:

text
Rollout (tasks + current skill) → Reflect (failures vs successes)
  → Bounded edit (add/delete/replace under budget) → Held-out gate (keep only if better)

Skill Steward manual adaptation (no GPU cluster required):

StepAction
1Baseline: held-out pass rate without skill
2With skill: same tasks, log pass rate
3Reflect: list 1–3 concrete failure modes
4Bounded edit: ≤10% line churn or one new section; no wholesale rewrite
5Re-run held-out only; keep edit only if improved
6Record outcome in references/evals.md + sources.md

Paper: https://arxiv.org/abs/2605.23904 · Site: https://microsoft.github.io/SkillOpt/

Related: SkillLens (model-generated skills study).

Show full SKILL.md (336 more words)Show less

Layer 4 — Ecosystem benchmarks (2026+)

ResourceUse
SkillsBenchInspiration for paired vanilla vs skill-augmented tasks
skillgradeRegression testing skill quality (mgechev)
Claude authoring best practicesEval-before-write workflow

Layer 5 — Runtime Dogfood Benchmarks

At 10,000x scale, NLP prompt evaluation fails because LLMs suffer Cognitive Overload navigating massive toolsets. Runtime dogfood should objectively assert their logical trajectory using deterministic traces.

  1. Capture compact traces: Store action IDs, tool counts, artifact digests, and redacted excerpts, not raw product traces.
  2. Define assertions: Expected action trajectory, declared surfaces used first, maximum tool calls, maximum repair/setup attempts, maximum unrelated tool calls, required return_to_goal_step, required artifacts, and negative checks for unrelated actions.
  3. Run dogfood benchmarks: Use steward benchmark --scenario <id> --json for runtime dogfood scenarios. Do not put product runtime scenarios under T1 behavior-critical skill evals.

The current steward eval --name registered-eval path is legacy/experimental. Skill quality remains pnpm run eval; runtime dogfood belongs to steward benchmark, where durability_blocked is valid blocked evidence when contract inputs are modified or untracked, not proof of runtime behavior.

Improve workflow (checklist)

- [ ] sources.md cites plugin-eval + SkillOpt if used
- [ ] pnpm run validate
- [ ] T1 behavior-critical: `pnpm run eval` + cases updated
- [ ] plugin-eval analyze (optional)
- [ ] 3+ prompt evals documented in references/evals.md
- [ ] Bounded edit applied; held-out improved
- [ ] skill-authoring-lifecycle checklist
- [ ] PR mentions eval delta

What to fix first (typical order)

  1. name / description (routing)—must include what + when
  2. Broken links / missing references/sources.md
  3. Delete or replace duplicated rules before adding a new section or eval case
  4. Move bulk to references/ (SKILL.md < 500 lines)
  5. Add error-handling / validation steps agents skip
  6. Token cost (description length, always-loaded content)

Anti-patterns

  • Rewriting entire SKILL.md from one failure (destroy working rules)
  • Self-editing without held-out prompts (overfit)
  • Adding skill rules, evals, or tools from one observed run when a smaller FAQ, error message, native command, observed-effect check, or deletion would solve the problem
  • Adding a new eval for duplicated guidance before trying to compress, delete, or replace the overlapping rule
  • Claims without references/sources.md rows
  • Evaluating only with static analyze—never running real prompts
  • LLM judge in CI (flake, cost) — offline only per ADR 0011
  • Passing pnpm run eval and claiming agent behavior is proven
SkillRole
skill-source-citationsSave research links
skill-authoring-lifecycleScaffold
skill-authoring-lifecyclePre-merge audit

Sources

See references/sources.md.

Install

bash
npx skills add arenukvern/skill_steward --skill skill-eval-improve

© Arenukvern, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (references) in .agents/skills/skill-eval-improve of Arenukvern/mcp_flutter.

  • SKILL.md
  • evals/cases/duplicated-guidance-compression-trigger.yaml
  • evals/cases/improve-routing-trigger.yaml
  • evals/cases/tool-loop-drift-trigger.yaml
  • evals/cases/validate-only-dormant.yaml
  • references/chrome-eval-design.md
  • references/eval-case-schema.md
  • references/evals-template.md
  • references/evals.md
  • references/plugin-eval.md
  • references/skillopt-and-research.md
  • references/sources.md

Open the folder on GitHubat commit 62f3ee1

Compare with similar skills

Skill Eval Improve next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Eval Improve compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Eval Improve this skillArenukvern/mcp_flutter386—~2.4kAutomated safety check: PassMIT
Aeon Skill EvalsBankrBot/skills1.2k—~660Automated safety check: PassNone
Waza Skill Evaluatormicrosoft/waza1.4k—~2kAutomated safety check: PassMIT
Waza Interactivemicrosoft/waza1.4k—~1.3kAutomated safety check: PassMIT
Continuous Agent Loopaffaan-m/ECC276k5 repos~298Automated safety check: PassMIT
Skill Release Gaterohitg00/ai-engineering-from-scratch66k—~1kAutomated safety check: PassMIT

Similar skills

  • Aeon Skill Evals

    BankrBot/skills

    Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation.

    1.2k GitHub stars~660 tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Waza Skill Evaluator

    microsoft/waza

    Official

    Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.

    1.4k GitHub stars~2k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Waza Interactive

    microsoft/waza

    Official

    Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

    1.4k GitHub stars~1.3k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Patterns for continuous autonomous agent loops with quality gates, evals, and recovery controls.

    276k GitHub starsUsed in 5 repos~298 tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    66k GitHub stars~1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Official

    Designs and verifies a deterministic grader that measures whether a GitHub Agentic Workflow run reached its real-world or repository outcome.

    5.4k GitHub stars~6.8k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from Arenukvern/mcp_flutter

All 23 skills in this repo
  • Harness Engineering Lifecycle

    Arenukvern/mcp_flutter

    Design, implement, and integrate generalized validation harnesses across a producer-consumer boundary after a local harness contract exists.

    386 GitHub stars~1.6k tokensUpdated 6 days ago
    Auto-check passed
  • Mixture Of Experts

    Arenukvern/mcp_flutter

    Run a Mixture of Experts (MoE) audit on any topic, plan, codebase, evidence archive, or process.

    386 GitHub stars~2.3k tokensUpdated 6 days ago
    Auto-check passed
  • Multi Agent Handoff

    Arenukvern/mcp_flutter

    Plan and document handoffs, parent lane contracts, and parallel batch contracts between specialized AI agents (foreman, workers, reviewers).

    386 GitHub stars~2.7k tokensUpdated 6 days ago
    Auto-check passed
  • Plugin Marketplace Setup

    Arenukvern/mcp_flutter

    Designs public or private Agent Skill and plugin marketplaces for Cursor, Claude Code, Codex, Zed, Open Plugin, and npx skills—manifest layout, install matrix, and Skill Steward vs product boundaries.

    386 GitHub stars~2.9k tokensUpdated 6 days ago
    Auto-check passed
  • Release Changelog Harness

    Arenukvern/mcp_flutter

    Chooses ecosystem-native release and changelog tooling (Changesets, Melos, release-plz) plus binary distribution (GitHub Release tarballs, install.sh) when the product is an executable.

    386 GitHub stars~2.6k tokensUpdated 6 days ago
    Auto-check passed
  • Repository Governance Lifecycle

    Arenukvern/mcp_flutter

    Master orchestration for repository governance, North Star impact, sub-Star boundaries, and repair-first or evidence-first drift checks.

    386 GitHub stars~1.5k tokensUpdated 6 days ago
    Auto-check passed

Works with

Questions about Skill Eval Improve

What does Skill Eval Improve do?

Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates. Skill Eval Improve is an agent skill from Arenukvern/mcp_flutter. Improves Agent Skills via validate → rule-based eval cases → plugin-eval → prompt evals → bounded edits with held-out gates.

When should I use Skill Eval Improve?

Skill Eval Improve fits situations like: tuning skill quality; adopting Chrome/Microsoft T-named quality gates—not for bulk validate-only; skillOpt automation.

How do I install Skill Eval Improve in Claude Code?

Run `npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a claude-code`. Or copy the skill folder (.agents/skills/skill-eval-improve in Arenukvern/mcp_flutter) into .claude/skills/skill-eval-improve in your project. Claude Code loads it when a task matches its description.

How do I install Skill Eval Improve in Codex?

Run `npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a codex`. Or copy the skill folder (.agents/skills/skill-eval-improve in Arenukvern/mcp_flutter) into .agents/skills/skill-eval-improve in your project. Codex loads it when a task matches its description.

Can I use Skill Eval Improve in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Arenukvern/mcp_flutter --skill skill-eval-improve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-eval-improve, .gemini/skills/skill-eval-improve, .github/skills/skill-eval-improve and .opencode/skills/skill-eval-improve in your project.

What does Skill Eval Improve need to run?

Going by SKILL.md and its folder, Skill Eval Improve needs the command-line tools its instructions call (pnpm and npx). Our summary lists: Node.js.

Does Skill Eval Improve access the network?

SKILL.md names 3 domains. As links in the text: microsoft.github.io, arxiv.org and github.com. This is read from the text; nothing was executed.

Is Skill Eval Improve safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skill Eval Improve use?

Skill Eval Improve is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Eval Improve use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Skill Eval Improve?

Skills that share tags, products or a category with Skill Eval Improve: Aeon Skill Evals (BankrBot/skills, 1.2k stars), Waza Skill Evaluator (microsoft/waza, 1.4k stars), Waza Interactive (microsoft/waza, 1.4k stars) and Continuous Agent Loop (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Eval Improve?

Arenukvern (a GitHub user) maintains it in Arenukvern/mcp_flutter, which has 386 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 3, 2026.

Source: Arenukvern/mcp_flutter on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.