Agent skill

Evaluate

by sharpdeveye in sharpdeveye/maestro

A skill your agent uses when the user wants a quality review, interaction audit, or to test the workflow against realistic scenarios.

MITAuto-check passed

Install Evaluate

skills CLI
$ npx skills add sharpdeveye/maestro --skill evaluate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sharpdeveye/maestro evaluate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sharpdeveye/maestro.git skills-src && mkdir -p .claude/skills && cp -r skills-src/source/skills/evaluate .claude/skills/evaluate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate
GitHub stars
592
Token cost
~776 tokens
SKILL.md length
362 words
Files
1
Skills in repo
25
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user wants a quality review, interaction audit, or to test the workflow against realistic scenarios.

  • Works in 4 steps: Overall quality grade (A-F) → Per-dimension scores with evidence → Specific scenario results → …
  • The user wants a quality review
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Interaction audit

What it does

Evaluate is an agent skill from sharpdeveye/maestro. Use when the user wants a quality review, interaction audit, or to test the workflow against realistic scenarios.

Its SKILL.md is about 780 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Workflow fluency for AI coding agents. 1 core skill · 25 commands · 7 domain references · memory layer · audit trail — works across Cursor, Claude Code, Gemini CLI, Copilot, and… The licence is MIT.

When your agent uses it

  • The user wants a quality review
  • Interaction audit
  • Test the workflow against realistic scenarios

Example prompts

  • “/evaluate”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Overall quality grade (A-F)
  2. Per-dimension scores with evidence
  3. Specific scenario results
  4. Priority improvements with recommended Maestro commands

What it can do on your machine

Read from SKILL.md and the folder at commit 00f9115. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluate loads about 776 tokens when it runs. Until then it costs about 31 tokens; SKILL.md has 362 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~31
When it runs · the whole SKILL.md, loaded when a task matches
~776

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sharpdeveye/maestro at commit 00f9115, republished under its MIT licence (© sharpdeveye). 362 words, ~776 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate/SKILL.md (or your agent's skills folder).
name
evaluate
description
Use when the user wants a quality review, interaction audit, or to test the workflow against realistic scenarios.
argument-hint
[workflow or scenario]
category
analysis
version
2.0.0
user-invocable
true

MANDATORY PREPARATION

Invoke /agent-workflow — it contains workflow principles, anti-patterns, and the Context Gathering Protocol. Follow the protocol before proceeding — if no workflow context exists yet, you MUST run /teach-maestro first. Consult the feedback-loops reference in the agent-workflow skill for evaluation patterns, golden test sets, and regression detection.


Evaluate the workflow's actual interaction quality by testing it against scenarios that represent real usage.

Evaluation Dimensions

1. Task Completion

  • Does the workflow actually accomplish what it's supposed to?
  • Does it handle the complete task or only the happy path?
  • Are edge cases addressed or silently dropped?

2. Output Quality

  • Is the output accurate, complete, and well-formatted?
  • Does it match the defined output schema (if any)?
  • Would a domain expert approve the output?

3. Error Behavior

  • What happens when input is malformed?
  • What happens when a tool fails?
  • What happens when the model is uncertain?
  • Is the error message useful or generic?

4. User Experience

  • Is the interaction natural and intuitive?
  • Are confirmations requested for destructive operations?
  • Is the response time acceptable?
  • Does the workflow communicate its limitations?

5. Consistency

  • Does the same input produce consistent output quality?
  • Are there random failures that aren't reproducible?
  • Does quality degrade over long conversations?
Show full SKILL.md (164 more words)Show less
Scenario Testing

Create and run test scenarios:

ScenarioInputExpectedActualGrade
Happy pathNormal inputCorrect output?A-F
Edge caseUnusual inputGraceful handling?A-F
Error caseBad inputHelpful error?A-F
Stress caseLarge/complex inputReasonable handling?A-F
AdversarialTricky/malicious inputSafe response?A-F
Evaluation Report

Produce a structured report with:

  1. Overall quality grade (A-F)
  2. Per-dimension scores with evidence
  3. Specific scenario results
  4. Priority improvements with recommended Maestro commands
Evaluation Checklist
  • All 5 dimensions tested with concrete scenarios
  • At least one edge case and one adversarial case tested
  • Results documented in the scenario table
  • Overall grade assigned with justification
  • Improvement actions reference specific Maestro commands

After evaluation, run /fortify to address error behavior gaps, /refine for output quality improvements, or /iterate to set up continuous quality monitoring.

NEVER:

  • Evaluate theoretically — run actual scenarios
  • Give an A grade unless the workflow handles all scenario types well
  • Skip adversarial testing for user-facing workflows
  • Evaluate only the happy path

© sharpdeveye, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in source/skills/evaluate of sharpdeveye/maestro.

Open the folder on GitHubat commit 00f9115

Compare with similar skills

Evaluate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate this skillsharpdeveye/maestro592—~776Automated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT
LLM Evaluationdavila7/claude-code-templates32k13 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT
EvaluatorsArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence
Agent Evaluation Reportingsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 2 repos~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 13 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check passed
  • Agent Evaluation Reporting

    sickn33/agentic-awesome-skills

    A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Official

    Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified.

    40k GitHub stars~6.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from sharpdeveye/maestro

All 25 skills in this repo
  • Accelerate

    sharpdeveye/maestro

    A skill your agent uses when the workflow is too slow, too expensive, or both and needs latency, cost, or token usage optimization.

    592 GitHub stars~745 tokensUpdated 5 mo ago
    Auto-check passed
  • Chain

    sharpdeveye/maestro

    A skill your agent uses when the workflow needs multi-step processing with sequential, parallel, or conditional tool compositions and proper data flow.

    592 GitHub stars~607 tokensUpdated 5 mo ago
    Auto-check passed
  • Compose

    sharpdeveye/maestro

    A skill your agent uses when a single agent demonstrably cannot handle the task and multi-agent coordination is justified.

    592 GitHub stars~720 tokensUpdated 5 mo ago
    Auto-check passed
  • Diagnose

    sharpdeveye/maestro

    A skill your agent uses when the user wants to find problems, audit workflow quality, or get a comprehensive health check on their AI workflow.

    592 GitHub stars~1.5k tokensUpdated 5 mo ago
    Auto-check passed
  • Extract Pattern

    sharpdeveye/maestro

    A skill your agent uses when the user wants to create templates, extract reusable patterns, document solutions, or build a pattern library from working workflows.

    592 GitHub stars~664 tokensUpdated 5 mo ago
    Auto-check passed
  • Fortify

    sharpdeveye/maestro

    A skill your agent uses when the workflow lacks error handling, has been failing in production, or needs retry logic, fallback strategies, and circuit breakers.

    592 GitHub stars~688 tokensUpdated 5 mo ago
    Auto-check passed

Questions about Evaluate

What does Evaluate do?

A skill your agent uses when the user wants a quality review, interaction audit, or to test the workflow against realistic scenarios. Evaluate is an agent skill from sharpdeveye/maestro. Use when the user wants a quality review, interaction audit, or to test the workflow against realistic scenarios.

When should I use Evaluate?

Evaluate fits situations like: the user wants a quality review; interaction audit; test the workflow against realistic scenarios.

How do I install Evaluate in Claude Code?

Run `npx skills add sharpdeveye/maestro --skill evaluate -a claude-code`. Or copy the skill folder (source/skills/evaluate in sharpdeveye/maestro) into .claude/skills/evaluate in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate in Codex?

Run `npx skills add sharpdeveye/maestro --skill evaluate -a codex`. Or copy the skill folder (source/skills/evaluate in sharpdeveye/maestro) into .agents/skills/evaluate in your project. Codex loads it when a task matches its description.

Can I use Evaluate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sharpdeveye/maestro --skill evaluate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate, .gemini/skills/evaluate, .github/skills/evaluate and .opencode/skills/evaluate in your project.

What does Evaluate need to run?

SKILL.md names no scripts, command-line tools or credentials: Evaluate is instructions for the agent only.

Does Evaluate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluate use?

Evaluate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate use?

About 776 tokens (SKILL.md is roughly 3.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate?

Skills that share tags, products or a category with Evaluate: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Evaluators (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate?

sharpdeveye (a GitHub user) maintains it in sharpdeveye/maestro, which has 592 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on April 29, 2026.

Source: sharpdeveye/maestro on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.