Agent skill

Report

by evo-hq in evo-hq/evo

Read-only evo run reporting. An agent skill from evo-hq/evo.

Apache-2.0Auto-check passedAgent Workflows

Install Report

skills CLI
$ npx skills add evo-hq/evo --skill report -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install evo-hq/evo report --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/evo-hq/evo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/evo/skills/report .claude/skills/report && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
report
GitHub stars
1.5k
Used in
1 other repo
Token cost
~1k tokens
SKILL.md length
530 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Read-only evo run reporting. An agent skill from evo-hq/evo.

  • Works in 4 steps: Run evo status, evo frontier, and evo… → Use evo show for the best node and any… → Use evo diff only to explain what… → …
  • The user invokes /evo:report
  • SKILL.md covers What it shows, How to invoke, When not to use and Overnight / Improvement Reports
  • Calls python

What it does

Report is an agent skill from evo-hq/evo. Read-only evo run reporting. Use when the user invokes /evo:report, asks what happened overnight, asks what improved recently, asks for the best/frontier candidates, asks for a quick score chart without opening the dashboard, or wants the scatter plot in chat output. Never run benchmarks, gates, Slurm commands, evo run, or ad-hoc verification scripts for report requests.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows. The repository describes itself as: turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents. The licence is Apache-2.0.

When your agent uses it

  • The user invokes /evo:report
  • Asks what happened overnight
  • Asks what improved recently
  • Asks for the best/frontier candidates

Example prompts

  • “/report”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Run evo status, evo frontier, and evo tree.
  2. Use evo show for the best node and any recent committed/evaluated
  3. Use evo diff only to explain what changed in a recorded experiment.
  4. If you need benchmark details, read the existing outcome.json,

What it can do on your machine

Read from SKILL.md and the folder at commit c70c04b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Report loads about 1k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 530 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from evo-hq/evo at commit c70c04b, republished under its Apache-2.0 licence (© evo-hq). 530 words, ~1,005 tokens.

Download SKILL.mdSave it as .claude/skills/report/SKILL.md (or your agent's skills folder).
name
report
description
Read-only evo run reporting. Use when the user invokes /evo:report, asks what happened overnight, asks what improved recently, asks for the best/frontier candidates, asks for a quick score chart without opening the dashboard, or wants the scatter plot in chat output. Never run benchmarks, gates, Slurm commands, evo run, or ad-hoc verification scripts for report requests.
evo_version
0.8.0

Report

Report the current evo workspace from recorded state only. A report request is read-only, even if the user phrases it casually as "what happened?", "what got better?", "what should I pay attention to?", or "I just woke up".

Do not spend compute while reporting:

  • Do not run evo run, evo gate check, benchmark commands, or project eval scripts.
  • Do not run python bench.py, python slurm_eval.py, sbatch, srun, squeue, sacct, or scancel to verify a result.
  • Do not create launcher, monitor, parsing, or analysis scripts.
  • Do not edit files.

Use stored evo state instead: evo report, evo status, evo tree, evo frontier, evo show <id>, evo diff <id>, and immutable artifacts under .evo/run_*/experiments/<exp>/attempts/<NNN>/.

For chart requests, render the dashboard's scatter plot as a colored terminal block, one chart per run, sized to the current terminal.

What it shows

Mirrors the web dashboard's score scatter (left rail of evo dashboard):

  • X = experiment creation order, Y = score
  • Dot color by status: green = committed valid result, red = failed, purple = active, grey = pending / evaluated / discarded / pruned
  • ★ marks the current best valid committed-result experiment. pruned with prune_kind=exhausted can still be best; prune_kind=invalid and its descendants cannot.
  • Yellow ring on dots that sit on the best-path spine (root → best)
  • Yellow stair line traces cumulative-best across valid committed-result experiments
  • ○ at the baseline for experiments that have no score yet (active / pending)

Every run in the workspace is rendered, stacked top-to-bottom, with a header line showing run_id · target · metric.

How to invoke

Run:

bash
evo report

That is it. Print the output verbatim in your reply so the user sees the chart. Do not summarize the chart in prose — the visual is the point.

Flags:

  • --color always|never|auto — force or suppress ANSI color. Default auto (color when stdout is a TTY). Pass --color always if you are piping through a host that strips TTY but renders ANSI in chat.
  • --watch [SECONDS] — live-refresh mode (like nvidia-smi -l). Re-reads the workspace every N seconds (default 2) and redraws in place. Ctrl-C to exit. Use this when you want to babysit a running optimization without manually re-invoking the report.
Show full SKILL.md (186 more words)Show less

When not to use

  • For one-off score lookups, evo status or evo show <id> is faster.
  • For navigating the tree shape, evo tree is the right command.
  • For interactive exploration (click a dot, open a drawer), point the user at evo dashboard instead.

Overnight / Improvement Reports

When the user asks what happened recently or what improved, summarize from recorded evo state:

  1. Run evo status, evo frontier, and evo tree.
  2. Use evo show <id> for the best node and any recent committed/evaluated nodes you mention.
  3. Use evo diff <id> only to explain what changed in a recorded experiment.
  4. If you need benchmark details, read the existing outcome.json, benchmark.log, or declared artifacts for that experiment. Treat missing artifacts as "not recorded", not as permission to rerun.

Report:

  • best current experiment and score;
  • score delta versus baseline or parent;
  • top candidates/frontier if relevant;
  • failed/evaluated nodes that need attention;
  • any caveats about gates, missing held-out checks, or tied candidates.

If the user wants fresh validation or reruns, ask them to explicitly start a new optimization or evaluation command. Do not infer that from a report request.

© evo-hq, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/evo/skills/report of evo-hq/evo.

Open the folder on GitHubat commit c70c04b

Used in 1 other repository

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in evo-hq/evo, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Report next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Report compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Report this skillevo-hq/evo1.5k1 repos~1kAutomated safety check: PassApache-2.0
MCP Server Builderanthropics/skills180k64 repos~2.3kAutomated safety check: PassApache-2.0
Hook Development for Claude Code Pluginsanthropics/claude-plugins-official38k11 repos~4.1kAutomated safety check: NotesApache-2.0
Using Superpowersfarm-fe/farm5.6k35 repos~1.4kAutomated safety check: PassMIT
Executing Plans Inlineobra/superpowers296k2 repos~5.1kAutomated safety check: PassMIT
Claude Code Agent Developmentanthropics/claude-plugins-official38k8 repos~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Hook Development for Claude Code Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.

    38k GitHub starsUsed in 11 repos~4.1k tokens
    Agent WorkflowsAuto-check: notes
  • Using Superpowers

    farm-fe/farm

    A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions

    5.6k GitHub starsUsed in 35 repos~1.4k tokens
    Agent WorkflowsAuto-check passed
  • Executing Plans Inline

    obra/superpowers

    Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.

    296k GitHub starsUsed in 2 repos~5.1k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 8 repos~2.8k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    795 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed

More from evo-hq/evo

  • Discover

    evo-hq/evo

    Initialize evo for the current repository by exploring the codebase, proposing unexplored optimization dimensions, constructing the benchmark inside a baseline worktree, and running the first…

    1.5k GitHub stars~12k tokensUpdated 3 days ago
    Auto-check: notes
  • Finetuning

    evo-hq/evo

    This skill should be used when picking or diagnosing a training move (SFT, LoRA, DPO/KTO/ORPO, RFT, GRPO/PPO/RLOO, RLHF), or when the user mentions fine-tuning, post-training, training recipe…

    1.5k GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check passed
  • Infra Setup

    evo-hq/evo

    Non-user-invocable provider/setup reference for evo backend switching, prerequisite checks, and auth/install guidance.

    1.5k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Ship

    evo-hq/evo

    Land the winning experiment from an evo run as a clean, mergeable change -- open a PR when the repo has a remote, otherwise merge into the working branch.

    1.5k GitHub stars~1.7k tokensUpdated 3 days ago
    Auto-check passed
  • Subagent

    evo-hq/evo

    Protocol that evo optimization subagents follow when dispatched from /optimize.

    1.5k GitHub starsUsed in 1 repo~7.1k tokens
    Auto-check: notes
  • Optimize

    evo-hq/evo

    Drive structured autoresearch iteration after evo:discover and the baseline commit.

    1.5k GitHub stars~13k tokensUpdated 3 days ago
    Auto-check passed

Categories

Questions about Report

What does Report do?

Read-only evo run reporting. An agent skill from evo-hq/evo. Report is an agent skill from evo-hq/evo. Read-only evo run reporting.

When should I use Report?

Report fits situations like: the user invokes /evo:report; asks what happened overnight; asks what improved recently; asks for the best/frontier candidates.

How do I install Report in Claude Code?

Run `npx skills add evo-hq/evo --skill report -a claude-code`. Or copy the skill folder (plugins/evo/skills/report in evo-hq/evo) into .claude/skills/report in your project. Claude Code loads it when a task matches its description.

How do I install Report in Codex?

Run `npx skills add evo-hq/evo --skill report -a codex`. Or copy the skill folder (plugins/evo/skills/report in evo-hq/evo) into .agents/skills/report in your project. Codex loads it when a task matches its description.

Can I use Report in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add evo-hq/evo --skill report -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/report, .gemini/skills/report, .github/skills/report and .opencode/skills/report in your project.

What does Report need to run?

Going by SKILL.md and its folder, Report needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Report access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Report safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Report use?

Report is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Report use?

About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Report?

Skills that share tags, products or a category with Report: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 296k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Report?

evo-hq (a GitHub organization) maintains it in evo-hq/evo, which has 1,466 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 5, 2026.

Source: evo-hq/evo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.