Agent skill

Benchmark And Docs Refresh

by open-edge-platform in open-edge-platform/anomalib

Run or continue model benchmarks, collect measured results, and refresh README/docs benchmark sections from generated artifacts.

Apache-2.0Auto-check passedDevelopment

Install Benchmark And Docs Refresh

skills CLI
$ npx skills add open-edge-platform/anomalib --skill benchmark-and-docs-refresh -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-edge-platform/anomalib benchmark-and-docs-refresh --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-edge-platform/anomalib.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchmark-and-docs-refresh .claude/skills/benchmark-and-docs-refresh && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-and-docs-refresh
GitHub stars
6.2k
Token cost
~808 tokens
SKILL.md length
392 words
Files
1
Skills in repo
17
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run or continue model benchmarks, collect measured results, and refresh README/docs benchmark sections from generated artifacts.

  • Works in 3 steps: derive a small helper script from the… → keep it model-specific unless multiple… → save measurable outputs such as CSV…
  • Benchmark tables in model docs need to be created
  • SKILL.md covers Scope, Request changes when, Preferred Benchmark Workflow and Required Evidence, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Benchmark And Docs Refresh is an agent skill from open-edge-platform/anomalib. Run or continue model benchmarks, collect measured results, and refresh README/docs benchmark sections from generated artifacts. Use when benchmark tables in model docs need to be created, updated, or corrected.

Its SKILL.md is about 810 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Technical documentation. The repository describes itself as: An anomaly detection library comprising state-of-the-art algorithms and features such as experiment management, hyper-parameter optimization, and edge inference. The licence is Apache-2.0.

When your agent uses it

  • Benchmark tables in model docs need to be created
  • Tasks that involve Technical documentation

Example prompts

  • “/benchmark-and-docs-refresh”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. derive a small helper script from the benchmark workflow
  2. keep it model-specific unless multiple models clearly need the same pattern
  3. save measurable outputs such as CSV files under results/

What it can do on your machine

Read from SKILL.md and the folder at commit 6f54715. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark And Docs Refresh loads about 808 tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 392 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~808

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-edge-platform/anomalib at commit 6f54715, republished under its Apache-2.0 licence (© open-edge-platform). 392 words, ~808 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-and-docs-refresh/SKILL.md (or your agent's skills folder).
name
benchmark-and-docs-refresh
description
Run or continue model benchmarks, collect measured results, and refresh README/docs benchmark sections from generated artifacts. Use when benchmark tables in model docs need to be created, updated, or corrected.

Benchmark and Docs Refresh

Use this skill to update benchmark sections in model documentation from real benchmark outputs.

Scope

This skill focuses on:

  • running or continuing benchmarks
  • collecting benchmark CSV results from results/
  • updating benchmark tables in model READMEs
  • updating matching docs pages when benchmark status changes

It does not own sample image export. Use model-sample-image-export for that.

Request changes when

  • incomplete benchmark coverage is presented;
  • README or docs benchmark status drifts from the actual run state.

Preferred Benchmark Workflow

Always prefer:

  • tools/experimental/benchmarking/benchmark.py

with an appropriate config file.

If the stock benchmark path is insufficient for a specific model:

  1. derive a small helper script from the benchmark workflow
  2. keep it model-specific unless multiple models clearly need the same pattern
  3. save measurable outputs such as CSV files under results/

Required Evidence

Only publish benchmark values when they come from actual artifacts, for example:

  • results/<model>_benchmark.csv
  • benchmark-generated CSV files under runs/ or results/
  • model-specific run outputs that clearly record the measured metrics

Never infer missing values.

Update Rules

When refreshing benchmark tables:

  1. Read the target README and matching docs page first.
  2. Read the benchmark artifact source.
  3. Fill only the shot-settings and metrics that actually exist.
  4. Leave unavailable rows blank or TODO.
  5. Update status wording if the benchmark is still partial or still running.

Table Conventions

Common sections to refresh:

  • ### Image-Level AUC
  • ### Pixel-Level AUC
  • ### Image F1 Score
  • ### Pixel F1 Score

If a README only contains placeholders, replace only the rows supported by measured results.

Show full SKILL.md (144 more words)Show less

Docs Synchronization Rules

If the README benchmark state changes, update the matching docs page under:

  • docs/source/markdown/guides/reference/models/image/<model>.md
  • docs/source/markdown/guides/reference/models/video/<model>.md

The docs page may stay shorter than the README, but it must not contradict it.

Quality Checks

Before finishing:

  1. Confirm the benchmark artifact still exists.
  2. Confirm copied values exactly match the artifact.
  3. Confirm averages are computed from measured values only.
  4. Confirm incomplete rows remain clearly incomplete.
  5. Confirm README/docs wording matches reality.

Reviewer checklist

  • Check that the artifact exists.
  • Check that every copied value matches.
  • Check that partial runs are labeled clearly.
  • Check README and docs wording for consistency.

Repo-Specific Notes

  • Some benchmark jobs in this repo may require derived helper scripts.
  • Some long runs are better continued in tmux/background sessions.
  • A benchmark can be complete enough to fill a subset of rows without justifying all rows.
  • Never replace TODOs with fabricated numbers.

© open-edge-platform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/benchmark-and-docs-refresh of open-edge-platform/anomalib.

Open the folder on GitHubat commit 6f54715

Compare with similar skills

Benchmark And Docs Refresh next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark And Docs Refresh compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark And Docs Refresh this skillopen-edge-platform/anomalib6.2k—~808Automated safety check: PassApache-2.0
Get API Docs with chubandrewyng/context-hub14k2 repos~775Automated safety check: PassMIT
Adk Sample Creatorgoogle/adk-python22k—~1.3kAutomated safety check: PassApache-2.0
Create SkillHyk260/PureChat5461 repos~823Automated safety check: PassMIT
Env And Assets Bootstraplllllllama/RigorPilot-Skills4971 repos~592Automated safety check: PassMIT
Glossarylightly-ai/lightly-studio896—~482Automated safety check: PassApache-2.0

Similar skills

  • Get API Docs with chub

    andrewyng/context-hub

    Fetches current documentation for third-party APIs and SDKs with the chub CLI before the agent writes code against them, instead of relying on remembered API shapes.

    14k GitHub starsUsed in 2 repos~775 tokens
    DevelopmentAuto-check passed
  • Adk Sample Creator

    google/adk-python

    Official

    Creates a new sample agent in the ADK Python repository — the sample directory, its agent.py, and its README.md — following the conventions the existing samples already use.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Create Skill

    Hyk260/PureChat

    Create a new skill in the current repository. An agent skill from Hyk260/PureChat.

    546 GitHub starsUsed in 1 repo~823 tokens
    DevelopmentAuto-check passed
  • Env And Assets Bootstrap

    lllllllama/RigorPilot-Skills

    Rigor Setup skill for README-first deep learning repo reproduction.

    497 GitHub starsUsed in 1 repo~592 tokens
    DevelopmentAuto-check passed
  • Glossary

    lightly-ai/lightly-studio

    Read when naming anything user-facing - GUI text, docs, public Python API names, arguments, docstrings, or error messages.

    896 GitHub stars~482 tokensUpdated today
    DevelopmentAuto-check passed
  • Add Model

    georg-wolflein/pathology-foundation-models

    Add a new pathology foundation model to the README (excluding magnification).

    207 GitHub stars~1.7k tokensUpdated 7 mo ago
    DevelopmentAuto-check passed

More from open-edge-platform/anomalib

All 17 skills in this repo
  • Anomalib Adding A Datamodule

    open-edge-platform/anomalib

    Adds a new dataset/datamodule to anomalib under src/anomalib/data/.

    6.2k GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Anomalib Adding A Model

    open-edge-platform/anomalib

    Adds a new anomaly-detection model to anomalib under src/anomalib/models/.

    6.2k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Anomalib Benchmarking

    open-edge-platform/anomalib

    Runs the anomalib benchmarking pipeline to train/evaluate a grid of model + dataset (+ category) combinations and collect metrics into a results CSV.

    6.2k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Anomalib Tiled Ensemble

    open-edge-platform/anomalib

    Runs and configures the anomalib tiled-ensemble pipeline, which trains/evaluates one model per image tile and merges results (with optional seam smoothing) for high-resolution anomaly detection.

    6.2k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Anomalib Training

    open-edge-platform/anomalib

    Trains an anomalib model on a dataset via the Python API or CLI, including training on a custom folder-structured dataset with the Folder datamodule.

    6.2k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Model Sample Image Export

    open-edge-platform/anomalib

    Export, validate, and publish model sample-result images into docs/source/images and reference them from README/docs pages.

    6.2k GitHub stars~874 tokensUpdated today
    Auto-check passed

Questions about Benchmark And Docs Refresh

What does Benchmark And Docs Refresh do?

Run or continue model benchmarks, collect measured results, and refresh README/docs benchmark sections from generated artifacts. Benchmark And Docs Refresh is an agent skill from open-edge-platform/anomalib. Run or continue model benchmarks, collect measured results, and refresh README/docs benchmark sections from generated artifacts.

When should I use Benchmark And Docs Refresh?

Benchmark And Docs Refresh fits situations like: benchmark tables in model docs need to be created; tasks that involve Technical documentation.

How do I install Benchmark And Docs Refresh in Claude Code?

Run `npx skills add open-edge-platform/anomalib --skill benchmark-and-docs-refresh -a claude-code`. Or copy the skill folder (.agents/skills/benchmark-and-docs-refresh in open-edge-platform/anomalib) into .claude/skills/benchmark-and-docs-refresh in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark And Docs Refresh in Codex?

Run `npx skills add open-edge-platform/anomalib --skill benchmark-and-docs-refresh -a codex`. Or copy the skill folder (.agents/skills/benchmark-and-docs-refresh in open-edge-platform/anomalib) into .agents/skills/benchmark-and-docs-refresh in your project. Codex loads it when a task matches its description.

Can I use Benchmark And Docs Refresh in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-edge-platform/anomalib --skill benchmark-and-docs-refresh -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-and-docs-refresh, .gemini/skills/benchmark-and-docs-refresh, .github/skills/benchmark-and-docs-refresh and .opencode/skills/benchmark-and-docs-refresh in your project.

What does Benchmark And Docs Refresh need to run?

SKILL.md names no scripts, command-line tools or credentials: Benchmark And Docs Refresh is instructions for the agent only.

Does Benchmark And Docs Refresh access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark And Docs Refresh safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark And Docs Refresh use?

Benchmark And Docs Refresh is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark And Docs Refresh use?

About 808 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark And Docs Refresh?

Skills that share tags, products or a category with Benchmark And Docs Refresh: Get API Docs with chub (andrewyng/context-hub, 14k stars), Adk Sample Creator (google/adk-python, 22k stars), Create Skill (Hyk260/PureChat, 546 stars) and Env And Assets Bootstrap (lllllllama/RigorPilot-Skills, 497 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark And Docs Refresh?

open-edge-platform (a GitHub organization) maintains it in open-edge-platform/anomalib, which has 6,226 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 7, 2026.

Source: open-edge-platform/anomalib on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.