Agent skill

Anomalib Benchmarking

by open-edge-platform in open-edge-platform/anomalib

Runs the anomalib benchmarking pipeline to train/evaluate a grid of model + dataset (+ category) combinations and collect metrics into a results CSV.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Anomalib Benchmarking

skills CLI
$ npx skills add open-edge-platform/anomalib --skill anomalib-benchmarking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-edge-platform/anomalib anomalib-benchmarking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-edge-platform/anomalib.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/anomalib-benchmarking .claude/skills/anomalib-benchmarking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
anomalib-benchmarking
GitHub stars
6.2k
Token cost
~1k tokens
SKILL.md length
352 words
Files
2
Skills in repo
17
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs the anomalib benchmarking pipeline to train/evaluate a grid of model + dataset (+ category) combinations and collect metrics into a results CSV.

  • Comparing multiple models/datasets/categories in one sweep
  • SKILL.md covers Code locations, Running it, Config structure and Where results go, plus 2 more sections
  • Calls python
  • Authoring/editing a benchmark config YAML

What it does

Anomalib Benchmarking is an agent skill from open-edge-platform/anomalib. Runs the anomalib benchmarking pipeline to train/evaluate a grid of model + dataset (+ category) combinations and collect metrics into a results CSV. Use when comparing multiple models/datasets/categories in one sweep, or authoring/editing a benchmark config YAML. Do not use for training a single model (see anomalib-training) or the tiled-ensemble pipeline (see anomalib-tiled-ensemble). For turning measured results into README/docs benchmark tables, see the benchmark-and-docs-refresh skill.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/evals.json`).

It sits in AI & LLM Engineering, covering CSV and tabular files. The repository describes itself as: An anomaly detection library comprising state-of-the-art algorithms and features such as experiment management, hyper-parameter optimization, and edge inference. The licence is Apache-2.0.

When your agent uses it

  • Comparing multiple models/datasets/categories in one sweep
  • Authoring/editing a benchmark config YAML
  • Training a single model (see anomalib-training)
  • The tiled-ensemble pipeline (see anomalib-tiled-ensemble)

Example prompts

  • “/anomalib-benchmarking”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit dc087d5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Anomalib Benchmarking loads about 1k tokens when it runs. Until then it costs about 129 tokens; SKILL.md has 352 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~129
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-edge-platform/anomalib at commit dc087d5, republished under its Apache-2.0 licence (© open-edge-platform). 352 words, ~1,046 tokens.

Download SKILL.mdSave it as .claude/skills/anomalib-benchmarking/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
anomalib-benchmarking
description
Runs the anomalib benchmarking pipeline to train/evaluate a grid of model + dataset (+ category) combinations and collect metrics into a results CSV. Use when comparing multiple models/datasets/categories in one sweep, or authoring/editing a benchmark config YAML. Do not use for training a single model (see anomalib-training) or the tiled-ensemble pipeline (see anomalib-tiled-ensemble). For turning measured results into README/docs benchmark tables, see the benchmark-and-docs-refresh skill.
license
Apache-2.0

Using the Benchmarking Pipeline

The benchmarking pipeline runs a grid of model/dataset/category combinations end-to-end (train + test) and writes measured metrics to a CSV — use it to produce real, reproducible numbers rather than hand-editing benchmark tables.

Code locations

  • src/anomalib/pipelines/benchmark/pipeline.py — Benchmark: top-level pipeline; picks SerialRunner or ParallelRunner based on configured accelerators and torch.cuda.device_count().
  • src/anomalib/pipelines/benchmark/generator.py — BenchmarkJobGenerator: expands the config (including grid: entries) into individual jobs.
  • src/anomalib/pipelines/benchmark/job.py — BenchmarkJob: runs one model/dataset combination, times it, and saves results.
  • tools/experimental/benchmarking/benchmark.py — thin CLI wrapper around Benchmark.
  • tools/experimental/benchmarking/sample.yaml — example config to copy from.

Running it

bash
# Via the tools wrapper
python tools/experimental/benchmarking/benchmark.py --config tools/experimental/benchmarking/sample.yaml

# Via the anomalib CLI (registered pipeline subcommand)
anomalib benchmark --config tools/experimental/benchmarking/sample.yaml

Config structure

yaml
accelerator:
  - cuda
  - cpu

benchmark:
  seed: 42
  model:
    class_path:
      grid: [Padim, Patchcore]
  data:
    class_path: MVTecAD
    init_args:
      category:
        grid:
          - bottle
          - capsule

Any field can use grid: [...] to sweep multiple values — the generator produces the Cartesian product of every grid field as separate jobs (here: 2 models × 2 categories = 4 jobs). Non-grid fields are held constant across all jobs. data.class_path / model.class_path follow the same anomalib.data.* / anomalib.models.* resolution as everywhere else in the repo (see anomalib-training).

Where results go

BenchmarkJob.save(...) writes one row per job into:

bash
runs/benchmark/<timestamp>/results.csv

(<timestamp> is generated when results are saved via BenchmarkJob.save(), e.g. 2026-08-24-10_30_00.) Each row includes the model/dataset/category combination and the measured metrics — this is the file to consume when building or refreshing README/docs benchmark tables.

There is also a separate, narrower helper tools/benchmark_mebin.py that writes to results/mebin_benchmark.csv for a specific benchmarking use case — prefer the pipeline above unless you specifically need that script's behavior.

Show full SKILL.md (137 more words)Show less

Gotchas

  • A grid sweep multiplies job count fast — check the Cartesian product size before launching a large sweep (e.g. 5 models × 10 categories = 50 full train+test runs).
  • accelerator: [cuda, cpu] creates one runner per entry, so every model/category combination runs once per accelerator (doubling the total job count). This is not a device-pool selector — if you only want to benchmark on GPU, use accelerator: [cuda].
  • Never hand-write or infer numbers into README/docs benchmark tables — always source them from a results.csv produced by an actual run of this pipeline.

Reviewer / self-check

  • Config's grid fields produce the intended, bounded set of jobs (no accidental huge sweep).
  • model.class_path / data.class_path values resolve to real exported classes.
  • Benchmark run completed and runs/benchmark/<timestamp>/results.csv exists before citing numbers anywhere else.
  • Test reference: tests/integration/pipelines/test_benchmark.py for how the pipeline is invoked programmatically if debugging job generation.

© open-edge-platform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/anomalib-benchmarking of open-edge-platform/anomalib.

  • SKILL.md
  • evals/evals.json

Open the folder on GitHubat commit dc087d5

Compare with similar skills

Anomalib Benchmarking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Anomalib Benchmarking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Anomalib Benchmarking this skillopen-edge-platform/anomalib6.2k—~1kAutomated safety check: PassApache-2.0
Perforatedai AnalyzePerforatedAI/PerforatedAI237—~5.1kAutomated safety check: PassApache-2.0
Codemie Analyticscodemie-ai/codemie-code294—~7.5kAutomated safety check: PassApache-2.0
Perforatedai PlotPerforatedAI/PerforatedAI237—~1.6kAutomated safety check: PassApache-2.0
Yolo Trainingfcakyon/claude-codex-settings1.2k—~1.4kAutomated safety check: PassApache-2.0
Extracting Clinical Entitiesmaziyarpanahi/openmed5.5k—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Perforatedai Analyze

    PerforatedAI/PerforatedAI

    Analyze PerforatedAI training results and provide optimization recommendations.

    237 GitHub stars~5.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Codemie Analytics

    codemie-ai/codemie-code

    CodeMie Analytics expert — use this skill whenever the user asks about CodeMie usage data, AI adoption metrics, user leaderboards, CLI insights, spending, LiteLLM costs, token usage, or wants to…

    294 GitHub stars~7.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Perforatedai Plot

    PerforatedAI/PerforatedAI

    Render a single-panel PAI figure of score versus parameter count from sweep CSVs, PAI run folders, or hand-supplied numbers.

    237 GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Yolo Training

    fcakyon/claude-codex-settings

    This skill should be used when user asks to "improve my mAP", "why is my model overfitting", "my training is diverging", "read my results.csv", "interpret my training curves", "my AP50 is good but…

    1.2k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Extracting Clinical Entities

    maziyarpanahi/openmed

    Run clinical and biomedical named-entity recognition on medical text with OpenMed's analyzetext.

    5.5k GitHub stars~1.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Weaviate

    weaviate/agent-skills

    Official

    Search, query, and manage Weaviate vector database collections.

    105 GitHub stars~1.8k tokensUpdated 2 days ago
    DatabasesAuto-check passed

More from open-edge-platform/anomalib

All 17 skills in this repo
  • Anomalib Adding A Model

    open-edge-platform/anomalib

    Adds a new anomaly-detection model to anomalib under src/anomalib/models/.

    6.2k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Anomalib Tiled Ensemble

    open-edge-platform/anomalib

    Runs and configures the anomalib tiled-ensemble pipeline, which trains/evaluates one model per image tile and merges results (with optional seam smoothing) for high-resolution anomaly detection.

    6.2k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Anomalib Training

    open-edge-platform/anomalib

    Trains an anomalib model on a dataset via the Python API or CLI, including training on a custom folder-structured dataset with the Folder datamodule.

    6.2k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Model Sample Image Export

    open-edge-platform/anomalib

    Export, validate, and publish model sample-result images into docs/source/images and reference them from README/docs pages.

    6.2k GitHub stars~874 tokensUpdated today
    Auto-check passed
  • UI Test Utils

    open-edge-platform/anomalib

    A skill your agent uses when writing or updating Anomalib Studio UI component or hook tests that need the shared render/renderHook helpers, React Router paths or parameters, React Query, theme…

    6.2k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Anomalib Adding A Datamodule

    open-edge-platform/anomalib

    Adds a new dataset/datamodule to anomalib under src/anomalib/data/.

    6.2k GitHub stars~5.1k tokensUpdated today
    Auto-check passed

Questions about Anomalib Benchmarking

What does Anomalib Benchmarking do?

Runs the anomalib benchmarking pipeline to train/evaluate a grid of model + dataset (+ category) combinations and collect metrics into a results CSV. Anomalib Benchmarking is an agent skill from open-edge-platform/anomalib. Runs the anomalib benchmarking pipeline to train/evaluate a grid of model + dataset (+ category) combinations and collect metrics into a results CSV.

When should I use Anomalib Benchmarking?

Anomalib Benchmarking fits situations like: comparing multiple models/datasets/categories in one sweep; authoring/editing a benchmark config YAML; training a single model (see anomalib-training); the tiled-ensemble pipeline (see anomalib-tiled-ensemble).

How do I install Anomalib Benchmarking in Claude Code?

Run `npx skills add open-edge-platform/anomalib --skill anomalib-benchmarking -a claude-code`. Or copy the skill folder (.agents/skills/anomalib-benchmarking in open-edge-platform/anomalib) into .claude/skills/anomalib-benchmarking in your project. Claude Code loads it when a task matches its description.

How do I install Anomalib Benchmarking in Codex?

Run `npx skills add open-edge-platform/anomalib --skill anomalib-benchmarking -a codex`. Or copy the skill folder (.agents/skills/anomalib-benchmarking in open-edge-platform/anomalib) into .agents/skills/anomalib-benchmarking in your project. Codex loads it when a task matches its description.

Can I use Anomalib Benchmarking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-edge-platform/anomalib --skill anomalib-benchmarking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/anomalib-benchmarking, .gemini/skills/anomalib-benchmarking, .github/skills/anomalib-benchmarking and .opencode/skills/anomalib-benchmarking in your project.

What does Anomalib Benchmarking need to run?

Going by SKILL.md and its folder, Anomalib Benchmarking needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Anomalib Benchmarking access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Anomalib Benchmarking safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Anomalib Benchmarking use?

Anomalib Benchmarking is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Anomalib Benchmarking use?

About 1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Anomalib Benchmarking?

Skills that share tags, products or a category with Anomalib Benchmarking: Perforatedai Analyze (PerforatedAI/PerforatedAI, 237 stars), Codemie Analytics (codemie-ai/codemie-code, 294 stars), Perforatedai Plot (PerforatedAI/PerforatedAI, 237 stars) and Yolo Training (fcakyon/claude-codex-settings, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Anomalib Benchmarking?

open-edge-platform (a GitHub organization) maintains it in open-edge-platform/anomalib, which has 6,230 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 8, 2026.

Source: open-edge-platform/anomalib on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.