Measure AlbumentationsX runtime changes with paired baseline and candidate benchmarks on the affected routes.

AGPL-3.0Auto-check passedAI & LLM Engineering

Install Benchmark

skills CLI
$ npx skills add albumentations-team/AlbumentationsX --skill benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install albumentations-team/AlbumentationsX benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/albumentations-team/AlbumentationsX.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/benchmark .claude/skills/benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark
GitHub stars
567
Token cost
~900 tokens
SKILL.md length
467 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Measure AlbumentationsX runtime changes with paired baseline and candidate benchmarks on the affected routes.

  • Works in 5 steps: Run baseline and candidate on the same… → Use at least 100 iterations for fast… → Verify correctness, seeded behavior, and… → …
  • Tasks that involve Computer vision
  • SKILL.md covers Select the workload, Use the existing catalog and Compare and report
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Benchmark is an agent skill from albumentations-team/AlbumentationsX. Measure AlbumentationsX runtime changes with paired baseline and candidate benchmarks on the affected routes.

Its SKILL.md is about 900 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Computer vision and Deep learning. The repository describes itself as: Image augmentation for computer vision. AGPL-3.0-only or commercial licensing. The licence is AGPL-3.0.

When your agent uses it

  • Tasks that involve Computer vision
  • Tasks that involve Deep learning

Example prompts

  • “/benchmark”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Run baseline and candidate on the same machine and environment, back-to-back, with warm-up and exactly one
  2. Use at least 100 iterations for fast functions. For slow functions, choose enough repetitions for stable timing,
  3. Verify correctness, seeded behavior, and aliasing before accepting a faster path.
  4. Report every before/after cell, speedup, exact revisions, command or ASV filter, environment, and rejected candidates.
  5. Investigate any regression above 5%; rework it or explain the measured trade-off.

What it can do on your machine

Read from SKILL.md and the folder at commit 1458043. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark loads about 900 tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 467 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~900

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from albumentations-team/AlbumentationsX at commit 1458043, republished under its AGPL-3.0 licence (© albumentations-team). 467 words, ~900 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark/SKILL.md (or your agent's skills folder).
name
benchmark
description
Measure AlbumentationsX runtime changes with paired baseline and candidate benchmarks on the affected routes.

Benchmark a runtime change

Benchmark when a change can alter executed work in a functional operation, sampling, target dispatch, or Compose. Documentation and mechanical refactors with no plausible runtime effect need no timing run.

Read Performance Optimization and its required reference before selecting candidates. Define the changed route and the workload dimension that could falsify the proposed speedup.

Select the workload

For pixel arithmetic, dtype conversion, layout, or backend routing, use the nine combinations of square sizes 256, 512, and 1024 with 1, 3, and 5 channels. Skip explicitly unsupported channel counts. Keep grayscale NumPy inputs (H, W, 1).

For other changes, vary the controlling axis: label count and density, random-output size, annotation count, or both sides of a routing threshold. Explain why the chosen matrix covers the claim.

  • Measure the direct functional call and the public Compose route when both are affected.
  • For batch changes, compare per-image dispatch with the batch route at batch sizes 4, 8, and 16.
  • For dtype routing, measure uint8 and float32. A change confined to one dtype may time that dtype and use correctness tests to verify the other dtype's wrapper round-trip.
  • For Compose changes, include root skip, no-op, probabilistic no-op, an always-applied cheap leaf, applied-parameter capture, trace, Tensor, processors, and concurrent calls. For annotation changes, include empty, single, and dense annotations alongside image-only inputs.
  • Construction time and retained allocations are optional context when RNG or graph ownership changes. Judge the result by repeated-call performance; a measured call-time gain can justify slower construction.
Show full SKILL.md (218 more words)Show less

Use the existing catalog

Start with benchmark/README.md for ASV commands and Performance Coverage for case ownership.

release-core supplies the fixed weekly and release profile. changed selects affected families for requested PR comparisons. Resolve these profiles with tools/select_benchmark_filters.py. Use ASV continuous with explicit baseline and candidate refs. An empty selection does not authorize a full-catalog run; choose an explicit regex for a larger investigation.

If the catalog cannot express the question, keep a focused local measurement under _internal/. Save matching before/after cells with shape, dtype, channels, parameters, allocation mode, elapsed time, and iteration count.

Compare and report

  1. Run baseline and candidate on the same machine and environment, back-to-back, with warm-up and exactly one CPU thread per process. Follow the thread controls in the required performance guide. Benchmark additional thread counts only when the user explicitly requests a thread-scaling experiment.
  2. Use at least 100 iterations for fast functions. For slow functions, choose enough repetitions for stable timing, aiming for more than one second per cell.
  3. Verify correctness, seeded behavior, and aliasing before accepting a faster path.
  4. Report every before/after cell, speedup, exact revisions, command or ASV filter, environment, and rejected candidates.
  5. Investigate any regression above 5%; rework it or explain the measured trade-off.

A one-revision timing is a baseline observation, not evidence of a speedup.

© albumentations-team, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .codex/skills/benchmark of albumentations-team/AlbumentationsX.

Open the folder on GitHubat commit 1458043

Compare with similar skills

Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark this skillalbumentations-team/AlbumentationsX567—~900Automated safety check: PassAGPL-3.0
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
Kaiming HeK-Dense-AI/mimeo282—~1.6kAutomated safety check: PassMIT
Matlab Analyze Spectral Imagesmatlab/matlab-agentic-toolkit1.1k—~3.7kAutomated safety check: PassCustom licence
Matlab Process Imagesmatlab/matlab-agentic-toolkit1.1k—~3.7kAutomated safety check: PassCustom licence

Similar skills

  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Kaiming He

    K-Dense-AI/mimeo

    Applies the reasoning style of Kaiming He, computer vision pioneer and creator of ResNet.

    282 GitHub stars~1.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Matlab Analyze Spectral Images

    matlab/matlab-agentic-toolkit

    Work with hyperspectral and multispectral images in MATLAB. An agent skill from matlab/matlab-agentic-toolkit.

    1.1k GitHub stars~3.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Matlab Process Images

    matlab/matlab-agentic-toolkit

    Load this first for any task involving images, pictures, photos, scans, frames, volumes, or visual data — including reading, writing, filtering, enhancing, denoising, sharpening, deblurring…

    1.1k GitHub stars~3.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Scholar Compute

    joshzyj/open-scholar-skill

    Design and execute computational social science analyses across 11 modules: text-as-data/NLP (STM, BERTopic, Wordfish, BERT, conText embedding regression, LLM annotation + DSL bias correction…

    168 GitHub stars~15k tokensUpdated 19 days ago
    AI & LLM EngineeringAuto-check passed

More from albumentations-team/AlbumentationsX

All 10 skills in this repo
  • Mixing Transforms

    albumentations-team/AlbumentationsX

    Policy for AlbumentationsX transforms that combine multiple images or objects.

    567 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Release Notes

    albumentations-team/AlbumentationsX

    Generate release notes for AlbumentationsX. An agent skill from albumentations-team/AlbumentationsX.

    567 GitHub stars~685 tokensUpdated today
    Auto-check passed
  • Performance Optimization

    albumentations-team/AlbumentationsX

    Systematic performance audit for AlbumentationsX runtime code.

    567 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Add Transform

    albumentations-team/AlbumentationsX

    Full checklist for adding a new transform to AlbumentationsX.

    567 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Docstring Deep Dive

    albumentations-team/AlbumentationsX

    Review public AlbumentationsX docstrings for useful descriptions, runnable examples, parameter semantics, and related transforms.

    567 GitHub stars~534 tokensUpdated today
    Auto-check passed
  • License Integrity

    albumentations-team/AlbumentationsX

    Maintain AlbumentationsX license, CLA, provenance notices, and packaged legal artifacts consistently.

    567 GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Questions about Benchmark

What does Benchmark do?

Measure AlbumentationsX runtime changes with paired baseline and candidate benchmarks on the affected routes. Benchmark is an agent skill from albumentations-team/AlbumentationsX. Measure AlbumentationsX runtime changes with paired baseline and candidate benchmarks on the affected routes.

When should I use Benchmark?

Benchmark fits situations like: tasks that involve Computer vision; tasks that involve Deep learning.

How do I install Benchmark in Claude Code?

Run `npx skills add albumentations-team/AlbumentationsX --skill benchmark -a claude-code`. Or copy the skill folder (.codex/skills/benchmark in albumentations-team/AlbumentationsX) into .claude/skills/benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark in Codex?

Run `npx skills add albumentations-team/AlbumentationsX --skill benchmark -a codex`. Or copy the skill folder (.codex/skills/benchmark in albumentations-team/AlbumentationsX) into .agents/skills/benchmark in your project. Codex loads it when a task matches its description.

Can I use Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add albumentations-team/AlbumentationsX --skill benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark, .gemini/skills/benchmark, .github/skills/benchmark and .opencode/skills/benchmark in your project.

What does Benchmark need to run?

SKILL.md names no scripts, command-line tools or credentials: Benchmark is instructions for the agent only.

Does Benchmark access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark use?

Benchmark is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark use?

About 900 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark?

Skills that share tags, products or a category with Benchmark: Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Kaiming He (K-Dense-AI/mimeo, 282 stars) and Matlab Analyze Spectral Images (matlab/matlab-agentic-toolkit, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark?

albumentations-team (a GitHub organization) maintains it in albumentations-team/AlbumentationsX, which has 567 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 7, 2026.

Source: albumentations-team/AlbumentationsX on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.