Agent skill

Offload Benchmark

by imbue-ai in imbue-ai/offload

Run local vs Offload benchmarks for mng and sculptor, then update the Benchmarks section of the Offload README.

MITAuto-check passedTesting & QA

Install Offload Benchmark

skills CLI
$ npx skills add imbue-ai/offload --skill offload-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install imbue-ai/offload offload-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/imbue-ai/offload.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/offload-benchmark .claude/skills/offload-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
offload-benchmark
GitHub stars
125
Token cost
~2.6k tokens
SKILL.md length
988 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Run local vs Offload benchmarks for mng and sculptor, then update the Benchmarks section of the Offload README.

  • Works in 6 steps: Read Current Test Commands → Run Benchmarks → Compute Metrics → …
  • Testing & QA work in your project
  • SKILL.md covers Prerequisites, Step 1: Read Current Test…, Step 2: Run Benchmarks and Step 3: Compute Metrics, plus 4 more sections
  • Calls just, cargo and modal

What it does

Offload Benchmark is an agent skill from imbue-ai/offload. Run local vs Offload benchmarks for mng and sculptor, then update the Benchmarks section of the Offload README.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA. The repository describes itself as: Offload your test computation to ephemeral compute. The licence is MIT.

When your agent uses it

  • Testing & QA work in your project

Example prompts

  • “/offload-benchmark”

Requirements

  • Docker

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Read Current Test Commands
  2. Run Benchmarks
  3. Compute Metrics
  4. Generate the Updated Benchmarks Section
  5. Update the README
  6. Report Results

What it can do on your machine

Read from SKILL.md and the folder at commit d3675da. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • just
    • cargo
    • modal

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Offload Benchmark loads about 2.6k tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 988 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from imbue-ai/offload at commit d3675da, republished under its MIT licence (© imbue-ai). 988 words, ~2,602 tokens.

Download SKILL.mdSave it as .claude/skills/offload-benchmark/SKILL.md (or your agent's skills folder).
name
offload-benchmark
description
Run local vs Offload benchmarks for mng and sculptor, then update the Benchmarks section of the Offload README.

Offload Benchmark

Run local and Offload test suites for mng and sculptor, collect timing data, and update the ## Benchmarks section of the Offload README with fresh numbers.

Prerequisites

  1. Offload must be installed: cargo install offload (or built from source in the offload repo).
  2. Modal must be authenticated: modal token new (if credentials are expired).
  3. Both repos must exist locally:
    • mng: ~/imbue/mng (or wherever the monorepo is checked out)
    • sculptor: ~/imbue/sculptor
  4. Run just install from the offload repo root to ensure the binary is up to date.

Verify all four before proceeding. If any prerequisite is missing, stop and tell the user.

Step 1: Read Current Test Commands

Do not hardcode commands. Read the justfile in each repo at invocation time, since commands change:

  • mng: Read ~/imbue/mng/justfile and locate:
    • test-integration -- the local baseline (unit + integration tests via pytest with xdist)
    • test-offload -- the Offload run on Modal
  • sculptor: Read ~/imbue/sculptor/justfile and locate:
    • test-integration -- the local baseline (Playwright integration tests via pytest with xdist)
    • test-integration-offload -- the Offload run on Modal

Record the exact commands from each justfile. If any target is missing or has changed significantly, stop and tell the user.

Also read each justfile to determine:

  • The default xdist worker count for the local baseline (look for -n auto --maxprocesses N or -n N)
  • The higher xdist worker count to use for a second local run (use double the default, capped at 8)

Step 2: Run Benchmarks

CRITICAL RULES:

  • NEVER run any two benchmark runs concurrently. Each run must have exclusive machine resources. Run them strictly sequentially.
  • Close all other heavy processes (browsers, IDEs, Docker containers) before starting, or at minimum warn the user to do so.
  • Use time (bash builtin) to measure wall-clock time for each run. Wrap each command in time (...) and record the real time.
  • If any tests fail in a run, flag the run as INVALID but still record the timing. Report the failure to the user at the end.
2a: mng benchmarks

Run these three commands sequentially from the mng repo root (~/imbue/mng):

  1. Local baseline (default xdist workers):

    time just test-integration

    Record the wall-clock time.

  2. Local high-xdist (override to higher worker count): Read the default -n value from the justfile. Run with double that value (capped at 8). For example, if the default is -n 4:

    time PYTEST_NUMPROCESSES=8 just test-integration

    If the justfile uses PYTEST_NUMPROCESSES or a similar env var for xdist workers, use that. Otherwise, pass the appropriate flag or env var. Read the justfile carefully to determine the correct override mechanism.

  3. Offload (warm cache): Run the offload command TWICE. Discard the first run's timing (it warms the image cache). Record only the second run's timing.

    just test-offload          # warm-up run (discard timing)
    time just test-offload     # benchmark run (record timing)

Record the number of tests collected (visible in pytest output) and the xdist -n values used.

2b: sculptor benchmarks

Run these three commands sequentially from the sculptor repo root (~/imbue/sculptor):

  1. Local baseline (default xdist workers):

    time just test-integration

    Record the wall-clock time.

  2. Local high-xdist (override to higher worker count): Read the default --maxprocesses value from the justfile (currently 3). Run with a higher value:

    time XDIST_WORKERS=8 just test-integration

    Use the override mechanism from the justfile (XDIST_WORKERS env var).

  3. Offload (warm cache): Run the offload command TWICE. Discard the first run's timing. Record only the second.

    just test-integration-offload          # warm-up run (discard timing)
    time just test-integration-offload     # benchmark run (record timing)

Record the number of tests collected and the xdist -n values used.

Step 3: Compute Metrics

For each project, compute:

MetricFormula
Time (s)Wall-clock seconds from time, rounded to 1 decimal
Time (%)(run_time / baseline_time) * 100, rounded to 1 decimal
Speedupbaseline_time / run_time, rounded to 2 decimals
Bar width (px)round((run_time / baseline_time) * 150) -- baseline is always 150px

The baseline is always the first run (default xdist) for each project.

Show full SKILL.md (388 more words)Show less

Step 4: Generate the Updated Benchmarks Section

Replace the entire ## Benchmarks section in the Offload README (from ## Benchmarks up to but not including the next ## heading) with the template below, filled in with measured values.

Use the correct labels for each project's test suite:

  • sculptor: "Integration Tests (Playwright)"
  • mng: "Unit + Integration Tests"

The bar chart images reference docs/bar-local.svg (gray, #444) and docs/bar-offload.svg (green, #22a355). These are 1x1 SVG rectangles scaled via the width attribute.

Template
markdown
## Benchmarks

Speedups measured on Imbue projects using Offload with the Modal provider. All local baselines were run on a {MACHINE_DESCRIPTION}.

### Sculptor Integration Tests (Playwright)

| Run Kind | Time (s) | Time (%) | Speedup |
|----------|----------|----------|---------|
| pytest with xdist, n={SCULPTOR_BASELINE_N} (baseline) | <img src="docs/bar-local.svg" width="150" height="4"> {SCULPTOR_BASELINE_TIME} | 100.0% | 1.00x |
| pytest with xdist, n={SCULPTOR_HIGH_N} | <img src="docs/bar-local.svg" width="{SCULPTOR_HIGH_BAR_WIDTH}" height="4"> {SCULPTOR_HIGH_TIME} | {SCULPTOR_HIGH_PCT}% | {SCULPTOR_HIGH_SPEEDUP}x |
| Offload (Modal, max {SCULPTOR_MAX_PARALLEL}) | <img src="docs/bar-offload.svg" width="{SCULPTOR_OFFLOAD_BAR_WIDTH}" height="4"> {SCULPTOR_OFFLOAD_TIME} | {SCULPTOR_OFFLOAD_PCT}% | **{SCULPTOR_OFFLOAD_SPEEDUP}x** |

<details>
<summary><strong>Notes</strong></summary>

{SCULPTOR_TEST_COUNT} Playwright integration tests (browser-based, each launching a full Sculptor instance).
Individual tests are heavyweight (Chromium + backend server per worker), so the default xdist cap is n={SCULPTOR_BASELINE_N}.
Offload bypasses xdist entirely, fanning out across up to {SCULPTOR_MAX_PARALLEL} isolated Modal sandboxes -- each running a single test against its own Sculptor instance. The high per-test cost makes Offload's per-sandbox overhead negligible, yielding a {SCULPTOR_OFFLOAD_SPEEDUP}x speedup.

</details>

### Mng Unit + Integration Tests

| Run Kind | Time (s) | Time (%) | Speedup |
|----------|----------|----------|---------|
| pytest with xdist, n={MNG_BASELINE_N} (baseline) | <img src="docs/bar-local.svg" width="150" height="4"> {MNG_BASELINE_TIME} | 100.0% | 1.00x |
| pytest with xdist, n={MNG_HIGH_N} | <img src="docs/bar-local.svg" width="{MNG_HIGH_BAR_WIDTH}" height="4"> {MNG_HIGH_TIME} | {MNG_HIGH_PCT}% | {MNG_HIGH_SPEEDUP}x |
| Offload (Modal, max {MNG_MAX_PARALLEL}) | <img src="docs/bar-offload.svg" width="{MNG_OFFLOAD_BAR_WIDTH}" height="4"> {MNG_OFFLOAD_TIME} | {MNG_OFFLOAD_PCT}% | **{MNG_OFFLOAD_SPEEDUP}x** |

<details>
<summary><strong>Notes</strong></summary>

{MNG_TEST_COUNT} tests collected (unit + integration, excluding acceptance and release).
Individual tests are lightweight and fast-running, so the default xdist cap is n={MNG_BASELINE_N}.
Offload bypasses xdist entirely, fanning out across up to {MNG_MAX_PARALLEL} isolated Modal sandboxes. The low per-test cost makes Offload's per-sandbox overhead proportionally larger, yielding a more modest {MNG_OFFLOAD_SPEEDUP}x speedup vs Sculptor's {SCULPTOR_OFFLOAD_SPEEDUP}x.

</details>
Placeholder Reference
PlaceholderSource
{MACHINE_DESCRIPTION}Ask the user, or read from the existing README intro paragraph
{*_BASELINE_N}Default xdist -n value from the justfile
{*_HIGH_N}The higher xdist count used in run 2
{*_BASELINE_TIME}Wall-clock seconds from run 1
{*_HIGH_TIME}Wall-clock seconds from run 2
{*_OFFLOAD_TIME}Wall-clock seconds from run 3 (second invocation only)
{*_HIGH_PCT}(high_time / baseline_time) * 100
{*_OFFLOAD_PCT}(offload_time / baseline_time) * 100
{*_HIGH_SPEEDUP}baseline_time / high_time
{*_OFFLOAD_SPEEDUP}baseline_time / offload_time
{*_HIGH_BAR_WIDTH}round((high_time / baseline_time) * 150)
{*_OFFLOAD_BAR_WIDTH}round((offload_time / baseline_time) * 150)
{*_MAX_PARALLEL}Read from the relevant offload*.toml config (max_parallel field)
{*_TEST_COUNT}Number of tests collected, from pytest output

Step 5: Update the README

  1. Read offload/README.md.
  2. Find the ## Benchmarks section (starts at ## Benchmarks, ends just before the next ## heading).
  3. Replace that entire section with the generated content from Step 4.
  4. Write the file back.
  5. Verify the replacement by reading the file again and confirming the new numbers appear.

Step 6: Report Results

Summarize to the user:

  1. A table of all six runs with their wall-clock times and pass/fail status.
  2. Any runs flagged as INVALID (test failures).
  3. The computed speedups.
  4. Confirmation that the README was updated (or not, if any run was INVALID -- in that case, ask the user whether to update anyway).

Rules

  • Never run benchmarks concurrently. Sequential execution only.
  • Run Offload twice, take the second time. The first run warms the Modal image cache.
  • If any tests fail, report results but mark as INVALID. Ask the user before updating the README with invalid data.
  • Maintain parallel phrasing in Notes sections. Both projects' Notes blocks should follow the same structure: test count and type, why the default xdist cap is what it is, how Offload bypasses xdist, and why the speedup is high or low.
  • Read justfiles at invocation time. Commands may have changed since this skill was written.
  • Preserve the rest of the README. Only replace the ## Benchmarks section.

© imbue-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/offload-benchmark of imbue-ai/offload.

Open the folder on GitHubat commit d3675da

Compare with similar skills

Offload Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Offload Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Offload Benchmark this skillimbue-ai/offload125—~2.6kAutomated safety check: PassMIT
Pester Failure AnalysisPowerShell/PowerShell56k—~5.1kAutomated safety check: PassMIT
Rust TDD Workflowrtk-ai/rtk83k—~753Automated safety check: NotesApache-2.0
Clawteam DevHKUDS/ClawTeam5.5k1 repos~1.1kAutomated safety check: PassMIT
Browser Testing with Chrome DevToolsaddyosmani/agent-skills102k4 repos~3.5kAutomated safety check: WarnMIT
Apple Container Test RunnerRustPython/RustPython22k—~467Automated safety check: PassMIT

Similar skills

  • Pester Failure Analysis

    PowerShell/PowerShell

    Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations.

    56k GitHub stars~5.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Enforces red-green-refactor for Rust work, with idiomatic test patterns, a naming convention and a pre-commit gate of cargo fmt, clippy and test.

    83k GitHub stars~753 tokensUpdated today
    Testing & QAAuto-check: notes
  • Clawteam Dev

    HKUDS/ClawTeam

    A skill your agent uses when working inside the ClawTeam repository itself: local development, debugging, reviewing, testing, validating multi-agent flows, or checking whether a code change actually…

    5.5k GitHub starsUsed in 1 repo~1.1k tokens
    Testing & QAAuto-check passed
  • Connects an agent to a real Chrome instance through the Chrome DevTools MCP server, so it can inspect the DOM, read console errors and profile performance directly.

    102k GitHub starsUsed in 4 repos~3.5k tokens
    Testing & QAAuto-check: warnings
  • Apple Container Test Runner

    RustPython/RustPython

    Runs RustPython tests inside a Linux container built with Apple's container CLI, so macOS users can compare Linux results with their local ones.

    22k GitHub stars~467 tokensUpdated today
    Testing & QAAuto-check passed
  • Codex Plugin QA

    code-yeongyu/oh-my-openagent

    Tests the omo Codex plugin in an isolated CODEX_HOME with a local mock model, proving hooks fired through app-server notifications without touching ~/.codex.

    70k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed

More from imbue-ai/offload

  • Offload

    imbue-ai/offload

    Activate when you see offload.toml in a repo, offload referenced in build targets (justfile, Makefile, scripts), or when you need to run a large test suite in parallel.

    125 GitHub stars~3.1k tokensUpdated 11 days ago
    Auto-check passed
  • Deburr Edge Cases

    imbue-ai/offload

    Strip AI-codegen cruft from a branch — defensive checks the type system should make unrepresentable, redundant validation, speculative abstraction, no-op wrappers — while preserving load-bearing…

    125 GitHub stars~2.5k tokensUpdated 11 days ago
    Auto-check passed
  • Release To Crates

    imbue-ai/offload

    Cut a release of a Rust crate and publish to crates.io, with externally-gated merge and publish steps.

    125 GitHub stars~2.3k tokensUpdated 11 days ago
    Auto-check passed
  • Offload Onboard

    imbue-ai/offload

    Onboard a repository to use Offload for parallel test execution on Modal.

    125 GitHub stars~7.2k tokensUpdated 11 days ago
    Auto-check: notes

Questions about Offload Benchmark

What does Offload Benchmark do?

Run local vs Offload benchmarks for mng and sculptor, then update the Benchmarks section of the Offload README. Offload Benchmark is an agent skill from imbue-ai/offload. Run local vs Offload benchmarks for mng and sculptor, then update the Benchmarks section of the Offload README.

When should I use Offload Benchmark?

Offload Benchmark fits situations like: testing & QA work in your project.

How do I install Offload Benchmark in Claude Code?

Run `npx skills add imbue-ai/offload --skill offload-benchmark -a claude-code`. Or copy the skill folder (skills/offload-benchmark in imbue-ai/offload) into .claude/skills/offload-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Offload Benchmark in Codex?

Run `npx skills add imbue-ai/offload --skill offload-benchmark -a codex`. Or copy the skill folder (skills/offload-benchmark in imbue-ai/offload) into .agents/skills/offload-benchmark in your project. Codex loads it when a task matches its description.

Can I use Offload Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add imbue-ai/offload --skill offload-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/offload-benchmark, .gemini/skills/offload-benchmark, .github/skills/offload-benchmark and .opencode/skills/offload-benchmark in your project.

What does Offload Benchmark need to run?

Going by SKILL.md and its folder, Offload Benchmark needs the command-line tools its instructions call (just, cargo and modal). Our summary lists: Docker.

Does Offload Benchmark access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Offload Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Offload Benchmark use?

Offload Benchmark is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Offload Benchmark use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Offload Benchmark?

Skills that share tags, products or a category with Offload Benchmark: Pester Failure Analysis (PowerShell/PowerShell, 56k stars), Rust TDD Workflow (rtk-ai/rtk, 83k stars), Clawteam Dev (HKUDS/ClawTeam, 5.5k stars) and Browser Testing with Chrome DevTools (addyosmani/agent-skills, 102k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Offload Benchmark?

imbue-ai (a GitHub organization) maintains it in imbue-ai/offload, which has 125 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 25, 2026.

Source: imbue-ai/offload on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.