Agent skill

Openclaw Test Performance

by openclaw in openclaw/openclaw

Benchmark, diagnose, and optimize OpenClaw test and plugin-suite runtime, import hotspots, CPU/RSS, heap growth, and slow coverage paths.

MITAuto-check passedTesting & QA

Install Openclaw Test Performance

skills CLI
$ npx skills add openclaw/openclaw --skill openclaw-test-performance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openclaw/openclaw openclaw-test-performance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openclaw/openclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/openclaw-test-performance .claude/skills/openclaw-test-performance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
openclaw-test-performance
GitHub stars
392k
Token cost
~2.9k tokens
SKILL.md length
1,137 words
Files
2
Skills in repo
93
Repo updated
First seen
Licence
MIT

At a glance

Benchmark, diagnose, and optimize OpenClaw test and plugin-suite runtime, import hotspots, CPU/RSS, heap growth, and slow coverage paths.

  • Works in 2 steps: Read the relevant local AGENTS.md files… → Establish a baseline before changing code
  • Testing & QA work in your project
  • SKILL.md covers Workflow, Plugin-Suite Workflow, Metric Collection and Common Root Causes, plus 3 more sections
  • Calls pnpm

What it does

Openclaw Test Performance is an agent skill from openclaw/openclaw. Benchmark, diagnose, and optimize OpenClaw test and plugin-suite runtime, import hotspots, CPU/RSS, heap growth, and slow coverage paths.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Testing & QA. It works with pnpm and Vitest. The repository describes itself as: The AI that really does things. Any OS. Any Platform. The lobster way. 🦞. The licence is MIT.

When your agent uses it

  • Testing & QA work in your project

Example prompts

  • “/openclaw-test-performance”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Read the relevant local AGENTS.md files before editing
  2. Establish a baseline before changing code

What it can do on your machine

Read from SKILL.md and the folder at commit 1eb5970. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Openclaw Test Performance loads about 2.9k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 1,137 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from openclaw/openclaw at commit 1eb5970, republished under its MIT licence (© openclaw). 1,137 words, ~2,877 tokens.

Download SKILL.mdSave it as .claude/skills/openclaw-test-performance/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
openclaw-test-performance
description
Benchmark, diagnose, and optimize OpenClaw test and plugin-suite runtime, import hotspots, CPU/RSS, heap growth, and slow coverage paths.

OpenClaw Test Performance

Use evidence first. The goal is real pnpm test, plugin-suite, and plugin-inspector speed/RSS improvement with coverage intact, not runner tuning by guesswork.

Workflow

  1. Read the relevant local AGENTS.md files before editing:
    • src/agents/AGENTS.md for agent/import hotspots.
    • src/channels/AGENTS.md and src/plugins/AGENTS.md for plugin/channel laziness.
    • src/gateway/AGENTS.md for server lifecycle tests.
    • test/helpers/AGENTS.md and src/channels/plugins/contracts/test-helpers/AGENTS.md for shared contract helpers.
    • src/infra/outbound/AGENTS.md for outbound/media/action tests.
  2. Establish a baseline before changing code:
    • Prefer pnpm test:perf:groups --full-suite --allow-failures --output <file> for full-suite ranking.
    • For bundled plugin breadth, run the smallest relevant pnpm test:extensions:batch <plugin[,plugin...]> or plugin-inspector command before jumping to the full extension sweep.
    • For a scoped hotspot use: /usr/bin/time -l pnpm test <file-or-files> --maxWorkers=1 --reporter=verbose
    • For import-heavy suspicion add: OPENCLAW_VITEST_IMPORT_DURATIONS=1 OPENCLAW_VITEST_PRINT_IMPORT_BREAKDOWN=1.
  3. Separate wall/runner noise from real file cost:
    • Compare Vitest duration, test body timing, import breakdown, wall time, and max RSS.
    • Re-run single files when grouped/full-suite numbers look stale or noisy.
    • If a full-suite grouped run reports a lane failure but JSON says tests passed, capture that as harness/noise and verify the suspect file directly.
  4. Pick the next attack by return and risk:
    • High return: one file/test dominates seconds or RSS and has a clear root.
    • High leverage: one plugin or SDK barrel causes every plugin-inspector or extension-batch run to load broad runtime.
    • Lower risk: static descriptors, target parsing, routing, auth bypass, setup hints, registry fixtures, or test server lifecycle.
    • Higher risk: real memory/runtime behavior, live providers, protocol contracts, or broad production refactors.
  5. Fix the root cause, not the symptom:
    • Move static metadata/parsing into narrow helpers or lightweight artifacts reused by full runtime and fast paths.
    • Prefer dependency injection, loaded-plugin-only lookup, explicit fixtures, and pure helpers over broad mocks.
    • Reuse suite-level servers/clients when a fresh handshake is irrelevant.
    • Keep schedulers/background loops off unless the test proves scheduling.
    • In plugin paths, move static metadata into manifest/lightweight artifacts and keep runtime plugin loads behind explicit execution boundaries.
  6. Preserve coverage shape:
    • Do not delete a slow integration proof unless the exact production composition is extracted into a named helper and tested.
    • Keep one cheap integration smoke when cross-component wiring matters.
    • State explicitly what incidental coverage was removed, if any.
  7. Re-benchmark the same command after the change and compute seconds plus percent gain.
  8. Update the running report when requested or when this thread is tracking one. Include before/after commands, artifacts, coverage notes, verification, and next attack order.
  9. Stage the intended paths, commit with standard Git, and push when the user asked for commits/pushes. Stage only files touched for this attack.

Plugin-Suite Workflow

Use this section when perf work involves bundled plugins, plugin-inspector, SDK barrels, package-boundary tests, or extension suites.

  1. Map the suite shape first:
    • source tests: pnpm test extensions/<id> or pnpm test:extensions:batch <id>
    • package boundaries: pnpm run test:extensions:package-boundary:canary and pnpm run test:extensions:package-boundary:compile
    • all bundled source tests: pnpm test:extensions
    • plugin import memory: pnpm test:extensions:memory -- --json .artifacts/test-perf/extensions-memory.json
    • plugin-inspector/report work: keep report primitives in plugin-inspector; keep wrappers thin and collect peak RSS when the command supports it.
  2. Start narrow, then widen:
    • one plugin changed: run that plugin's tests and plugin-inspector slice.
    • SDK/public barrel changed: add representative provider, channel, memory, and feature plugins.
    • loader/runtime mirror changed: add package-boundary checks and build/package proof as needed.
    • unknown shared plugin behavior: run test:extensions:batch groups before pnpm test:extensions.
  3. Treat plugin-inspector failures as product signals:
    • JSON must parse.
    • warnings/errors must be classified, not hidden.
    • runtime capture should be quiet and config-tolerant.
    • command output should include wall time, exit code, and peak RSS when available.
  4. Follow $openclaw-testing for host selection. Trusted source benchmarks can run locally with comparable machine/load conditions. Use $crabbox when clean packaging, Linux/platform behavior, isolation, or an explicit remote request is part of the proof; reuse and clean up only the owned lease.
  5. If plugin performance is package-artifact sensitive, switch to release-openclaw-plugin-testing and Package Acceptance rather than trusting source-only timing.
Show full SKILL.md (500 more words)Show less

Metric Collection

Collect at least one stable metric before and after. Prefer the same machine and same command. For Testbox comparisons, use the same tbx_... id when possible.

MetricUse forPreferred source
wall timeuser-visible suite cost/usr/bin/time -l, test wrapper duration, Testbox run time
Vitest durationtest body/import costVitest output per file/shard
import durationbroad barrel/runtime loadsOPENCLAW_VITEST_IMPORT_DURATIONS=1
max RSSmemory pressure and OOM risk/usr/bin/time -l, pnpm test:extensions:memory, wrapper memory summaries
CPU/user/sysCPU-bound vs wait-bound split/usr/bin/time -l locally, Testbox job timing when local CPU is noisy
heap evidencereal leak vs retained module graphopenclaw-test-heap-leaks workflow

Local scoped command with CPU/RSS:

bash
timeout 240 /usr/bin/time -l pnpm test <file> --maxWorkers=1 --reporter=verbose

Plugin import memory profile:

bash
pnpm build
pnpm test:extensions:memory -- --top 20 --json .artifacts/test-perf/extensions-memory.json

Targeted plugin import memory:

bash
pnpm test:extensions:memory -- --extension discord --extension telegram --skip-combined

Heap/RSS escalation:

bash
pnpm test:perf:groups \
  --config test/vitest/vitest.unit-fast.config.ts \
  --allow-failures \
  --output .artifacts/test-perf/unit-fast-memory.json
pnpm test:perf:profile:runner -- \
  --output-dir .artifacts/test-perf/vitest-runner-profile -- <file>

Use openclaw-test-heap-leaks when RSS keeps growing across intervals, workers OOM, or the suspect command has app-object retention. Do not call RSS growth a leak until snapshots or retainers support it.

Common Root Causes

  • Full bundled channel/plugin runtime loaded for static data.
  • getChannelPlugin() fallback used when an already-loaded fixture or pure parser would suffice.
  • Broad api.ts, runtime-api.ts, test-api.ts, or plugin-sdk barrels pulled into hot tests.
  • SDK root aliases or package barrels pulling focused subpaths back into a broad plugin graph.
  • Plugin-inspector loading runtime code just to render metadata, reports, or CI policy scores.
  • Bundled plugin capture reusing real config/home state instead of synthetic, redacted, isolated state.
  • Partial-real mocks using importActual() around broad modules.
  • vi.resetModules() plus fresh imports in per-test loops.
  • Test plugin registry seeded in beforeAll while runtime state resets in afterEach.
  • Per-test gateway/server/client startup when state reset would suffice.
  • Runtime/default model/auth selection paid by idle snapshots or fixtures.
  • Plugin-owned media/action discovery triggered before checking whether args contain plugin-owned fields.
  • Parallel Vitest runs sharing node_modules/.experimental-vitest-cache without distinct OPENCLAW_VITEST_FS_MODULE_CACHE_PATH values.

Benchmark Commands

Scoped file:

bash
timeout 240 /usr/bin/time -l pnpm test <file> --maxWorkers=1 --reporter=verbose

Scoped file with import breakdown:

bash
timeout 240 /usr/bin/time -l env \
  OPENCLAW_VITEST_IMPORT_DURATIONS=1 \
  OPENCLAW_VITEST_PRINT_IMPORT_BREAKDOWN=1 \
  pnpm test <file> --maxWorkers=1 --reporter=verbose

Grouped suite:

bash
pnpm test:perf:groups --full-suite --allow-failures \
  --output .artifacts/test-perf/<name>.json

Extension batch:

bash
pnpm test:extensions:batch <plugin[,plugin...]> -- --reporter=verbose

All extension tests:

bash
pnpm test:extensions

Package-boundary plugin checks:

bash
pnpm run test:extensions:package-boundary:canary
pnpm run test:extensions:package-boundary:compile

Reuse an existing Vitest JSON report:

bash
pnpm test:perf:groups --report <vitest-json> \
  --output .artifacts/test-perf/<name>.json

Verification

  • Always run the targeted test surface that proves the change.
  • For source changes, run pnpm check:changed before push; in maintainer Testbox mode run it in the warmed Testbox.
  • For test-only changes, run pnpm test:changed or the exact edited tests.
  • Run pnpm build when touching lazy-loading, bundled artifacts, package boundaries, dynamic imports, build output, or public surfaces.
  • For plugin SDK/barrel/runtime changes, compare exact commits with pnpm plugin-sdk:api:diff -- --base <base-sha> --head <head-sha> when the public API surface may drift. For PR-local proof, use the branch merge base as <base-sha> and the exact tested head commit as <head-sha>.
  • For plugin-suite perf fixes, verify at least one representative plugin batch plus the changed gate; use Package Acceptance if the bug only exists in a packed artifact.
  • If deps are missing/stale, run pnpm install and retry the exact failed command once.
  • Use the report format:
markdown
| Metric         | Before |  After |          Gain |
| -------------- | -----: | -----: | ------------: |
| File wall time |   `Xs` |   `Ys` |  `-Zs` (`P%`) |
| Max RSS        |  `XMB` |  `YMB` | `-ZMB` (`P%`) |
| CPU user/sys   | `X/Ys` | `A/Bs` |       explain |

Handoff

Keep the final concise:

  • Root cause.
  • Suite/plugin scope.
  • Files changed.
  • Before/after wall, Vitest/import, CPU, and RSS numbers where available.
  • Leak classification if memory was involved: real leak, retained module graph, or inconclusive.
  • Coverage retained.
  • Verification commands.
  • Testbox ID or workflow URL for remote proof.
  • Commit hash and push status.

© openclaw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/openclaw-test-performance of openclaw/openclaw.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 1eb5970

Compare with similar skills

Openclaw Test Performance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Openclaw Test Performance compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Openclaw Test Performance this skillopenclaw/openclaw392k—~2.9kAutomated safety check: PassMIT
Ya Runkzahel/yepanywhere533—~957Automated safety check: PassMIT
Ckeditor5 TestingTriliumNext/Trilium38k—~3.3kAutomated safety check: PassAGPL-3.0
Ha Frontend Testinghome-assistant/frontend5.7k—~1.7kAutomated safety check: PassApache-2.0
Odc TestingDouglasNeuroInformatics/OpenDataCapture119—~1.5kAutomated safety check: NotesApache-2.0
Next QA Loopbreaking-brake/cc-wf-studio5.4k—~2.5kAutomated safety check: PassCustom licence

Similar skills

  • Ya Run

    kzahel/yepanywhere

    Launch and drive Yep Anywhere — an isolated dev server plus real browser interaction — and run the repository's check suite.

    533 GitHub stars~957 tokensUpdated today
    Testing & QAAuto-check passed
  • Ckeditor5 Testing

    TriliumNext/Trilium

    Testing CKEditor 5 plugins in the Trilium monorepo. An agent skill from TriliumNext/Trilium.

    38k GitHub stars~3.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Ha Frontend Testing

    home-assistant/frontend

    Home Assistant frontend testing and validation workflow. An agent skill from home-assistant/frontend.

    5.7k GitHub stars~1.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Odc Testing

    DouglasNeuroInformatics/OpenDataCapture

    Test a change in Open Data Capture. An agent skill from DouglasNeuroInformatics/OpenDataCapture.

    119 GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check: notes
  • Next QA Loop

    breaking-brake/cc-wf-studio

    Runs one unattended iteration of a QA loop: tend any open QA pull request, then build one queued qa issue as tests or test tooling only and open a PR.

    5.4k GitHub stars~2.5k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Openclaw Test Heap Leaks

    SafeAI-Lab-X/ClawKeeper

    Investigate pnpm test memory growth, Vitest worker OOMs, and suspicious RSS increases in OpenClaw using the scripts/test-parallel.mjs heap snapshot tooling.

    1k GitHub stars~1.2k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from openclaw/openclaw

All 93 skills in this repo
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Tmux

    openclaw/openclaw

    Control tmux sessions/panes for interactive CLIs: list, capture output, send keys, paste text, monitor prompts.

    392k GitHub starsUsed in 2 repos~640 tokens
    Auto-check passed
  • Feishu Doc

    openclaw/openclaw

    Feishu document read/write workflows. An agent skill from openclaw/openclaw.

    392k GitHub stars~516 tokensUpdated today
    Auto-check passed
  • Openclaw PR Maintainer

    openclaw/openclaw

    Review, triage, repair, or land OpenClaw issues and pull requests with current-source evidence and the native maintainer workflow.

    392k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Browser Automation

    openclaw/openclaw

    A skill your agent uses when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

    392k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Clawsweeper

    openclaw/openclaw

    A skill your agent uses for all ClawSweeper work: OpenClaw issue/PR sweep reports, repair jobs, cloud fix PRs, @clawsweeper maintainer mention commands, trusted ClawSweeper-reviewed…

    392k GitHub stars~3k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Openclaw Test Performance

What does Openclaw Test Performance do?

Benchmark, diagnose, and optimize OpenClaw test and plugin-suite runtime, import hotspots, CPU/RSS, heap growth, and slow coverage paths. Openclaw Test Performance is an agent skill from openclaw/openclaw. Benchmark, diagnose, and optimize OpenClaw test and plugin-suite runtime, import hotspots, CPU/RSS, heap growth, and slow coverage paths.

When should I use Openclaw Test Performance?

Openclaw Test Performance fits situations like: testing & QA work in your project.

How do I install Openclaw Test Performance in Claude Code?

Run `npx skills add openclaw/openclaw --skill openclaw-test-performance -a claude-code`. Or copy the skill folder (.agents/skills/openclaw-test-performance in openclaw/openclaw) into .claude/skills/openclaw-test-performance in your project. Claude Code loads it when a task matches its description.

How do I install Openclaw Test Performance in Codex?

Run `npx skills add openclaw/openclaw --skill openclaw-test-performance -a codex`. Or copy the skill folder (.agents/skills/openclaw-test-performance in openclaw/openclaw) into .agents/skills/openclaw-test-performance in your project. Codex loads it when a task matches its description.

Can I use Openclaw Test Performance in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openclaw/openclaw --skill openclaw-test-performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/openclaw-test-performance, .gemini/skills/openclaw-test-performance, .github/skills/openclaw-test-performance and .opencode/skills/openclaw-test-performance in your project.

What does Openclaw Test Performance need to run?

Going by SKILL.md and its folder, Openclaw Test Performance needs the command-line tools its instructions call (pnpm).

Does Openclaw Test Performance access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Openclaw Test Performance safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Openclaw Test Performance use?

Openclaw Test Performance is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Openclaw Test Performance use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Openclaw Test Performance?

Skills that share tags, products or a category with Openclaw Test Performance: Ya Run (kzahel/yepanywhere, 533 stars), Ckeditor5 Testing (TriliumNext/Trilium, 38k stars), Ha Frontend Testing (home-assistant/frontend, 5.7k stars) and Odc Testing (DouglasNeuroInformatics/OpenDataCapture, 119 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Openclaw Test Performance?

openclaw (a GitHub organization) maintains it in openclaw/openclaw, which has 391,610 GitHub stars. The repository holds 93 skills in this directory. The repository was last updated on October 8, 2026.

Source: openclaw/openclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.