Memory And Disposal
andymai/brepjs
This skill should be used when managing WASM handle lifetimes or hunting memory leaks in brepjs — when a task mentions "createHandle() without using keyword risks WASM memory leak"…
Analyze and improve the runtime speed of a Wado program's compiled guest Wasm — profile hot functions, read the generated WIR for allocations and copies, reason about the WasmGC cost model, and…
$ npx skills add wado-lang/wado --skill wado-performance -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wado-lang/wado wado-performance --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wado-lang/wado.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/wado-performance .claude/skills/wado-performance && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "wado-performance" agent skill from https://github.com/wado-lang/wado/tree/main/.claude/skills/wado-performance into .claude/skills/wado-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wado-performance", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wado-lang/wado/tree/main/.claude/skills/wado-performanceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wado-lang/wado --skill wado-performance -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wado-lang/wado wado-performance --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wado-lang/wado.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/wado-performance .agents/skills/wado-performance && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "wado-performance" agent skill from https://github.com/wado-lang/wado/tree/main/.claude/skills/wado-performance into .agents/skills/wado-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wado-performance", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wado-lang/wado --skill wado-performance -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wado-lang/wado wado-performance --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wado-lang/wado.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/wado-performance .cursor/skills/wado-performance && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "wado-performance" agent skill from https://github.com/wado-lang/wado/tree/main/.claude/skills/wado-performance into .cursor/skills/wado-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wado-performance", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wado-lang/wado.git --path .claude/skills/wado-performance--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wado-lang/wado --skill wado-performance -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wado-lang/wado wado-performance --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wado-lang/wado.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/wado-performance .gemini/skills/wado-performance && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "wado-performance" agent skill from https://github.com/wado-lang/wado/tree/main/.claude/skills/wado-performance into .gemini/skills/wado-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wado-performance", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wado-lang/wado wado-performanceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wado-lang/wado --skill wado-performance -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wado-lang/wado.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/wado-performance .github/skills/wado-performance && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "wado-performance" agent skill from https://github.com/wado-lang/wado/tree/main/.claude/skills/wado-performance into .github/skills/wado-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wado-performance", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wado-lang/wado --skill wado-performance -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wado-lang/wado wado-performance --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wado-lang/wado.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/wado-performance .opencode/skills/wado-performance && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "wado-performance" agent skill from https://github.com/wado-lang/wado/tree/main/.claude/skills/wado-performance into .opencode/skills/wado-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wado-performance", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
wado-performanceAnalyze and improve the runtime speed of a Wado program's compiled guest Wasm — profile hot functions, read the generated WIR for allocations and copies, reason about the WasmGC cost model, and…
Wado Performance is an agent skill from wado-lang/wado. Analyze and improve the runtime speed of a Wado program's compiled guest Wasm — profile hot functions, read the generated WIR for allocations and copies, reason about the WasmGC cost model, and A/B-measure a fix. Use for any guest-side speed question, whatever the program does. For host-side native compiler profiling see profiling-wado-compiler; for wrong code out of an optimizer pass see optimizer-debug.
Its SKILL.md is about 7.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `dead-ends.md` and `scripts/analyze_guest_profile.ts`).
It sits in Development, covering Performance optimization. It works with WebAssembly. The repository describes itself as: The Wado Programming Language. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b1f1e2a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (TypeScript), which the agent can run.
Shell commands in SKILL.md call:
misecargogitnodeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Wado Performance loads about 7.5k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 4,379 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from wado-lang/wado at commit b1f1e2a, republished under its MIT licence (© wado-lang). 4,379 words, ~7,515 tokens.
.claude/skills/wado-performance/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Speed of the compiled guest Wasm (what wasmtime runs), not the native
compiler — that is profiling-wado-compiler.
Loop: profile the hot function → read its WIR for what it allocates/copies per iteration → change one thing → A/B both arms in one session, plus the WIR diff of the hot function → keep or revert (§5 says which evidence decides).
Re-profile before choosing a target: the percentages a WEP or an older note quotes predate whatever has landed since.
A speedup lands in the compiler, not in Wado source. Editing .wado files,
the stdlib included, is fine as an experiment: an ablation that prices a piece,
or a hand-written shape that shows what the optimizer should emit. Shipping
such an edit as the speedup is forbidden, unless the user approves it or asks
for it. A fast iteration loop on the stdlib is not a reason to make an
exception. §2 and §4 say where the fix goes instead.
wado run --profile guest,profile.json,1 prog.wado # interval 1 short runs, 0 exhaustive
wado test --profile guest,profile.json,1 file.wado # same, over one file's test blocks
node .claude/skills/wado-performance/scripts/analyze_guest_profile.ts profile.json [--top N]The script reports self (leaf) and inclusive counts per function; names keep
monomorphization detail, so each instantiation is separate. Loop a one-shot hot
phase N times so it clears the fixed setup (aim ≥ ~200 samples). Firefox
Profiler (profiler.firefox.com) gives a flame graph; perf + --profile jitdump gives instruction-level (store- vs compute-bound), see
docs/jitdump-profiling.md.
Dev-profile inflation: a cargo run wado JITs guest code near-release but
runs the wasmtime runtime / GC / allocator at dev speed (~4–5× slower), so
profiles over-weight allocation/GC frames — read percentages as relative and
size any GC or allocation win by its release number, not the dev multiple. A
flat-CST rewrite that cut a benchmark ~3× on dev gained ~1.47× on release,
because release GC was only ~⅓ of wall-clock to begin with. Pure compute does not
inflate, so a compute-bound win carries over intact.
A sample lands at the next epoch check, not where the time went. The guest
profiler samples on an epoch deadline and wasmtime checks the epoch at function
entries and loop headers, so straight-line code is charged to whichever it
reaches next. A derived deserializer's field-dispatch chain reported as 73%
self in deserialize_i32 on an 80-field struct, and 19% in deserialize_bool
on cbor-twitter — neither function is more than a bounds check and two compares.
A hot small leaf is telling you how often it is entered, so go read its caller.
The profile ranks candidates; it does not locate them.
A low self-percentage rules out one dataset, not the function.
FieldSchema::lookup read 0.71% on json-catalog, whose widest struct is 16
fields. Rewriting that same function cut 44% off cbor-twitter's decode, where
User has 40. Profile a per-item cost on the input whose items are widest.
Rule out a super-linear pass before blaming GC — that same inflation makes an algorithmic blow-up read as GC-bound; sweep input size to tell them apart. Faster-than-linear growth is a hypothesis and not a verdict, since a live set can grow that way too, so the WIR is what settles which.
Sweeping a shape dimension is sharper than size, and the one to vary is the one the suspect is indexed by. Decoding 1000 CBOR records, holding that count fixed so no per-record term is left for the growth to be:
i32 fields per record | 5 | 10 | 20 | 40 | 80 |
|---|---|---|---|---|---|
| ns per field | 85 | 80 | 87 | 129 | 221 |
Hold everything but the dimension under test — a sweep that varies two answers about neither.
wado dump -O2 prog.wado # final WIR
wado dump --tir-monomorphized prog.wado # how `?`, for-of, … desugarThree villains, each a heap alloc or deep copy, and a bug when one lands per element in a loop:
struct.new / Box<…> — a heap object. for x of &list boxes every
element (WasmGC has no interior references, so a by-ref iterator materializes
&T as a box). A tuple is a GC struct too, so retyping a two-field struct as
[u64, u64] allocates the same. Two things remove the allocation: multivalue
on a return the caller destructures, and SROA on a literal whose fields the
optimizer can split into locals. A merge point defeats both preconditions,
which is a pass to fix rather than a call site to rewrite (below).array.new / array.new_default — a fresh GC array (_default
zero-fills); watch for one per call where a buffer could be reused.$value_copy$T… — a value-semantics deep copy of a value-typed binding/arg
unless the source is fresh (a call / literal / variant result, or a fresh
value's payload). x? desugars to match f() {…}, so freshness must see
through the match; a missed copy shows up here and is removable.Count the copies, not only the time: a cut count is a result even when the benchmark is flat. The entry package's copies are its remarks:
wado compile -O2 --log-level info prog.wado 2>&1 | grep -c 'remark: a copy of'Remarks skip the stdlib, so read a stdlib copy off wado dump -O2.
Moving a copy from a callee into its call sites multiplies it by the number of sites.
Also: a Trait::method(…) call left in a hot loop (the inliner declined it), and
array_set_u8 / array_get_value (bounds-checked; one per element is the store floor
for Array<T>-backed String / List).
An optimizer fix reaches every Wado program. The stdlib edit that routes around one reaches a single call site, and hides the gap that produced it. So when the shape you are about to rewrite by hand is one a pass exists for, name the pass and read its precondition first. That is where the fix belongs.
Three turned up this way while cutting fts. All three are also live in
short, the path every ${x} on a float takes, which is why fixing the first
paid on a benchmark fts never touched.
sroa matches only a direct literal binding. An inlined get_pow10 left
let pm = <block with two exits>, one building the struct and one calling out
for it, with pm.hi / pm.lo the only uses. The PmHiLo was a heap object
per conversion. Fixed by extending slot_temp_sroa, which already scalarized
the [tag, slots…] shape, to a struct literal and to an exit that hands over
the aggregate: json-canada de +12.7%, ser +7.2%, and uscale_pow10
deleted.multi_value_return is all-or-nothing per callee. It wants every call
site to be let $tmp = Call(f) whose only uses are field accesses.
mul_pow10 had seven sites and six were exactly that; the one yielding the
call as a block value disqualified all seven. Fixing the first gap retired
this one, since the offending site is now a let. The precondition is
unchanged, so the next callee to hit it pays the same way.cold_outline refuses a region containing a return (control_escapes),
which is every rare slow path there is. Leaving fixed_width_for_prec's
out-of-range tail inline cost json-canada ser 6.5% against a byte-identical
serialize path: growing a hot function moves everything downstream of it in
the module. Hand-splitting it into a function restored the row, and
fixed_width_out_of_range in fpfmt.wado is that split.copying collector
traces what survives a cycle; an object that dies before the next one is never
copied, however many there were. Cutting transient allocations therefore moves
nothing — a compiler pass that removed thousands of per-token Box<i32> allocs
measured within noise under copying (and −0.7 ms/iter under null). Chase the
footprint, not the volume. The same rule retires "iterate by index to stop
for x of &list boxing": the boxes die immediately.List<i32> per state (~7.4K permanently live objects) made
identical hot-function wasm run 3–6× slower purely from the resident graph;
flattening it to offset/count columns fixed it. A resident 160 KB flat
List<i32> costs ~+0.9 ms/parse, the 7,400-list shape ~+2.4 ms. Prefer flat
columns over nested lists, and don't build what nothing reads. Measure the GC
share with --collector null (it leaks, so drive a fixed iteration count) vs
--collector copying.wado starts every guest
with 256 MiB (--gc-heap-initial), which the copying collector splits in two,
so a benchmark allocates through half a heap rather than climbing to one. A
raw wasmtime invocation starts at zero and doubles its way up to the working
set, paying a full trace at every rung — and there the ranking of two
compilers flips with the heap size rather than with the code. Even a fixed
size is one sample: a change that allocates less moves where the collections
land, and gale-gen read 4% slower at 512m where the sweep's geometric mean
read it faster.
On an allocation-heavy row, sweep --gc-heap-initial and compare geometric
means (dead-ends.md).with_capacity zero-fills. List::with_capacity(n) is an
array.new_default, so an over-sized arena pays for every slot it never uses —
once badly enough to turn a 2× faster build into a 4× slower one. Growing from
[] by doubling is not the fix either: it zero-fills ~2.4× more than a
reasonable pre-size. Size it about right, or grow.array.get costs ~20 machine instructions; gets sharing a block cost
~8. wasmtime re-derives the object's null check, its length load and the
overflow-checked element address per get, and neither hoists them out of a
loop nor shares them across blocks — only across gets in one block. Read the
actual sequence with wasmtime explore -W gc,function-references f.wat; a
byte-at-a-time loop is 22 instructions and 6 branches per byte, four gets in
one block are 18 + 4×8. So a scan reads several bytes per bounds check and
then tests them: peek_after_whitespace_run in core:json is that shape,
worth 12.6% on json-catalog deserialize. It pays in proportion to the run it
covers, against the one partial block it always wastes — under ~16 bytes per
run it is a loss. Four is where the sharing stops: a wider block only adds
lone gets, and measures worse the wider it gets (dead-ends.md).
array.set shares nothing: a store may write the header as far as
Cranelift knows, so four adjacent sets reload the length four times. Only
array.copy / array.fill amortise a write.struct.new per struct and
costs 6.5% on cbor-twitter: past the register file, forty ref locals live
across a call-heavy loop are forty spill slots reloaded at every call boundary,
plus a ref.null init apiece at entry. "The allocation is gone" says nothing
about which side won (dead-ends.md).array.copy is fast; leave it alone. It beats a hand-written loop from a
couple of bytes on — the loop pays the bounds check above on both the get and
the set of every byte. Neither hand-roll it nor contort an algorithm to avoid
it (dead-ends.md). Its length decides the cost: a constant one compiles to
inline loads and stores, a run-time one calls wasmtime_builtin_memory_copy.
So a small copy repeated per item pays to reach the copy with a constant
length, which is what param_spec cloning on name.used gave the JSON key
writer./ and % are cheap (Cranelift magic-multiply, x/k and x%k
fused) — don't trade a divide for extra multiplies.else if chain competitively, and a match over it (a br_table) adds an
indirect branch: two separate rewrites to jump tables measured flat and
slightly slower. Such a frame is usually call-frequency-bound, not
dispatch-bound — cut the calls, not the branch. What does answer to dispatch is
a cascade long enough to pay for that branch, or one that is not a cascade at
all: independent ifs no arm leaves test every key whatever matched, which
nir/if_chain_to_match is what fixes. A set membership test,
k matches { A | B | 'x'..='z' | … }, is neither. match_to_bitset lowers
it to a branch-free mask test when the scrutinee is 32 bits or narrower and
the members span at most 256 values, so write it as the set rather than
hand-rolling a range compare or a table. Past that span it is a br_table.`{v}` allocates a
throwaway String and copies it in, per value; buf.push_display(&v) skips
both. A run of adjacent push / push_str calls is fused into one capacity
check by nir/string_push, so write the appends plainly and let it batch them.internal_raw_data() / returning Array<T> by value is a copy API — for a
single read use get_unchecked / set_byte_unchecked.assert of a caller-guaranteed precondition is free; the same test as a
guard is not. is_json_ws in core:json needs b < 64 (Wasm masks a shift
count mod 64) and every call site short-circuits on b > b' ' first. Writing
the precondition as assert b < 64 measures flat on json-catalog deserialize;
writing it as b < 64 && in the returned expression costs 6%, and dropping
the bitset for four compares costs 19%. So a hot leaf whose precondition the
callers establish keeps both the assert and the fast body. Measure the assert
and the guard as separate arms: folded into one they read as a single cost,
and the assert takes the blame for what the guard spent.An inline hint or a hand-inlined body is almost never worth it: it reaches one
call site, and it silently outlives the measurement that justified it. The
stdlib carries no #[inline] / #[inline(never)], because a hint that makes
code faster marks a case the optimizer misses.
Improving the inliner is what pays, and it keeps being improved. A better price
reaches every program at once. When it declines a callee that should be spliced,
that is a cost-model bug: find the price with WADO_TRACE=inline and fix it in
optimize/inline.rs.
Never bend source to the inliner's current prices. Collapsing lets into one
expression, choosing an operation width, or splitting a function only to get
under the threshold are all banned. It is §2's rule applied to the inliner: the
contortion reaches one call site and hides the gap from every other. zlib's
read_u32_le is the case that set this rule. Widening each byte to u64
priced it one instruction over the threshold, and doing the arithmetic in u32
with one widening at the end got it inlined. But the widening costs nothing:
Cranelift folds an i64.extend_i32_u into the load or arithmetic producing it. The
price was wrong, and is_zero_extension in optimize/inline.rs is its fix.
Raising the threshold wholesale is not the lever: it bloats hot loops and measured slower. The lever is a price that matches what the spliced code costs.
A split that is right on its own stays. A slow path that is not rare is one: a
width > 0 branch runs every time a width is set, so it is hot when taken and
no cold_path() marker should claim otherwise. A rare heavy sub-case needs no
hand-split at all: nir/cold_outline moves what the marker opens into a
function of its own, so the leaf inlines at its hot-path size. A marker
cold_outline cannot take (mid-loop-body, see that pass's module doc) is a gap
in that pass, per §2.
Report an improvement as +%, a regression as −%. Compute it as speedup,
base / head − 1 on ms/iter (or head / base − 1 on throughput), so faster is
always positive. A report that writes one faster row as "+5%" and another as
"−5% ms" leaves the reader to work out which way each number points.
Only relative numbers carry signal. A/B both arms in the same session, best of
three or four, alternating and with the order swapped once — the first run of a
session reads high, so a fixed order silently taxes whichever arm goes second.
Run on an idle host, nothing else building: an A/B taken beside a compiling
test suite has put both arms inside each other's spread and flipped their
ranking. Check ps and free as well as uptime — a load average lags a
session that just started and says nothing about memory, and another agent's on
this box put the same test target at 17× its idle time before OOM-killing the
command after it. Nothing in a number says whether its host was idle, so
benchmark/README.md is a sanity check on the arm you just built, never the
control for it — even on the machine that produced it; a HEAD build has measured
615 MB/s against its own recorded 656 in the same afternoon. Isolate the phase —
A/B a float-format change on fts, not on a serialize benchmark that dilutes it.
core:json's inputs pull opposite ways: citm_catalog.json is mostly
pretty-printing whitespace, canada.json minified floats. A scan win on one is
no evidence about the other, so measure both, and json-twitter for strings.
A CI "Performance Alert" on rows the diff cannot reach is the shared runner. Confirm it locally before treating it as a regression.
While the user is iterating, the reading is 4–5 back-to-back pairs on the target row plus the wasm hash; the whole-suite A/B below is for the wrap-up.
A dev-build A/B is only valid where the dev build is. The inflation §1 describes flips A/B verdicts, not only profile weights. Dev runs the wasmtime runtime, GC and allocator at dev speed, so a row bound by allocation reads a different winner. json-canada is store- and compute-bound, so it matched release to under 2% and made a fast stdlib loop possible. On the same change dev called syntax-highlight -1.2% where release said +1.4%, and cbor-canada and cbor-twitter deserialize -2.9% and -1.2% where release said +0.3%. Every row that moved is a deserialize or a CST build, which is what allocates. Iterate on dev, then settle any row whose work is building an object graph on release.
Scratch files below go in scratchpad/. A path handed to the harness, which
runs from benchmark/, must be absolute.
Build the head arm first: every command below compares against
target/release/wado, and a stale one compares main with itself.
cargo build --release --bin wadoRun every benchmark at -O0 through -O3 before trusting an optimizer change:
both suites have passed a miscompile only the large bodies gale_gen produces
reach, where it reads as ERROR task failed. The last -O flag wins, so
WADO_BENCH_FLAGS=-O1 mise run benchmark-all runs the suite at -O1.
A change to the compiler needs two compilers. benchmark-baseline builds
origin/main's once and caches it under that commit; WADO_BIN then runs it
through this tree's harness, so only the compiler differs — the baseline's own
benchmark/ would put the branch's harness changes inside the comparison too.
The task fetches origin/main each time it runs and deletes the baseline of an
older main, so copy it out once and compare against the copy: a moved main is
another compiler.
cp "$(mise run benchmark-baseline)" scratchpad/wado-main # slow the first time# alternate, so neither arm always goes second
WADO_BIN="$PWD/scratchpad/wado-main" mise run benchmark-all > scratchpad/b1.log 2>&1
mise run benchmark-all > scratchpad/h1.log 2>&1 # …and so on, 3 each
node benchmark/ab.ts --base scratchpad/b{1,2,3}.log --head scratchpad/h{1,2,3}.logTime the Wado rows alone, with mise run all-wado. The reference arms (C,
Rust, JavaScript, the Java ones) run the same binary whatever the compiler does,
so re-timing them buys nothing and stretches a round many times over.
That is the gap the host drifts across: a three-arm benchmark-all comparison
came back with ANTLR4 (Java) at -2.2%, count-prime / JavaScript at +1.4% and
a prime sieve 4.3% "faster" from a string-append change, all unreadable. The
same arms over all-wado, six rounds back to back, settled every row. Keep
sieve in the selection as the in-band control, and run benchmark-all once at
the end for the record.
It is a task of benchmark/mise.toml, so it runs from benchmark/ only.
cd benchmark
mise run all-wado # every Wado row
mise run all-wado json_catalog sieve # those, by nameIts log feeds ab.ts and pick.ts like any other, one row per benchmark file.
Hash the wasm before you time anything. Compile every benchmark under both
compilers and compare. A row whose bytes are identical cannot have moved, so
whatever the suite says about it is the host. That is a stronger check than
timing, and it leaves only the few rows that differ to measure. A
field_scalarize fix came out byte-identical on all but three benchmarks. The
suite had meanwhile called fts 6.9% SLOWER with non-overlapping ranges; the
identical SHA-256 retired that reading outright.
Gate the compare on each compiler's exit status, not on its output file. A
compile that failed leaves the previous round's file in place, and comparing
those reads as "identical". That is the one answer this check must never give by
accident. Drop the schema modules, which are no world entry point, and give
http_routing the world it targets. Every remaining benchmark must compile, so
a FAILED row is one to go and read, and it carries the diagnostic explaining
it. wado compile reports on stderr even when it succeeds, so its output is
held back and printed with the failure, which keeps the sweep's own lines
readable.
Keep the world one --world=… token: zsh does not word-split an expansion, so
two words in a variable reach wado as a single flag it rejects.
for f in benchmark/*/*.wado; do
case "$f" in *_schema.wado) continue ;; esac
world=
case "$f" in */http_routing/*) world=--world=wasi:http/service ;; esac
scratchpad/wado-main compile -O2 ${world:+"$world"} -o scratchpad/b.wasm "$f" > scratchpad/cc.log 2>&1 \
|| { echo "FAILED $f"; cat scratchpad/cc.log; continue; }
target/release/wado compile -O2 ${world:+"$world"} -o scratchpad/h.wasm "$f" > scratchpad/cc.log 2>&1 \
|| { echo "FAILED $f"; cat scratchpad/cc.log; continue; }
cmp -s scratchpad/b.wasm scratchpad/h.wasm || echo "DIFFERS $f"
doneGive the four wasm-size programs the same pass at -Os: no benchmark covers
sqlite_highlight, the largest generated program in the tree.
For each row that differs, diff the two wado dump --wir -O2 outputs function by
function, which names what moved. A correctness fix is held to byte-identical
output on every program it does not fix. Then time only the rows whose hot path moved, back to back, and read the
rest as unmoved.
ab.ts decides each row by whether the arms' [min, max] overlap, not by the
delta: on a 5 ms benchmark a 6% gap between bests sits inside one arm's own
spread. Read the reference rows first — C, Rust and JavaScript run the same
binary in both arms, so a SLOWER among them is the host drifting and no Wado
row can be read either.
Confirm a surviving row before believing it: the whole-suite arms are minutes apart, and the reference rows only catch drift big enough to cross a range. Loop that one benchmark back to back and check the ranking holds pair by pair.
for i in 1 2 3 4 5; do
scratchpad/wado-main run -O2 benchmark/sieve/sieve.wado
target/release/wado run -O2 benchmark/sieve/sieve.wado
doneA change to lib/core/*.wado needs two compilers as well, which the source tree
hides: a release build embeds the stdlib where a dev build reads it from disk.
benchmark/wado.sh falls back to cargo run --release whenever WADO_BIN is
unset, so swapping an arm's .wado files into the tree invalidates
wado-compiler and rebuilds it under lto=thin / codegen-units=1. That is
two full rebuilds per alternating round, and they are the wall clock rather than
the benchmark.
Build one binary per arm first, from the same compiler source with lib/ at the
branch's fork point and at HEAD, so lib/ is all that differs. git restore
removes a file the source tree lacks, so neither binary embeds a stdlib
belonging to neither arm. It overwrites uncommitted edits under lib/, so
commit them first.
set -e # a failed build would leave an earlier A/B's binary as its arm
rm -f scratchpad/wado-base scratchpad/wado-head
fork=$(git merge-base origin/main HEAD)
# however the block ends, `lib/` goes back to HEAD rather than stay reverted
trap 'git restore --source=HEAD --worktree -- wado-compiler/lib' EXIT
git restore --source="$fork" --worktree -- wado-compiler/lib
cargo build --release --bin wado --quiet
cp target/release/wado scratchpad/wado-base
git restore --source=HEAD --worktree -- wado-compiler/lib
cargo build --release --bin wado --quiet
cp target/release/wado scratchpad/wado-head
for r in 1 2 3; do
for arm in base head; do
WADO_BIN="$PWD/scratchpad/wado-$arm" mise run benchmark-json-catalog
done
doneThe tree's sources stop mattering once the binaries exist, so a round costs what the benchmark costs. Rounds are cheap enough then to run six or ten of them, which is what it takes to resolve a delta near 1% out of this row's spread.
WADO_SKIP_PASS=<pass> is a third arm off the same binary, which is how a
regression is attributed to one pass without a third build. WADO_BENCH_FLAGS
sweeps a knob the same way; the harness appends it to every wado compile and
wado run it issues, so only a knob both accept can be swept.
Give a threshold a temporary env override and sweep it, rather than
rebuilding per value — and reach for it the moment a change looks like it only
pays above some size, because that shape usually means two rewrites are riding
one knob. if_chain_to_match appeared to need a 12-arm floor; overriding its
threshold and match_to_switch's separately showed the fusion was never the
cost at any width and the br_table past it was the whole of it on the row that
regressed, turning 3.6% down on cbor-catalog into 2.1% up. Delete the overrides before
committing: read per node visit, std::env::var is itself a compile-time cost.
A runtime setting one benchmark wants becomes a CLI option with a conservative
default, and that benchmark opts in through gc_heap_flags in
benchmark/wado.sh. Sweep a candidate below the default as well as above it.
What decides adoption, in priority order:
dead-ends.md.Only a WIR diff decides case 2. Diff the two wado dump -O2 outputs and read
what the hot function issues per iteration — a run of N capacity checks collapsed
to one, a call gone from a loop body. Nothing else establishes "fewer
instructions": not the dump's line count, and not the wasm byte count.
Neither wasm size nor dump size correlates with speed. Smaller output is
routinely slower and larger output routinely faster — the bytes are mostly code
that never runs, and what does run is priced by what the loop executes. The three
quantities move independently: the append fusion grew wado dump -O2 on
syntax-highlight 8.3% (a fused write unparses its offset as an expression) and
shrank the -Os binary 1.5%, while the thing that justified it was a +8%
benchmark and a diff showing one less capacity check per key. As evidence about
speed, size is only the tiebreaker at rank 3. Where size and speed trade, take
the speed and state the size cost. A fold of an idiom wasmtime already matches (shift-or into rotl)
buys size only, so judge it by bytes.
dead-ends.md (next to this file) is the record: every optimization measured
and dropped, with the A/B that killed it and what it generalizes to. Read it
before starting, and add an entry whenever an A/B comes back flat or negative
— a dead end nobody wrote down is one somebody re-measures.
Stop when the floor is the representation — a store-bound loop on an
Array<T>-backed String is near-optimal short of leaving GC arrays.
dead-ends.md — what has already been measured and dropped.profiling-wado-compiler — the native wado binary (host side).benchmark — run the suite / wasm-size report.optimizer-debug — a NIR/WIR pass producing wrong code, not just slow.© wado-lang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts) in .claude/skills/wado-performance of wado-lang/wado.
Open the folder on GitHubat commit b1f1e2a
Wado Performance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Wado Performance this skillwado-lang/wado | 117 | — | ~7.5k | Automated safety check: Pass | MIT | |
| Memory And Disposalandymai/brepjs | 115 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Code Review ChecklistshareAI-lab/learn-claude-code | 78k | 5 repos | ~1.1k | Automated safety check: Pass | MIT | |
| LLM Torch Profiler Analysissgl-project/sglang | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | |
| Pycrazyguitar/pysheeet | 8.2k | — | ~886 | Automated safety check: Pass | MIT | |
| Cmux Debugging Guidemanaflow-ai/cmux | 28k | 1 repos | ~1.1k | Automated safety check: Pass | Custom licence |
andymai/brepjs
This skill should be used when managing WASM handle lifetimes or hunting memory leaks in brepjs — when a task mentions "createHandle() without using keyword risks WASM memory leak"…
shareAI-lab/learn-claude-code
Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
crazyguitar/pysheeet
Comprehensive Python programming reference covering syntax, concurrency, networking, databases, ML/LLM development, and HPC.
manaflow-ai/cmux
Covers debug logging, the Debug menu, profiling rules and runtime pitfalls for working on the cmux macOS terminal app.
keybase/client
Analyzes V8, Chrome and Electron .heapsnapshot files with Node scripts to find memory leaks, detached DOM nodes and the retainer paths that keep objects alive.
wado-lang/wado
Profile the native Rust wado binary (compile/serve/run) for host-side bottlenecks — CPU with a sampling profiler, memory with the span trace's RSS and valgrind DHAT.
wado-lang/wado
Measure how long GitHub Actions jobs and steps took across past runs, and find the pull request that made CI slower.
wado-lang/wado
Investigate and improve code coverage for the wado-compiler crate.
wado-lang/wado
The only way to merge origin/main into a branch, conflicts or not.
wado-lang/wado
Transpile Wado Wasm components to JS with jco, then run, debug, and benchmark them on Node.
wado-lang/wado
Task-completion flow: /cr (a /code-review answered with /code-review-response, which ends with /distill), update docs (spec/cheatsheet/compiler/optimizer), then run mise run on-task-done (build…
Works with
Categories
Analyze and improve the runtime speed of a Wado program's compiled guest Wasm — profile hot functions, read the generated WIR for allocations and copies, reason about the WasmGC cost model, and…. Wado Performance is an agent skill from wado-lang/wado. Analyze and improve the runtime speed of a Wado program's compiled guest Wasm — profile hot functions, read the generated WIR for allocations and copies, reason about the WasmGC cost model, and A/B-measure a fix.
Wado Performance fits situations like: any guest-side speed question; whatever the program does.
Run `npx skills add wado-lang/wado --skill wado-performance -a claude-code`. Or copy the skill folder (.claude/skills/wado-performance in wado-lang/wado) into .claude/skills/wado-performance in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wado-lang/wado --skill wado-performance -a codex`. Or copy the skill folder (.claude/skills/wado-performance in wado-lang/wado) into .agents/skills/wado-performance in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wado-lang/wado --skill wado-performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wado-performance, .gemini/skills/wado-performance, .github/skills/wado-performance and .opencode/skills/wado-performance in your project.
Going by SKILL.md and its folder, Wado Performance needs TypeScript for the scripts in its folder and the command-line tools its instructions call (mise, cargo, git and node). Our summary lists: Node.js.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Wado Performance is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.5k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Wado Performance: Memory And Disposal (andymai/brepjs, 115 stars), Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars) and Py (crazyguitar/pysheeet, 8.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wado-lang (a GitHub organization) maintains it in wado-lang/wado, which has 117 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.
Source: wado-lang/wado on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.