Update V8 Version
openinterpreter/openinterpreter
Bumps the pinned v8 and rusty_v8 versions in Codex, validates the release-candidate path with the v8-canary check, and traces failures to upstream build changes.
Optimize Rust code until nothing left to improve. An agent skill from katopz/katgpt-rs.
$ npx skills add katopz/katgpt-rs --skill rust-optimize -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install katopz/katgpt-rs rust-optimize --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/katopz/katgpt-rs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/rust-optimize .claude/skills/rust-optimize && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "rust-optimize" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/rust-optimize into .claude/skills/rust-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rust-optimize", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/rust-optimizeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add katopz/katgpt-rs --skill rust-optimize -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install katopz/katgpt-rs rust-optimize --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/katopz/katgpt-rs.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/rust-optimize .agents/skills/rust-optimize && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "rust-optimize" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/rust-optimize into .agents/skills/rust-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rust-optimize", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add katopz/katgpt-rs --skill rust-optimize -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install katopz/katgpt-rs rust-optimize --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/katopz/katgpt-rs.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/rust-optimize .cursor/skills/rust-optimize && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "rust-optimize" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/rust-optimize into .cursor/skills/rust-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rust-optimize", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/katopz/katgpt-rs.git --path .agents/skills/rust-optimize--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add katopz/katgpt-rs --skill rust-optimize -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install katopz/katgpt-rs rust-optimize --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/katopz/katgpt-rs.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/rust-optimize .gemini/skills/rust-optimize && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "rust-optimize" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/rust-optimize into .gemini/skills/rust-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rust-optimize", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install katopz/katgpt-rs rust-optimizeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add katopz/katgpt-rs --skill rust-optimize -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/katopz/katgpt-rs.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/rust-optimize .github/skills/rust-optimize && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "rust-optimize" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/rust-optimize into .github/skills/rust-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rust-optimize", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add katopz/katgpt-rs --skill rust-optimize -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install katopz/katgpt-rs rust-optimize --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/katopz/katgpt-rs.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/rust-optimize .opencode/skills/rust-optimize && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "rust-optimize" agent skill from https://github.com/katopz/katgpt-rs/tree/develop/.agents/skills/rust-optimize into .opencode/skills/rust-optimize/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rust-optimize", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
rust-optimizeOptimize Rust code until nothing left to improve. An agent skill from katopz/katgpt-rs.
Rust Optimize is an agent skill from katopz/katgpt-rs. Optimize Rust code until nothing left to improve. Loops automatically. Use when the user says "optimize" or "/optimize".
Its SKILL.md is about 7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files.
It works with Rust. The repository describes itself as: A neuro-symbolic micro-Transformer with speculative decoding, constraint pruning, recurrent attention, and adaptive test-time scaling — built in Rust. The licence is MIT.
11 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d0b32e2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Rust), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
mcyoung.xyzFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Rust Optimize loads about 7k tokens when it runs. Until then it costs about 34 tokens; SKILL.md has 2,478 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from katopz/katgpt-rs at commit d0b32e2, republished under its MIT licence (© katopz). 2,478 words, ~6,989 tokens.
.claude/skills/rust-optimize/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Optimize *.rs files in the target scope using the checklist below.
This is a LOOP. You keep optimizing until you cannot find anything to optimize.
After each turn, end your response with exactly one of:
Continue optimizing remaining files. — you made code changes this turnNo optimizations this pass. — you read files but changed nothingIf you only summarized what you read without changing code, say "no optimizations this pass".
O(n^{(d-1)/d}) instead of interior O(n) for low-dim d≤3 region mass queries — curse-of-dim caps it at d≤3)String → &'static str, pre-allocate, eliminate .to_string()#[repr(u8)] enums, remove #[repr(C)]f32 counters → u32/u64chunks_exact(), use index arithmeticArc<RwLock<HashMap>> → papaya, Mutex<u32> → AtomicU32unwrap() → ?, let _ = → .log_err()WinitSettings::default() (Continuous) pegs CPU at 100% even idle — switch to UpdateMode::reactive_low_power(max_wait) matching tick ratePresentMode::Fifo controls GPU only, independent of CPU loop spinLineList / TriangleList mesh once, NEVER redraw via gizmos.line() per-frame (each call is CPU + per-frame buffer grow)Entity does NOT release its Assets<Mesh> / Assets<StandardMaterial> handles — call meshes.remove(handle.0.id()) + materials.remove(...) on despawn or leak indefinitely (regrow-on-new-EntityId patterns leak ~N assets per cycle)bevy_egui sets set_request_repaint_callback that fires bevy_winit::WakeUp on delay.is_zero() — bypasses reactive max_wait and spins the loop. Diagnose via in-process perf overlay before assuming the engine itself is the hog.Last-schedule std::thread::sleep system capped at e.g. 60 FPS is a belt-and-suspenders defense when vsync + reactive mode somehow don't throttlesysinfo for CPU/RAM + Bevy's Time::delta() for FPS) inside the GUI itself — external ps sampling misses focused-window swings and requires keeping a terminal attachedperf: or refactor: prefix).Hot-path Rust optimization patterns. Apply to any microsecond-sensitive code.
std::hint::black_box() to prevent dead-code elimination--nocapture test harness[T; N] when domain is boundedVec::with_capacity() once, clear() + reuse across calls&mut [T] parameters instead of allocating inside hot loopsFor code operating on cell complexes / meshes / grid manifolds (DEC, FEM, game maps):
∂M (surface area, O(n^{(d-1)/d}) cells) instead of the interior M (volume, O(n) cells). Valid when the field is curl-free / exact; reconstruction error is bounded by the harmonic component (compute via Hodge decomposition). Win shrinks fast as dimension d grows — practical only for d ≤ 3 (2D game maps, 3D belief regions, KG embeddings). For d ≥ 8 (HLA state) or d ≥ 64 (style weights) the boundary is larger than the interior, so boundary-only is a loss.curl(grad)=0 / div(curl)=0 enforced by DEC operator construction (not a soft penalty) gives mass-conservation invariants for free. Use as a modelless validator: if div(flow) > τ, mass leaked/created = anomaly.exterior_derivative, codifferential, hodge_decompose) depend only on the cell complex topology, not on the field values. Compute once on map/complex load, cache, invalidate only when topology changes — zero per-tick DEC op cost on a stable map.Auto-vectorization (let LLVM do the work):
usize → u64 slices (same layout on 64-bit) for wider SIMD lanesu64 equality comparison — compiler maps to _mm256_cmpeq_epi64 on AVX2bool as usize instead of if)Portable std::simd recipes (when auto-vec isn't enough):
Reach for std::simd only after profiling shows auto-vectorization failing. Patterns below are distilled from mcyoung's vb64 writeup (https://mcyoung.xyz/2023/11/27/simd-base64/). Full skeleton in recipes/swizzle_lookup.rs.
match on byte ranges with simd_ge/simd_le masks + mask.select(splat_a, splat_b). 1 select beats N branches.swizzle_dyn: if (byte >> 4) - (byte == c) distinguishes all ranges, build an 8-entry offset table and do 1 shuffle instead of N compares. Index vector must be same width as lookup table.sextets.cast::<u16>() << Simd::from([2,4,6,8]), then split into lo = v.cast::<u8>() and hi = (v >> 8).cast::<u8>(), OR them after rotating hi by 1 lane. Lets bits cross byte boundaries without per-bit ops.|i| i + i/(k-1) to skip those lanes. Compile-time indices = single vpshufb.out.reserve(final_len + N/4), write full SIMD vectors via ptr.cast::<Simd<u8,N>>().write_unaligned(), only call set_len() after success. On error, never commit — garbage writes vanish.error |= !ok from each iteration, return Err once after the loop. Errors are rare; don't pay branch cost per chunk.chunks_exact(N) for the hot path + Simd::from_slice(). For the remainder, load u64 from p and p+len-8 (overlap by 1 byte), OR them — 2 loads cover any 8–15 byte tail.Decision rules:
swizzle_dyn requires index vector length == lookup table length — pad the table if neededLaneCount<N>: SupportedLaneCountN by benchmark; on x86-64 with AVX2, N = 32 (one YMM) is usually optimalrayon::join(|| left(), || right()) for recursive divide-and-conquer — the primitive that powers Rayon; work-stealing ensures threads don't sit idle.par_sort() and .par_extend() instead of manual .par_iter().collect() — Rayon provides optimized parallel versions of stdlib algorithmsThreadPool to isolate core usage (e.g., reserve CPU for a web server):let pool = rayon::ThreadPoolBuilder::new().num_threads(4).build().unwrap();
pool.install(|| { /* parallel code here */ });Vec, slices) — Rayon splits chunks efficiently; LinkedList forces traversal before splittingcriterion — warm-up cost and orchestration overhead may negate gains| Feature | std::iter | rayon::par_iter | tokio::spawn |
|---|---|---|---|
| Best For | Small data / Simple logic | Big data / CPU-heavy | I/O / Networking |
| Overhead | Zero | Medium (task splitting) | High (runtime / context switch) |
| Execution | Sequential | Multi-threaded (parallel) | Concurrent (event loop) |
#[repr(u8)] on field-less enums to guarantee 1-byte size[u8; 1024]) instead of Vec::with_capacity() — eliminates allocator overhead in tight loops (~3.6× faster)papaya::HashMap<ThreadId, T> for per-thread WASM stores — lock-free reads on existing entries, uncontended Mutex per thread. Better than a single global Mutex for multi-threaded servers.(id, x, y) array alongside action indices and output results buffer.wasmtime::TypedFunc is cheap to clone (handle index). Clone it to release &self borrow before calling mutable Store methods — avoids borrow-checker conflicts with zero cost.Real-time render loops fail in ways the standard CPU profile doesn't catch. The hot CPU path can be 0% algorithmic work — it can be the event loop itself spinning between frames. Patterns below are distilled from the bevy_egui_orchard_demo 100%-CPU saga (commits d6fb415 → 565524f → 01b4687 in riir-game-sdk).
Bevy's default WinitSettings uses UpdateMode::Continuous, which spins the winit event loop as fast as possible. Even an empty Bevy app sits at ~50% CPU on macOS (Apple Silicon). This is a known Bevy issue — not user code.
Fix: switch to reactive mode so the loop parks between events:
use bevy::winit::{UpdateMode, WinitSettings};
use std::time::Duration;
app.insert_resource(WinitSettings {
// max_wait = sim tick rate. Scene can't visibly change faster than the
// sim produces state, so waiting match-tick-rate between idle redraws is free.
focused_mode: UpdateMode::reactive_low_power(Duration::from_millis(50)), // 20 Hz sim
unfocused_mode: UpdateMode::reactive_low_power(Duration::from_secs(1)),
});Vsync is independent. PresentMode::Fifo controls GPU-side blocking on the swapchain; it does NOT prevent the CPU event loop from spinning. You need BOTH PresentMode::Fifo (GPU vsync) AND WinitSettings::reactive_low_power (CPU loop).
NEVER redraw static line/mesh data via per-frame APIs:
// BAD: ~4k gizmos.line() calls per frame, each pushing to a per-frame CPU buffer
fn draw_overlays_system(mut gizmos: Gizmos, overlays: Res<Overlays>) {
for &[a, b] in &overlays.grid_lines { gizmos.line(a, b, COLOR); } // 728 calls
for &[a, b] in &overlays.isolines { gizmos.line(a, b, COLOR); } // 3450 calls
}
// GOOD: bake into a LineList mesh once at Startup, render cost is GPU-only
fn setup_scene(mut commands: Commands, mut meshes: ResMut<Assets<Mesh>>, overlays: Res<Overlays>) {
let mesh = build_line_list_mesh(&overlays.grid_lines, &overlays.isolines);
commands.spawn((
Mesh3d(meshes.add(mesh)),
MeshMaterial3d(materials.add(StandardMaterial { unlit: true, ..default() })),
));
}Despawning a Bevy Entity does NOT release its asset handles. The Assets<T> cache is keyed by AssetId and only grows.
// BAD: leaks mesh + material forever. Regrow-on-new-EntityId patterns
// (apples harvested + respawned) compound: each cycle adds N leaked assets.
for (be, _mirror) in &mirrors {
if !present.contains(&mirror.snap_id) {
commands.entity(be).despawn(); // entity gone, mesh + material still in cache
}
}
// GOOD: release assets alongside despawn
for (be, _mirror, mesh_h, mat_h) in &mirrors {
if !present.contains(&mirror.snap_id) {
meshes.remove(mesh_h.0.id());
materials.remove(mat_h.0.id());
commands.entity(be).despawn();
}
}Symptom: RSS grows linearly over a long session (e.g. 200MB → 900MB over an hour). The bevy_egui_orchard_demo saga hit exactly this.
bevy_egui (and similar immediate-mode GUI integrations) install a set_request_repaint_callback that fires bevy_winit::WakeUp events on delay.is_zero(). These bypass the reactive max_wait — the loop wakes immediately, processes the egui frame, egui repaints again, fires another WakeUp, infinite 60+ Hz loop.
Important nuance (verified in egui 0.33 source): plain ui.label(format!(...)) does NOT call Context::request_repaint(). Repaints originate only from: hover/click animations (only while in-progress), tooltips, scroll areas, CollapsingHeader, SidePanel/TopBottomPanel resize handles, explicit ctx.request_repaint() calls. A format!("frames {}", n) label whose value changes every frame does NOT itself drive repaints — but it IS wasteful (the displayed number flickers). Throttle per-frame-changing displayed values to 1 Hz for readability, not for perf.
Diagnose via in-process perf overlay before assuming the engine itself is the hog. If ORCHARD_NO_EGUI=1 drops CPU dramatically, the culprit is egui-side. Use finer-grained bisection (ORCHARD_NO_PERF_OVERLAY, ORCHARD_NO_MINIMAP, ORCHARD_NO_DEBUG_PANEL) to identify WHICH panel — typically it's a resize handle the cursor is parked over, or a hover-state animation that hasn't settled.
Mitigations:
ctx.set_request_repaint_callback to enforce a min delayRunningMode::Reactive if exposedA Last-schedule sleep is a brute-force defense when vsync + reactive mode somehow don't throttle (driver quirks, repaint storms, etc.):
fn frame_rate_cap_system(mut last_frame_end: Local<Option<Instant>>) {
const TARGET_FPS: u64 = 60;
let target = Duration::from_secs_f32(1.0 / TARGET_FPS as f32);
let now = Instant::now();
if let Some(last) = *last_frame_end {
let elapsed = now.duration_since(last);
if elapsed < target { std::thread::sleep(target - elapsed); }
}
*last_frame_end = Some(Instant::now());
}
app.add_systems(Last, frame_rate_cap_system);External ps -p $PID -o %cpu,rss sampling misses per-frame swings and requires a terminal attached. Ship a perf overlay inside the GUI itself so the user can see what's burning CPU while they interact:
use sysinfo::{Pid, ProcessRefreshKind, ProcessesToUpdate, System};
#[derive(Resource)]
struct PerfMonitor {
sys: System,
pid: Pid,
last_refresh: Instant,
cpu_pct: f32,
rss_bytes: u64,
fps_history: Vec<f32>,
cpu_history: Vec<f32>,
}
impl PerfMonitor {
fn refresh(&mut self) {
let now = Instant::now();
if now.duration_since(self.last_refresh) < Duration::from_millis(250) { return; }
self.sys.refresh_processes_specifics(
ProcessesToUpdate::Some(&[self.pid]), false, ProcessRefreshKind::everything(),
);
if let Some(p) = self.sys.process(self.pid) {
self.cpu_pct = p.cpu_usage(); // 100.0 = 1 core fully used
self.rss_bytes = p.memory(); // bytes on macOS, bytes on Linux
}
self.last_refresh = now;
}
}Note sysinfo MSRV: 0.32 supports Rust 1.93; 0.39+ requires 1.95. The sysinfo API changed between versions (refresh_process_specifics in 0.39 → refresh_processes_specifics(ProcessesToUpdate::Some(&[pid]), ...) in 0.32).
When a perf issue only manifests under specific conditions (focused window, mouse hover, etc.), ship env-var-driven kill switches so the user can isolate the culprit without code changes:
struct RunFlags {
no_egui: bool, // ORCHARD_NO_EGUI=1 — kills entire EguiPlugin
no_perf_overlay: bool, // ORCHARD_NO_PERF_OVERLAY=1 — hides only the 📊 perf window
no_minimap: bool, // ORCHARD_NO_MINIMAP=1 — hides only the left-panel minimap
no_debug_panel: bool, // ORCHARD_NO_DEBUG_PANEL=1 — hides only the bottom debug panel
no_shadows: bool, // ORCHARD_NO_SHADOWS=1 — disables shadow-map rendering
no_sim: bool, // ORCHARD_NO_SIM=1 — skips the per-tick sim step
no_overlays: bool, // ORCHARD_NO_OVERLAYS=1 — skips the static LineList overlay mesh
bench_secs: Option<f32>, // ORCHARD_BENCH=<n> auto-exits after n seconds
}The user runs each variant for ~10s, notes CPU%, and the env var that drops CPU is the culprit subsystem. Start coarse (NO_EGUI) then go fine (NO_PERF_OVERLAY / NO_MINIMAP / NO_DEBUG_PANEL) — this distinguishes "egui as a whole" from "a specific panel" and avoids mis-attributing the cause to the wrong surface.
// BAD: rayon overhead (~5μs) >> computation (~0.1μs per row)
let counts: Vec<_> = (0..10).into_par_iter()
.map(|i| compute_row(variants, i))
.collect();
// GOOD: serial for small m
let counts: Vec<_> = (0..10)
.map(|i| compute_row(variants, i))
.collect();Threshold: rayon wins only at m ≥ 64 with μs/row work, or m ≥ 1000 with ns/row work.
GPU kernel launch overhead is ~50μs. If your computation is 2-5μs, GPU is a net negative. GPU wins only for: batched matmul, large tensor ops, or when you can amortize launch across many ops.
// BAD: allocates every call, every sample, every position
for sample in 0..10 {
let support = rule.support(vocab_size); // Vec allocation!
}
// GOOD: pre-compute once, reuse
let config = Config::default().with_cached_data(size);
for sample in 0..10 {
let support = config.data_for(rule); // &[T] — zero alloc
}// BAD: O(n) scan per query
fn query(&self, key: usize) -> f32 {
for item in &self.items {
if item.key == key { ... }
}
}
// GOOD: O(1) precomputed index
struct Store {
stats: [SlotStats; MAX_SLOTS], // updated on insert/evict
}
fn query(&self, key: usize) -> f32 {
self.stats[key].rate() // O(1)
}// BAD: same value recomputed N×M times
for sample in 0..N {
for &pos in &positions {
let h = expensive_calc(data[pos]); // SAME value every sample!
}
}
// GOOD: compute once per position
let cache: Vec<f32> = positions.iter()
.map(|&pos| expensive_calc(data[pos]))
.collect(); // M calls instead of N×MAlways benchmark before AND after adding parallelism. If the serial version is faster, keep serial. Parallel overhead: thread wake (~2μs) + work stealing (~3μs) + synchronization. If your total work is < 10μs, parallelism will make it slower.
Mutex introduces contention — 16 threads fighting for one lock effectively run sequentially (or slower).
Prefer atomic types or reduce/fold patterns:
// BAD: shared Mutex — threads serialize on lock
let results = Mutex::new(Vec::new());
(0..1000).into_par_iter().for_each(|i| {
results.lock().unwrap().push(compute(i)); // contention!
});
// GOOD: map + collect — threads work independently, merge at end
let results: Vec<_> = (0..1000).into_par_iter()
.map(|i| compute(i))
.collect();If a closure inside a Rayon thread panics, Rayon propagates that panic to the calling thread.
This can crash your entire application if not handled at the top level. Wrap parallel closures
in catch_unwind or ensure invariants are validated before entering Rayon.
Splitting work too finely loses CPU cache benefits. Processing contiguous chunks is faster than jumping across memory addresses in parallel. Prefer chunk-based splitting over per-element parallelism when data is large but per-element work is small.
Adding code behind a feature flag still affects the entire binary when enabled:
Mitigation:
[[bin]]) or test filesWASM fuel limits prevent infinite loops but can silently trap legitimate computation. Complex BFS/graph algorithms with N entities on bounded domains can spike well above average:
// BAD: fuel based on average case, traps on worst case
const FUEL_PER_CALL: u64 = 10_000; // sufficient for 1–2 bombs
// BFS with 4+ bombs × 4 directions × range × 169 cells = ~40K ops → SILENT TRAP
// GOOD: fuel based on worst-case analysis + headroom
const FUEL_PER_CALL: u64 = 50_000; // 16 bombs × 4 dirs × range 3 × 169 cells ≈ 40K + marginSymptom: WASM returns false for valid inputs that should return true. Only manifests with complex inputs. Batch APIs may mask this if they use higher fuel multipliers. Fuzz-test with maximum entity counts to catch fuel traps.
When validating N items against the same state (e.g., N players on one game grid), serializing the state N times wastes both allocation and FFI overhead:
// BAD: 24 × (serialize + FFI + compute) = ~12µs/tick
for player in 0..4 {
for action in 0..6 {
let state = serialize(grid, player, action); // 24 serializations!
wasm.is_valid(state); // 24 FFI calls!
}
}
// GOOD: 1 × (serialize + FFI + batch compute) = ~1.7µs/tick
let state = serialize_grid(grid, bombs); // 1 serialization
wasm.batch_validate(state, players, actions, results); // 1 FFI callThe batch API turns N×M individual calls into 1 call. The WASM module internally loops over all combinations, reusing the parsed state. For 4 players × 6 actions, this gives ~5.8× speedup.
Laptop CPUs throttle aggressively. A 30% "regression" may just be heat. Always compare same-commit, back-to-back runs to isolate feature impact from system noise.
// tests/prof_bench.rs — run with: cargo test --features X prof_bench -- --nocapture
#[cfg(feature = "X")]
#[test]
fn prof_components() {
let warmup = 100;
let iters = 10000;
for _ in 0..warmup { black_box(component_a()); }
let start = Instant::now();
for _ in 0..iters { black_box(component_a()); }
let t_a = start.elapsed();
// ... same pattern for component_b, component_c ...
println!(" Component A: {:.2} μs", t_a.as_micros() as f64 / iters as f64);
println!(" Total Δ: {:.2} μs", total.as_micros() as f64 / iters as f64);
}// Pattern: batch validate N items × M actions in one FFI call
//
// Memory layout written to WASM:
// [0..state_end) shared state (grid + bombs, no per-entity data)
// [players_off..+N×12) entity array: N × (id, x, y) as u32 LE
// [actions_off..+M×4) action indices as u32 LE
// [results_off..+N×M×4) output: u32 LE results (0/1 or Q16.16)
//
// WASM export signature:
// batch_is_valid(state_ptr, state_len, players_ptr, player_count,
// actions_ptr, action_count, results_ptr) -> u32
const MAX_ENTITIES: usize = 4;
const ACTION_COUNT: usize = 6;
const ACTIONS_BYTES: [u8; ACTION_COUNT * 4] = [0,0,0,0, 1,0,0,0, 2,0,0,0, 3,0,0,0, 4,0,0,0, 5,0,0,0];
fn batch_validate(&self, grid: &Grid, players: &[(u8,i32,i32)], bombs: &[Bomb]) -> BatchResult {
self.with_inner(|inner| {
// 1. Serialize shared state once (zero-copy stack buffer)
let (state_bytes, state_tokens) = inner.state_buf.serialize_grid(grid, bombs);
let mut tmp = [0u8; 1024];
tmp[..state_bytes].copy_from_slice(inner.state_buf.as_bytes(state_bytes));
// 2. Compute aligned offsets
let players_off = (state_bytes + 7) & !7; // align8
let actions_off = players_off + players.len() * 12;
let results_off = actions_off + ACTION_COUNT * 4;
// 3. Write to WASM memory
inner.write_memory(0, &tmp[..state_bytes])?;
inner.write_memory(players_off, &players_to_bytes(players))?;
inner.write_memory(actions_off, &ACTIONS_BYTES)?;
// 4. Call batch export
let batch_fn = inner.batch_fn.as_ref()?.clone();
batch_fn.call(&mut inner.store, (0, state_tokens, players_off as u32,
players.len() as u32, actions_off as u32, ACTION_COUNT as u32,
results_off as u32))?;
// 5. Read results
Some(BatchResult::from_memory(inner, results_off, players.len(), ACTION_COUNT))
})
}© katopz, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in .agents/skills/rust-optimize of katopz/katgpt-rs.
Open the folder on GitHubat commit d0b32e2
Rust Optimize next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Rust Optimize this skillkatopz/katgpt-rs | 134 | — | ~7k | Automated safety check: Pass | MIT | |
| Update V8 Versionopeninterpreter/openinterpreter | 69k | 2 repos | ~845 | Automated safety check: Pass | Apache-2.0 | |
| Firecrawl Page Scrape Integrationfirecrawl/firecrawl | 189k | 1 repos | ~944 | Automated safety check: Pass | ISC | |
| Migrate Core Code to Submodulestinyhumansai/openhuman | 41k | — | ~2.6k | Automated safety check: Pass | GPL-3.0 | |
| Rust TDD Workflowrtk-ai/rtk | 83k | — | ~753 | Automated safety check: Notes | Apache-2.0 | |
| Rust Best Practicesfarm-fe/farm | 5.6k | 3 repos | ~1.1k | Automated safety check: Pass | MIT |
openinterpreter/openinterpreter
Bumps the pinned v8 and rusty_v8 versions in Codex, validates the release-candidate path with the v8-canary check, and traces failures to upstream build changes.
firecrawl/firecrawl
Adds Firecrawl's /scrape endpoint to application code to pull markdown, HTML, links, screenshots or structured data from a single known URL.
tinyhumansai/openhuman
Plans and carries out moving non-host-specific code and its tests from the OpenHuman core into vendored tiny submodule libraries, then releases the submodule and re-pins the host.
rtk-ai/rtk
Enforces red-green-refactor for Rust work, with idiomatic test patterns, a naming convention and a pre-commit gate of cargo fmt, clippy and test.
farm-fe/farm
Guide for writing idiomatic Rust code based on Apollo GraphQL's best practices handbook.
AprilNEA/OpenLogi
Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.
katopz/katgpt-rs
Write a reasoned architectural proposal (.proposals/NNN.md) grounded in focused codebase grep + prior-art paper search.
katopz/katgpt-rs
Audit + enforce game-stack boundary rules across the multi-repo workspace.
katopz/katgpt-rs
Audit feature-gate status claims across the multi-repo stack.
katopz/katgpt-rs
Audit cross-repo GOAT/gain primitive cherry-pick status across the multi-repo stack (katgpt-rs upstream, riir- consumers).
katopz/katgpt-rs
Research workflow for distilling ML/AI papers into modelless inference primitives, freeze/thaw runtime patterns, latent-space operations, AND model-based training plans across the multi-repo stack.
katopz/katgpt-rs
Pre-implementation DRY gate + existing-code drift audit for the multi-repo workspace.
Works with
Optimize Rust code until nothing left to improve. An agent skill from katopz/katgpt-rs. Rust Optimize is an agent skill from katopz/katgpt-rs. Optimize Rust code until nothing left to improve.
Rust Optimize fits situations like: the user says optimize.
Run `npx skills add katopz/katgpt-rs --skill rust-optimize -a claude-code`. Or copy the skill folder (.agents/skills/rust-optimize in katopz/katgpt-rs) into .claude/skills/rust-optimize in your project. Claude Code loads it when a task matches its description.
Run `npx skills add katopz/katgpt-rs --skill rust-optimize -a codex`. Or copy the skill folder (.agents/skills/rust-optimize in katopz/katgpt-rs) into .agents/skills/rust-optimize in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add katopz/katgpt-rs --skill rust-optimize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rust-optimize, .gemini/skills/rust-optimize, .github/skills/rust-optimize and .opencode/skills/rust-optimize in your project.
Going by SKILL.md and its folder, Rust Optimize needs Rust for the scripts in its folder.
SKILL.md names 1 domain. As links in the text: mcyoung.xyz. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Rust Optimize is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Rust Optimize: Update V8 Version (openinterpreter/openinterpreter, 69k stars), Firecrawl Page Scrape Integration (firecrawl/firecrawl, 189k stars), Migrate Core Code to Submodules (tinyhumansai/openhuman, 41k stars) and Rust TDD Workflow (rtk-ai/rtk, 83k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
katopz (a GitHub user) maintains it in katopz/katgpt-rs, which has 134 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.
Source: katopz/katgpt-rs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.