Optimize or review performance for the Static Web Server (SWS) project — profiling, bottlenecks, resource usage, compression, and caching

Apache-2.0Auto-check passedBackend & APIs

Install Performance

skills CLI
$ npx skills add static-web-server/static-web-server --skill performance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install static-web-server/static-web-server performance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/static-web-server/static-web-server.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/performance .claude/skills/performance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
performance
GitHub stars
2.4k
Token cost
~3.1k tokens
SKILL.md length
1,529 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Optimize or review performance for the Static Web Server (SWS) project — profiling, bottlenecks, resource usage, compression, and caching

  • Tasks that involve Caching
  • SKILL.md covers General Approach, Rust Performance, HTTP Performance and File I/O Performance, plus 6 more sections
  • Calls cargo

What it does

Performance is an agent skill from static-web-server/static-web-server. Optimize or review performance for the Static Web Server (SWS) project — profiling, bottlenecks, resource usage, compression, and caching

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Backend & APIs, covering Caching. It works with Rust and Linux. The repository describes itself as: A cross-platform, high-performance and asynchronous web server for static files-serving. ⚡. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Caching

Example prompts

  • “/performance”

What it can do on your machine

Read from SKILL.md and the folder at commit 32ec4aa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • valgrind.org
    • crates.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Performance loads about 3.1k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 1,529 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from static-web-server/static-web-server at commit 32ec4aa, republished under its Apache-2.0 licence (© static-web-server). 1,529 words, ~3,110 tokens.

Download SKILL.mdSave it as .claude/skills/performance/SKILL.md (or your agent's skills folder).
name
performance
description
Optimize or review performance for the Static Web Server (SWS) project — profiling, bottlenecks, resource usage, compression, and caching

Performance Optimization

Load this skill when profiling, optimizing, or reviewing code for performance — latency, throughput, memory, or CPU.

When to load: a change touches the request hot path (handler.rs, static_files/, compression.rs, fs/stream.rs), a benchmark is being added or interpreted, a regression in requests/sec or latency is suspected, or a new dependency may affect binary size or runtime cost.

General Approach

  • Measure before optimizing: Use profiling tools (perf, flamegraph) to identify bottlenecks. Never optimize based on intuition
  • Set a target: Define acceptable latency/throughput before starting. Stop optimizing when the target is met
  • Optimize the hot path: Focus on code that runs on every request. Startup code and config parsing are low priority

Rust Performance

  • Profile with perf and flamegraph: cargo flamegraph --bin static-web-server for CPU profiles. Profile under load (e.g., wrk or bombardier)
  • Profile heap allocations with DHAT: Use DHAT or dhat-rs to find hot allocation sites. Reducing 10 allocations per million instructions can have measurable impact
  • Avoid unnecessary allocations: Prefer &Path over &PathBuf, &[u8] over Vec<u8>, pass by reference where ownership is not needed
  • Pre-compute at startup: Canonicalize paths, parse config, compile regex patterns, build Aho-Corasick automata once. Never on the request path
  • Pre-allocate collections: Use Vec::with_capacity, String::with_capacity when the size is known
  • Stream large responses: Use tokio::fs::File + tokio::io::copy for file serving. Never buffer the full file in memory
  • Inline small hot functions: Use #[inline] on small functions called on every request (e.g., header name normalization, MIME lookups). Use #[cold] on error-path functions to guide branch prediction away from the hot path
  • Prefer filter_map over filter().map(): Avoids an intermediate layer in hot iterator chains
  • Use iter().copied() for small types: When iterating over &u8, &u32, etc., .copied() lets LLVM generate better code than receiving references
  • Use chunks_exact when chunk size evenly divides length: Faster than chunks because it eliminates a remainder check per iteration
  • Prefer ok_or_else over ok_or: ok_or(expensive()) always evaluates its argument. ok_or_else(|| expensive()) is lazy and only evaluates on None
  • Eliminate bounds checks in hot loops: Use iteration instead of index-based access, or add an upfront assertion on the range to let the compiler prove bounds are safe

HTTP Performance

Connection Handling
  • HTTP/1.1 keep-alive: Enabled by default via Hyper. Reduces connection setup overhead for subsequent requests
  • HTTP/2 multiplexing: Enable with --http2 --tls. Multiple concurrent streams over a single TCP connection
  • Worker threads: Default is num_cpus * 1. Increase --threads-multiplier for workloads with mixed CPU and I/O blocking (e.g., many concurrent clients with dynamic compression enabled — compression per-request is CPU-bound but high concurrency adds I/O wait interleaving). For pure CPU-bound workloads with minimal I/O, increasing threads beyond CPU count rarely helps.
  • Max blocking threads: Default 512. For I/O-heavy patterns (large file serving), this is sufficient
  • Graceful shutdown: Use --grace-period to allow in-flight requests to complete before shutdown
Compression Tradeoffs
  • Static compression is free: Pre-compressed .br/.gz/.zst files are served with zero CPU. Always prefer this for production
  • Dynamic compression overhead: On-the-fly compression trades CPU for bandwidth. Use --compression-level fastest for high-traffic sites
  • Minimum size threshold: Responses below 860 bytes skip dynamic compression entirely — the overhead exceeds any bandwidth savings
  • Compression algorithm priority (by compression ratio × speed): zstd > brotli > gzip > deflate. zstd offers the best ratio-speed tradeoff
Caching Headers
  • Cache-Control is enabled by default: SWS sets max-age based on file extension:
    • 1 year for static assets (.css, .js, .png, .woff2, etc.)
    • 1 hour for feeds/API (.json, .xml, .rss, .atom)
    • 1 hour fallback for unknown extensions

You can override these defaults per file or extension using the configuration file. The above values are defaults, not hardcoded limits.

  • Conditional requests: SWS supports If-Modified-Since and If-Unmodified-Since via ConditionalHeaders. Returns 304 when the file hasn't changed
  • ETag not implemented: SWS uses Last-Modified instead. For byte-level cache validation, put SWS behind a CDN or reverse proxy

File I/O Performance

Buffering
  • Optimal buffer size: optimal_buf_size() selects the best buffer size based on file metadata (uses std::fs::Metadata::blksize() when available)
  • BufReader with take(): For byte-range requests, a BufReader wraps the file handle and limits bytes read to the requested range
  • Streaming avoids full-file buffering: FileStream reads in chunks. The response body is a stream, not a byte buffer
Path Operations
  • Canonicalize once at startup: The root directory is canonicalized in server/opts.rs. Per-request path resolution reuses this
  • try_metadata() caches nothing: Each call (in src/fs/meta.rs) is a filesystem syscall. The experimental memory cache feature (mini-moka, in src/mem_cache/) caches file metadata and content
  • Avoid clone() in the hot path: static_files.rs avoids cloning file paths for non-directory requests
Pre-compressed Static Files
  • Zero-CPU serving: SWS detects .br/.gz/.zst variants via Accept-Encoding and serves them directly. No compression step runs
  • Build-time pre-compression: Generate variants with maximum quality: brotli -q 11, gzip -9, zstd -19. SWS serves them as-is
  • Vary header: Vary: Accept-Encoding is appended so caches know to store multiple variants

Memory

  • Minimal per-connection state: SWS stores only the remote address and handler opts (shared via Arc). No per-connection buffers
  • Response body is a stream: File contents are streamed, not buffered. Exception: small generated responses (health endpoint, error pages, directory listing HTML)
  • Experimental in-memory cache: mini-moka (in src/mem_cache/) caches hot files in memory with LFU admission and LRU eviction. Configurable capacity (default 100 entries), ttl (default 1800s), tti, and max_file_size. Keys use CompactString to reduce allocation
Show full SKILL.md (671 more words)Show less

Allocation Patterns

  • Prefer clone_from over reassign-and-clone: a.clone_from(&b) reuses a's existing heap allocation when possible, avoiding an extra alloc/free. Especially valuable for Vec and String in hot loops
  • Reuse collections across iterations: Declare the collection outside the loop, call .clear() at the end of each iteration. Avoids repeated alloc/free while keeping the heap allocation alive
  • Use Cow<'_, str> / Cow<'_, Path> for mixed borrowed/owned data: Avoids allocating a String/PathBuf when the data is already a static literal or an existing slice that won't be modified
  • Use SmallVec<[T; N]> for short, stack-like sequences: When most allocations hold ≤ N elements (e.g., header value lists, index file candidates), smallvec avoids heap allocation entirely for the common case
  • Convert finalized Vec to Box<[T]> with into_boxed_slice(): Drops the unused capacity word, shrinking the type from 3 words to 2. Good for config-time data that is built once and never grown
  • Return impl Iterator<Item=T> instead of Vec<T> from helpers: Avoids an allocation when the caller only needs to iterate
  • Avoid format! when a literal or write! suffices: Every format! call allocates a String. Write directly to a &mut String or use std::fmt::Write instead

Type Sizes

  • Keep hot types under 128 bytes: The compiler emits memcpy for values larger than 128 bytes. If a hot type exceeds this, check its layout with RUSTFLAGS=-Zprint-type-sizes cargo +nightly build --release
  • Box large enum variants: If one variant is much larger than the others, box its fields to bring all variants to a similar small size. Reduces stack pressure and cache churn
  • Use smaller integer types for index/count fields: Prefer u32 over usize for counts and offsets stored in frequently instantiated structs (e.g., header tables, path segments). Cast to usize at use sites
  • Assert type sizes in tests: Add static_assertions::assert_eq_size!(HotType, [u8; N]); for performance-critical types so that accidental size regressions cause a compile error

Release Build Configuration

The default cargo build --release profile is a good starting point, but the following options can improve throughput for production SWS builds:

OptionEffectCargo.toml
lto = "thin"Cross-crate inlining, 5–15% speedup, moderate compile cost[profile.release]
codegen-units = 1Single codegen unit, enables more optimizations, slower compile[profile.release]
panic = "abort"Removes unwinding machinery, smaller binary, slight speedup[profile.release]

For a custom server build where broad CPU compatibility is not required:

RUSTFLAGS="-C target-cpu=native" cargo build --release

This emits AVX/SSE instructions optimal for the build machine, which can improve compression throughput.

Note: target-cpu=native produces a non-portable binary. Do not use for distributed release artifacts.

Benchmarking

Micro-benchmarks (CodSpeed)

The benches/ directory is a standalone crate with Criterion benchmarks for hot-path functions (path sanitization, header handling, redirects, basic auth). They run on every pull request via the codspeed workflow and are tracked on CodSpeed.

bash
cd benches

# Plain Criterion run (local timings only)
cargo bench

# Same suite measured with CodSpeed CPU simulation (deterministic, flamegraphs)
cargo codspeed build
codspeed run --mode simulation -- cargo codspeed run

Add a bench when a change touches the request hot path so a future regression is caught by CI instead of by users.

Tools
  • HTTP load generators: wrk, bombardier, oha, hey
  • CPU profiling: perf record + flamegraph, cargo flamegraph, cargo instruments (macOS), samply (cross-platform)
  • Allocation profiling: dhat-rs (all platforms), DHAT via Valgrind (Linux) — identifies hot allocation sites
  • Monitoring: Prometheus metrics via --metrics + /metrics endpoint
What to Measure
  • Requests per second at concurrency levels: 1, 10, 100, 1000
  • Latency percentiles: p50, p95, p99
  • Memory usage: RSS before and under load
  • CPU utilization: Per-core usage during load test
  • Allocation rate: Use DHAT to confirm per-request allocations are not growing unexpectedly

Checklist

  • Is there a benchmark or profile showing the bottleneck?
  • Are paths canonicalized once, not per-request?
  • Is static compression used where possible (zero CPU)?
  • Is dynamic compression size-threshold applied (860 bytes)?
  • Are large files streamed, not buffered?
  • Are Cache-Control headers set appropriately for the content type?
  • Are worker threads configured for the workload?
  • Is keep-alive or HTTP/2 enabled for connection reuse?
  • Are hot allocation sites identified with DHAT or similar?
  • Are collections pre-allocated or reused rather than recreated per request?
  • Are clone() calls on the hot path justified — or replaceable with clone_from, Cow, or a reference?
  • Are hot types under 128 bytes (no unintended memcpy)?
  • Are #[inline] / #[cold] attributes applied where profiling shows they help?
  • Are ok_or_else / lazy combinators used instead of eager alternatives in hot paths?

© static-web-server, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/performance of static-web-server/static-web-server.

Open the folder on GitHubat commit 32ec4aa

Compare with similar skills

Performance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Performance compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Performance this skillstatic-web-server/static-web-server2.4k—~3.1kAutomated safety check: PassApache-2.0
Gpg SigningProrise-cool/Claude-Code-Multi-Agent305—~3.9kAutomated safety check: WarnNone
Golem Mark Read Only Rustgolemcloud/golem1.5k—~1.6kAutomated safety check: PassCustom licence
Caching Architecturemajiayu000/litellm-rs116—~2kAutomated safety check: PassMIT
macOS Cleanerdaymade/claude-code-skills1.4k—~8.7kAutomated safety check: PassMIT
Apple Container Test RunnerRustPython/RustPython22k—~467Automated safety check: PassMIT

Similar skills

  • Gpg Signing

    Prorise-cool/Claude-Code-Multi-Agent

    Comprehensive guide to GPG commit signing. An agent skill from Prorise-cool/Claude-Code-Multi-Agent.

    305 GitHub stars~3.9k tokensUpdated 21 days ago
    Backend & APIsAuto-check: warnings
  • Marking Rust agent methods as read-only for a side-effect-free guarantee and result caching.

    1.5k GitHub stars~1.6k tokensUpdated today
    Backend & APIsAuto-check passed
  • Caching Architecture

    majiayu000/litellm-rs

    LiteLLM-RS response caching architecture. An agent skill from majiayu000/litellm-rs.

    116 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • macOS Cleaner

    daymade/claude-code-skills

    Diagnoses and safely reclaims macOS disk space: caches, logs, app remnants, large/duplicate files, Docker/OrbStack, Chromium code-sign clones, developer caches.

    1.4k GitHub stars~8.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Apple Container Test Runner

    RustPython/RustPython

    Runs RustPython tests inside a Linux container built with Apple's container CLI, so macOS users can compare Linux results with their local ones.

    22k GitHub stars~467 tokensUpdated today
    Testing & QAAuto-check passed
  • CI Pipeline Synthesizer

    kajisho5/ffmpeg-skill

    Generate GitHub Actions CI/CD pipeline configurations for automated building and testing of library and package projects.

    1.9k GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed

More from static-web-server/static-web-server

All 9 skills in this repo
  • Design

    static-web-server/static-web-server

    Design or review software architecture, API contracts, data models, and module boundaries for the Static Web Server (SWS) project

    2.4k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Issue Tracking

    static-web-server/static-web-server

    Triage, debug, fix, and document issues for the Static Web Server (SWS) project — bug reports, root cause analysis, fix implementation, and regression prevention

    2.4k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Prose

    static-web-server/static-web-server

    Author or edit any prose for the Static Web Server (SWS) project — documentation, design docs, READMEs, PR descriptions, issue bodies, commit message bodies, or other human-readable text — following…

    2.4k GitHub stars~971 tokensUpdated yesterday
    Auto-check passed
  • Rust Backend

    static-web-server/static-web-server

    Write or review Rust backend code for the Static Web Server (SWS) project — crates, modules, functions, types, error handling, and async code

    2.4k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Static File Serving

    static-web-server/static-web-server

    Serve static files and web assets with optimal headers, MIME types, compression, and caching for the Static Web Server (SWS) project

    2.4k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Testing

    static-web-server/static-web-server

    Write or review tests for the Static Web Server (SWS) project — unit tests, integration tests, test fixtures, and mocking strategies

    2.4k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Performance

What does Performance do?

Optimize or review performance for the Static Web Server (SWS) project — profiling, bottlenecks, resource usage, compression, and caching. Performance is an agent skill from static-web-server/static-web-server.

When should I use Performance?

Performance fits situations like: tasks that involve Caching.

How do I install Performance in Claude Code?

Run `npx skills add static-web-server/static-web-server --skill performance -a claude-code`. Or copy the skill folder (.agents/skills/performance in static-web-server/static-web-server) into .claude/skills/performance in your project. Claude Code loads it when a task matches its description.

How do I install Performance in Codex?

Run `npx skills add static-web-server/static-web-server --skill performance -a codex`. Or copy the skill folder (.agents/skills/performance in static-web-server/static-web-server) into .agents/skills/performance in your project. Codex loads it when a task matches its description.

Can I use Performance in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add static-web-server/static-web-server --skill performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/performance, .gemini/skills/performance, .github/skills/performance and .opencode/skills/performance in your project.

What does Performance need to run?

Going by SKILL.md and its folder, Performance needs the command-line tools its instructions call (cargo).

Does Performance access the network?

SKILL.md names 2 domains. As links in the text: valgrind.org and crates.io. This is read from the text; nothing was executed.

Is Performance safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Performance use?

Performance is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Performance use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Performance?

Skills that share tags, products or a category with Performance: Gpg Signing (Prorise-cool/Claude-Code-Multi-Agent, 305 stars), Golem Mark Read Only Rust (golemcloud/golem, 1.5k stars), Caching Architecture (majiayu000/litellm-rs, 116 stars) and macOS Cleaner (daymade/claude-code-skills, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Performance?

static-web-server (a GitHub organization) maintains it in static-web-server/static-web-server, which has 2,373 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 5, 2026.

Source: static-web-server/static-web-server on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.