Agent skill

Compute Mamba Ratio

by sgl-project in sgl-project/sglang

Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

Apache-2.0Auto-check passed

Install Compute Mamba Ratio

skills CLI
$ npx skills add sgl-project/sglang --skill compute-mamba-ratio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang compute-mamba-ratio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/compute-mamba-ratio .claude/skills/compute-mamba-ratio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
compute-mamba-ratio
GitHub stars
37k
Used in
2 other repos
Token cost
~2.9k tokens
SKILL.md length
1,466 words
Files
1
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

  • Works in 7 steps: L — average context (input + output) in… → The two per-GPU byte constants — one of → S — from --mamba-radix-cache-strategy,… → …
  • A user asks what ratio to set
  • SKILL.md covers The formula, Inputs to collect from the user, Compute and Procedure, plus 2 more sections
  • Calls mamba

What it does

Compute Mamba Ratio is an agent skill from sgl-project/sglang. Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and serving config. Use when a user asks what ratio to set, why concurrency is clamped, or how to size the state vs KV pools for a hybrid model.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • A user asks what ratio to set
  • Why concurrency is clamped
  • How to size the state vs KV pools for a hybrid model

Example prompts

  • “/compute-mamba-ratio”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. L — average context (input + output) in tokens.
  2. The two per-GPU byte constants — one of
  3. S — from --mamba-radix-cache-strategy, the overlap scheduler, and SGLANG_OPT_MAMBA_SKIP_DECODE_LOCK (table below).
  4. D — --speculative-num-draft-tokens (0 if NOSPEC; also 0 when ReplaySSM spec-verify is enabled — see caveats).
  5. dcp_size — --dcp-size (1 if no DCP). If DCP and spec with a replicated draft KV, apply the draft caveat (below).
  6. KV dtype (bf16 / fp8) and ssm dtype (fp32 / bf16) — they set the two byte constants (see the token_equiv 2×2).
  7. (only to also predict the clamp, not just r) rest = per-GPU memory − weights at the chosen --mem-fraction-static. Read avail mem after…

What it can do on your machine

Read from SKILL.md and the folder at commit f620d73. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • mamba

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Compute Mamba Ratio loads about 2.9k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 1,466 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit f620d73, republished under its Apache-2.0 licence (© sgl-project). 1,466 words, ~2,945 tokens.

Download SKILL.mdSave it as .claude/skills/compute-mamba-ratio/SKILL.md (or your agent's skills folder).
name
compute-mamba-ratio
description
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and serving config. Use when a user asks what ratio to set, why concurrency is clamped, or how to size the state vs KV pools for a hybrid model.

Optimal hybrid dual-pool ratio (--mamba-full-memory-ratio)

A hybrid model (attention layers + linear-attention layers — the recurrent-state family: Mamba/SSM, GDN, KDA, etc.) splits serving memory into two independently-budgeted pools, fixed once at startup:

  • state pool (the linear-attention recurrent state) → caps concurrency (hard: whole slots, worst-case reserved, fail-loud)
  • full-KV pool (attention KV) → caps context × concurrency (soft: paged, over-committable via retraction)

--mamba-full-memory-ratio r splits the post-weight budget: mamba_budget = rest · r/(1+r), i.e. mamba_budget : kv_budget = r. This skill picks the r (or the pin---max-mamba-cache-size alternative) at which neither pool bottlenecks first for the user's workload.

The formula

r*  =  (S + D) · token_equiv · dcp_size / L
token_equiv  =  state_bytes_per_slot / kv_bytes_per_token
  • L = average context length per request (input + output tokens)
  • token_equiv = full-KV token-equivalent of one state slot
  • S = state slots per running request (cache-strategy dependent, table below)
  • D = --speculative-num-draft-tokens (0 if NOSPEC); each running req carries D extra intermediate states
  • dcp_size = --dcp-size (1 without DCP). DCP shards the per-rank KV by dcp_size, so KV gets ~dcp_size× cheaper per request → the balance shifts that much toward the state pool.

r is dimensionless (just the split). To also predict the actual concurrency you need rest (below).

Inputs to collect from the user

  1. L — average context (input + output) in tokens.
  2. The two per-GPU byte constants — one of:
    • (a) measured (preferred, exact) — from one boot log at any ratio:
      • Mamba Cache is allocated. ... ssm_state size X GB with max_mamba_cache_size: N → state_bytes_per_slot = X / N
      • KV Cache is allocated. #tokens: M, KV size: Y GB → kv_bytes_per_token = Y / M
    • (b) derived — model arch (linear-layer count + state dims d_state/d_conv/heads/head_dim; attention type + KV dims: MLA latent dim, or GQA kv_heads·head_dim·layers) × the dtypes below.
  3. S — from --mamba-radix-cache-strategy, the overlap scheduler, and SGLANG_OPT_MAMBA_SKIP_DECODE_LOCK (table below).
  4. D — --speculative-num-draft-tokens (0 if NOSPEC; also 0 when ReplaySSM spec-verify is enabled — see caveats).
  5. dcp_size — --dcp-size (1 if no DCP). If DCP and spec with a replicated draft KV, apply the draft caveat (below).
  6. KV dtype (bf16 / fp8) and ssm dtype (fp32 / bf16) — they set the two byte constants (see the token_equiv 2×2).
  7. (only to also predict the clamp, not just r) rest = per-GPU memory − weights at the chosen --mem-fraction-static. Read avail mem after Load weight end, or Memory pool end. avail mem + pool sizes, from the boot log.
S (state slots per running request), set by --mamba-radix-cache-strategy

All strategies keep prefix caching on; they differ in the track buffer that snapshots chunk-boundary state under the overlap scheduler. S = base + ping-pong, where base is 3 (live state + radix retention/COW headroom), and ping-pong = 2 (overlap on, non-lazy) / 1 (lazy, or overlap off) / 0 (no track buffer at all). This mirrors kv_cache_configurator._calculate_mamba_ratio; read it there if a release moves the constants.

SGLANG_OPT_MAMBA_SKIP_DECODE_LOCK=1 takes the base to 2, since the decode-time skip frees one resident slot per running request. no_buffer is the exception and stays at an effective 3: it adds that slot straight back, because its binding limit is the prefill-to-decode peak, which a decode-time change does not shrink.

strategySS with decode-lock skipnotes
--disable-radix-cache11prefix cache off entirely; most concurrency, no reuse
no_buffer33no track buffer; requires overlap off (+ page_size 1) → lower decode throughput
extra_buffer (default), overlap on54track buffer reserved per running request
extra_buffer, overlap off43same buffer, but ping-pong costs 1 slot instead of 2
extra_buffer_lazy (recommended)43track buffer allocated lazily at the boundary → 1 fewer slot/req; requires overlap on

The skip is a real concurrency lever, not a rounding detail: at the default strategy it takes S from 5 to 4, so the same state budget admits 25% more requests and r* drops by the same factor.

Overlap counts as off when --disable-overlap-schedule is passed or --pp-size > 1 (pipeline parallelism turns it off for you). So a PP deployment left on the default strategy sits on the S = 4 row, not S = 5 — using 5 there overstates the state cost and starves KV.

token_equiv — compute it from the two byte constants

token_equiv = state_bytes_per_slot / kv_bytes_per_token. Both bytes are per-model, per-GPU — read them from a boot log (or derive from the arch + dtypes) and divide. That's the whole thing; the dtype effects below are just for a quick mental estimate.

The recurrent state has two parts: state_bytes = SSM_bytes + conv_bytes. The SSM part scales with --mamba-ssm-dtype; the conv part is always bf16 (SGLANG_MAMBA_CONV_DTYPE, fixed). How the dtype knobs move token_equiv:

  • fp8 KV → ×2 (universal, exact): fp8 is 1 byte vs bf16's 2, so kv_bytes_per_token halves.
  • fp32 ssm vs bf16 ssm → ×≈2: only the SSM tensor doubles, conv stays fixed, so the exact factor is (2·SSM + conv)/(SSM + conv) — just under 2. conv is usually small relative to SSM, so estimate ≈2 (e.g. one measured model: conv ≈7% of state → factor ≈1.9). Only bother with the exact split if a model's conv is unusually large.
  • fp32 ssm + fp8 KV → ×≈4.

For an actual number, don't apply factors — read state_bytes_per_slot and kv_bytes_per_token for your model+dtype from the boot log and divide.

The GB the boot log prints for both pools is GiB (1024³). Convert both the same way and token_equiv is unaffected either way, since it is a ratio — but get it right before comparing an absolute byte figure against your own log.

Worked example (one measured model, TP8, fp32 ssm + fp8 KV), read off one boot log:

Mamba Cache is allocated. max_mamba_cache_size: 257, conv_state size: 0.46GB, ssm_state size: 13.04GB
KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 1167552, KV size: 15.03 GB

state_bytes_per_slot = (0.46 + 13.04) GiB / 257 ≈ 56.4 MB, kv_bytes_per_token = 15.03 GiB / 1167552 ≈ 13.8 KB → token_equiv ≈ **4080**.

Show full SKILL.md (585 more words)Show less

Compute

python
def optimal_ratio(L, state_bytes_per_slot, kv_bytes_per_token, S, D=0, dcp_size=1):
    token_equiv = state_bytes_per_slot / kv_bytes_per_token
    r = (S + D) * token_equiv * dcp_size / L
    return r  # value of --mamba-full-memory-ratio (>1 is legal)

def predict_clamp(rest_bytes, r, state_bytes_per_slot, S, D=0):
    mamba_budget = rest_bytes * r / (1 + r)
    slots = mamba_budget / state_bytes_per_slot
    # spec: each running req reserves (S+D) worth; non-spec just S
    return int(slots // S)  # NOSPEC; with spec the budget joint-solves for (S+D)·per_req per req

Then state the result three ways: the r value, the predicted clamp (if rest given), and which pool binds (min(mamba_clamp, KV_cap), where KV_cap = kv_tokens · dcp_size / L).

Procedure

  1. Free lever first: if the boot log shows large idle in Memory pool end. avail mem (e.g. 30–40 GB at mem-frac 0.85), raise --mem-fraction-static (→0.92) before touching the split — it grows rest for both pools at no cost. Validate graph-capture headroom once.
  2. Collect the inputs. Prefer a real boot log for the two byte constants.
  3. Compute r*. If r* would drive the state pool below one request's worth (mamba_budget < S · per_req, happens at very long L), switch to pinning --max-mamba-cache-size = target_concurrency · S and let the rest go to KV — a sub-0.15 r is fragile.
  4. Report r, predicted clamp, binding pool, and any dtype accuracy gate that applies.

Worked examples (TP8, B300, validated against measured clamps)

Constants (from boot logs): state_bytes_per_slot ≈ 56.4 MB (fp32 ssm), kv_bytes_per_token ≈ 13.8 KB (fp8) → token_equiv ≈ 4080. Cache strategy extra_buffer_lazy → S = 4. Workload L = 9216 (8192 in + 1024 out). NOSPEC → D = 0.

  • TP (dcp_size=1): r = 4 · 4080 · 1 / 9216 ≈ 1.8. Measured: clamp 88, KV cap 87 → balanced. ✅
  • DCP8 (dcp_size=8): r = 4 · 4080 · 8 / 9216 ≈ 14. Measured at r=14: clamp 125, KV cap 129 → balanced. ✅ (At the naive r=1.8 the DCP KV pool is ~8× over-provisioned — cap 683 vs clamp 86 — wasting budget that should go to the state pool.)

predict_clamp on the same model at the default extra_buffer, overlap on and the decode-lock skip off (S = 5), three boot logs:

configmax_mamba_cache_sizemmcs // Smeasured max_running_requests
TP8, no DCP2575151
DCP8, r = 5.974519090
DCP8, r = 94749494

Reference r for this example model (token_equiv ≈ 4080, i.e. fp32 ssm + fp8 KV; S=5 default; multiply by (S+D)/S for spec, by dcp_size for DCP) — recompute with your own token_equiv for a different model:

L2K4K8K32K64K128K
r10.05.02.50.620.310.16

Caveats

  • DCP + spec: a replicated (non-DCP-sharded) draft KV does not shard by dcp_size; at long L it dominates per-token cost, so the clean ×dcp_size overstates DCP's advantage — fall back to a per-token direct-solve (KV term /dcp + an un-sharded draft-KV term) when spec is on. (NOSPEC → clean ×dcp_size holds.)
  • Spec + ReplaySSM → use D=0, not the draft-token count. When ReplaySSM spec-verify is enabled, the D intermediate SSM states move off the per-request slot budget onto a fixed ring (a one-time deduction from rest, not a per-req term). So the mamba-slot cost per running request drops back to S (clamp = mmcs / S, not /(S+D)), and the balance ratio returns to the NOSPEC value (r* ≈ S·token_equiv·dcp/L). Measured example (TP8, D=8, L≈9K): applying D=8 computes r*≈2.6 but the true optimum is r≈1.0 — D=8 lands KV-bound at ~45% below the achievable peak concurrency. Plain spec (no replayssm) keeps D = the draft-token count.
  • Under DCP the balance point sits beyond any realistic L (the state pool binds first for almost everything), so the practical recommendation is to pin --max-mamba-cache-size = target_concurrency · S directly rather than dial a large r.
  • --max-mamba-cache-size overrides r; bytes beyond the r budget come out of the KV pool one-for-one.
  • Precision changes are behind accuracy gates: --kv-cache-dtype fp8_e4m3 (doubles token_equiv → doubles r) and --mamba-ssm-dtype bfloat16 (~halves token_equiv; also silently switches the linear-attention decode backend on SM100+ — pin --linear-attn-decode-backend triton) shift outputs; validate accuracy for the workload before production.
  • Asymmetry: the state pool is worst-case-reserved and fail-loud; KV degrades gracefully (retraction). When L is spiky/uncertain, bias r up rather than starve the state pool.

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/compute-mamba-ratio of sgl-project/sglang.

Open the folder on GitHubat commit f620d73

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Compute Mamba Ratio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Compute Mamba Ratio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Compute Mamba Ratio this skillsgl-project/sglang37k2 repos~2.9kAutomated safety check: PassApache-2.0
SQL Optimizationgithub/awesome-copilot40k2 repos~2.3kAutomated safety check: PassMIT
Maxsickn33/agentic-awesome-skills47k1 repos~1.4kAutomated safety check: PassMIT
Agent Performance Optimizerruvnet/ruflo74k2 repos~3.6kAutomated safety check: PassMIT
Database Optimizerdavila7/claude-code-templates32k8 repos~2.5kAutomated safety check: PassMIT
Prompt Optimizeraffaan-m/ECC276k2 repos~2.4kAutomated safety check: PassMIT

Similar skills

  • SQL Optimization

    github/awesome-copilot

    Official

    Universal SQL performance optimization assistant for comprehensive query tuning, indexing strategies, and database performance analysis across all SQL databases (MySQL, PostgreSQL, SQL Server…

    40k GitHub starsUsed in 2 repos~2.3k tokens
    DatabasesAuto-check passed
  • Max

    sickn33/agentic-awesome-skills

    Cleans up and improves existing code without changing behavior.

    47k GitHub starsUsed in 1 repo~1.4k tokens
    DevelopmentAuto-check passed
  • Agent skill for performance-optimizer - invoke with $agent-performance-optimizer

    74k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Database Optimizer

    davila7/claude-code-templates

    Expert database optimizer specializing in modern performance tuning, query optimization, and scalable architectures.

    32k GitHub starsUsed in 8 repos~2.5k tokens
    DatabasesAuto-check passed
  • Prompt Optimizer

    affaan-m/ECC

    分析原始提示,识别意图和差距,匹配ECC组件(技能/命令/代理/钩子),并输出一个可直接粘贴的优化提示。仅提供咨询角色——绝不自行执行任务。触发时机:当用户说“优化提示”、“改进我的提示”、“如何编写提示”、“帮我优化这个指令”或明确要求提高提示质量时。中文等效表达同样触发:“优化prompt”、“改进prompt”、“怎么写prompt”、“帮我优化这个指令”。不触发时机:当用户希望直接执行任…

    276k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Ito Compute

    affaan-m/ECC

    Query live GPU inventory, submit an authenticated Itô fixed-rate RFQ, inspect RFQ or procurement status, revoke device credentials, and run explicitly gated node qualification through the separately…

    276k GitHub starsUsed in 1 repo~1.7k tokens
    Business, Finance & HRAuto-check passed

More from sgl-project/sglang

All 32 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Kl Consistency Test

    sgl-project/sglang

    Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths…

    37k GitHub starsUsed in 2 repos~3.7k tokens
    Auto-check passed

Questions about Compute Mamba Ratio

What does Compute Mamba Ratio do?

Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…. Compute Mamba Ratio is an agent skill from sgl-project/sglang. Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and serving config.

When should I use Compute Mamba Ratio?

Compute Mamba Ratio fits situations like: A user asks what ratio to set; why concurrency is clamped; how to size the state vs KV pools for a hybrid model.

How do I install Compute Mamba Ratio in Claude Code?

Run `npx skills add sgl-project/sglang --skill compute-mamba-ratio -a claude-code`. Or copy the skill folder (.agents/skills/compute-mamba-ratio in sgl-project/sglang) into .claude/skills/compute-mamba-ratio in your project. Claude Code loads it when a task matches its description.

How do I install Compute Mamba Ratio in Codex?

Run `npx skills add sgl-project/sglang --skill compute-mamba-ratio -a codex`. Or copy the skill folder (.agents/skills/compute-mamba-ratio in sgl-project/sglang) into .agents/skills/compute-mamba-ratio in your project. Codex loads it when a task matches its description.

Can I use Compute Mamba Ratio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill compute-mamba-ratio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compute-mamba-ratio, .gemini/skills/compute-mamba-ratio, .github/skills/compute-mamba-ratio and .opencode/skills/compute-mamba-ratio in your project.

What does Compute Mamba Ratio need to run?

Going by SKILL.md and its folder, Compute Mamba Ratio needs the command-line tools its instructions call (mamba). Our summary lists: Python 3.

Does Compute Mamba Ratio access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Compute Mamba Ratio safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Compute Mamba Ratio use?

Compute Mamba Ratio is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Compute Mamba Ratio use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Compute Mamba Ratio?

Skills that share tags, products or a category with Compute Mamba Ratio: SQL Optimization (github/awesome-copilot, 40k stars), Max (sickn33/agentic-awesome-skills, 47k stars), Agent Performance Optimizer (ruvnet/ruflo, 74k stars) and Database Optimizer (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Compute Mamba Ratio?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,907 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.