SQL Optimization
github/awesome-copilot
Universal SQL performance optimization assistant for comprehensive query tuning, indexing strategies, and database performance analysis across all SQL databases (MySQL, PostgreSQL, SQL Server…
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…
$ npx skills add sgl-project/sglang --skill compute-mamba-ratio -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sgl-project/sglang compute-mamba-ratio --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/compute-mamba-ratio .claude/skills/compute-mamba-ratio && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "compute-mamba-ratio" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/compute-mamba-ratio into .claude/skills/compute-mamba-ratio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compute-mamba-ratio", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sgl-project/sglang/tree/main/.agents/skills/compute-mamba-ratioType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sgl-project/sglang --skill compute-mamba-ratio -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sgl-project/sglang compute-mamba-ratio --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/compute-mamba-ratio .agents/skills/compute-mamba-ratio && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "compute-mamba-ratio" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/compute-mamba-ratio into .agents/skills/compute-mamba-ratio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compute-mamba-ratio", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill compute-mamba-ratio -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sgl-project/sglang compute-mamba-ratio --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/compute-mamba-ratio .cursor/skills/compute-mamba-ratio && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "compute-mamba-ratio" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/compute-mamba-ratio into .cursor/skills/compute-mamba-ratio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compute-mamba-ratio", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sgl-project/sglang.git --path .agents/skills/compute-mamba-ratio--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sgl-project/sglang --skill compute-mamba-ratio -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sgl-project/sglang compute-mamba-ratio --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/compute-mamba-ratio .gemini/skills/compute-mamba-ratio && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "compute-mamba-ratio" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/compute-mamba-ratio into .gemini/skills/compute-mamba-ratio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compute-mamba-ratio", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sgl-project/sglang compute-mamba-ratioInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sgl-project/sglang --skill compute-mamba-ratio -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/compute-mamba-ratio .github/skills/compute-mamba-ratio && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "compute-mamba-ratio" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/compute-mamba-ratio into .github/skills/compute-mamba-ratio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compute-mamba-ratio", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill compute-mamba-ratio -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sgl-project/sglang compute-mamba-ratio --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/compute-mamba-ratio .opencode/skills/compute-mamba-ratio && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "compute-mamba-ratio" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/compute-mamba-ratio into .opencode/skills/compute-mamba-ratio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compute-mamba-ratio", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
compute-mamba-ratioCompute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…
Compute Mamba Ratio is an agent skill from sgl-project/sglang. Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and serving config. Use when a user asks what ratio to set, why concurrency is clamped, or how to size the state vs KV pools for a hybrid model.
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f620d73. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
mambaFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Compute Mamba Ratio loads about 2.9k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 1,466 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sgl-project/sglang at commit f620d73, republished under its Apache-2.0 licence (© sgl-project). 1,466 words, ~2,945 tokens.
.claude/skills/compute-mamba-ratio/SKILL.md (or your agent's skills folder).--mamba-full-memory-ratio)A hybrid model (attention layers + linear-attention layers — the recurrent-state family: Mamba/SSM, GDN, KDA, etc.) splits serving memory into two independently-budgeted pools, fixed once at startup:
--mamba-full-memory-ratio r splits the post-weight budget: mamba_budget = rest · r/(1+r), i.e. mamba_budget : kv_budget = r. This skill picks the r (or the pin---max-mamba-cache-size alternative) at which neither pool bottlenecks first for the user's workload.
r* = (S + D) · token_equiv · dcp_size / L
token_equiv = state_bytes_per_slot / kv_bytes_per_tokenL = average context length per request (input + output tokens)token_equiv = full-KV token-equivalent of one state slotS = state slots per running request (cache-strategy dependent, table below)D = --speculative-num-draft-tokens (0 if NOSPEC); each running req carries D extra intermediate statesdcp_size = --dcp-size (1 without DCP). DCP shards the per-rank KV by dcp_size, so KV gets ~dcp_size× cheaper per request → the balance shifts that much toward the state pool.r is dimensionless (just the split). To also predict the actual concurrency you need rest (below).
L — average context (input + output) in tokens.Mamba Cache is allocated. ... ssm_state size X GB with max_mamba_cache_size: N → state_bytes_per_slot = X / NKV Cache is allocated. #tokens: M, KV size: Y GB → kv_bytes_per_token = Y / Md_state/d_conv/heads/head_dim; attention type + KV dims: MLA latent dim, or GQA kv_heads·head_dim·layers) × the dtypes below.S — from --mamba-radix-cache-strategy, the overlap scheduler, and SGLANG_OPT_MAMBA_SKIP_DECODE_LOCK (table below).D — --speculative-num-draft-tokens (0 if NOSPEC; also 0 when ReplaySSM spec-verify is enabled — see caveats).dcp_size — --dcp-size (1 if no DCP). If DCP and spec with a replicated draft KV, apply the draft caveat (below).token_equiv 2×2).r) rest = per-GPU memory − weights at the chosen --mem-fraction-static. Read avail mem after Load weight end, or Memory pool end. avail mem + pool sizes, from the boot log.S (state slots per running request), set by --mamba-radix-cache-strategyAll strategies keep prefix caching on; they differ in the track buffer that snapshots chunk-boundary state under the overlap scheduler. S = base + ping-pong, where base is 3 (live state + radix retention/COW headroom), and ping-pong = 2 (overlap on, non-lazy) / 1 (lazy, or overlap off) / 0 (no track buffer at all). This mirrors kv_cache_configurator._calculate_mamba_ratio; read it there if a release moves the constants.
SGLANG_OPT_MAMBA_SKIP_DECODE_LOCK=1 takes the base to 2, since the decode-time skip frees one resident slot per running request. no_buffer is the exception and stays at an effective 3: it adds that slot straight back, because its binding limit is the prefill-to-decode peak, which a decode-time change does not shrink.
| strategy | S | S with decode-lock skip | notes |
|---|---|---|---|
--disable-radix-cache | 1 | 1 | prefix cache off entirely; most concurrency, no reuse |
no_buffer | 3 | 3 | no track buffer; requires overlap off (+ page_size 1) → lower decode throughput |
extra_buffer (default), overlap on | 5 | 4 | track buffer reserved per running request |
extra_buffer, overlap off | 4 | 3 | same buffer, but ping-pong costs 1 slot instead of 2 |
extra_buffer_lazy (recommended) | 4 | 3 | track buffer allocated lazily at the boundary → 1 fewer slot/req; requires overlap on |
The skip is a real concurrency lever, not a rounding detail: at the default strategy it takes S from 5 to 4, so the same state budget admits 25% more requests and r* drops by the same factor.
Overlap counts as off when --disable-overlap-schedule is passed or --pp-size > 1 (pipeline parallelism turns it off for you). So a PP deployment left on the default strategy sits on the S = 4 row, not S = 5 — using 5 there overstates the state cost and starves KV.
token_equiv — compute it from the two byte constantstoken_equiv = state_bytes_per_slot / kv_bytes_per_token. Both bytes are per-model, per-GPU — read them from a boot log (or derive from the arch + dtypes) and divide. That's the whole thing; the dtype effects below are just for a quick mental estimate.
The recurrent state has two parts: state_bytes = SSM_bytes + conv_bytes. The SSM part scales with --mamba-ssm-dtype; the conv part is always bf16 (SGLANG_MAMBA_CONV_DTYPE, fixed). How the dtype knobs move token_equiv:
kv_bytes_per_token halves.(2·SSM + conv)/(SSM + conv) — just under 2. conv is usually small relative to SSM, so estimate ≈2 (e.g. one measured model: conv ≈7% of state → factor ≈1.9). Only bother with the exact split if a model's conv is unusually large.For an actual number, don't apply factors — read state_bytes_per_slot and kv_bytes_per_token for your model+dtype from the boot log and divide.
The GB the boot log prints for both pools is GiB (1024³). Convert both the same way and token_equiv is unaffected either way, since it is a ratio — but get it right before comparing an absolute byte figure against your own log.
Worked example (one measured model, TP8, fp32 ssm + fp8 KV), read off one boot log:
Mamba Cache is allocated. max_mamba_cache_size: 257, conv_state size: 0.46GB, ssm_state size: 13.04GB
KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 1167552, KV size: 15.03 GBstate_bytes_per_slot = (0.46 + 13.04) GiB / 257 ≈ 56.4 MB, kv_bytes_per_token = 15.03 GiB / 1167552 ≈ 13.8 KB → token_equiv ≈ **4080**.
def optimal_ratio(L, state_bytes_per_slot, kv_bytes_per_token, S, D=0, dcp_size=1):
token_equiv = state_bytes_per_slot / kv_bytes_per_token
r = (S + D) * token_equiv * dcp_size / L
return r # value of --mamba-full-memory-ratio (>1 is legal)
def predict_clamp(rest_bytes, r, state_bytes_per_slot, S, D=0):
mamba_budget = rest_bytes * r / (1 + r)
slots = mamba_budget / state_bytes_per_slot
# spec: each running req reserves (S+D) worth; non-spec just S
return int(slots // S) # NOSPEC; with spec the budget joint-solves for (S+D)·per_req per reqThen state the result three ways: the r value, the predicted clamp (if rest given), and which pool binds (min(mamba_clamp, KV_cap), where KV_cap = kv_tokens · dcp_size / L).
Memory pool end. avail mem (e.g. 30–40 GB at mem-frac 0.85), raise --mem-fraction-static (→0.92) before touching the split — it grows rest for both pools at no cost. Validate graph-capture headroom once.r*. If r* would drive the state pool below one request's worth (mamba_budget < S · per_req, happens at very long L), switch to pinning --max-mamba-cache-size = target_concurrency · S and let the rest go to KV — a sub-0.15 r is fragile.r, predicted clamp, binding pool, and any dtype accuracy gate that applies.Constants (from boot logs): state_bytes_per_slot ≈ 56.4 MB (fp32 ssm), kv_bytes_per_token ≈ 13.8 KB (fp8) → token_equiv ≈ 4080. Cache strategy extra_buffer_lazy → S = 4. Workload L = 9216 (8192 in + 1024 out). NOSPEC → D = 0.
r = 4 · 4080 · 1 / 9216 ≈ 1.8. Measured: clamp 88, KV cap 87 → balanced. ✅r = 4 · 4080 · 8 / 9216 ≈ 14. Measured at r=14: clamp 125, KV cap 129 → balanced. ✅ (At the naive r=1.8 the DCP KV pool is ~8× over-provisioned — cap 683 vs clamp 86 — wasting budget that should go to the state pool.)predict_clamp on the same model at the default extra_buffer, overlap on and the decode-lock skip off (S = 5), three boot logs:
| config | max_mamba_cache_size | mmcs // S | measured max_running_requests |
|---|---|---|---|
| TP8, no DCP | 257 | 51 | 51 |
DCP8, r = 5.97 | 451 | 90 | 90 |
DCP8, r = 9 | 474 | 94 | 94 |
Reference r for this example model (token_equiv ≈ 4080, i.e. fp32 ssm + fp8 KV; S=5 default; multiply by (S+D)/S for spec, by dcp_size for DCP) — recompute with your own token_equiv for a different model:
| L | 2K | 4K | 8K | 32K | 64K | 128K |
|---|---|---|---|---|---|---|
| r | 10.0 | 5.0 | 2.5 | 0.62 | 0.31 | 0.16 |
dcp_size; at long L it dominates per-token cost, so the clean ×dcp_size overstates DCP's advantage — fall back to a per-token direct-solve (KV term /dcp + an un-sharded draft-KV term) when spec is on. (NOSPEC → clean ×dcp_size holds.)D=0, not the draft-token count. When ReplaySSM spec-verify is enabled, the D intermediate SSM states move off the per-request slot budget onto a fixed ring (a one-time deduction from rest, not a per-req term). So the mamba-slot cost per running request drops back to S (clamp = mmcs / S, not /(S+D)), and the balance ratio returns to the NOSPEC value (r* ≈ S·token_equiv·dcp/L). Measured example (TP8, D=8, L≈9K): applying D=8 computes r*≈2.6 but the true optimum is r≈1.0 — D=8 lands KV-bound at ~45% below the achievable peak concurrency. Plain spec (no replayssm) keeps D = the draft-token count.L (the state pool binds first for almost everything), so the practical recommendation is to pin --max-mamba-cache-size = target_concurrency · S directly rather than dial a large r.--max-mamba-cache-size overrides r; bytes beyond the r budget come out of the KV pool one-for-one.--kv-cache-dtype fp8_e4m3 (doubles token_equiv → doubles r) and --mamba-ssm-dtype bfloat16 (~halves token_equiv; also silently switches the linear-attention decode backend on SM100+ — pin --linear-attn-decode-backend triton) shift outputs; validate accuracy for the workload before production.L is spiky/uncertain, bias r up rather than starve the state pool.© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/compute-mamba-ratio of sgl-project/sglang.
Open the folder on GitHubat commit f620d73
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.
Compute Mamba Ratio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Compute Mamba Ratio this skillsgl-project/sglang | 37k | 2 repos | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| SQL Optimizationgithub/awesome-copilot | 40k | 2 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Maxsickn33/agentic-awesome-skills | 47k | 1 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Agent Performance Optimizerruvnet/ruflo | 74k | 2 repos | ~3.6k | Automated safety check: Pass | MIT | |
| Database Optimizerdavila7/claude-code-templates | 32k | 8 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Prompt Optimizeraffaan-m/ECC | 276k | 2 repos | ~2.4k | Automated safety check: Pass | MIT |
github/awesome-copilot
Universal SQL performance optimization assistant for comprehensive query tuning, indexing strategies, and database performance analysis across all SQL databases (MySQL, PostgreSQL, SQL Server…
sickn33/agentic-awesome-skills
Cleans up and improves existing code without changing behavior.
ruvnet/ruflo
Agent skill for performance-optimizer - invoke with $agent-performance-optimizer
davila7/claude-code-templates
Expert database optimizer specializing in modern performance tuning, query optimization, and scalable architectures.
affaan-m/ECC
分析原始提示,识别意图和差距,匹配ECC组件(技能/命令/代理/钩子),并输出一个可直接粘贴的优化提示。仅提供咨询角色——绝不自行执行任务。触发时机:当用户说“优化提示”、“改进我的提示”、“如何编写提示”、“帮我优化这个指令”或明确要求提高提示质量时。中文等效表达同样触发:“优化prompt”、“改进prompt”、“怎么写prompt”、“帮我优化这个指令”。不触发时机:当用户希望直接执行任…
affaan-m/ECC
Query live GPU inventory, submit an authenticated Itô fixed-rate RFQ, inspect RFQ or procurement status, revoke device credentials, and run explicitly gated node qualification through the separately…
sgl-project/sglang
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
sgl-project/sglang
Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
sgl-project/sglang
Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.
sgl-project/sglang
Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths…
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…. Compute Mamba Ratio is an agent skill from sgl-project/sglang. Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and serving config.
Compute Mamba Ratio fits situations like: A user asks what ratio to set; why concurrency is clamped; how to size the state vs KV pools for a hybrid model.
Run `npx skills add sgl-project/sglang --skill compute-mamba-ratio -a claude-code`. Or copy the skill folder (.agents/skills/compute-mamba-ratio in sgl-project/sglang) into .claude/skills/compute-mamba-ratio in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sgl-project/sglang --skill compute-mamba-ratio -a codex`. Or copy the skill folder (.agents/skills/compute-mamba-ratio in sgl-project/sglang) into .agents/skills/compute-mamba-ratio in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill compute-mamba-ratio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compute-mamba-ratio, .gemini/skills/compute-mamba-ratio, .github/skills/compute-mamba-ratio and .opencode/skills/compute-mamba-ratio in your project.
Going by SKILL.md and its folder, Compute Mamba Ratio needs the command-line tools its instructions call (mamba). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Compute Mamba Ratio is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Compute Mamba Ratio: SQL Optimization (github/awesome-copilot, 40k stars), Max (sickn33/agentic-awesome-skills, 47k stars), Agent Performance Optimizer (ruvnet/ruflo, 74k stars) and Database Optimizer (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,907 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.
Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.