Agent skill

Kl Consistency Test

by sgl-project in sgl-project/sglang

Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Kl Consistency Test

skills CLI
$ npx skills add sgl-project/sglang --skill kl-consistency-test -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang kl-consistency-test --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/kl-consistency-test .claude/skills/kl-consistency-test && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kl-consistency-test
GitHub stars
37k
Used in
2 other repos
Token cost
~3.7k tokens
SKILL.md length
2,126 words
Files
1
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths…

  • Works in 2 steps: Every operator on the path is… → The two paths compute the same function.…
  • Adding a KL test to a model
  • SKILL.md covers What the test is for, Two independent conditions…, The three helpers differ in… and Run it the way CI runs it, plus 6 more sections
  • Calls python3 and curl

What it does

Kl Consistency Test is an agent skill from sgl-project/sglang. Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths computing the same function), which helper separates them, how to pick a threshold once they hold, and how to localize a divergence to a single operator. Use when adding a KL test to a model, picking or defending a kldiv threshold, or investigating a KL number that is too high.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with SGLang. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • Adding a KL test to a model
  • Defending a kldiv threshold
  • Investigating a KL number that is too high

Example prompts

  • “/kl-consistency-test”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Every operator on the path is batch-invariant. A token's result must not
  2. The two paths compute the same function. Decode's context and state at a

What it can do on your machine

Read from SKILL.md and the folder at commit f620d73. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • thinkingmachines.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Kl Consistency Test loads about 3.7k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 2,126 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit f620d73, republished under its Apache-2.0 licence (© sgl-project). 2,126 words, ~3,698 tokens.

Download SKILL.mdSave it as .claude/skills/kl-consistency-test/SKILL.md (or your agent's skills folder).
name
kl-consistency-test
description
Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths computing the same function), which helper separates them, how to pick a threshold once they hold, and how to localize a divergence to a single operator. Use when adding a KL test to a model, picking or defending a kl_div threshold, or investigating a KL number that is too high.

KL Consistency Tests

What the test is for

kl_test_utils scores the same token twice -- once as a prefill input logprob, once as a decode output logprob -- and compares. The two paths run different kernels over different shapes, so agreement is a statement about state, not about answer quality: it catches a radix-cache prefix that does not reproduce a fresh prefill, a stale conv/mamba checkpoint, a SWA pool that evicted something it still needed.

gsm8k passing says nothing about this. Accuracy is insensitive to a handful of corrupted tokens; the KL check is not.

Two independent conditions produce a zero

Reaching bit-identity needs both, and they fail for unrelated reasons. Knowing which one a nonzero belongs to is most of the debugging.

  1. Every operator on the path is batch-invariant. A token's result must not depend on how many tokens share its forward. Note this is a property across the two paths, not a property of each: a kernel can be perfectly reproducible at M=1 and again at M=N while disagreeing between them, which is exactly what a tile-size switch or a message-size-dependent reduction does.

  2. The two paths compute the same function. Decode's context and state at a position must equal what a fresh prefill computes there -- the same KV set, the same sliding window, the same conv/mamba state, a restored cache prefix that reproduces a recomputed one. This is logic, not arithmetic, and it survives any amount of numerical hygiene.

The conditions are independent, and one measurement separates them: with (1) satisfied, match and decode_cache_hit read exactly 0 while prefill_cache_hit stays nonzero when a prefix restore is wrong. Same server, same prompts -- float noise cannot pick a code path, so a helper-specific divergence is (2).

Order the work accordingly. Settle (1) first: until it holds, its noise is orders of magnitude above anything (2) produces and hides it completely.

The three helpers differ in what touches the cache

KLDivergenceMixin runs the last two. Pick deliberately -- they are not interchangeable, and only the cache-hit pair exercises prefix reuse.

HelperCache involvement
..._match_helperboth sides flush; no cache at all
..._match_prefill_cache_hit_helperprompt is prefilled once to warm the cache, then the generation prefill restores from it
..._match_decode_cache_hit_helperdecode side runs on a warmed cache

A divergence confined to one helper is diagnostic. match clean but prefill_cache_hit dirty means the restore path is wrong, not the arithmetic -- float noise does not pick a code path.

Run it the way CI runs it

KLDivergenceMixin defaults: max_samples=32, max_new_tokens=512. Do not characterize with fewer.

avg_kl_div is the k3 estimator, exp(logr) - 1 - logr, applied to the sampled token's logprob. It is exponentially sensitive to the tail, so the mean is carried by a handful of tokens. At 4 samples the same config measured 0.049 to 0.158 -- a 3x spread that invalidates any A/B comparison drawn from it.

When characterizing rather than gating, report tail statistics -- the fraction of tokens past a threshold, and the max -- rather than the mean.

Generate past the sliding window if the model has one, so decode carries the window through the handover from prompt tokens to generated ones.

Condition 1: determinism is not batch-invariance

This distinction decides whether a threshold means anything.

  • Deterministic: same input, same shape, same result on every run.
  • Batch-invariant: a token's result does not depend on how many other tokens share its batch.

The KL check compares a prefill of thousands of tokens against decode steps of one, so it measures the second. --enable-deterministic-inference buys both -- it swaps the aten kernels for fixed-reduction versions and pins the NCCL algorithm and channel count -- but only for kernels it covers. Custom kernels that never reach an aten op are outside batch_invariant_ops and stay shape-dependent.

The consequence: a nonzero KL under deterministic inference that appears in every helper alike means some kernel on the path is still batch-dependent. Localize it (below) rather than widening the threshold.

Background, and the source of the fixed-reduction approach the aten overrides take: Defeating nondeterminism in LLM inference.

How much batch-invariance an operator needs

For a token-wise operator -- GEMM, norm, activation, the router's linear -- a token's output depends only on that token's row, so pinning the reduction order is the whole requirement. Once its result is independent of how many rows share the launch, it is done.

Two kinds need more than that, and they are where the remaining nonzero usually lives:

  • Operators that reduce across tokens -- attention over a KV range, and any collective. Fixing the arithmetic order is not enough if the extent still varies: an all-reduce whose tree shape follows the message size, or an attention split whose block boundary follows the query count, gives a token a different reduction depending on its batch. Pin the shape, not just the order.
  • Operators that carry state across calls -- conv windows, SSM checkpoints. These are batch-invariant per call and still diverge, because what they store is reused by a later request. That is condition 2, and no amount of reduction-order work reaches it.

So "make everything batch-invariant" closes condition 1 for the token-wise majority, and the residual after that is concentrated in these two classes.

MoE amplifies this to a degree dense models do not

Top-k routing is a discrete decision over near-tied scores. A 1e-8 difference in gate weights flips which experts a token is routed to, the outputs diverge completely, and 42 layers compound it. Measured on one MoE checkpoint: a gate GEMM that switched tiling between M=8 and M=16 produced a 1.6e-5 logits difference, which became 20-37 nat on individual high-confidence tokens and a KL of 0.177.

A dense model of comparable size shows the same root cause as ~1e-4. So a KL in the hundredths is not evidence of a worse bug on a MoE model -- it is the same class of numerical difference, amplified. Do not calibrate a MoE threshold by analogy to a dense one.

Condition 2: the two paths must compute the same function

Once condition 1 holds, whatever remains is a state bug, and the helper it appears in names the path. A restore that does not reproduce a recomputed prefix shows up in prefill_cache_hit alone; the other two stay at exactly 0.

What the signature looks like, and how to read it:

  • Which sequences. Divergence concentrated in a couple of requests out of a batch, with the rest bit-identical, is a condition triggered by those requests -- not a systematic offset. Compare their prompt lengths, cached_tokens, and page and checkpoint-interval remainders against the ones that pass.
  • Where in the generation. Contiguous from the first generated token means the state was already wrong when generation began, so the fault is in the prefix restore rather than in decode. Divergence starting mid-generation points instead at something that happens during decode -- a window handover, a checkpoint rotation.
  • Whether it is a race. Re-run under different configurations that should not matter (page size, TP degree, buffer strategy). Bit-identical numbers across them mean a deterministic logic fault, which is far cheaper to chase than a race.

Generate past the sliding window if the model has one: the handover from prompt tokens to generated ones inside the window is where eviction and checkpoint rotation actually run.

Choosing a threshold

Once every kernel on the path is batch-invariant, prefill and decode agree bit for bit and the honest assertion is a stray-ulp floor, not a tolerance:

python
KL_DIV_THRESHOLD = 1e-9   # measured 0; anything a state bug produces is orders above

A loose threshold tolerates float noise and small logic errors alike, which is how a state-reuse bug hides. Prefer running the KL case on its own deterministic server and asserting near-zero, and keep the accuracy case on the production numerics -- one server cannot serve both.

Thresholds are per (model, tp). A value calibrated at tp=1 does not transfer: tp=1 has no all-reduce, so it never exercises the source that dominates at tp>1.

Show full SKILL.md (831 more words)Show less

Localizing a divergence

Ablations answer "does it change" but never "where". The forward-hook dumper points at the operator directly, and has done so reliably: run it once and read off the first layer whose output differs while its inputs are bit-identical.

bash
DUMPER_ENABLE=0 DUMPER_SERVER_PORT=reuse DUMPER_NON_INTRUSIVE_MODE=all \
DUMPER_DIR=/path/to/dumps python3 -m sglang.launch_server ... \
  --disable-cuda-graph --disable-prefill-cuda-graph
curl -X POST localhost:PORT/dumper/configure -d '{"enable": true, "exp_name": "dec"}'

Five settings that are each required, and each fails silently if wrong:

  • DUMPER_ENABLE=0 plus DUMPER_SERVER_PORT=reuse. The port sentinel makes may_enable true so the hooks register, while enable=0 keeps warmup from dumping. Enabling at boot dumps every warmup prefill -- that is how a run wrote 1.8T and filled a shared disk. Add a watchdog that kills the run below a free-space floor.
  • DUMPER_NON_INTRUSIVE_MODE=all. The default core writes only positions, seq_lens, req_pool_indices, input_ids, rids -- no module tensors, and no error to tell you.
  • DUMPER_SERVER_PORT=reuse is a literal sentinel, not a port number; the /dumper/{method} route only registers for that exact value.
  • --disable-prefill-cuda-graph on top of --disable-cuda-graph. Some models default prefill onto a CUDA graph, and Python forward hooks do not run inside a replay -- the prefill pass then dumps the embedding and nothing else.
  • Prefer dumper.py over --debug-tensor-dump-*: the latter asserts on a top-level module named model, which multimodal wrappers do not have.

Prove the alignment before reading any diff. Decode pass k and prefill row plen + k consume the same token, so the embedding output must be bit-identical. If it is not, the rows are misaligned and every downstream number is meaningless. Getting this wrong once produced a confident, entirely wrong root cause.

Read the result as: the first layer where a module's inputs are bit-identical and its output is not is the operator. Everything after it inherits.

When the divergence needs a CUDA graph

A divergence that only appears with a captured graph defeats both usual probes, and the failure is silent in each case:

  • The dumper's hooks do not run during replay — the graph replays kernels, not Python. Disabling the graph to collect a dump also removes the divergence, so a clean layer-by-layer diff means nothing. Confirm the bug still reproduces under the exact flags you dump with.
  • Anything that syncs to host dies during capture (.item(), float(), .tolist()). Guard probes with torch.cuda.is_current_stream_capturing() or the server will not boot.
  • The Python wrapper around a captured kernel is not called at replay. Instrumenting it logs only the phases that stayed eager. Read that as evidence, not as a broken probe: it means the kernel runs with the arguments bound at capture, so any tensor handed in fresh per replay is invisible to it — a bug shape in its own right.

What works instead is to probe the state that gets reused, outside the graph: at the moment a request donates its checkpoint, log the slot id, the length it claims to have checkpointed at, and abs().max() over the stored state. Run it twice with the graph on and off and diff per slot. A handful of slots whose content differs, with claimed lengths matching the prefixes of the requests that go wrong, localizes the write in one round — where a dozen ablations only bound the trigger.

Make the probe prove it fired. A probe on a code path that is not taken prints nothing, which is indistinguishable from "measured, no difference". Assert a minimum hit count, or log unconditionally at entry. Instrument the single choke point every caller reaches rather than one call site.

Confirm the mechanism, do not infer it

Two failure modes cost the most time, both avoidable:

  • A flag that changes nothing. Bit-identical results before and after a toggle mean the flag did not take effect -- a dispatch guarded on a hidden condition, a path never taken for that config. Check the guard before concluding the component is innocent.
  • A harness that measures something else. Capture through the helper's own functions rather than reconstructing its inputs. Reconstructing them once appended a generation twice and produced a plausible, wrong conclusion; another time a different num_samples silently selected a different prompt set through the get_input_ids cache key.

The logprob arrays are indexed by absolute position: with logprob_start_len=0, input_token_logprobs carries one entry per input token, the first is None, and entry k scores input_ids[k]. The helpers slice the tail, which lands on the generated span; analysis that indexes absolutely has to agree with that. An off-by-one here reads a neighbouring token, whose logprob is usually close enough to look like a real signal.

For an isolated claim, reduce to a standalone repro. A ten-line script calling the suspect op at M=1 and M=288 settles batch-invariance in seconds, and belongs in the PR ahead of any end-to-end number.

Reading code to find a suspect is the slowest of these. One investigation refuted eight successive code-derived hypotheses, each internally consistent, before a direct measurement of the reused state found the defect in a single round. Prefer, in order: a single-variable A/B that isolates the trigger, asking what the wrong output is the correct answer to, probing the reused state itself, and only then reading for a mechanism to explain what was measured.

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/kl-consistency-test of sgl-project/sglang.

Open the folder on GitHubat commit f620d73

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in sgl-project/sglang, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Kl Consistency Test next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kl Consistency Test compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kl Consistency Test this skillsgl-project/sglang37k2 repos~3.7kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Clean Startup Logguqiong96/Lsglang1431 repos~4.5kAutomated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
slime RL Post-TrainingOrchestra-Research/AI-Research-SKILLs13k4 repos~2.8kAutomated safety check: PassMIT
SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs13k2 repos~2.9kAutomated safety check: PassMIT

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Clean Startup Log

    guqiong96/Lsglang

    Clean up noisy startup warnings and spurious prints in SGLang server logs.

    143 GitHub starsUsed in 1 repo~4.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • slime RL Post-Training

    Orchestra-Research/AI-Research-SKILLs

    Guides reinforcement-learning post-training of LLMs with slime, which pairs Megatron-LM training with SGLang rollouts, including GRPO runs on GLM, Qwen3 and Llama 3 models.

    13k GitHub starsUsed in 4 repos~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 2 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • verl RL Training

    Orchestra-Research/AI-Research-SKILLs

    Trains LLMs with reinforcement learning using verl, from ByteDance's Seed team, with GRPO, PPO and other algorithms and swappable training and rollout backends.

    13k GitHub starsUsed in 2 repos~2.4k tokens
    AI & LLM EngineeringAuto-check passed

More from sgl-project/sglang

All 32 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

Works with

Questions about Kl Consistency Test

What does Kl Consistency Test do?

Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths…. Kl Consistency Test is an agent skill from sgl-project/sglang. Write, calibrate, and debug the prefill-vs-decode logprob (KL) consistency tests in sglang -- the two independent conditions a zero requires (every operator batch-invariant, and the two paths computing the same function), which helper separates them, how to pick a threshold once they hold, and how to localize a divergence to a single operator.

When should I use Kl Consistency Test?

Kl Consistency Test fits situations like: adding a KL test to a model; defending a kldiv threshold; investigating a KL number that is too high.

How do I install Kl Consistency Test in Claude Code?

Run `npx skills add sgl-project/sglang --skill kl-consistency-test -a claude-code`. Or copy the skill folder (.agents/skills/kl-consistency-test in sgl-project/sglang) into .claude/skills/kl-consistency-test in your project. Claude Code loads it when a task matches its description.

How do I install Kl Consistency Test in Codex?

Run `npx skills add sgl-project/sglang --skill kl-consistency-test -a codex`. Or copy the skill folder (.agents/skills/kl-consistency-test in sgl-project/sglang) into .agents/skills/kl-consistency-test in your project. Codex loads it when a task matches its description.

Can I use Kl Consistency Test in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill kl-consistency-test -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kl-consistency-test, .gemini/skills/kl-consistency-test, .github/skills/kl-consistency-test and .opencode/skills/kl-consistency-test in your project.

What does Kl Consistency Test need to run?

Going by SKILL.md and its folder, Kl Consistency Test needs the command-line tools its instructions call (python3 and curl). Our summary lists: Python 3.

Does Kl Consistency Test access the network?

SKILL.md names 1 domain. As links in the text: thinkingmachines.ai. This is read from the text; nothing was executed.

Is Kl Consistency Test safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Kl Consistency Test use?

Kl Consistency Test is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kl Consistency Test use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kl Consistency Test?

Skills that share tags, products or a category with Kl Consistency Test: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Clean Startup Log (guqiong96/Lsglang, 143 stars), Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars) and slime RL Post-Training (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kl Consistency Test?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,907 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.