Hugging Face Tokenizers
Orchestra-Research/AI-Research-SKILLs
Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.
Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids).
$ npx skills add sgl-project/sglang --skill processor-model-parity -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sgl-project/sglang processor-model-parity --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/processor-model-parity .claude/skills/processor-model-parity && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "processor-model-parity" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/processor-model-parity into .claude/skills/processor-model-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processor-model-parity", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sgl-project/sglang/tree/main/.agents/skills/processor-model-parityType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sgl-project/sglang --skill processor-model-parity -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sgl-project/sglang processor-model-parity --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/processor-model-parity .agents/skills/processor-model-parity && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "processor-model-parity" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/processor-model-parity into .agents/skills/processor-model-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processor-model-parity", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill processor-model-parity -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sgl-project/sglang processor-model-parity --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/processor-model-parity .cursor/skills/processor-model-parity && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "processor-model-parity" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/processor-model-parity into .cursor/skills/processor-model-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processor-model-parity", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sgl-project/sglang.git --path .agents/skills/processor-model-parity--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sgl-project/sglang --skill processor-model-parity -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sgl-project/sglang processor-model-parity --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/processor-model-parity .gemini/skills/processor-model-parity && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "processor-model-parity" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/processor-model-parity into .gemini/skills/processor-model-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processor-model-parity", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sgl-project/sglang processor-model-parityInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sgl-project/sglang --skill processor-model-parity -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/processor-model-parity .github/skills/processor-model-parity && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "processor-model-parity" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/processor-model-parity into .github/skills/processor-model-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processor-model-parity", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sgl-project/sglang --skill processor-model-parity -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sgl-project/sglang processor-model-parity --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/processor-model-parity .opencode/skills/processor-model-parity && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "processor-model-parity" agent skill from https://github.com/sgl-project/sglang/tree/main/.agents/skills/processor-model-parity into .opencode/skills/processor-model-parity/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "processor-model-parity", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
processor-model-parityAdd or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids).
Processor Model Parity is an agent skill from sgl-project/sglang. Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids). Use when adding a model to sglang-processor, porting a model's rendering from sgl-router, bumping Dynamo's renderer, or debugging a processor-vs-Python prompt mismatch.
Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Development, covering Natural language processing. It works with SGLang, Python, Rust and DeepSeek. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f620d73. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
cargopythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Processor Model Parity loads about 3.1k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 1,362 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sgl-project/sglang at commit f620d73, republished under its Apache-2.0 licence (© sgl-project). 1,362 words, ~3,089 tokens.
.claude/skills/processor-model-parity/SKILL.md (or your agent's skills folder).rust/sglang-processor renders chat requests for SGLang's Rust hosts (the Rust
server, sgl-router, the renderer). A model is supported only when, for the same
request body, the processor produces the same prompt text and the same prompt
token ids as Python's OpenAIServingChat. Python is the reference, Dynamo is
the implementation, and SGLang code exists only where the two differ.
Scope: the request -> prompt -> token ids path. Output parsing (reasoning and tool calls) and multimodal preprocessing are not covered here.
The worked example is DeepSeek-V4-Flash-0731: src/render/models/deepseek_v4.rs
plus tests/fixtures/parity/deepseek-v4-flash-0731.json. Paths below are relative
to rust/sglang-processor/ unless they start with python/.
| Piece | File | Needed when |
|---|---|---|
| Identity -> formatter | src/render/selection.rs (select_chat_formatter) | always; mirror the order of chat_encoding.resolve_chat_encoding_spec |
| Formatter variant | src/render/mod.rs (ChatFormatter, its render_request arm, plus the render_prompt, stop_strs and resolve_thinking arms) | always |
| SGLang-specific steps | src/render/models/<model>.rs, re-exported in models/mod.rs | only where Dynamo diverges from Python |
| Parity cases | tests/fixtures/parity/<model-id>.json | always |
| README row | the render table in README.md | always |
No new test code: tests/parity.rs checks every fixture in the directory.
render_request(Value, ..) takes the SGLang request body, not Dynamo's
OAIChatLikeRequest, because Python reads fields the trait lacks (task,
continue_final_message, the reasoning object). Pick the shape by how Python
renders the model:
None spec such as dsv4 or dsv32):
give the model a dedicated ChatFormatter variant and render_request arm, like
DeepSeek-V4. If selection.rs already wires the model to a Dynamo
PromptFormatter, replace that branch.apply_chat_template (Jinja, spec None): the first such model
adds one generic arm that adapts the body into OAIChatLikeRequest, in its own
PR. sgl-router's ChatRequest in
experimental/sgl-router/src/tokenizer/chat_formatter.rs is a working adapter.Pick the checkpoint. Use the one the model's cookbook page recommends
(docs/cookbook/), pin its Hugging Face commit as revision, and cache it.
Note whether its tokenizer ships a chat_template; that can change the spec.
Trace the Python path for the model. Write down every transformation
between the HTTP body and prompt_ids, in order:
protocol.py::ChatCompletionRequest.normalize_reasoning_inputs: reasoning
and reasoning_effort set the default thinking / enable_thinking.
Reuse render/reasoning.rs for this rather than porting it again.serving_chat._convert_to_internal_request: chat_template_kwargs.reasoning_effort
replaces the request effort; server --default-chat-template-kwargs merge in.chat_encoding.resolve_chat_encoding_spec: which branch renders the model
(a spec such as dsv4, dsv32, kimi_k3, inkling, or None for Jinja).
It also reads the server's --tool-call-parser and whether the tokenizer has
a chat template. If the deployment depends on the parser, set
tool_call_parser in the fixture, and add it to ChatFormatterOptions in the
same PR. Note any per-checkpoint state the server resolves once (DeepSeek-V4's
effort profile reads encoding/encoding_dsv4.py).serving_chat._apply_jinja_template / _encode_messages:
message dumping, content flattening, _handle_last_assistant_message, system
insertion, tools, model-specific fields, the encoder or apply_chat_template
call, and how the prompt is tokenized.encoding_dsv4.py), line by line. History
rules live there: dropped turns, </think> on empty reasoning, argument
formatting, and the errors it raises.This list is the divergence inventory.
Find Dynamo's counterpart. Check dynamo_renderer::native_formatter_for,
the model's low-level encoder (for example deepseek::v4::encode_messages_with_options),
or the Jinja template. Prefer the lowest-level Dynamo entry that lets SGLang own
request semantics. Dynamo's OpenAI-level formatters apply their own defaults
(they filter tools by tool_choice, inject response_format, and map effort),
which often differ from SGLang's. Diff the low-level encoder against Python's
for the rules in step 1.
Write one request case per inventory item in the fixture (name +
request only), then generate the expected outputs from Python (below). Cover
the checklist; reuse the request shapes from the DeepSeek-V4 fixture where they apply.
Implement. Start from a plain Dynamo call and run tests/parity.rs. For
each failing case, add the smallest SGLang step that fixes it, with a one-line
doc comment naming the Python source it mirrors. Never patch the fixture to
match Rust.
Verify text offline, then token ids with the checkpoint cached (commands below). If sgl-router or the renderer uses the formatter, run their suites too.
user
reduced to role and content; null content becomes "". (ignored_fields, tool_results)parts)tool_choice allows; field order
and defaults of Function.model_dump; message-level tools; empty lists.
(tools, tools_none, tools_named, message_tools, empty_message_tools)strict: 1 and defer_loading: "false" become booleans, and
strict: null is rejected. Probe the pydantic model in Python to get its exact
rules; don't guess them. (tool_bool_coercion, tool_strict_null, continuation_coerced)serde_json preserve_order is declared in Cargo.toml). Floats must print as
Python's json.dumps prints them (1e-06, 1e+16, 100.0).
(tool_results, agentic_thinking, tool_float_values, history_float_arguments)thinking > request effort (!= "none") > reasoning.enabledserver default kwargs >
SGLANG_DEFAULT_THINKING.enabledfollows Python truthiness (1is on); strings use Python's yes-word set. (thinking_false,thinking_none,drop_thinking_ignored,reasoning_enabled_numeric,reasoning_enable_string)
reasoning, reasoning_effort and env;
the checkpoint's profile mapping. (effort_high, effort_max, kwargs_effort, effort_conflict, reasoning_object)continue_final_message
a separately tokenized prefix. Check whether content is flattened before the
split (it is for DeepSeek-V4). Run each check where Python runs it: a final
assistant turn's tool calls are discarded, so object checks on their arguments
belong after the split. Hosts reach the model through render_prompt too, so
check the renderer's continuation handling (sglang-renderer render()) as well.
(final_assistant, continuation_parts, continuation_bos, continuation_only, final_tool_call_no_arguments, continuation_tool_call_null_arguments)task placement, a system turn mid-conversation,
inserted empty system turns. (task_action, task_after_developer, consecutive_task, mid_system)multi_turn, thinking_multi_turn)tokenizer.encode adds special tokens by default; the
continuation prefix is encoded alone and loses a leading BOS
(_append_assistant_prefix_to_prompt_ids). tests/parity.rs does both, using the
fixture's bos_token_id; render_prompt returns the prefix as its own segment so
hosts keep that boundary. (continuation_bos, continuation_after_text)error, and tests/parity.rs then requires render_request to fail. That
includes pydantic rejections and Python crashes (a 500 is still a rejection).
(invalid_tool_arguments, continuation_only) If Dynamo rejects a shape Python accepts, say in the
PR that hosts fall back to the engine for it.known_gap with the reason. The test
reports the case without failing, and fails once it matches so the marker gets
removed. (tool_float_values: Dynamo's deepseek::common::to_json)tests/scripts/generate_parity.py loads the pinned snapshot and resolves the
spec and per-checkpoint state with SGLang's own resolvers. It then renders every
case through the real OpenAIServingChat._apply_jinja_template and writes:
prompt, token_count and token_sha256 for each case, or error when
Python raises (name, request and known_gap are kept as written). An
AttributeError aborts instead: the stub server lacks something, so extend it;config (the config.json fields the processor reads, plus
resolved overrides such as the effort profile) and bos_token_id.cd rust/sglang-processor
PYTHONPATH=$REPO/python HF_HUB_CACHE=<hub cache> HF_HUB_OFFLINE=1 \
python tests/scripts/generate_parity.py tests/fixtures/parity/<model-id>.jsonWhen a new spec needs more, extend serving_chat() in the script (one place):
__init__ resolves, for
example _dsv41_default_reasoning_effort.apply_chat_template(tokenize=True): the recorder sees no
text. Record the template's text output instead.Pin revision to a commit hash and regenerate only on purpose.
cd rust
cargo test -p sglang-processor --locked # text, offline
HF_HUB_CACHE=<hub cache> cargo test -p sglang-processor --test parity --locked # + token ids
cargo clippy -p sglang-processor --all-targets --locked -- -D warnings
cargo check -p sglang-processor --no-default-features --features render,tokenizer --lockedThe token check loads the pinned commit's snapshot from the cache. When the
snapshot is missing, the test prints token ids not checked (visible with
-- --nocapture), so confirm that line is absent before claiming token parity.
With the snapshot cached, the test also checks that the checkpoint resolves the
fixture's recorded DeepSeek-V4 profile.
dynamo-renderer or dynamo-tokenizers.SGLANG_DEFAULT_THINKING, SGLANG_DSV4_REASONING_EFFORT) are read per
request, as in Python. The generator pins them and tests/parity.rs clears them,
so fixtures do not depend on the machine.OAIChatLikeRequest path) cannot carry every field,
such as task and message-level tools. Parity is defined on render_request;
report host-adapter gaps separately rather than bending the model code.render/mod.rs
beyond the dispatch arm.© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/processor-model-parity of sgl-project/sglang.
Open the folder on GitHubat commit f620d73
Processor Model Parity next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Processor Model Parity this skillsgl-project/sglang | 37k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Stellar DevVelaPayments/vela-payments | 131 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Debug Sessionai-dynamo/dynamo | 8.3k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Release Skillsnexmoe/eve | 421 | 3 repos | ~3.3k | Automated safety check: Pass | None | |
| RustPython C-API ExpansionRustPython/RustPython | 22k | — | ~831 | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.
VelaPayments/vela-payments
End-to-end Stellar development playbook. An agent skill from VelaPayments/vela-payments.
ai-dynamo/dynamo
Sets up a structured debugging session for a Dynamo bug — pull the report from a Linear ticket, GitHub issue, or pasted text, capture the environment, create a persistent worklog markdown file, and…
nexmoe/eve
Universal release workflow. An agent skill from nexmoe/eve.
RustPython/RustPython
Implements missing CPython C-API functions in RustPython's crates/capi, mapping each header to its module with the pyo3-ffi header split.
kucherenko/jscpd
Measures a code port between languages or frameworks with jscpd's function-level comparison, porting tests before code and tracking what is left unmatched.
sgl-project/sglang
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
sgl-project/sglang
Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.
sgl-project/sglang
Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
sgl-project/sglang
Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.
Categories
Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids). Processor Model Parity is an agent skill from sgl-project/sglang. Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids).
Processor Model Parity fits situations like: adding a model to sglang-processor; porting a models rendering from sgl-router; bumping Dynamos renderer; debugging a processor-vs-Python prompt mismatch.
Run `npx skills add sgl-project/sglang --skill processor-model-parity -a claude-code`. Or copy the skill folder (.agents/skills/processor-model-parity in sgl-project/sglang) into .claude/skills/processor-model-parity in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sgl-project/sglang --skill processor-model-parity -a codex`. Or copy the skill folder (.agents/skills/processor-model-parity in sgl-project/sglang) into .agents/skills/processor-model-parity in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill processor-model-parity -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/processor-model-parity, .gemini/skills/processor-model-parity, .github/skills/processor-model-parity and .opencode/skills/processor-model-parity in your project.
Going by SKILL.md and its folder, Processor Model Parity needs the command-line tools its instructions call (cargo and python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Processor Model Parity is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Processor Model Parity: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Stellar Dev (VelaPayments/vela-payments, 131 stars), Debug Session (ai-dynamo/dynamo, 8.3k stars) and Release Skills (nexmoe/eve, 421 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,907 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.
Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.