Agent skill

Processor Model Parity

by sgl-project in sgl-project/sglang

Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids).

Apache-2.0Auto-check passedDevelopment

Install Processor Model Parity

skills CLI
$ npx skills add sgl-project/sglang --skill processor-model-parity -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sgl-project/sglang processor-model-parity --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sgl-project/sglang.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/processor-model-parity .claude/skills/processor-model-parity && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
processor-model-parity
GitHub stars
37k
Token cost
~3.1k tokens
SKILL.md length
1,362 words
Files
1
Skills in repo
32
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids).

  • Works in 6 steps: Pick the checkpoint. Use the one the… → Trace the Python path for the model.… → Find Dynamo's counterpart. Check… → …
  • Adding a model to sglang-processor
  • SKILL.md covers Where a model's code goes, Method, Checklist (DeepSeek-V4 case… and Generating fixtures, plus 2 more sections
  • Calls cargo and python

What it does

Processor Model Parity is an agent skill from sgl-project/sglang. Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids). Use when adding a model to sglang-processor, porting a model's rendering from sgl-router, bumping Dynamo's renderer, or debugging a processor-vs-Python prompt mismatch.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Natural language processing. It works with SGLang, Python, Rust and DeepSeek. The repository describes itself as: SGLang is a high-performance serving framework for large language models and multimodal models. The licence is Apache-2.0.

When your agent uses it

  • Adding a model to sglang-processor
  • Porting a models rendering from sgl-router
  • Bumping Dynamos renderer
  • Debugging a processor-vs-Python prompt mismatch

Example prompts

  • “s rendering from sgl-router, bumping Dynamo”
  • “/processor-model-parity”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Pick the checkpoint. Use the one the model's cookbook page recommends
  2. Trace the Python path for the model. Write down every transformation
  3. Find Dynamo's counterpart. Check dynamo_renderer::native_formatter_for,
  4. Write one request case per inventory item in the fixture (name +
  5. Implement. Start from a plain Dynamo call and run tests/parity.rs. For
  6. Verify text offline, then token ids with the checkpoint cached (commands

What it can do on your machine

Read from SKILL.md and the folder at commit f620d73. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Processor Model Parity loads about 3.1k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 1,362 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sgl-project/sglang at commit f620d73, republished under its Apache-2.0 licence (© sgl-project). 1,362 words, ~3,089 tokens.

Download SKILL.mdSave it as .claude/skills/processor-model-parity/SKILL.md (or your agent's skills folder).
name
processor-model-parity
description
Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids). Use when adding a model to sglang-processor, porting a model's rendering from sgl-router, bumping Dynamo's renderer, or debugging a processor-vs-Python prompt mismatch.

Processor model parity

rust/sglang-processor renders chat requests for SGLang's Rust hosts (the Rust server, sgl-router, the renderer). A model is supported only when, for the same request body, the processor produces the same prompt text and the same prompt token ids as Python's OpenAIServingChat. Python is the reference, Dynamo is the implementation, and SGLang code exists only where the two differ.

Scope: the request -> prompt -> token ids path. Output parsing (reasoning and tool calls) and multimodal preprocessing are not covered here.

The worked example is DeepSeek-V4-Flash-0731: src/render/models/deepseek_v4.rs plus tests/fixtures/parity/deepseek-v4-flash-0731.json. Paths below are relative to rust/sglang-processor/ unless they start with python/.

Where a model's code goes

PieceFileNeeded when
Identity -> formattersrc/render/selection.rs (select_chat_formatter)always; mirror the order of chat_encoding.resolve_chat_encoding_spec
Formatter variantsrc/render/mod.rs (ChatFormatter, its render_request arm, plus the render_prompt, stop_strs and resolve_thinking arms)always
SGLang-specific stepssrc/render/models/<model>.rs, re-exported in models/mod.rsonly where Dynamo diverges from Python
Parity casestests/fixtures/parity/<model-id>.jsonalways
README rowthe render table in README.mdalways

No new test code: tests/parity.rs checks every fixture in the directory.

render_request(Value, ..) takes the SGLang request body, not Dynamo's OAIChatLikeRequest, because Python reads fields the trait lacks (task, continue_final_message, the reasoning object). Pick the shape by how Python renders the model:

  • Python calls its own encoder (a non-None spec such as dsv4 or dsv32): give the model a dedicated ChatFormatter variant and render_request arm, like DeepSeek-V4. If selection.rs already wires the model to a Dynamo PromptFormatter, replace that branch.
  • Python calls apply_chat_template (Jinja, spec None): the first such model adds one generic arm that adapts the body into OAIChatLikeRequest, in its own PR. sgl-router's ChatRequest in experimental/sgl-router/src/tokenizer/chat_formatter.rs is a working adapter.

Method

  1. Pick the checkpoint. Use the one the model's cookbook page recommends (docs/cookbook/), pin its Hugging Face commit as revision, and cache it. Note whether its tokenizer ships a chat_template; that can change the spec.

  2. Trace the Python path for the model. Write down every transformation between the HTTP body and prompt_ids, in order:

    • protocol.py::ChatCompletionRequest.normalize_reasoning_inputs: reasoning and reasoning_effort set the default thinking / enable_thinking. Reuse render/reasoning.rs for this rather than porting it again.
    • serving_chat._convert_to_internal_request: chat_template_kwargs.reasoning_effort replaces the request effort; server --default-chat-template-kwargs merge in.
    • chat_encoding.resolve_chat_encoding_spec: which branch renders the model (a spec such as dsv4, dsv32, kimi_k3, inkling, or None for Jinja). It also reads the server's --tool-call-parser and whether the tokenizer has a chat template. If the deployment depends on the parser, set tool_call_parser in the fixture, and add it to ChatFormatterOptions in the same PR. Note any per-checkpoint state the server resolves once (DeepSeek-V4's effort profile reads encoding/encoding_dsv4.py).
    • That spec's branch in serving_chat._apply_jinja_template / _encode_messages: message dumping, content flattening, _handle_last_assistant_message, system insertion, tools, model-specific fields, the encoder or apply_chat_template call, and how the prompt is tokenized.
    • The encoder itself (for example encoding_dsv4.py), line by line. History rules live there: dropped turns, </think> on empty reasoning, argument formatting, and the errors it raises.

    This list is the divergence inventory.

  3. Find Dynamo's counterpart. Check dynamo_renderer::native_formatter_for, the model's low-level encoder (for example deepseek::v4::encode_messages_with_options), or the Jinja template. Prefer the lowest-level Dynamo entry that lets SGLang own request semantics. Dynamo's OpenAI-level formatters apply their own defaults (they filter tools by tool_choice, inject response_format, and map effort), which often differ from SGLang's. Diff the low-level encoder against Python's for the rules in step 1.

  4. Write one request case per inventory item in the fixture (name + request only), then generate the expected outputs from Python (below). Cover the checklist; reuse the request shapes from the DeepSeek-V4 fixture where they apply.

  5. Implement. Start from a plain Dynamo call and run tests/parity.rs. For each failing case, add the smallest SGLang step that fixes it, with a one-line doc comment naming the Python source it mirrors. Never patch the fixture to match Rust.

  6. Verify text offline, then token ids with the checkpoint cached (commands below). If sgl-router or the renderer uses the formatter, run their suites too.

Show full SKILL.md (720 more words)Show less

Checklist (DeepSeek-V4 case that covers each)

  • Message dump: roles lowercased; unknown and null fields dropped; user reduced to role and content; null content becomes "". (ignored_fields, tool_results)
  • Content parts: which part types survive and the join separator. (parts)
  • Tools: all request tools or only those tool_choice allows; field order and defaults of Function.model_dump; message-level tools; empty lists. (tools, tools_none, tools_named, message_tools, empty_message_tools)
  • Pydantic coercion: typed request fields are coerced before rendering. For example, strict: 1 and defer_loading: "false" become booleans, and strict: null is rejected. Probe the pydantic model in Python to get its exact rules; don't guess them. (tool_bool_coercion, tool_strict_null, continuation_coerced)
  • Tool-call arguments: how the encoder wants them. DeepSeek-V4 takes a compact JSON string; other encoders format them their own way. Keep key order (serde_json preserve_order is declared in Cargo.toml). Floats must print as Python's json.dumps prints them (1e-06, 1e+16, 100.0). (tool_results, agentic_thinking, tool_float_values, history_float_arguments)
  • Thinking: kwargs thinking > request effort (!= "none") > reasoning.enabled

    server default kwargs > SGLANG_DEFAULT_THINKING. enabled follows Python truthiness (1 is on); strings use Python's yes-word set. (thinking_false, thinking_none, drop_thinking_ignored, reasoning_enabled_numeric, reasoning_enable_string)

  • Effort: precedence between kwargs, reasoning, reasoning_effort and env; the checkpoint's profile mapping. (effort_high, effort_max, kwargs_effort, effort_conflict, reasoning_object)
  • Final assistant turn: becomes a user turn, or with continue_final_message a separately tokenized prefix. Check whether content is flattened before the split (it is for DeepSeek-V4). Run each check where Python runs it: a final assistant turn's tool calls are discarded, so object checks on their arguments belong after the split. Hosts reach the model through render_prompt too, so check the renderer's continuation handling (sglang-renderer render()) as well. (final_assistant, continuation_parts, continuation_bos, continuation_only, final_tool_call_no_arguments, continuation_tool_call_null_arguments)
  • Model-specific fields and turns: task placement, a system turn mid-conversation, inserted empty system turns. (task_action, task_after_developer, consecutive_task, mid_system)
  • History: the encoder's rules for earlier turns: reasoning kept or dropped, whole turns dropped, closing tags on empty reasoning. (multi_turn, thinking_multi_turn)
  • Tokens: Python's tokenizer.encode adds special tokens by default; the continuation prefix is encoded alone and loses a leading BOS (_append_assistant_prefix_to_prompt_ids). tests/parity.rs does both, using the fixture's bos_token_id; render_prompt returns the prefix as its own segment so hosts keep that boundary. (continuation_bos, continuation_after_text)
  • Errors: reject what Python rejects. A case Python raises on is recorded with error, and tests/parity.rs then requires render_request to fail. That includes pydantic rejections and Python crashes (a 500 is still a rejection). (invalid_tool_arguments, continuation_only) If Dynamo rejects a shape Python accepts, say in the PR that hosts fall back to the engine for it.
  • Known gaps: when a mismatch is inside Dynamo and out of the processor's reach, fix it upstream and mark the case known_gap with the reason. The test reports the case without failing, and fails once it matches so the marker gets removed. (tool_float_values: Dynamo's deepseek::common::to_json)

Generating fixtures

tests/scripts/generate_parity.py loads the pinned snapshot and resolves the spec and per-checkpoint state with SGLang's own resolvers. It then renders every case through the real OpenAIServingChat._apply_jinja_template and writes:

  • prompt, token_count and token_sha256 for each case, or error when Python raises (name, request and known_gap are kept as written). An AttributeError aborts instead: the stub server lacks something, so extend it;
  • the fixture-level config (the config.json fields the processor reads, plus resolved overrides such as the effort profile) and bos_token_id.
sh
cd rust/sglang-processor
PYTHONPATH=$REPO/python HF_HUB_CACHE=<hub cache> HF_HUB_OFFLINE=1 \
    python tests/scripts/generate_parity.py tests/fixtures/parity/<model-id>.json

When a new spec needs more, extend serving_chat() in the script (one place):

  • More server state: set the extra attribute that __init__ resolves, for example _dsv41_default_reasoning_effort.
  • Tokenizes through apply_chat_template(tokenize=True): the recorder sees no text. Record the template's text output instead.

Pin revision to a commit hash and regenerate only on purpose.

Verify

sh
cd rust
cargo test -p sglang-processor --locked                                       # text, offline
HF_HUB_CACHE=<hub cache> cargo test -p sglang-processor --test parity --locked  # + token ids
cargo clippy -p sglang-processor --all-targets --locked -- -D warnings
cargo check -p sglang-processor --no-default-features --features render,tokenizer --locked

The token check loads the pinned commit's snapshot from the cache. When the snapshot is missing, the test prints token ids not checked (visible with -- --nocapture), so confirm that line is absent before claiming token parity. With the snapshot cached, the test also checks that the checkpoint resolves the fixture's recorded DeepSeek-V4 profile.

Pitfalls

  • Dynamo version bumps change rendering silently: rerun every fixture after bumping dynamo-renderer or dynamo-tokenizers.
  • Env vars (SGLANG_DEFAULT_THINKING, SGLANG_DSV4_REASONING_EFFORT) are read per request, as in Python. The generator pins them and tests/parity.rs clears them, so fixtures do not depend on the machine.
  • Typed hosts (the renderer's OAIChatLikeRequest path) cannot carry every field, such as task and message-level tools. Parity is defined on render_request; report host-adapter gaps separately rather than bending the model code.
  • Keep model code in its own file: no model-specific branches in render/mod.rs beyond the dispatch arm.

© sgl-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/processor-model-parity of sgl-project/sglang.

Open the folder on GitHubat commit f620d73

Compare with similar skills

Processor Model Parity next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Processor Model Parity compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Processor Model Parity this skillsgl-project/sglang37k—~3.1kAutomated safety check: PassApache-2.0
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.4kAutomated safety check: PassMIT
Stellar DevVelaPayments/vela-payments131—~1.8kAutomated safety check: PassMIT
Debug Sessionai-dynamo/dynamo8.3k—~1.2kAutomated safety check: PassApache-2.0
Release Skillsnexmoe/eve4213 repos~3.3kAutomated safety check: PassNone
RustPython C-API ExpansionRustPython/RustPython22k—~831Automated safety check: PassMIT

Similar skills

  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 6 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Stellar Dev

    VelaPayments/vela-payments

    End-to-end Stellar development playbook. An agent skill from VelaPayments/vela-payments.

    131 GitHub stars~1.8k tokensUpdated 3 days ago
    Backend & APIsAuto-check passed
  • Debug Session

    ai-dynamo/dynamo

    Sets up a structured debugging session for a Dynamo bug — pull the report from a Linear ticket, GitHub issue, or pasted text, capture the environment, create a persistent worklog markdown file, and…

    8.3k GitHub stars~1.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Release Skills

    nexmoe/eve

    Universal release workflow. An agent skill from nexmoe/eve.

    421 GitHub starsUsed in 3 repos~3.3k tokens
    DevelopmentAuto-check passed
  • RustPython C-API Expansion

    RustPython/RustPython

    Implements missing CPython C-API functions in RustPython's crates/capi, mapping each header to its module with the pyo3-ffi header split.

    22k GitHub stars~831 tokensUpdated today
    DevelopmentAuto-check passed
  • Measures a code port between languages or frameworks with jscpd's function-level comparison, porting tests before code and tracking what is left unmatched.

    6.4k GitHub stars~5k tokensUpdated today
    DevelopmentAuto-check passed

More from sgl-project/sglang

All 32 skills in this repo
  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    Auto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    Auto-check passed
  • Compute Mamba Ratio

    sgl-project/sglang

    Compute the optimal --mamba-full-memory-ratio (or --max-mamba-cache-size pin) for a hybrid attention + linear-attention (Mamba / GDN / KDA) model's two serving memory pools, from the workload and…

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Env Var Conventions

    sgl-project/sglang

    Conventions for SGLang environment variables — where to define, how to access, how to name, and how to deprecate.

    37k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed

Questions about Processor Model Parity

What does Processor Model Parity do?

Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids). Processor Model Parity is an agent skill from sgl-project/sglang. Add or verify a model's chat prompt rendering and tokenization in rust/sglang-processor so it matches SGLang's Python serving path exactly (same prompt text, same token ids).

When should I use Processor Model Parity?

Processor Model Parity fits situations like: adding a model to sglang-processor; porting a models rendering from sgl-router; bumping Dynamos renderer; debugging a processor-vs-Python prompt mismatch.

How do I install Processor Model Parity in Claude Code?

Run `npx skills add sgl-project/sglang --skill processor-model-parity -a claude-code`. Or copy the skill folder (.agents/skills/processor-model-parity in sgl-project/sglang) into .claude/skills/processor-model-parity in your project. Claude Code loads it when a task matches its description.

How do I install Processor Model Parity in Codex?

Run `npx skills add sgl-project/sglang --skill processor-model-parity -a codex`. Or copy the skill folder (.agents/skills/processor-model-parity in sgl-project/sglang) into .agents/skills/processor-model-parity in your project. Codex loads it when a task matches its description.

Can I use Processor Model Parity in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sgl-project/sglang --skill processor-model-parity -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/processor-model-parity, .gemini/skills/processor-model-parity, .github/skills/processor-model-parity and .opencode/skills/processor-model-parity in your project.

What does Processor Model Parity need to run?

Going by SKILL.md and its folder, Processor Model Parity needs the command-line tools its instructions call (cargo and python). Our summary lists: Python 3.

Does Processor Model Parity access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Processor Model Parity safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Processor Model Parity use?

Processor Model Parity is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Processor Model Parity use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Processor Model Parity?

Skills that share tags, products or a category with Processor Model Parity: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Stellar Dev (VelaPayments/vela-payments, 131 stars), Debug Session (ai-dynamo/dynamo, 8.3k stars) and Release Skills (nexmoe/eve, 421 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Processor Model Parity?

sgl-project (a GitHub organization) maintains it in sgl-project/sglang, which has 36,907 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 9, 2026.

Source: sgl-project/sglang on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.