Vllm Deploy Docker
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration.
$ npx skills add vllm-project/vllm-omni --skill add-tts-model -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vllm-project/vllm-omni add-tts-model --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/add-tts-model .claude/skills/add-tts-model && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "add-tts-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/add-tts-model into .claude/skills/add-tts-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-tts-model", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/add-tts-modelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vllm-project/vllm-omni --skill add-tts-model -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vllm-project/vllm-omni add-tts-model --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/add-tts-model .agents/skills/add-tts-model && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "add-tts-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/add-tts-model into .agents/skills/add-tts-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-tts-model", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-omni --skill add-tts-model -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vllm-project/vllm-omni add-tts-model --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/add-tts-model .cursor/skills/add-tts-model && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "add-tts-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/add-tts-model into .cursor/skills/add-tts-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-tts-model", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vllm-project/vllm-omni.git --path .claude/skills/add-tts-model--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vllm-project/vllm-omni --skill add-tts-model -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vllm-project/vllm-omni add-tts-model --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/add-tts-model .gemini/skills/add-tts-model && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "add-tts-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/add-tts-model into .gemini/skills/add-tts-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-tts-model", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vllm-project/vllm-omni add-tts-modelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vllm-project/vllm-omni --skill add-tts-model -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/add-tts-model .github/skills/add-tts-model && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "add-tts-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/add-tts-model into .github/skills/add-tts-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-tts-model", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-omni --skill add-tts-model -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vllm-project/vllm-omni add-tts-model --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/add-tts-model .opencode/skills/add-tts-model && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "add-tts-model" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/add-tts-model into .opencode/skills/add-tts-model/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "add-tts-model", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
add-tts-modelIntegrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration.
Add Tts Model is an agent skill from vllm-project/vllm-omni. Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. Use when adding a new TTS model, wiring stage separation for speech synthesis, enabling online voice generation serving, debugging TTS integration behavior, or building audio output pipelines.
Its SKILL.md is about 8.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/cuda-graph-example.md`, `references/optional-deps.md` and `references/precommit-dco.md`).
It sits in Media & Creative, covering Text to speech and voice. It works with vLLM, Hugging Face and CUDA. The repository describes itself as: A framework for efficient model inference with omni-modality models. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 88a35c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitruffhfFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Add Tts Model loads about 8.7k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 94 tokens; SKILL.md has 3,679 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vllm-project/vllm-omni at commit 88a35c0, republished under its Apache-2.0 licence (© vllm-project). 3,679 words, ~8,732 tokens.
.claude/skills/add-tts-model/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.HF Reference -> Stage Separation -> Online Serving -> Async Chunk -> CUDA Graph -> Pre-commit/DCO
(Phase 1) (Phase 2) (Phase 3) (Phase 4) (Phase 5) (Phase 6)Three architecture patterns are supported:
inference_stream() generator. Use this when the upstream model bundles ARThe single-stage variants skip Phase 4 (async_chunk) but Phase 5 (CUDA graph) is still encouraged for the inner AR loop.
These rules apply to every TTS model regardless of architecture (AR vs AR+diffusion, single-stage vs two-stage, codec-based vs VAE-based). They surface repeatedly across PRs — check them at the end of every phase.
Pick exactly one per-step semantics for forward() and document it in the docstring:
If you choose delta, verify the full emit→consolidate→consume chain:
forward() returns {"model_outputs": <new_chunk_only>, ...}_consolidate_multimodal_tensors() in vllm_omni/engine/output_processor.py concatenates the audio key into one tensor at finish. If it skips the key (continue), offline consumers receive only the final chunk. See output_processor.py for the concrete list of handled modality keys.engine.generate()) receive a single concatenated tensor.Cumulative-vs-delta mismatch is the most common silent bug — offline RTF benchmarks pass, but users hear replays or truncation.
outputs[0].outputs[0].multimodal_output[<key>] can be any of Tensor, list[Tensor] (pre-consolidation snapshot), np.ndarray, or scalar. When writing tests, examples, and benchmarks:
dict.get("a") or dict.get("b") on tensor values — Python evaluates the tensor's boolean, raising RuntimeError: Boolean value of Tensor with more than one value is ambiguous. Use explicit if x is None chains.if isinstance(x, list): x = torch.cat([t.reshape(-1) for t in x], dim=0).shape / dtype / duration explicitly; do not rely on truthiness for presence checks.Inside any per-step model loop (AR decode, diffusion solver, CFM Euler, vocoder block loop):
tensor.item(), .cpu(), or .tolist() — each triggers a GPU→CPU sync; at 10 steps × 60 frames × 4 ops that is 2400 syncs per request.dst.copy_(src) over dst.fill_(src.item()) when writing a scalar tensor into a buffer.torch.compile(Model.forward, fullgraph=False) on the whole forward over per-submodule compile — fewer dispatch boundaries, larger fusion regions. Measure before choosing granularity.torch.where / masking instead.Profile first, optimize second. See the profiling docs / project memory for the trace-analysis workflow.
Offline RTF alone is necessary but not sufficient. Every new TTS model must pass all three:
| Layer | Catches | Tool |
|---|---|---|
| Offline RTF / duration check | Throughput regressions, missing audio, wrong sample rate | end2end.py, pytest e2e |
| Browser streaming playback | Delta/cumulative bugs, chunk boundary glitches, TTFP regressions | Gradio demo over /v1/audio/speech?stream=true |
| Concurrent requests | Per-request state leaks, codec window round-robin gaps | max_num_seqs>1 smoke test with 4+ parallel prompts |
Declaring a model "done" without all three has shipped regressions more than once.
If the model caches anything across forward() calls (streaming generators, codec buffers, sliding-window pads, CUDA graph state), key it by request ID:
self._state: dict[str, YourState] = {} # request_key → state
# fetch: request_key = str(info.get("_omni_req_id", "0"))
# free on finish: del self._state[request_key]A shared buffer silently corrupts audio across concurrent requests — the symptom is crosstalk or truncation only under load.
Goal: Understand the reference implementation and verify it produces correct audio.
config.json fields, model_type, sub-model configs<|voice|>, <|audio_start|>, <|im_end|>, etc.)Goal: Split the model into vLLM-Omni stages and get offline inference working.
vllm_omni/model_executor/models/registry.pyconfiguration_<model>.py) with model_type registrationforward() for autoregressive token generationforward(): codec codes -> audio waveformOmniOutput with multimodal_outputsvllm_omni/model_executor/models/<model>/pipeline.pyvllm_omni/deploy/ for placement,
memory sizing, connectors, and runtime overrides| Parameter | Impact if Wrong |
|---|---|
| Hop length | Audio duration wrong, streaming noise |
| Token ID mapping | Garbage codes -> noise output |
| Codebook count/size | Shape mismatch crashes |
| Stop token | Generation never stops or stops too early |
| dtype / autocast | Numerical issues, silent quality degradation |
| Repetition penalty | Must match reference (often 1.0 for TTS) |
When audio output is wrong, check in this order:
These bugs appear in almost every new TTS PR. Check all before the first push. See also the cross-cutting invariants I1 (output contract) and I5 (per-request state) above — the rules below are the Phase 2-specific instances of those invariants:
forward() appends new codes; do not reset between steps or audio will be truncated (fish speech: fix: accumulate audio_codes across steps)fix: emit delta audio not full waveform)model_outputs — if any early-return branch skips setting model_outputs, the serving layer silently drops that step's audio (fish speech: fix: ensure ALL return paths emit model_outputs)fix: per-request vocode + delta emission)fix: use model device for CUDA stream)max_num_seqs — set to at least 4 in production deploy configs; for single-stage models this is the only stage. For two-stage models, Stage 0 (AR) needs max_num_seqs ≥ 4 to pipeline concurrent requests; Stage 1 (codec decoder) is model-specific and may intentionally use max_num_seqs: 1. Defaulting the AR stage to 1 causes audio gaps under concurrency because the codec window round-robins across requests (RFC #2568)Patch optional dependencies (torchaudio / torchcodec / soundfile) at
the top of load_weights(), not at module import. Failures to do so cause
cryptic errors only on environments missing the optional package — after
the model is already deployed. See
references/optional-deps.md for the full
pattern, signature constraints, and MOSS-TTS-Nano reference.
When the upstream model cannot be cleanly split into an AR stage and a
separate decoder, run the full pipeline inside a single AR worker and
stream audio through a per-request inference_stream() generator keyed by
_omni_req_id. Define one StagePipelineConfig with
execution_type=StageExecutionType.LLM_AR, engine_output_type="audio",
final_output=True, and owns_tokenizer=True. Set async_chunk: false in
the deploy YAML. Only extract params from
additional_information that you actually forward, or pre-commit fails
ruff F841.
Full walkthrough with the complete forward() / _create_stream_gen()
skeleton plus pipeline/deploy definitions:
references/single-stage-ar.md. For an
in-tree reference, look for any single-stage AR model under
vllm_omni/model_executor/models/, such as MOSS-TTS-Nano.
VoxCPM2 is a different pattern and should not reuse this skeleton — it
runs the base LM under vLLM PagedAttention with external side-computation.
See plan/voxcpm2_native_ar_design.md.
vllm_omni/model_executor/models/<model_name>/pipeline.py topologyvllm_omni/deploy/end2end.py at examples/offline_inference/text_to_speech/<model>/end2end.pyexamples/offline_inference/text_to_speech/README.md (table row + per-model section). Do not create a top-level examples/offline_inference/<model>/ dir or a per-model README.md inside text_to_speech/<model>/ — the hub README is the documented surface and the mkdocs generate_examples hook only descends one level into examples/<category>/.Goal: Expose the model via /v1/audio/speech API endpoint.
Write one adapter under vllm_omni/entrypoints/openai/tts_adapters/.
serving_speech.py should not need an edit — detection, stage discovery and
dispatch are all derived from what the adapter declares.
Create vllm_omni/entrypoints/openai/tts_adapters/your_model.py:
@register_tts_adapter
class YourModelAdapter(ARTTSAdapter):
name = "your_model" # registry key + log label
stage_keys = frozenset({"your_stage_key"}) # the deploy yaml's model_stage
def validate(self, request) -> str | None:
if not request.input or not request.input.strip():
return "Input text cannot be empty"
return None
async def build(self, request, sampling_params_list, has_inline_ref_audio):
params = {"text": [request.input]}
if request.voice is not None:
params["voice"] = [request.voice]
return PreparedRequest(
prompt={"prompt": request.input},
tts_params=params,
model_type=self.name,
)Then add the module to the import block at the bottom of
tts_adapters/__init__.py so it registers. That import line is the only shared
file a new model touches — which is also why the old rebase-conflict hotspot is
gone.
Pure-diffusion TTS does not go through adapters yet. Under
for_diffusion(),create_speech()routes straight to_create_diffusion_speech()and never callsvalidate()/build().DiffusionTTSAdapteris scaffolding with no production subclass, so logic placed in one would silently never run. Diffusion-engine models follow the existing diffusion path; wiring it through adapters is open work (#4855).
If a stage key alone cannot identify the model, declare
model_archs = frozenset({"YourModelForConditionalGeneration"}); add
arch_identifies_entry_stage = True when the model owns no stage key at all
(Ming dense). For a rule that is not set membership, override matches() (see
covo_audio.py). For a genuine overlap with another adapter, give one an
explicit detect_priority — test_tts_detection.py fails on an unordered
overlap. Models that only serve speech in some topologies override
stage_serves_speech() (see audex.py).
Reuse shared helpers via self.ctx.server rather than reimplementing them:
_resolve_ref_audio, _apply_uploaded_speaker, _validate_ref_audio_format,
_max_instructions_length. Read a comparable adapter first — fish_speech.py
(voice cloning), higgs_audio_v3.py (parameter-heavy), moss_tts.py (family
sharing a base class).
Do not add
self._tts_model_type == ...branches toserving_speech.py. Older models predate the adapter framework and still have them; they are being migrated out (RFC #4327, #4855).tools/pre_commit/check_tts_adapter.pyis a ratchet on the remaining count and fails the commit if it grows. Behaviour that no adapter hook can express is a missing hook — propose it on the RFC.
Unused variable rule: only extract fields in
build()that are actually forwarded to the model. Unused extractions failruff F841. For voice-cloning fields (ref_audio->prompt_audio_path,ref_text->prompt_text), add them to the params and verify they reach the model call.
Handle model-specific parameters:
ref_audio encoding and prompt injectionmax_new_tokens override in sampling paramsCreate client scripts: speech_client.py, run_server.sh
Test all response formats: wav, mp3, flac, pcm
Add Gradio demo: Interactive web UI with streaming support
import base64
from pathlib import Path
def build_voice_clone_prompt(ref_audio_path: str, text: str, codec) -> list:
"""Build prompt with reference audio for voice cloning, called from the adapter."""
audio_bytes = Path(ref_audio_path).read_bytes()
codes = codec.encode(audio_bytes) # Encode on CPU using model's codec (e.g., DAC)
token_ids = [code + codec.vocab_offset for code in codes.flatten().tolist()]
return [
{"role": "system", "content": f"<|voice|>{''.join(chr(t) for t in token_ids)}"},
{"role": "user", "content": text},
]Follow the vllm-omni-test skill for markers, file naming (test_{slug}.py / test_{slug}_expansion.py), Buildkite wiring, and copy-paste run commands. Also read test_system_overview.md and test_writing_guide.md.
Classify the model's CI priority first (high / medium / low). High-priority TTS models are typically those on the integration hot path or listed in tracking issues such as #1832; medium and low tiers cover the long tail. When unsure, ask the reviewer which tier applies.
| Priority | Required test levels | Files & markers |
|---|---|---|
| High | L1 unit/logic · L2 online smoke · L3 online + offline integration · L4 feature + performance | See table below |
| Medium | L3 online + offline · L4 feature only | Skip dedicated L1/L2 unless fixing a logic bug |
| Low | L4 feature only | One or two *_expansion.py parametrized cases |
Per-level deliverables (TTS / pytest.mark.tts):
| Level | Location | Marker | CI pipeline | Notes |
|---|---|---|---|---|
| L1 | tests/model_executor/…, tests/entrypoints/openai_api/…, stage-processor tests | core_model + cpu | test-ready.yml | Prompt assembly, async_chunk helpers, adapter validation — no GPU |
| L2 | tests/e2e/online_serving/test_{slug}.py | core_model + advanced_model (both on baseline smoke) + tts + @hardware_test(...) | test-ready.yml (ready label) | Default deploy smoke: single /v1/audio/speech or offline OmniRunner path |
| L3 | tests/e2e/online_serving/test_{slug}.py and tests/e2e/offline_inference/test_{slug}.py | Baseline smoke: core_model + advanced_model; heavier cases: advanced_model only (+ tts) | test-merge.yml or merged into nightly TTS function job | Streaming, voice clone, batch/queue, async_chunk |
| L4 | tests/e2e/online_serving/test_{slug}_expansion.py, optional offline expansion | full_model + tts | test-nightly.yml (:full_moon: TTS · Function Test with L4) | Feature matrix; perf → tests/dfx/perf/tests/test_tts.json |
L2 & L3 online — same file, dual marks on the baseline smoke: The first / simplest case in test_{slug}.py (default deploy, single non-streaming /v1/audio/speech or equivalent offline path) should carry both @pytest.mark.core_model and @pytest.mark.advanced_model on the same function so it runs in L2 (test-ready.yml, --run-level core_model, basic validation) and L3 (test-merge.yml, --run-level advanced_model, deeper validation) without duplicating the test. In-tree examples: test_voxcpm2_tts.py::test_text_to_audio_001, test_qwen3_tts_customvoice.py::test_text_to_audio_001.
Heavier scenarios in the same file use advanced_model only (streaming, extra languages, concurrency, async_chunk, batch). Example: test_voice_clone_en_streaming_001 → advanced_model only. When migrating L3 to nightly, move those heavier cases into test_{slug}_expansion.py with full_model and drop the dedicated merge job (see test_ming_tts_expansion.py, test_glm_tts_expansion.py).
@pytest.mark.core_model
@pytest.mark.advanced_model
@pytest.mark.tts
@hardware_test(res={"cuda": "L4"}, num_cards=1)
@pytest.mark.parametrize("omni_server", tts_server_params, indirect=True)
def test_voice_clone_en_non_streaming_001(omni_server, online_client) -> None:
online_client.send_audio_speech_request({...})L4 consolidation: Prefer parametrized OmniServerParams rows (default, async_chunk, feature flags) in one expansion module rather than many merge-only files (#1832).
L4 performance (high-priority models): Add latency / throughput / stress rows in tests/dfx/perf/tests/test_tts.json, or a dedicated tests/dfx/perf/tests/test_{slug}.json when the model must not join the shared nightly server matrix before integration lands (see VoxCPM2 / Coqui XTTS pattern). Register the model in benchmarks/tts/model_configs.yaml for local bench_tts.py. Wire a separate test-nightly.yml Perf Test step when the JSON is not merged into test_tts.json yet.
Keep model-specific code inside test modules — not tests/helpers/{slug}.py:
MODEL, deploy path, vendored REF_AUDIO_URL, get_prompt(), and inline request_config dicts in each test_{slug}.py, test_{slug}_expansion.py, offline test_{slug}.py, and L1 test_{slug}_*.py as needed.tests/helpers/{slug}.py (or tests/helpers/{model_name}.py) to deduplicate constants or request builders across those files. A little duplication is intentional; follow in-tree references such as tests/e2e/online_serving/test_glm_tts.py and tests/e2e/online_serving/test_cosyvoice3_tts_expansion.py.tests/helpers/ is for repo-wide harness code only (mark.py, media.py, runtime.py, stage_config.py, assertions.py, fixtures/). Import those; do not extend the tree with per-model modules.Runtime send helpers (tests/helpers/runtime.py) — online and offline e2e:
| Path | Fixture | Call |
|---|---|---|
Online /v1/* | online_client | online_client.send_*_request(request_config) |
| Offline inference | offline_client | offline_client.send_*_request(request_config) |
runtime.py first — reuse send_omni_request, send_diffusion_request, send_audio_speech_request (online + offline Qwen-style TTS), send_single_stage_tts_request (Coqui XTTS / MOSS-TTS-Nano offline), etc.send_<feature>_request (or send_<route>_http_request for negative/dfx) in runtime.py with general assert_* bundled inside, then call it from the test.request_config dicts only — not omni.generate, not _collect_audio(), not raw HTTP/SDK.tests/helpers/runtime.py into Buildkite source_file_dependencies when you add helpers.See vllm-omni-test skill § Runtime send helpers for full tables and exceptions.
tts_adapters/ plus its line in the package import blockexamples/online_serving/text_to_speech/<model>/test-ready.yml (L1/L2), test-merge.yml or nightly function job (L3), test-nightly.yml (L4) — see vllm-omni-test skillexamples/online_serving/text_to_speech/README.md (table row + per-model section). Do not create a top-level examples/online_serving/<model>/ dir or a per-model README.md inside text_to_speech/<model>/.OmniServerParams set per file. omni_server is module-scoped; a second
id in the same file forces mid-module teardown/restart and exposes startup
races (APIConnectionError on the first request post-restart). Split variants
into separate files instead.raw.githubusercontent.com over TLS. Inline ref audio as
data:audio/wav;base64,...; the serving layer accepts both URL and data URL./health; don't add time.sleep in tests. If warmup is incomplete, make
/health return non-200 until you're actually ready.core_model + advanced_model; heavier cases: advanced_model only; L4 expansion: full_modeltests/helpers/{slug}.py; keep constants and request_config payloads in the test fileruntime.py — online_client.send_audio_speech_request (online); offline_client.send_audio_speech_request (Qwen-style offline) or send_single_stage_tts_request (single-stage offline). Add a new send_*_request in runtime.py when none fits; do not embed omni.generate or HTTP in testsGoal: Enable inter-stage streaming so audio chunks are produced while AR generation continues.
async_chunk_process_next_stage_input_func.async_chunk: true
connectors:
connector_of_shared_memory:
name: SharedMemoryConnector
extra:
codec_streaming: true
codec_chunk_frames: 25
codec_left_context_frames: 25OmniOutputstream=true with PCM outputStage 0 (AR) Stage 1 (Decoder)
| |
|-- chunk 0 (25 frames) ------> decode -> audio chunk 0 -> client
|-- chunk 1 (25 frames) ------> decode -> audio chunk 1 -> client
|-- chunk 2 (25 frames) ------> decode -> audio chunk 2 -> client
...context_audio_samples = context_frames * hop_lengthasync_chunk: trueGoal: Capture the AR loop as a CUDA graph for significant speedup.
torch.argmax instead of torch.multinomial (graph-safe)See references/cuda-graph-example.md for a worked skeleton (Qwen3-TTS code predictor, 16-step AR loop), performance expectations (3–5× on the graphed component for fixed batch_size=1), and the graph-safety constraints you must honor inside the captured region.
Goal: Every commit passes pre-commit lint and carries a DCO
Signed-off-by line that matches the author email.
pre-commit install.pre-commit run --files <changed-files> before every push; accept any
auto-fixes, stage, re-commit.git commit -s. DCO checks that author email and
Signed-off-by email match — git config user.email must match your
GitHub account email.Common pre-commit failures, recovery commands for missing sign-off, and the
full pre-commit run invocation for a TTS model:
references/precommit-dco.md.
Follow the add-recipe skill to add or update the
in-repository model-family recipe, the recipes/README.md index, the supported
models table, TTS example documentation, and the speech API contract using
validated evidence.
Use this checklist when integrating a new TTS model:
forward() docstring states cumulative vs delta; consolidation path audited end-to-enddict.get(a) or dict.get(b) on tensor values; list form handled.item() / .cpu() / Python branch on tensor values inside per-step loops_omni_req_id; entries freed when the request finishesregistry.pymodel_type registrationmax_num_seqs ≥ 4 in the production deploy config unless the model has a tested lower limitload_weights() time (torchaudio/soundfile/etc.)forward() return paths emit model_outputspipeline.py and registered in OMNI_PIPELINESvllm_omni/deploy/end2end.py produces audio matching reference qualitytts_adapters/ and registered in the import blockbuild() that are forwarded to the model call (ruff F841)test-ready.yml / test-merge.yml or nightly TTS job / test-nightly.ymlrecipes/README.md row added via the add-recipe skillasync_chunk: true and connector chunk parametersstream=true) workspre-commit run --files <changed> passes before every pushSigned-off-by matching the author email (git commit -s)git config user.email matches the email registered on your GitHub accountIn-skill references (details split out of the main body):
forward() / generator skeleton for the MOSS-TTS-Nano-style patternProject docs and adjacent skills:
plan/voxcpm2_native_ar_design.md — VoxCPM2's vLLM-native AR + side-computation pattern (distinct from the generator-based single-stage described above)© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in .claude/skills/add-tts-model of vllm-project/vllm-omni.
Open the folder on GitHubat commit 88a35c0
Add Tts Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Add Tts Model this skillvllm-project/vllm-omni | 7.1k | — | ~8.7k | Automated safety check: Pass | Apache-2.0 | |
| Vllm Deploy Dockervllm-project/vllm-skills | 103 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Esmfold2JimLiu/science-skills | 227 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Production Defaultshuggingface/skills | 11k | 1 repos | ~6.9k | Automated safety check: Pass | Apache-2.0 |
vllm-project/vllm-skills
Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
huggingface/skills
Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.
guqiong96/Lvllm
Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…
vllm-project/vllm-omni
Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.
vllm-project/vllm-omni
Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness.
vllm-project/vllm-omni
Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.
vllm-project/vllm-omni
Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings.
vllm-project/vllm-omni
Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA.
vllm-project/vllm-omni
Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…
Works with
Categories
Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. Add Tts Model is an agent skill from vllm-project/vllm-omni. Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration.
Add Tts Model fits situations like: adding a new TTS model; wiring stage separation for speech synthesis; enabling online voice generation serving; debugging TTS integration behavior.
Run `npx skills add vllm-project/vllm-omni --skill add-tts-model -a claude-code`. Or copy the skill folder (.claude/skills/add-tts-model in vllm-project/vllm-omni) into .claude/skills/add-tts-model in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vllm-project/vllm-omni --skill add-tts-model -a codex`. Or copy the skill folder (.claude/skills/add-tts-model in vllm-project/vllm-omni) into .agents/skills/add-tts-model in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-omni --skill add-tts-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-tts-model, .gemini/skills/add-tts-model, .github/skills/add-tts-model and .opencode/skills/add-tts-model in your project.
Going by SKILL.md and its folder, Add Tts Model needs the command-line tools its instructions call (git, ruff and hf). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Add Tts Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.7k tokens (SKILL.md is roughly 35k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Add Tts Model: Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Esmfold2 (JimLiu/science-skills, 227 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vllm-project (a GitHub organization) maintains it in vllm-project/vllm-omni, which has 7,097 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 9, 2026.
Source: vllm-project/vllm-omni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.