SGLang Structured Serving
Orchestra-Research/AI-Research-SKILLs
Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
$ npx skills add dstackai/dstack --skill dstack-prototyping -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install dstackai/dstack dstack-prototyping --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/dstackai/dstack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dstack-prototyping .claude/skills/dstack-prototyping && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dstack-prototyping" agent skill from https://github.com/dstackai/dstack/tree/master/skills/dstack-prototyping into .claude/skills/dstack-prototyping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dstack-prototyping", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/dstackai/dstack/tree/master/skills/dstack-prototypingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add dstackai/dstack --skill dstack-prototyping -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install dstackai/dstack dstack-prototyping --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dstackai/dstack.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/dstack-prototyping .agents/skills/dstack-prototyping && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dstack-prototyping" agent skill from https://github.com/dstackai/dstack/tree/master/skills/dstack-prototyping into .agents/skills/dstack-prototyping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dstack-prototyping", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dstackai/dstack --skill dstack-prototyping -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install dstackai/dstack dstack-prototyping --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dstackai/dstack.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/dstack-prototyping .cursor/skills/dstack-prototyping && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dstack-prototyping" agent skill from https://github.com/dstackai/dstack/tree/master/skills/dstack-prototyping into .cursor/skills/dstack-prototyping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dstack-prototyping", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/dstackai/dstack.git --path skills/dstack-prototyping--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add dstackai/dstack --skill dstack-prototyping -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install dstackai/dstack dstack-prototyping --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dstackai/dstack.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/dstack-prototyping .gemini/skills/dstack-prototyping && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dstack-prototyping" agent skill from https://github.com/dstackai/dstack/tree/master/skills/dstack-prototyping into .gemini/skills/dstack-prototyping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dstack-prototyping", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install dstackai/dstack dstack-prototypingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add dstackai/dstack --skill dstack-prototyping -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/dstackai/dstack.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/dstack-prototyping .github/skills/dstack-prototyping && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dstack-prototyping" agent skill from https://github.com/dstackai/dstack/tree/master/skills/dstack-prototyping into .github/skills/dstack-prototyping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dstack-prototyping", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dstackai/dstack --skill dstack-prototyping -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install dstackai/dstack dstack-prototyping --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dstackai/dstack.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/dstack-prototyping .opencode/skills/dstack-prototyping && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dstack-prototyping" agent skill from https://github.com/dstackai/dstack/tree/master/skills/dstack-prototyping into .opencode/skills/dstack-prototyping/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dstack-prototyping", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dstack-prototypingUse with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
Dstack Prototyping is an agent skill from dstackai/dstack. Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. Guides task-first prototyping on real hardware, choosing fleets/backends that can reuse idle instances and caches, checking vLLM/SGLang sources, and verifying the final dstack service with a model request.
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Prototyping, LLM inference and serving and Container orchestration. It works with SGLang, vLLM, Kubernetes and NVIDIA AI Platform. The repository describes itself as: A unified orchestration layer for heterogeneous AI compute. It standardizes how to manage compute and run training and inference on GPU clouds, Kubernetes, VMs, or bare-metal… The licence is MPL-2.0.
Read from SKILL.md and the folder at commit 0d578c8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and bash).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
dstack.airecipes.vllm.aidocs.sglang.iogithub.comlmsys.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Dstack Prototyping loads about 1.6k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 829 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from dstackai/dstack at commit 0d578c8, republished under its MPL-2.0 licence (© dstackai). 829 words, ~1,608 tokens.
.claude/skills/dstack-prototyping/SKILL.md (or your agent's skills folder).Use /dstack for CLI commands, YAML fields, apply/attach behavior, service URLs,
and other dstack syntax. This skill explains how to use dstack runs while the
model-serving configuration is still unknown.
Find a working dstack service configuration for the requested model.
Before submitting a service, use a task on real hardware to test the serving image, install/runtime assumptions, model download, cache path, command, port, launch flags, resources, env vars, backend/fleet choice, and local model request. Then submit the same configuration as a service and verify the model through the dstack service URL.
Pick the offer whose hardware best fits the goal at hand. Only when several offers fit comparably, choose a VM-based backend, an SSH fleet, or a Kubernetes fleet: they support idle instances and/or instance volumes, so later runs reuse the provisioned/idle instance or instance volumes for caching model weights (and possibly other writes), while container-based backends start clean on every run.
Fetch https://dstack.ai/docs/concepts/backends.md and classify backends
from the fetched document, not from memory.
If the intention is to use PD disaggregation, the fleet must use placement: cluster. Since PD disaggregation implies running a router, unlike workers that must run on GPUs, the router normally should run on a CPU instance. Use dstack fleet to see existing fleets and dstack fleet get <fleet name> --json to inspect a specific fleet.
Check serving-framework sources early enough to choose the image, command, launch flags, resources, cache paths, request format, and expected model behavior.
For vLLM and SGLang, use these as credible sources:
https://recipes.vllm.ai/ and
https://recipes.vllm.ai/models.jsonhttps://docs.sglang.io/ (fetch /llms.txt for the page
index)https://docs.sglang.io/cookbook/autoregressive/introhttps://github.com/vllm-project/vllm/releases and
https://github.com/sgl-project/sglang/releaseshttps://www.lmsys.org/blog/2026-07-02-agent-assisted-sglang-developmentBefore submitting a service, start a long-lived task:
commands:
- sleep infinityor an equivalent idle command.
Submit the task detached, attach or SSH into it when available, and run commands inside the live environment. Test the image, installs, model download and cache path, serving command, port, launch flags, local model request, and expected model behavior.
When starting a long-running command in the background from a non-interactive
SSH command, use nohup, redirect stdin from /dev/null, and redirect
stdout/stderr to a log file so the SSH command returns while the process keeps
running. For example (the command can be any long-running command):
nohup vllm serve ... </dev/null > /tmp/vllm.log 2>&1 &If the image, hardware choice, or major install path changes, submit another task so the changed setup is tested before service verification.
Do not move to a service after checking only GPU visibility, imports, logs, or a health endpoint. Start the server inside the task and send a request that uses the requested model. For a chat or reasoning model, check the response behavior the endpoint is expected to support, such as reasoning output when that model is supposed to expose it.
Follow /dstack structured status guidance when polling task or service status.
After requesting a task or service stop before another submission, wait until
that run reaches a terminal status. This allows dstack to reuse its instance or
instance volumes when available.
Submit the service after the task has verified the configuration: image, command, port, resources, env vars, cache mounts if used, backend/fleet choice, and model request.
Use the service as a duplicate check of the same configuration under dstack service runtime. The model request that worked locally in the task must also work through the dstack service URL.
If service verification fails because the image, install, model download, command, resources, cache, or model behavior needs to change, go back to a task. If the tested serving setup is still right and only the dstack service configuration is wrong, fix the configuration and submit the service again.
If a fleet has placement: cluster and a CPU-only instance, you must use a
configuration with the router on the CPU-only instance, regardless of whether
the workers are aggregated or PD disaggregated.
Whenever possible, connect the workers over gRPC, not HTTP: with a gRPC
router, request parsing, serialization, and tokenization move from the
serving engine to the router, so latency improves just by introducing it.
When using a router:
sleep infinity even when using groups (set it in
each group's commands; top-level commands is not allowed with groups),
and run the actual commands on each node interactively over SSH.https://dstack.ai/docs/concepts/tasks.md
and "Router" in https://dstack.ai/docs/concepts/services.md.If the intention is to use PD disaggregation:
## Router: the router and the prefill/decode workers run as
separate groups, and the fleet needs an interconnect (placement: cluster).https://dstack.ai/docs/concepts/tasks.md and "Replica groups" and
"PD disaggregation" in https://dstack.ai/docs/concepts/services.md.© dstackai, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/dstack-prototyping of dstackai/dstack.
Open the folder on GitHubat commit 0d578c8
Dstack Prototyping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dstack Prototyping this skilldstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Hyperloom SetupAMD-AGI/Hyperloom | 216 | — | ~7.2k | Automated safety check: Notes | Custom licence | |
| Tao Launch WorkflowNVIDIA/skills | 3.5k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| One EvalOpenDCAI/One-Eval | 165 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.
AMD-AGI/Hyperloom
Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.
NVIDIA/skills
The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
OpenDCAI/One-Eval
驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
dstackai/dstack
Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.
dstackai/dstack
dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.
Categories
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. Dstack Prototyping is an agent skill from dstackai/dstack. Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
Dstack Prototyping fits situations like: tasks that involve Prototyping; tasks that involve LLM inference and serving; tasks that involve Container orchestration.
Run `npx skills add dstackai/dstack --skill dstack-prototyping -a claude-code`. Or copy the skill folder (skills/dstack-prototyping in dstackai/dstack) into .claude/skills/dstack-prototyping in your project. Claude Code loads it when a task matches its description.
Run `npx skills add dstackai/dstack --skill dstack-prototyping -a codex`. Or copy the skill folder (skills/dstack-prototyping in dstackai/dstack) into .agents/skills/dstack-prototyping in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dstackai/dstack --skill dstack-prototyping -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dstack-prototyping, .gemini/skills/dstack-prototyping, .github/skills/dstack-prototyping and .opencode/skills/dstack-prototyping in your project.
SKILL.md names no scripts, command-line tools or credentials: Dstack Prototyping is instructions for the agent only.
SKILL.md names 5 domains. In commands or code: dstack.ai, recipes.vllm.ai, docs.sglang.io, github.com and lmsys.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Dstack Prototyping is published under the MPL-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Dstack Prototyping: SGLang Structured Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hyperloom Setup (AMD-AGI/Hyperloom, 216 stars), Tao Launch Workflow (NVIDIA/skills, 3.5k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
dstackai (a GitHub organization) maintains it in dstackai/dstack, which has 2,272 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.
Source: dstackai/dstack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.