Agent skill

Dstack Prototyping

by dstackai in dstackai/dstack

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

MPL-2.0Auto-check passedAI & LLM Engineering

Install Dstack Prototyping

skills CLI
$ npx skills add dstackai/dstack --skill dstack-prototyping -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dstackai/dstack dstack-prototyping --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dstackai/dstack.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dstack-prototyping .claude/skills/dstack-prototyping && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dstack-prototyping
GitHub stars
2.3k
Token cost
~1.6k tokens
SKILL.md length
829 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MPL-2.0

At a glance

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

  • Tasks that involve Prototyping
  • SKILL.md covers Goal, Choose Where To Run, Check Serving Sources and Use A Task Before Service, plus 3 more sections
  • Reaches dstack.ai and recipes.vllm.ai
  • Tasks that involve LLM inference and serving

What it does

Dstack Prototyping is an agent skill from dstackai/dstack. Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. Guides task-first prototyping on real hardware, choosing fleets/backends that can reuse idle instances and caches, checking vLLM/SGLang sources, and verifying the final dstack service with a model request.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Prototyping, LLM inference and serving and Container orchestration. It works with SGLang, vLLM, Kubernetes and NVIDIA AI Platform. The repository describes itself as: A unified orchestration layer for heterogeneous AI compute. It standardizes how to manage compute and run training and inference on GPU clouds, Kubernetes, VMs, or bare-metal… The licence is MPL-2.0.

When your agent uses it

  • Tasks that involve Prototyping
  • Tasks that involve LLM inference and serving
  • Tasks that involve Container orchestration

Example prompts

  • “/dstack-prototyping”

What it can do on your machine

Read from SKILL.md and the folder at commit 0d578c8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • dstack.ai
    • recipes.vllm.ai
    • docs.sglang.io
    • github.com
    • lmsys.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dstack Prototyping loads about 1.6k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 829 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dstackai/dstack at commit 0d578c8, republished under its MPL-2.0 licence (© dstackai). 829 words, ~1,608 tokens.

Download SKILL.mdSave it as .claude/skills/dstack-prototyping/SKILL.md (or your agent's skills folder).
name
dstack-prototyping
description
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. Guides task-first prototyping on real hardware, choosing fleets/backends that can reuse idle instances and caches, checking vLLM/SGLang sources, and verifying the final dstack service with a model request.

dstack Prototyping

Use /dstack for CLI commands, YAML fields, apply/attach behavior, service URLs, and other dstack syntax. This skill explains how to use dstack runs while the model-serving configuration is still unknown.

Goal

Find a working dstack service configuration for the requested model.

Before submitting a service, use a task on real hardware to test the serving image, install/runtime assumptions, model download, cache path, command, port, launch flags, resources, env vars, backend/fleet choice, and local model request. Then submit the same configuration as a service and verify the model through the dstack service URL.

Choose Where To Run

Pick the offer whose hardware best fits the goal at hand. Only when several offers fit comparably, choose a VM-based backend, an SSH fleet, or a Kubernetes fleet: they support idle instances and/or instance volumes, so later runs reuse the provisioned/idle instance or instance volumes for caching model weights (and possibly other writes), while container-based backends start clean on every run.

Fetch https://dstack.ai/docs/concepts/backends.md and classify backends from the fetched document, not from memory.

If the intention is to use PD disaggregation, the fleet must use placement: cluster. Since PD disaggregation implies running a router, unlike workers that must run on GPUs, the router normally should run on a CPU instance. Use dstack fleet to see existing fleets and dstack fleet get <fleet name> --json to inspect a specific fleet.

Check Serving Sources

Check serving-framework sources early enough to choose the image, command, launch flags, resources, cache paths, request format, and expected model behavior.

For vLLM and SGLang, use these as credible sources:

  • vLLM recipes and model index: https://recipes.vllm.ai/ and https://recipes.vllm.ai/models.json
  • SGLang docs: https://docs.sglang.io/ (fetch /llms.txt for the page index)
  • SGLang model recipes: https://docs.sglang.io/cookbook/autoregressive/intro
  • Release notes: https://github.com/vllm-project/vllm/releases and https://github.com/sgl-project/sglang/releases
  • Performance-loop methodology (profiling, benchmark contracts): https://www.lmsys.org/blog/2026-07-02-agent-assisted-sglang-development

Use A Task Before Service

Before submitting a service, start a long-lived task:

yaml
commands:
  - sleep infinity

or an equivalent idle command.

Submit the task detached, attach or SSH into it when available, and run commands inside the live environment. Test the image, installs, model download and cache path, serving command, port, launch flags, local model request, and expected model behavior.

When starting a long-running command in the background from a non-interactive SSH command, use nohup, redirect stdin from /dev/null, and redirect stdout/stderr to a log file so the SSH command returns while the process keeps running. For example (the command can be any long-running command):

shell
nohup vllm serve ... </dev/null > /tmp/vllm.log 2>&1 &

If the image, hardware choice, or major install path changes, submit another task so the changed setup is tested before service verification.

Do not move to a service after checking only GPU visibility, imports, logs, or a health endpoint. Start the server inside the task and send a request that uses the requested model. For a chat or reasoning model, check the response behavior the endpoint is expected to support, such as reasoning output when that model is supposed to expose it.

Follow /dstack structured status guidance when polling task or service status. After requesting a task or service stop before another submission, wait until that run reaches a terminal status. This allows dstack to reuse its instance or instance volumes when available.

Show full SKILL.md (312 more words)Show less

Verify As A Service

Submit the service after the task has verified the configuration: image, command, port, resources, env vars, cache mounts if used, backend/fleet choice, and model request.

Use the service as a duplicate check of the same configuration under dstack service runtime. The model request that worked locally in the task must also work through the dstack service URL.

If service verification fails because the image, install, model download, command, resources, cache, or model behavior needs to change, go back to a task. If the tested serving setup is still right and only the dstack service configuration is wrong, fix the configuration and submit the service again.

Router

If a fleet has placement: cluster and a CPU-only instance, you must use a configuration with the router on the CPU-only instance, regardless of whether the workers are aggregated or PD disaggregated. Whenever possible, connect the workers over gRPC, not HTTP: with a gRPC router, request parsing, serialization, and tokenization move from the serving engine to the router, so latency improves just by introducing it.

When using a router:

  • Use node groups for the task and replica groups for the service: tasks' node groups are the equivalent of services' replica groups.
  • With tasks, still use sleep infinity even when using groups (set it in each group's commands; top-level commands is not allowed with groups), and run the actual commands on each node interactively over SSH.
  • When testing inference, call the router endpoint, not the workers directly (unless you want to test if they are alive).
  • Look for "Prototyping services" in https://dstack.ai/docs/concepts/tasks.md and "Router" in https://dstack.ai/docs/concepts/services.md.

PD disaggregation

If the intention is to use PD disaggregation:

  • Follow ## Router: the router and the prefill/decode workers run as separate groups, and the fleet needs an interconnect (placement: cluster).
  • Look for "Node groups" and "PD disaggregation" in https://dstack.ai/docs/concepts/tasks.md and "Replica groups" and "PD disaggregation" in https://dstack.ai/docs/concepts/services.md.

© dstackai, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/dstack-prototyping of dstackai/dstack.

Open the folder on GitHubat commit 0d578c8

Compare with similar skills

Dstack Prototyping next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dstack Prototyping compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dstack Prototyping this skilldstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.9kAutomated safety check: PassMIT
Hyperloom SetupAMD-AGI/Hyperloom216—~7.2kAutomated safety check: NotesCustom licence
Tao Launch WorkflowNVIDIA/skills3.5k—~4.5kAutomated safety check: NotesApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
One EvalOpenDCAI/One-Eval165—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 3 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Hyperloom Setup

    AMD-AGI/Hyperloom

    Configures Hyperloom after pip install --target . An agent skill from AMD-AGI/Hyperloom.

    216 GitHub stars~7.2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Tao Launch Workflow

    NVIDIA/skills

    Official

    The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action.

    3.5k GitHub stars~4.5k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • One Eval

    OpenDCAI/One-Eval

    驱动 One-Eval 对 API 或本地模型做端到端评测,覆盖纯文本、多模态、代码生成、函数调用和 Agent benchmark。当用户想评测模型在一个或多个 benchmark 上的表现、比较分数、补充 metric,或生成图文评测报告时使用本 skill。

    165 GitHub stars~2.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    900 GitHub stars~2.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from dstackai/dstack

  • Dstack Presets

    dstackai/dstack

    Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

    2.3k GitHub stars~403 tokensUpdated today
    Auto-check passed
  • Dstack

    dstackai/dstack

    dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.

    2.3k GitHub stars~6.2k tokensUpdated today
    Auto-check: warnings

Questions about Dstack Prototyping

What does Dstack Prototyping do?

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. Dstack Prototyping is an agent skill from dstackai/dstack. Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

When should I use Dstack Prototyping?

Dstack Prototyping fits situations like: tasks that involve Prototyping; tasks that involve LLM inference and serving; tasks that involve Container orchestration.

How do I install Dstack Prototyping in Claude Code?

Run `npx skills add dstackai/dstack --skill dstack-prototyping -a claude-code`. Or copy the skill folder (skills/dstack-prototyping in dstackai/dstack) into .claude/skills/dstack-prototyping in your project. Claude Code loads it when a task matches its description.

How do I install Dstack Prototyping in Codex?

Run `npx skills add dstackai/dstack --skill dstack-prototyping -a codex`. Or copy the skill folder (skills/dstack-prototyping in dstackai/dstack) into .agents/skills/dstack-prototyping in your project. Codex loads it when a task matches its description.

Can I use Dstack Prototyping in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dstackai/dstack --skill dstack-prototyping -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dstack-prototyping, .gemini/skills/dstack-prototyping, .github/skills/dstack-prototyping and .opencode/skills/dstack-prototyping in your project.

What does Dstack Prototyping need to run?

SKILL.md names no scripts, command-line tools or credentials: Dstack Prototyping is instructions for the agent only.

Does Dstack Prototyping access the network?

SKILL.md names 5 domains. In commands or code: dstack.ai, recipes.vllm.ai, docs.sglang.io, github.com and lmsys.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Dstack Prototyping safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dstack Prototyping use?

Dstack Prototyping is published under the MPL-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dstack Prototyping use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dstack Prototyping?

Skills that share tags, products or a category with Dstack Prototyping: SGLang Structured Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hyperloom Setup (AMD-AGI/Hyperloom, 216 stars), Tao Launch Workflow (NVIDIA/skills, 3.5k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dstack Prototyping?

dstackai (a GitHub organization) maintains it in dstackai/dstack, which has 2,272 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.

Source: dstackai/dstack on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.