vLLM Model Serving
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor).
$ npx skills add pproenca/dot-skills --skill ray-llm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install pproenca/dot-skills ray-llm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/ray-llm .claude/skills/ray-llm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ray-llm" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/ray-llm into .claude/skills/ray-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-llm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/ray-llmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add pproenca/dot-skills --skill ray-llm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install pproenca/dot-skills ray-llm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.experimental/ray-llm .agents/skills/ray-llm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ray-llm" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/ray-llm into .agents/skills/ray-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-llm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pproenca/dot-skills --skill ray-llm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install pproenca/dot-skills ray-llm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.experimental/ray-llm .cursor/skills/ray-llm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ray-llm" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/ray-llm into .cursor/skills/ray-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-llm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/pproenca/dot-skills.git --path skills/.experimental/ray-llm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add pproenca/dot-skills --skill ray-llm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install pproenca/dot-skills ray-llm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.experimental/ray-llm .gemini/skills/ray-llm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ray-llm" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/ray-llm into .gemini/skills/ray-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-llm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install pproenca/dot-skills ray-llmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add pproenca/dot-skills --skill ray-llm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.experimental/ray-llm .github/skills/ray-llm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ray-llm" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/ray-llm into .github/skills/ray-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-llm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pproenca/dot-skills --skill ray-llm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install pproenca/dot-skills ray-llm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.experimental/ray-llm .opencode/skills/ray-llm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ray-llm" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/ray-llm into .opencode/skills/ray-llm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ray-llm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ray-llmLLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor).
Ray LLM is an agent skill from pproenca/dot-skills. LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor). Corrects the stale defaults a model produces — the archived ray-llm repo and its YAML configs, hand-rolled vLLM engines inside plain Serve deployments, the removed buildllmprocessor name, deprecated boolean stage flags, top-level LLMServer/LLMRouter imports, free-form accelerator strings, one-deployment-per-LoRA-adapter…
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 20 other files, including reference files and assets (for example `AGENTS.md`, `assets/templates/_template.md` and `metadata.json`).
It sits in AI & LLM Engineering, covering Deployment, LLM inference and serving and Fine-tuning. It works with OpenAI and vLLM. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ray LLM loads about 1.7k tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 228 tokens; SKILL.md has 573 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 573 words, ~1,705 tokens.
.claude/skills/ray-llm/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.Library-reference skill for LLM workloads on open-source Ray — 13 rules across 5 categories covering ray.serve.llm (OpenAI-compatible, vLLM-backed serving) and ray.data.llm (batch inference). This surface churned faster than any other part of Ray — a standalone repo was absorbed and archived, entry points were renamed, and config shapes restructured — so the examples a model learned from mostly no longer run. Each rule names the wrong default it corrects; there is no rule for things a capable model already gets right.
Scope is the LLM-specific layer. Generic Serve/Data/cluster decisions (deployment lifecycle, autoscaling semantics, KubeRay) are the sibling ray skill — the two compose.
Pinned to ray 2.57.0 (ray[llm] extra, which pins its matching vLLM). API claims were verified against the unpacked 2.57.0 wheel and the installed package source, and every config example in the rules was constructed under CPU-only pydantic validation (including the traps, which fail exactly as described); engine/GPU runtime behavior is source-verified only — no model was actually served.
ray.data.llm names| # | Category | Prefix | Covers |
|---|---|---|---|
| 1 | Serving Setup | serve- | LLMConfig + build_openai_app over hand-rolled engines and the archived repo; model_id vs model_source; the ray[llm]↔vLLM version pin; relocated LLMServer/OpenAiIngress imports |
| 2 | Batch Inference | batch- | build_processor (old name removed), stage configs over boolean flags, CPU-default accelerator_type and autoscaling concurrency, HTTP/Serve processor alternatives |
| 3 | Placement & Accelerators | place- | The engine's own TP×PP placement group (and when to override its strategy), validated accelerator_type names |
| 4 | Autoscaling & Routing | scale- | deployment_config.autoscaling_config with engine-sized replicas, ingress replica sizing, prefix-cache-affinity routing |
| 5 | LoRA & API Surface | api- | Dynamic LoRA multiplexing over per-adapter deployments; the full OpenAI endpoint surface; GPU-free config validation |
serve-builtin-not-handrolled — LLMConfig + build_openai_app; the ray-llm repo is archived, hand-rolled engines re-implement lessserve-model-id-vs-source — model_id is the client-facing name; model_source is where weights liveserve-ray-llm-extra-pins-vllm — ray[llm] pins its exact vLLM; don't mix independently chosen versionsserve-deprecated-server-router — LLMServer/LLMRouter moved; the ingress class is OpenAiIngress (exact casing)batch-build-processor-renamed — build_llm_processor is removed; build_processor + vLLMEngineProcessorConfigbatch-stage-configs-not-flags — boolean apply_chat_template/tokenize/detokenize gave way to stage configsbatch-concurrency-and-alternatives — accelerator_type=None means CPU; (min, max) concurrency; HTTP/Serve processorsplace-engine-builds-pg — the engine builds its TP×PP placement group; override strategy, don't hand-rollplace-accelerator-type-validated — canonical accelerator constants; "A10" normalizes, most typos raisescale-autoscaling-deployment-config — standard Serve autoscaling nested in deployment_config; replicas are whole enginesscale-prefix-cache-routing — PrefixCacheAffinityRouter keeps same-prefix requests on warm KV cachesapi-lora-dynamic-multiplexing — dynamic adapter loading from cloud storage, not one deployment per adapterapi-endpoint-surface-cpu-validation — embeddings/transcription/score/tokenize are served too; configs validate without GPUsRead a reference file when its decision comes up. Each rule names the wrong default it corrects, then shows the canonical way (with an incorrect/correct contrast only where the wrong way is a real trap).
ray — the sibling rule pack for classic-ML Ray (Train, Tune, Data, Serve, Core, KubeRay production topology); generic Serve and cluster decisions live there| File | Description |
|---|---|
| references/_sections.md | Category definitions and ordering |
| assets/templates/_template.md | Template for new rules |
| metadata.json | Version and source references |
© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 17 other files (references, assets) in skills/.experimental/ray-llm of pproenca/dot-skills.
Open the folder on GitHubat commit cf93c57
Ray LLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ray LLM this skillpproenca/dot-skills | 214 | — | ~1.7k | Automated safety check: Pass | MIT | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Aqua Deploymentoracle/accelerated-data-science | 125 | — | ~2.4k | Automated safety check: Pass | UPL-1.0 | |
| Vllm Deploy K8svllm-project/vllm-skills | 103 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Rtvi Vlm Customize ModelNVIDIA/skills | 3.5k | — | ~5k | Automated safety check: Notes | Apache-2.0 | |
| LLM App Builderrevfactory/harness-100 | 1.3k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
oracle/accelerated-data-science
Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.
vllm-project/vllm-skills
Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.
NVIDIA/skills
How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.
revfactory/harness-100
Full pipeline where an agent team collaborates to develop an LLM app.
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
pproenca/dot-skills
Audio forensics and voice recovery guidelines for CSI-level audio analysis.
pproenca/dot-skills
Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.
pproenca/dot-skills
Create well-structured RFCs and technical proposals for software projects.
pproenca/dot-skills
Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.
pproenca/dot-skills
Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…
pproenca/dot-skills
Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…
Categories
LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor). Ray LLM is an agent skill from pproenca/dot-skills.llm (buildprocessor).
Ray LLM fits situations like: productionizing LLM serving; batch inference on Ray.
Run `npx skills add pproenca/dot-skills --skill ray-llm -a claude-code`. Or copy the skill folder (skills/.experimental/ray-llm in pproenca/dot-skills) into .claude/skills/ray-llm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add pproenca/dot-skills --skill ray-llm -a codex`. Or copy the skill folder (skills/.experimental/ray-llm in pproenca/dot-skills) into .agents/skills/ray-llm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill ray-llm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ray-llm, .gemini/skills/ray-llm, .github/skills/ray-llm and .opencode/skills/ray-llm in your project.
SKILL.md names no scripts, command-line tools or credentials: Ray LLM is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ray LLM is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ray LLM: vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Aqua Deployment (oracle/accelerated-data-science, 125 stars), Vllm Deploy K8s (vllm-project/vllm-skills, 103 stars) and Rtvi Vlm Customize Model (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 214 GitHub stars. The repository holds 182 skills in this directory. The repository was last updated on August 15, 2026.
Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.