Agent skill

Ray LLM

by pproenca in pproenca/dot-skills

LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor).

MITAuto-check passedAI & LLM Engineering

Install Ray LLM

skills CLI
$ npx skills add pproenca/dot-skills --skill ray-llm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pproenca/dot-skills ray-llm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/ray-llm .claude/skills/ray-llm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ray-llm
GitHub stars
214
Token cost
~1.7k tokens
SKILL.md length
573 words
Files
18 (incl. references, assets)
Skills in repo
182
Repo updated
First seen
Licence
MIT

At a glance

LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor).

  • Works in 5 steps: Serving Setup → Batch Inference → Placement & Accelerators → …
  • Productionizing LLM serving
  • SKILL.md covers When to Apply, Rule Categories, Quick Reference and How to Use, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Ray LLM is an agent skill from pproenca/dot-skills. LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor). Corrects the stale defaults a model produces — the archived ray-llm repo and its YAML configs, hand-rolled vLLM engines inside plain Serve deployments, the removed buildllmprocessor name, deprecated boolean stage flags, top-level LLMServer/LLMRouter imports, free-form accelerator strings, one-deployment-per-LoRA-adapter…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 20 other files, including reference files and assets (for example `AGENTS.md`, `assets/templates/_template.md` and `metadata.json`).

It sits in AI & LLM Engineering, covering Deployment, LLM inference and serving and Fine-tuning. It works with OpenAI and vLLM. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.

When your agent uses it

  • Productionizing LLM serving
  • Batch inference on Ray

Example prompts

  • “/ray-llm”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Serving Setup
  2. Batch Inference
  3. Placement & Accelerators
  4. Autoscaling & Routing
  5. LoRA & API Surface

What it can do on your machine

Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ray LLM loads about 1.7k tokens when it runs, and up to ~7k if it reads all its reference files. Until then it costs about 228 tokens; SKILL.md has 573 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~228
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 573 words, ~1,705 tokens.

Download SKILL.mdSave it as .claude/skills/ray-llm/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
ray-llm
description
LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + build_openai_app) and batch inference with ray.data.llm (build_processor). Corrects the stale defaults a model produces — the archived ray-llm repo and its YAML configs, hand-rolled vLLM engines inside plain Serve deployments, the removed build_llm_processor name, deprecated boolean stage flags, top-level LLMServer/LLMRouter imports, free-form accelerator strings, one-deployment-per-LoRA-adapter designs — with the 2.57 idioms (stage configs, placement_group_config over hand-rolled PGs, deployment_config autoscaling, dynamic LoRA multiplexing, prefix-cache-affinity routing, the full OpenAI endpoint surface). Use when writing, reviewing, or productionizing LLM serving or batch inference on Ray. Classic-ML Ray (Train/Tune/Data/Serve/clusters) lives in the sibling ray skill.

Ray LLM

Library-reference skill for LLM workloads on open-source Ray — 13 rules across 5 categories covering ray.serve.llm (OpenAI-compatible, vLLM-backed serving) and ray.data.llm (batch inference). This surface churned faster than any other part of Ray — a standalone repo was absorbed and archived, entry points were renamed, and config shapes restructured — so the examples a model learned from mostly no longer run. Each rule names the wrong default it corrects; there is no rule for things a capable model already gets right.

Scope is the LLM-specific layer. Generic Serve/Data/cluster decisions (deployment lifecycle, autoscaling semantics, KubeRay) are the sibling ray skill — the two compose.

Pinned to ray 2.57.0 (ray[llm] extra, which pins its matching vLLM). API claims were verified against the unpacked 2.57.0 wheel and the installed package source, and every config example in the rules was constructed under CPU-only pydantic validation (including the traps, which fail exactly as described); engine/GPU runtime behavior is source-verified only — no model was actually served.

When to Apply

  • Standing up or reviewing an OpenAI-compatible LLM serving deployment on Ray
  • Writing batch LLM inference over datasets — summarization, embedding, scoring at scale
  • Sizing or placing multi-GPU models — tensor/pipeline parallelism, accelerator selection
  • Scaling LLM deployments — replica autoscaling, ingress sizing, request routing
  • Serving families of LoRA fine-tunes of a shared base model
  • Migrating code that uses the archived ray-llm repo, hand-rolled vLLM engines, or pre-2.5x ray.data.llm names

Rule Categories

#CategoryPrefixCovers
1Serving Setupserve-LLMConfig + build_openai_app over hand-rolled engines and the archived repo; model_id vs model_source; the ray[llm]↔vLLM version pin; relocated LLMServer/OpenAiIngress imports
2Batch Inferencebatch-build_processor (old name removed), stage configs over boolean flags, CPU-default accelerator_type and autoscaling concurrency, HTTP/Serve processor alternatives
3Placement & Acceleratorsplace-The engine's own TP×PP placement group (and when to override its strategy), validated accelerator_type names
4Autoscaling & Routingscale-deployment_config.autoscaling_config with engine-sized replicas, ingress replica sizing, prefix-cache-affinity routing
5LoRA & API Surfaceapi-Dynamic LoRA multiplexing over per-adapter deployments; the full OpenAI endpoint surface; GPU-free config validation

Quick Reference

Show full SKILL.md (245 more words)Show less
1. Serving Setup
2. Batch Inference
3. Placement & Accelerators
4. Autoscaling & Routing
5. LoRA & API Surface

How to Use

Read a reference file when its decision comes up. Each rule names the wrong default it corrects, then shows the canonical way (with an incorrect/correct contrast only where the wrong way is a real trap).

  • ray — the sibling rule pack for classic-ML Ray (Train, Tune, Data, Serve, Core, KubeRay production topology); generic Serve and cluster decisions live there

Reference Files

FileDescription
references/_sections.mdCategory definitions and ordering
assets/templates/_template.mdTemplate for new rules
metadata.jsonVersion and source references

© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (references, assets) in skills/.experimental/ray-llm of pproenca/dot-skills.

  • SKILL.md
  • AGENTS.md
  • assets/templates/_template.md
  • metadata.json
  • references/_sections.md
  • references/api-endpoint-surface-cpu-validation.md
  • references/api-lora-dynamic-multiplexing.md
  • references/batch-build-processor-renamed.md
  • references/batch-concurrency-and-alternatives.md
  • references/batch-stage-configs-not-flags.md
  • references/place-accelerator-type-validated.md
  • references/place-engine-builds-pg.md
  • references/scale-autoscaling-deployment-config.md
  • references/scale-prefix-cache-routing.md
  • references/serve-builtin-not-handrolled.md
  • references/serve-deprecated-server-router.md
  • references/serve-model-id-vs-source.md
  • references/serve-ray-llm-extra-pins-vllm.md

Open the folder on GitHubat commit cf93c57

Compare with similar skills

Ray LLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ray LLM compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ray LLM this skillpproenca/dot-skills214—~1.7kAutomated safety check: PassMIT
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT
Aqua Deploymentoracle/accelerated-data-science125—~2.4kAutomated safety check: PassUPL-1.0
Vllm Deploy K8svllm-project/vllm-skills103—~2kAutomated safety check: PassApache-2.0
Rtvi Vlm Customize ModelNVIDIA/skills3.5k—~5kAutomated safety check: NotesApache-2.0
LLM App Builderrevfactory/harness-1001.3k—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Aqua Deployment

    oracle/accelerated-data-science

    Official

    Deploy LLM models on OCI using AI Quick Actions (AQUA) - single model, multi-model, stacked (LoRA), with GPU shape selection, vLLM configuration, streaming, and tool calling.

    125 GitHub stars~2.4k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    103 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Backend & APIsAuto-check: notes
  • LLM App Builder

    revfactory/harness-100

    Full pipeline where an agent team collaborates to develop an LLM app.

    1.3k GitHub stars~1.9k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed

More from pproenca/dot-skills

All 182 skills in this repo
  • Audio Voice Recovery

    pproenca/dot-skills

    Audio forensics and voice recovery guidelines for CSI-level audio analysis.

    214 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Codemod React Pipeline

    pproenca/dot-skills

    Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.

    214 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Dev Rfc

    pproenca/dot-skills

    Create well-structured RFCs and technical proposals for software projects.

    214 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Dx Harness

    pproenca/dot-skills

    Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.

    214 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Language Spec Author

    pproenca/dot-skills

    Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…

    214 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Python Pep Author

    pproenca/dot-skills

    Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…

    214 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Ray LLM

What does Ray LLM do?

LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor). Ray LLM is an agent skill from pproenca/dot-skills.llm (buildprocessor).

When should I use Ray LLM?

Ray LLM fits situations like: productionizing LLM serving; batch inference on Ray.

How do I install Ray LLM in Claude Code?

Run `npx skills add pproenca/dot-skills --skill ray-llm -a claude-code`. Or copy the skill folder (skills/.experimental/ray-llm in pproenca/dot-skills) into .claude/skills/ray-llm in your project. Claude Code loads it when a task matches its description.

How do I install Ray LLM in Codex?

Run `npx skills add pproenca/dot-skills --skill ray-llm -a codex`. Or copy the skill folder (skills/.experimental/ray-llm in pproenca/dot-skills) into .agents/skills/ray-llm in your project. Codex loads it when a task matches its description.

Can I use Ray LLM in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill ray-llm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ray-llm, .gemini/skills/ray-llm, .github/skills/ray-llm and .opencode/skills/ray-llm in your project.

What does Ray LLM need to run?

SKILL.md names no scripts, command-line tools or credentials: Ray LLM is instructions for the agent only.

Does Ray LLM access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ray LLM safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ray LLM use?

Ray LLM is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ray LLM use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.

What are the alternatives to Ray LLM?

Skills that share tags, products or a category with Ray LLM: vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Aqua Deployment (oracle/accelerated-data-science, 125 stars), Vllm Deploy K8s (vllm-project/vllm-skills, 103 stars) and Rtvi Vlm Customize Model (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ray LLM?

pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 214 GitHub stars. The repository holds 182 skills in this directory. The repository was last updated on August 15, 2026.

Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.