Agent skill

Engine Vllm

by autonomous-ai in autonomous-ai/openharness

Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…

MITAuto-check passedAI & LLM Engineering

Install Engine Vllm

skills CLI
$ npx skills add autonomous-ai/openharness --skill engine-vllm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/openharness engine-vllm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/store/agents/autonomous-grid/skills/engine-vllm .claude/skills/engine-vllm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
engine-vllm
GitHub stars
1.1k
Token cost
~1.9k tokens
SKILL.md length
847 words
Files
1
Skills in repo
99
Repo updated
First seen
Licence
MIT

At a glance

Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…

  • Works in 2 steps: The recipe — "$GRID_FLEET" recipe vllm… → No recipe — read the model:…
  • Tasks that involve LLM inference and serving
  • SKILL.md covers When to use it, 1. The recipe — "$GRID_FLEET"…, 2. No recipe — read the model:… and Install and version, plus 5 more sections
  • Calls uv and docker; reaches raw.githubusercontent.com and wheels.vllm.ai

What it does

Engine Vllm is an agent skill from autonomous-ai/openharness. Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it to the person's fleet. Load before installing, configuring, starting or stopping vLLM.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving and Model hubs and datasets. It works with vLLM, NVIDIA AI Platform, Linux and Hugging Face. The repository describes itself as: The ultimate harness for coding agents and beyond. All your agents. All your machines. One command center. Start with code, then follow your curiosity and build across… The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve Model hubs and datasets

Example prompts

  • “s official vLLM recipe — or, when it has none, from the model”
  • “/engine-vllm”

Requirements

  • Python 3
  • Docker

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. The recipe — "$GRID_FLEET" recipe vllm ORG/NAME
  2. No recipe — read the model: "$GRID_FLEET" model-facts ORG/NAME

What it can do on your machine

Read from SKILL.md and the folder at commit 54a1f1b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • raw.githubusercontent.com
    • wheels.vllm.ai

    Also links to:

    • docs.vllm.ai
    • recipes.vllm.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Engine Vllm loads about 1.9k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 847 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~74
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from autonomous-ai/openharness at commit 54a1f1b, republished under its MIT licence (© autonomous-ai). 847 words, ~1,931 tokens.

Download SKILL.mdSave it as .claude/skills/engine-vllm/SKILL.md (or your agent's skills folder).
name
engine-vllm
description
Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it to the person's fleet. Load before installing, configuring, starting or stopping vLLM.

vLLM (Linux + NVIDIA CUDA or AMD ROCm)

Official sources, read 2026-09-29. Agent index: recipes.vllm.ai/llms.txt. Reference docs as raw markdown under https://raw.githubusercontent.com/vllm-project/vllm/main/docs/: OpenAI-compatible server (serving/online_serving/openai_compatible_server.md) · Tool calling (features/tool_calling.md) · Reasoning outputs (features/reasoning_outputs.md) · Supported models (models/supported_models.md) · GPU install · Docker · Optimization · Conserving memory Tested: only vLLM 0.11.0 on an Apple-silicon Mac (CPU backend), 2026-09-29 — no GPU run yet. Tags: [doc] official source above, [run] seen on the tested machine, [?] unverified.

When to use it

  • Linux with an NVIDIA GPU of compute capability 7.5 or newer [doc], or an AMD GPU with ROCm [doc], and a Hugging Face model (safetensors, FP8, AWQ…). vLLM is built for many requests at once and for models split across several GPUs.
  • A GGUF file for one person on one GPU: Grid's own engine serves it with nothing to install.
  • Not on a Mac. There it runs on the CPU only, slowly (26 tokens in 2.3 s for a 0.5B model), and needed pinned transformers<5, --dtype float32 for JSON output and a batched-token limit to start [run].

1. The recipe — "$GRID_FLEET" recipe vllm ORG/NAME

Reads recipes.vllm.ai live (models.json, then <ORG>/<NAME>.json) and prints:

FieldUse
minVersionthe installed vllm --version must be at least this
baseArgs, baseEnvalways
features.tool_calling.args, features.reasoning.argsalways for an agent
features.* with optIn: trueonly on purpose (faster decoding, text only to free memory, longer context)
variants.<name>pick by vramMinimumGb against the GPU; a variant can carry its own modelId and extraArgs
recommendedthe site's reference command and Docker image for its reference hardware
installthe recipe's own install steps, including any pinned extras
guideprose; its Troubleshooting holds model-specific fixes

The recipe beats the generic docs: a family's generic parser entry and a model's recipe can name different parsers [doc]. Exit code 3 means there is no recipe: go to step 2.

2. No recipe — read the model: "$GRID_FLEET" model-facts ORG/NAME

  1. support.vllm.listed must be true (its architecture is in vLLM's supported-models list). false: the Transformers backend (--model-impl transformers) may still run it [doc] — a test, not a promise.
  2. modelCard.serveCommands / toolCallParsers / reasoningParsers: the model authors' own vLLM command. Use it when present.
  3. Otherwise match chatTemplate.toolSyntax against the parser descriptions in features/tool_calling.md (each parser documents the output format it reads) and, when chatTemplate.thinking, pick the reasoning parser from the table in features/reasoning_outputs.md [doc]. Quote the doc line you matched.
  4. contextLength must be ≥ 65536. Sampling defaults come from the model's generation_config.json automatically [doc]; leave them.
  5. No match in the docs: say so and stop — do not cycle through parsers until one passes.

Install and version

  • Slow step, ask first. Venv: uv venv && uv pip install -U vllm --torch-backend=auto [doc]; pin with "vllm==X.Y.Z" at or above the recipe's minVersion. Wheels target CUDA 12.9 by default, builds for 12.8 and 13.0 exist, Blackwell needs CUDA ≥ 12.8 [doc]. AMD: --extra-index-url https://wheels.vllm.ai/rocm (Python 3.12, ROCm 7.0, glibc ≥ 2.35) [doc]. Check the GPU and driver with nvidia-smi.
  • Docker: vllm/vllm-openai (NVIDIA), vllm/vllm-openai-rocm (AMD) [doc], with --gpus all --ipc=host -p 127.0.0.1:P:8000 -v ~/.cache/huggingface:/root/.cache/huggingface [doc].
Show full SKILL.md (339 more words)Show less

Common configs (the recipe's flags win)

Each adds the recipe's (or step 2's) tool and reasoning flags, --host 127.0.0.1 --port P and a --max-model-len of at least 65536.

CaseAdd
One GPU, one agent--max-model-len 131072 --max-num-seqs 4 --enable-prefix-caching
One GPU, many people--max-model-len 65536 --enable-prefix-caching (max-num-seqs left to vLLM)
Several GPUs, one machine--tensor-parallel-size <GPU count> [doc]
Memory is tightthe recipe's FP8 variant, --kv-cache-dtype fp8, fewer --max-num-seqs [doc]

Start, ready, join

"$GRID_FLEET" serve vllm-P --env HF_HUB_OFFLINE=1 -- ~/.grid/envs/vllm/bin/vllm serve <model id or snapshot dir> \
  --served-model-name <id> <flags>

(Install into ~/.grid/envs/vllm so fleet models finds it; fleet serve keeps it alive after your shell returns; log run/vllm-P.log.)

  • Verify: "$GRID_FLEET" verify --at http://127.0.0.1:P/v1 --model <id> --kind vllm (bounded, narrated).
  • Ready: PID alive; log shows Application startup complete. [run] (minutes for a big model); GET /health → 200 and GET /v1/models lists <id> [run]; one bounded answer; one tool call.
  • Join: "$GRID_FLEET" run -- join GRID --at http://127.0.0.1:P/v1 -m <id> --advertise-as ALIAS. /v1 is required — without it /models and /chat/completions answer 404 [run]. --api-key guards only /v1 routes [doc].
  • Thinking off for everyday use: the recipe's guide names the flag when there is one (--default-chat-template-kwargs '{"enable_thinking": false}' in the recipes that document it [doc]).

Knobs

--gpu-memory-utilization (share of GPU memory pre-allocated for weights plus cache), --max-model-len, --max-num-seqs, --max-num-batched-tokens, --kv-cache-dtype fp8, --enforce-eager (fastest start, slower serving), --tensor-parallel-size [doc].

Stop

"$GRID_FLEET" stop vllm-P (or docker stop the container you started); confirm the port is free.

Known failures → what to do

SignDo
"preempted … not enough KV cache space"raise --gpu-memory-utilization, lower --max-num-seqs, or more tensor parallel [doc]
CUDA out of memory at startFP8 variant, --kv-cache-dtype fp8, fewer sequences — context stays ≥ 65536
"max_num_batched_tokens … smaller than max_model_len"raise --max-num-batched-tokens to at least the context [doc][run]
an error the recipe's guide namesapply the guide's fix exactly [doc]
tokenizer AttributeError after installtransformers newer than that vLLM supports (0.11.0 needed transformers<5 [run]); use a vLLM at or above the recipe minimum
no tool_calls--enable-auto-tool-choice plus the recipe's parser [doc][run]
404 on /modelsthe --at URL lacks /v1 [run]
"CUDA error: an illegal memory access"read vLLM's own maintainer skill .agents/skills/debug-ima/SKILL.md at https://raw.githubusercontent.com/vllm-project/vllm/main/

© autonomous-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in store/agents/autonomous-grid/skills/engine-vllm of autonomous-ai/openharness.

Open the folder on GitHubat commit 54a1f1b

Compare with similar skills

Engine Vllm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Engine Vllm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Engine Vllm this skillautonomous-ai/openharness1.1k—~1.9kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Add Modelguoqingbao/xinfer333—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Check Modelguoqingbao/xinfer333—~3.8kAutomated safety check: PassMIT

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Check Model

    guoqingbao/xinfer

    Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

    333 GitHub stars~3.8k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check passed
  • Test Model

    guoqingbao/xinfer

    Test LLM models served by xinfer for correctness, output quality, and performance.

    333 GitHub stars~2.6k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check passed

More from autonomous-ai/openharness

All 99 skills in this repo
  • G-code Slicer Tool

    autonomous-ai/openharness

    Slices 3D mesh files into printer-profiled plain G-code through real slicer CLIs, with backend discovery, input inspection, dry runs and static validation.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Home Assistant Automation Builder

    autonomous-ai/openharness

    Turns a home-automation request into standard, testable automations.yaml, run against Home Assistant Core's real triggers and verified with its own trace tool.

    1.1k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Score Music Composer

    autonomous-ai/openharness

    Turns a musical brief into LilyPond concert-pitch music, checked parts for each instrument and a playable practice pack.

    1.1k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • OrcaSlicer 3MF and G-code Workflow

    autonomous-ai/openharness

    Turns an STL and explicit printer and material requirements into compared OrcaSlicer plans, an editable 3MF project, checked G-code and a portable handoff.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Sheets and Docs Report Builder

    autonomous-ai/openharness

    Builds an editable DOCX report, a formula-driven XLSX workbook and a fresh LibreOffice PDF preview from one structured source file, then checks them together.

    1.1k GitHub stars~708 tokensUpdated today
    Auto-check passed
  • Bambu Labs

    autonomous-ai/openharness

    Dry-run, upload, and cautiously initiate local Bambu Lab print jobs from validated plain .gcode, using Bambu LAN FTPS/MQTT handoffs.

    1.1k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check: warnings

Questions about Engine Vllm

What does Engine Vllm do?

Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…. Engine Vllm is an agent skill from autonomous-ai/openharness. Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it to the person's fleet.

When should I use Engine Vllm?

Engine Vllm fits situations like: tasks that involve LLM inference and serving; tasks that involve Model hubs and datasets.

How do I install Engine Vllm in Claude Code?

Run `npx skills add autonomous-ai/openharness --skill engine-vllm -a claude-code`. Or copy the skill folder (store/agents/autonomous-grid/skills/engine-vllm in autonomous-ai/openharness) into .claude/skills/engine-vllm in your project. Claude Code loads it when a task matches its description.

How do I install Engine Vllm in Codex?

Run `npx skills add autonomous-ai/openharness --skill engine-vllm -a codex`. Or copy the skill folder (store/agents/autonomous-grid/skills/engine-vllm in autonomous-ai/openharness) into .agents/skills/engine-vllm in your project. Codex loads it when a task matches its description.

Can I use Engine Vllm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/openharness --skill engine-vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/engine-vllm, .gemini/skills/engine-vllm, .github/skills/engine-vllm and .opencode/skills/engine-vllm in your project.

What does Engine Vllm need to run?

Going by SKILL.md and its folder, Engine Vllm needs the command-line tools its instructions call (uv and docker). Our summary lists: Python 3; Docker.

Does Engine Vllm access the network?

SKILL.md names 4 domains. In commands or code: raw.githubusercontent.com and wheels.vllm.ai; the agent is likely to contact these when it follows the instructions. As links in the text: docs.vllm.ai and recipes.vllm.ai. This is read from the text; nothing was executed.

Is Engine Vllm safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Engine Vllm use?

Engine Vllm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Engine Vllm use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Engine Vllm?

Skills that share tags, products or a category with Engine Vllm: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Add Model (guoqingbao/xinfer, 333 stars) and Resolve (alexziskind1/model-shelf, 130 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Engine Vllm?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/openharness, which has 1,137 GitHub stars. The repository holds 99 skills in this directory. The repository was last updated on October 7, 2026.

Source: autonomous-ai/openharness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.