SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…
$ npx skills add autonomous-ai/openharness --skill engine-vllm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install autonomous-ai/openharness engine-vllm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/store/agents/autonomous-grid/skills/engine-vllm .claude/skills/engine-vllm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "engine-vllm" agent skill from https://github.com/autonomous-ai/openharness/tree/main/store/agents/autonomous-grid/skills/engine-vllm into .claude/skills/engine-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-vllm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/autonomous-ai/openharness/tree/main/store/agents/autonomous-grid/skills/engine-vllmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add autonomous-ai/openharness --skill engine-vllm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install autonomous-ai/openharness engine-vllm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/store/agents/autonomous-grid/skills/engine-vllm .agents/skills/engine-vllm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "engine-vllm" agent skill from https://github.com/autonomous-ai/openharness/tree/main/store/agents/autonomous-grid/skills/engine-vllm into .agents/skills/engine-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-vllm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add autonomous-ai/openharness --skill engine-vllm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install autonomous-ai/openharness engine-vllm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/store/agents/autonomous-grid/skills/engine-vllm .cursor/skills/engine-vllm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "engine-vllm" agent skill from https://github.com/autonomous-ai/openharness/tree/main/store/agents/autonomous-grid/skills/engine-vllm into .cursor/skills/engine-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-vllm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/autonomous-ai/openharness.git --path store/agents/autonomous-grid/skills/engine-vllm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add autonomous-ai/openharness --skill engine-vllm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install autonomous-ai/openharness engine-vllm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/store/agents/autonomous-grid/skills/engine-vllm .gemini/skills/engine-vllm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "engine-vllm" agent skill from https://github.com/autonomous-ai/openharness/tree/main/store/agents/autonomous-grid/skills/engine-vllm into .gemini/skills/engine-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-vllm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install autonomous-ai/openharness engine-vllmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add autonomous-ai/openharness --skill engine-vllm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .github/skills && cp -r skills-src/store/agents/autonomous-grid/skills/engine-vllm .github/skills/engine-vllm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "engine-vllm" agent skill from https://github.com/autonomous-ai/openharness/tree/main/store/agents/autonomous-grid/skills/engine-vllm into .github/skills/engine-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-vllm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add autonomous-ai/openharness --skill engine-vllm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install autonomous-ai/openharness engine-vllm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/store/agents/autonomous-grid/skills/engine-vllm .opencode/skills/engine-vllm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "engine-vllm" agent skill from https://github.com/autonomous-ai/openharness/tree/main/store/agents/autonomous-grid/skills/engine-vllm into .opencode/skills/engine-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "engine-vllm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
engine-vllmServe a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…
Engine Vllm is an agent skill from autonomous-ai/openharness. Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it to the person's fleet. Load before installing, configuring, starting or stopping vLLM.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving and Model hubs and datasets. It works with vLLM, NVIDIA AI Platform, Linux and Hugging Face. The repository describes itself as: The ultimate harness for coding agents and beyond. All your agents. All your machines. One command center. Start with code, then follow your curiosity and build across… The licence is MIT.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 54a1f1b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvdockerFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
raw.githubusercontent.comwheels.vllm.aiAlso links to:
docs.vllm.airecipes.vllm.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Engine Vllm loads about 1.9k tokens when it runs. Until then it costs about 74 tokens; SKILL.md has 847 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from autonomous-ai/openharness at commit 54a1f1b, republished under its MIT licence (© autonomous-ai). 847 words, ~1,931 tokens.
.claude/skills/engine-vllm/SKILL.md (or your agent's skills folder).Official sources, read 2026-09-29. Agent index: recipes.vllm.ai/llms.txt.
Reference docs as raw markdown under https://raw.githubusercontent.com/vllm-project/vllm/main/docs/:
OpenAI-compatible server (serving/online_serving/openai_compatible_server.md) ·
Tool calling (features/tool_calling.md) ·
Reasoning outputs (features/reasoning_outputs.md) ·
Supported models (models/supported_models.md) ·
GPU install · Docker ·
Optimization · Conserving memory
Tested: only vLLM 0.11.0 on an Apple-silicon Mac (CPU backend), 2026-09-29 — no GPU run yet.
Tags: [doc] official source above, [run] seen on the tested machine, [?] unverified.
transformers<5, --dtype float32 for JSON output and a batched-token limit to start [run]."$GRID_FLEET" recipe vllm ORG/NAMEReads recipes.vllm.ai live (models.json, then <ORG>/<NAME>.json) and prints:
| Field | Use |
|---|---|
minVersion | the installed vllm --version must be at least this |
baseArgs, baseEnv | always |
features.tool_calling.args, features.reasoning.args | always for an agent |
features.* with optIn: true | only on purpose (faster decoding, text only to free memory, longer context) |
variants.<name> | pick by vramMinimumGb against the GPU; a variant can carry its own modelId and extraArgs |
recommended | the site's reference command and Docker image for its reference hardware |
install | the recipe's own install steps, including any pinned extras |
guide | prose; its Troubleshooting holds model-specific fixes |
The recipe beats the generic docs: a family's generic parser entry and a model's recipe can name different parsers [doc]. Exit code 3 means there is no recipe: go to step 2.
"$GRID_FLEET" model-facts ORG/NAMEsupport.vllm.listed must be true (its architecture is in vLLM's supported-models list). false:
the Transformers backend (--model-impl transformers) may still run it [doc] — a test, not a promise.modelCard.serveCommands / toolCallParsers / reasoningParsers: the model authors' own vLLM command.
Use it when present.chatTemplate.toolSyntax against the parser descriptions in features/tool_calling.md
(each parser documents the output format it reads) and, when chatTemplate.thinking, pick the reasoning
parser from the table in features/reasoning_outputs.md [doc]. Quote the doc line you matched.contextLength must be ≥ 65536. Sampling defaults come from the model's generation_config.json
automatically [doc]; leave them.uv venv && uv pip install -U vllm --torch-backend=auto [doc]; pin with
"vllm==X.Y.Z" at or above the recipe's minVersion. Wheels target CUDA 12.9 by default, builds for 12.8
and 13.0 exist, Blackwell needs CUDA ≥ 12.8 [doc]. AMD: --extra-index-url https://wheels.vllm.ai/rocm
(Python 3.12, ROCm 7.0, glibc ≥ 2.35) [doc]. Check the GPU and driver with nvidia-smi.vllm/vllm-openai (NVIDIA), vllm/vllm-openai-rocm (AMD) [doc], with
--gpus all --ipc=host -p 127.0.0.1:P:8000 -v ~/.cache/huggingface:/root/.cache/huggingface [doc].Each adds the recipe's (or step 2's) tool and reasoning flags, --host 127.0.0.1 --port P and a
--max-model-len of at least 65536.
| Case | Add |
|---|---|
| One GPU, one agent | --max-model-len 131072 --max-num-seqs 4 --enable-prefix-caching |
| One GPU, many people | --max-model-len 65536 --enable-prefix-caching (max-num-seqs left to vLLM) |
| Several GPUs, one machine | --tensor-parallel-size <GPU count> [doc] |
| Memory is tight | the recipe's FP8 variant, --kv-cache-dtype fp8, fewer --max-num-seqs [doc] |
"$GRID_FLEET" serve vllm-P --env HF_HUB_OFFLINE=1 -- ~/.grid/envs/vllm/bin/vllm serve <model id or snapshot dir> \
--served-model-name <id> <flags>(Install into ~/.grid/envs/vllm so fleet models finds it; fleet serve keeps it alive after your shell
returns; log run/vllm-P.log.)
"$GRID_FLEET" verify --at http://127.0.0.1:P/v1 --model <id> --kind vllm (bounded, narrated).Application startup complete. [run] (minutes for a big model); GET /health
→ 200 and GET /v1/models lists <id> [run]; one bounded answer; one tool call."$GRID_FLEET" run -- join GRID --at http://127.0.0.1:P/v1 -m <id> --advertise-as ALIAS. /v1 is required —
without it /models and /chat/completions answer 404 [run]. --api-key guards only /v1 routes [doc].--default-chat-template-kwargs '{"enable_thinking": false}' in the recipes that document it [doc]).--gpu-memory-utilization (share of GPU memory pre-allocated for weights plus cache), --max-model-len,
--max-num-seqs, --max-num-batched-tokens, --kv-cache-dtype fp8, --enforce-eager (fastest start,
slower serving), --tensor-parallel-size [doc].
"$GRID_FLEET" stop vllm-P (or docker stop the container you started); confirm the port is free.
| Sign | Do |
|---|---|
| "preempted … not enough KV cache space" | raise --gpu-memory-utilization, lower --max-num-seqs, or more tensor parallel [doc] |
| CUDA out of memory at start | FP8 variant, --kv-cache-dtype fp8, fewer sequences — context stays ≥ 65536 |
| "max_num_batched_tokens … smaller than max_model_len" | raise --max-num-batched-tokens to at least the context [doc][run] |
| an error the recipe's guide names | apply the guide's fix exactly [doc] |
tokenizer AttributeError after install | transformers newer than that vLLM supports (0.11.0 needed transformers<5 [run]); use a vLLM at or above the recipe minimum |
no tool_calls | --enable-auto-tool-choice plus the recipe's parser [doc][run] |
404 on /models | the --at URL lacks /v1 [run] |
| "CUDA error: an illegal memory access" | read vLLM's own maintainer skill .agents/skills/debug-ima/SKILL.md at https://raw.githubusercontent.com/vllm-project/vllm/main/ |
© autonomous-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in store/agents/autonomous-grid/skills/engine-vllm of autonomous-ai/openharness.
Open the folder on GitHubat commit 54a1f1b
Engine Vllm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Engine Vllm this skillautonomous-ai/openharness | 1.1k | — | ~1.9k | Automated safety check: Pass | MIT | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Add Modelguoqingbao/xinfer | 333 | — | ~4.2k | Automated safety check: Notes | MIT | |
| Resolvealexziskind1/model-shelf | 130 | — | ~792 | Automated safety check: Pass | MIT | |
| Check Modelguoqingbao/xinfer | 333 | — | ~3.8k | Automated safety check: Pass | MIT |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
guoqingbao/xinfer
Adapt and port new LLM model architectures to this xinfer project.
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
guoqingbao/xinfer
Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.
guoqingbao/xinfer
Test LLM models served by xinfer for correctness, output quality, and performance.
autonomous-ai/openharness
Slices 3D mesh files into printer-profiled plain G-code through real slicer CLIs, with backend discovery, input inspection, dry runs and static validation.
autonomous-ai/openharness
Turns a home-automation request into standard, testable automations.yaml, run against Home Assistant Core's real triggers and verified with its own trace tool.
autonomous-ai/openharness
Turns a musical brief into LilyPond concert-pitch music, checked parts for each instrument and a playable practice pack.
autonomous-ai/openharness
Turns an STL and explicit printer and material requirements into compared OrcaSlicer plans, an editable 3MF project, checked G-code and a portable handoff.
autonomous-ai/openharness
Builds an editable DOCX report, a formula-driven XLSX workbook and a fresh LibreOffice PDF preview from one structured source file, then checks them together.
autonomous-ai/openharness
Dry-run, upload, and cautiously initiate local Bambu Lab print jobs from validated plain .gcode, using Bambu LAN FTPS/MQTT handoffs.
Works with
Categories
Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it…. Engine Vllm is an agent skill from autonomous-ai/openharness. Serve a Hugging Face model with vLLM on a Linux machine with an NVIDIA or AMD GPU, configured from the model's official vLLM recipe — or, when it has none, from the model's own files — and join it to the person's fleet.
Engine Vllm fits situations like: tasks that involve LLM inference and serving; tasks that involve Model hubs and datasets.
Run `npx skills add autonomous-ai/openharness --skill engine-vllm -a claude-code`. Or copy the skill folder (store/agents/autonomous-grid/skills/engine-vllm in autonomous-ai/openharness) into .claude/skills/engine-vllm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add autonomous-ai/openharness --skill engine-vllm -a codex`. Or copy the skill folder (store/agents/autonomous-grid/skills/engine-vllm in autonomous-ai/openharness) into .agents/skills/engine-vllm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/openharness --skill engine-vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/engine-vllm, .gemini/skills/engine-vllm, .github/skills/engine-vllm and .opencode/skills/engine-vllm in your project.
Going by SKILL.md and its folder, Engine Vllm needs the command-line tools its instructions call (uv and docker). Our summary lists: Python 3; Docker.
SKILL.md names 4 domains. In commands or code: raw.githubusercontent.com and wheels.vllm.ai; the agent is likely to contact these when it follows the instructions. As links in the text: docs.vllm.ai and recipes.vllm.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Engine Vllm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Engine Vllm: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Add Model (guoqingbao/xinfer, 333 stars) and Resolve (alexziskind1/model-shelf, 130 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
autonomous-ai (a GitHub organization) maintains it in autonomous-ai/openharness, which has 1,137 GitHub stars. The repository holds 99 skills in this directory. The repository was last updated on October 7, 2026.
Source: autonomous-ai/openharness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.