Aider Delegate
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.
$ npx skills add Prism-Shadow/penguin-harness --skill vllm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Prism-Shadow/penguin-harness vllm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/model-development/skills/vllm .claude/skills/vllm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vllm" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/model-development/skills/vllm into .claude/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/model-development/skills/vllmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Prism-Shadow/penguin-harness --skill vllm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Prism-Shadow/penguin-harness vllm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/model-development/skills/vllm .agents/skills/vllm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vllm" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/model-development/skills/vllm into .agents/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Prism-Shadow/penguin-harness --skill vllm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Prism-Shadow/penguin-harness vllm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/model-development/skills/vllm .cursor/skills/vllm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vllm" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/model-development/skills/vllm into .cursor/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Prism-Shadow/penguin-harness.git --path plugins/model-development/skills/vllm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Prism-Shadow/penguin-harness --skill vllm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Prism-Shadow/penguin-harness vllm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/model-development/skills/vllm .gemini/skills/vllm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vllm" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/model-development/skills/vllm into .gemini/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Prism-Shadow/penguin-harness vllmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Prism-Shadow/penguin-harness --skill vllm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/model-development/skills/vllm .github/skills/vllm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vllm" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/model-development/skills/vllm into .github/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Prism-Shadow/penguin-harness --skill vllm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Prism-Shadow/penguin-harness vllm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/model-development/skills/vllm .opencode/skills/vllm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vllm" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/model-development/skills/vllm into .opencode/skills/vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vllmDeploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.
Vllm is an agent skill from Prism-Shadow/penguin-harness. Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.
Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving and Structured output and tool calling. It works with vLLM, OpenAI, Qwen and Ollama. The repository describes itself as: 🐧 Unified and Stable RSI Platform. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d56d9ce. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3curlpipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
huggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vllm loads about 1k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 461 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Prism-Shadow/penguin-harness at commit d56d9ce, republished under its Apache-2.0 licence (© Prism-Shadow). 461 words, ~1,026 tokens.
.claude/skills/vllm/SKILL.md (or your agent's skills folder).vLLM serves open-weight LLMs on local GPUs with high-throughput inference behind an OpenAI-compatible API, ready for chat and agent workloads.
If the user's message only invokes this skill (e.g. "use vllm skill") without a concrete request, ask the user what they want. Do not run any command until the goal is clear.
Ask the user which model to serve; if they have no preference, recommend the small default Qwen/Qwen3.5-0.8B. Also ask what context length the workload needs.
vLLM needs an NVIDIA or AMD GPU. Engine choice follows the user's preference: Ollama also runs on GPUs and is the simpler default — pick vLLM for high-throughput serving, and Ollama on macOS or CPU-only machines, which vLLM does not serve. Confirm the hardware first:
nvidia-smi # NVIDIA: GPU model and free VRAM (AMD ROCm: rocm-smi)
python3 --version # a recent Python is requiredThe model must fit the available VRAM — model size and context length drive the serve flags below.
curl http://localhost:8000/v1/models.penguin config model add ... --client-type openai-chat --base-url http://localhost:8000/v1 — a served model is not visible to Penguin until added.penguin config model list.Use a fresh virtual environment (or uv):
python3 -m venv .venv && source .venv/bin/activate
pip install vllmvllm serve Qwen/Qwen3.5-0.8B --port 8000This exposes an OpenAI-compatible API at http://localhost:8000/v1. Key flags:
--served-model-name <name> — the model id clients request (defaults to the model path).--api-key <key> — require this bearer token on every request.--max-model-len <n> — context window; agent sessions need a large one.--gpu-memory-utilization <0..1> — fraction of VRAM to claim (default 0.9).--tensor-parallel-size <n> — shard across n GPUs.--dtype <auto|bfloat16|float16> and --quantization <awq|gptq|fp8> — precision and quantized weights.If the port is taken, pick a free one — never kill a process already listening on it.
Agent harnesses (PenguinHarness included) send tools with their requests. vLLM must opt in at startup:
vllm serve Qwen/Qwen3.5-0.8B --enable-auto-tool-choice --tool-call-parser hermesChoose the parser for the model family — e.g. hermes for Qwen models, llama3_json for Llama models. Without these flags, requests that set tool_choice fail with 400 "auto" tool choice requires --enable-auto-tool-choice and --tool-call-parser to be set.
curl http://localhost:8000/v1/modelsModel configuration is the penguin CLI's job — penguin config model add registers an endpoint and penguin config model list shows what has been registered. A served model is not visible to Penguin until you add it:
penguin config model add --provider custom --client-type openai-chat \
--base-url http://localhost:8000/v1 --model-id <served-model-name> --api-key <key>
penguin config model list # the new entry should now be listed--gpu-memory-utilization or --max-model-len, or serve a quantized model.--max-model-len (bounded by VRAM).400 on tool calls: restart the server with the tool-calling flags above.© Prism-Shadow, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/model-development/skills/vllm of Prism-Shadow/penguin-harness.
Open the folder on GitHubat commit d56d9ce
Vllm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vllm this skillPrism-Shadow/penguin-harness | 2.5k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Aider DelegateamElnagdy/delegate-skills | 2.3k | 3 repos | ~3k | Automated safety check: Pass | MIT | |
| Perfupraullenchai/Rapid-MLX | 3.9k | — | ~1.6k | Automated safety check: Notes | Custom licence | |
| Resolvealexziskind1/model-shelf | 130 | — | ~792 | Automated safety check: Pass | MIT | |
| Tanstack AIsecondsky/claude-skills | 227 | 1 repos | ~3.6k | Automated safety check: Notes | MIT | |
| Vllm Ascendascend-ai-coding/awesome-ascend-skills | 174 | — | ~2.7k | Automated safety check: Pass | None |
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
raullenchai/Rapid-MLX
Autonomous performance optimization: research, PoC, benchmark, implement, review, PR
alexziskind1/model-shelf
Always resolve Hugging Face models via model-shelf before any download.
secondsky/claude-skills
TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama.
ascend-ai-coding/awesome-ascend-skills
vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU.
majiayu000/claude-skill-registry
A skill your agent uses when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or…
Prism-Shadow/penguin-harness
Make a reply easier to read and act on with rich blocks inside ordinary Markdown — a choice the user picks from, a form that collects several answers, a procedure as steps with warnings in place, a…
Prism-Shadow/penguin-harness
A skill your agent uses when developing PenguinHarness itself — changing packages/{core,server,web,cli,desktop,landing,docs,skills}, the built-in model catalog, the installers or the release…
Prism-Shadow/penguin-harness
Create and edit Bento presentations — self-contained .bento.html decks whose document is JSON.
Prism-Shadow/penguin-harness
A skill your agent uses when standing PenguinHarness up to try a change by hand — launching the Web App, the desktop shell, the landing page, the docs site or the component gallery to click through…
Prism-Shadow/penguin-harness
A skill your agent uses when changing the PenguinHarness Web App (packages/web) or the shared UI package — adding or restyling any UI, picking a status colour, adding an icon, laying out a row or a…
Prism-Shadow/penguin-harness
Drive the PenguinHarness agent browser — the desktop app's built-in browser or the user's own Chrome — from the shell with penguin browser: open pages, read them as simplified HTML or text, act with…
Categories
Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads. Vllm is an agent skill from Prism-Shadow/penguin-harness. Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.
Vllm fits situations like: tasks that involve LLM inference and serving; tasks that involve Structured output and tool calling.
Run `npx skills add Prism-Shadow/penguin-harness --skill vllm -a claude-code`. Or copy the skill folder (plugins/model-development/skills/vllm in Prism-Shadow/penguin-harness) into .claude/skills/vllm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Prism-Shadow/penguin-harness --skill vllm -a codex`. Or copy the skill folder (plugins/model-development/skills/vllm in Prism-Shadow/penguin-harness) into .agents/skills/vllm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Prism-Shadow/penguin-harness --skill vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm, .gemini/skills/vllm, .github/skills/vllm and .opencode/skills/vllm in your project.
Going by SKILL.md and its folder, Vllm needs the command-line tools its instructions call (python3, curl and pip). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: huggingface.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Vllm is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vllm: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Perfup (raullenchai/Rapid-MLX, 3.9k stars), Resolve (alexziskind1/model-shelf, 130 stars) and Tanstack AI (secondsky/claude-skills, 227 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Prism-Shadow (a GitHub organization) maintains it in Prism-Shadow/penguin-harness, which has 2,455 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.
Source: Prism-Shadow/penguin-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.