Jetson Inference Mem Tune
NVIDIA/skills
Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.
Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-vllm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-vllm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-vllm .claude/skills/agentsop-vllm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agentsop-vllm" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-vllm into .claude/skills/agentsop-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-vllm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-vllmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-vllm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-vllm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/agentsop-vllm .agents/skills/agentsop-vllm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agentsop-vllm" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-vllm into .agents/skills/agentsop-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-vllm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-vllm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-vllm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/agentsop-vllm .cursor/skills/agentsop-vllm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agentsop-vllm" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-vllm into .cursor/skills/agentsop-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-vllm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentsope/SkillAlchemy.git --path skills/agentsop-vllm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-vllm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-vllm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/agentsop-vllm .gemini/skills/agentsop-vllm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agentsop-vllm" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-vllm into .gemini/skills/agentsop-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-vllm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentsope/SkillAlchemy agentsop-vllmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentsope/SkillAlchemy --skill agentsop-vllm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/agentsop-vllm .github/skills/agentsop-vllm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-vllm" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-vllm into .github/skills/agentsop-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-vllm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentsope/SkillAlchemy --skill agentsop-vllm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentsope/SkillAlchemy agentsop-vllm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/agentsop-vllm .opencode/skills/agentsop-vllm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agentsop-vllm" agent skill from https://github.com/agentsope/SkillAlchemy/tree/master/skills/agentsop-vllm into .opencode/skills/agentsop-vllm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agentsop-vllm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agentsop-vllmDecision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy.
Agentsop Vllm is an agent skill from agentsope/SkillAlchemy. Decision SOP for serving LLMs with vLLM. Covers PagedAttention mental model, quantization/parallelism/batching tradeoffs, OOM triage, and when NOT to use vLLM. Activates when a coder-agent is choosing or tuning an inference engine, debugging vLLM throughput/latency/OOM, or comparing vLLM against TGI/SGLang/TensorRT-LLM/llama.cpp.
Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-architecture.md`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM, llama.cpp, NVIDIA AI Platform and SGLang. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d0f0355. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
arxiv.orgdevelopers.redhat.comdocs.vllm.aigithub.comyottalabs.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agentsop Vllm loads about 6.1k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 2,504 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentsope/SkillAlchemy at commit d0f0355, republished under its MIT licence (© agentsope). 2,504 words, ~6,066 tokens.
.claude/skills/agentsop-vllm/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Activate this skill when any of the following hold:
Do NOT activate for: training/fine-tuning (use accelerate/deepspeed/trl), CPU-only edge inference (use llama.cpp/Ollama), Apple Silicon production (vLLM Metal/MPS is experimental, not production-ready as of 2026) [aimadetools.com 2026], API-only consumption of hosted models (just call the OpenAI/Anthropic SDK).
vLLM's defining insight (Kwon et al., SOSP 2023) is that LLM serving's bottleneck was not compute — it was KV-cache memory fragmentation. Pre-vLLM systems pre-allocated a contiguous KV-cache slot per request, sized for the maximum possible output length; in early 2023, inference engines used only 20–40% of available GPU memory because of internal+external fragmentation [arxiv.org/abs/2309.06180; zilliz.com/learn].
PagedAttention applies classic OS paging to KV cache:
Result: near-zero memory waste → larger batch sizes → 2–4× throughput vs FasterTransformer/Orca at equal latency [arxiv.org/abs/2309.06180]; 14–24× vs vanilla HuggingFace Transformers [yottalabs.ai 2026].
vLLM inherits Orca's iteration-level scheduling (OSDI 2022, 36.9× over FasterTransformer [medium.com/byte-sized-ai]). Instead of waiting for a static batch to finish, the scheduler reassigns batch slots every decode step: a request that finishes early frees its slot to a waiting request. Static batching is dead; continuous batching is table stakes.
Takeaway: when tuning, separate TTFT (time-to-first-token, gated by prefill+queue) from ITL (gated by decode bandwidth and batch interference).
Per Red Hat's tuning hierarchy [developers.redhat.com 2026]:
[Step 0] Confirm vLLM is the right tool
├─ Production, GPU-backed, concurrent users? → continue
└─ Else → see §7 (ecosystem) and stop
[Step 1] Pick the model + precision
├─ Model fits in single-GPU VRAM at BF16? → keep BF16, TP=1
├─ Need 50% VRAM cut, ~zero quality loss? → FP8 (Hopper/Ada+) [arxiv 2411.02355]
├─ Need 4× VRAM cut, tolerate ~1.6pt avg drop? → AWQ-4 or GPTQ-4
└─ Reasoning-heavy / coding workload? → favor FP8 > AWQ; verify on eval set
[Step 2] Choose parallelism
├─ Fits 1 GPU → TP=1, PP=1
├─ Fits 1 node, NVLink present → TP=#GPUs/node
├─ Fits 1 node, only PCIe (e.g. L40S) → PP within node (TP-only over PCIe collapses)
├─ Multi-node → TP=GPUs/node, PP=#nodes
└─ MoE model (Mixtral, DSv3) → DP attention + EP/TP for MoE layers
[Step 3] Set memory/batch envelope
├─ --gpu-memory-utilization 0.90 (default; 0.85 if sharing GPU)
├─ --max-model-len = (longest realistic prompt + output) — NOT model max!
├─ --max-num-seqs (start 256; lower if preemption logs appear)
└─ --max-num-batched-tokens (raise for TTFT; lower for ITL)
[Step 4] Turn on the free wins
├─ enable_prefix_caching=True → if any system-prompt/few-shot reuse
├─ enable_chunked_prefill=True (V1: default on) → tame long-prompt HoL blocking
└─ kv_cache_dtype="fp8" → +KV headroom, Ampere+ only
[Step 5] Optional: speculative decoding
├─ Low QPS, latency-bound, have draft/EAGLE weights? → EAGLE-3 or MTP (high gain)
├─ No draft model, zero setup cost? → n-gram (modest gain, ~1.17×)
└─ High QPS / large batch? → skip; gains shrink, complexity rises
[Step 6] Benchmark on YOUR workload
├─ Replay representative ISL/OSL distribution
├─ Watch Prometheus: num_requests_waiting, KV cache occupancy, preemption count
└─ Tune in this order: §3 Step 3 → Step 4 → Step 5 → reconsider Step 2
[Step 7] Scale out
├─ Latency SLA violated under load? → add a replica (data parallelism across pods)
├─ Tail TTFT high? → prefix-aware routing; pin prefix to replica
└─ Cost too high? → revisit quantization, smaller model, draft modeltorch.OutOfMemoryError: CUDA out of memory during engine init or warmup.--max-model-len to the realistic max (prompt + output), not the model's architectural max [markaicode.com 2026].--gpu-memory-utilization to 0.85 if other processes share the GPU; raise to 0.95 if vLLM is alone and KV cache is too small.--kv-cache-dtype fp8 (Ampere+) or --enforce-eager (skip CUDA-graph reservation).--max-num-seqs (e.g. 256 → 64 → 16).--tensor-parallel-size to shard the model.--quantization {fp8|awq|gptq|...} flag; documented expected accuracy delta.--tensor-parallel-size=<GPUs in node>, PP=1.--pipeline-parallel-size instead of large TP — all-reduce over PCIe will tank TP throughput [docs.vllm.ai parallelism_scaling].TP=GPUs/node, PP=#nodes. Use InfiniBand if possible.--enable-prefix-caching (in V1, often on by default).num_requests_waiting > 0 sustained → engine queue-bound; check num_requests_running.--max-num-seqs or quantize KV to FP8.--speculative-config JSON, with measured tokens/s delta on representative traffic.Situation: A team serves Llama-3-70B-FP8 on 2×H100. P50 TTFT is fine at 250 ms, but P99 spikes to 4 s when traffic bursts. They consider raising --max-num-seqs from 64 to 256 to handle bursts.
Tension:
max_num_seqs → more concurrency → higher throughput, but also more contention for KV cache → preemption + head-of-line blocking on prefills.max_num_batched_tokens → better TTFT (more prefill per step) → worse ITL for streaming requests already mid-decode.Resolution heuristic:
num_requests_waiting > 0 and KV occupancy < 80%: raise max_num_seqs.max_model_len.max_num_batched_tokens (e.g. 2048) to favor decode/ITL; for batch/offline: raise (≥8192) to favor TTFT/throughput [docs.vllm.ai optimization; anyscale.com].Evidence: [developers.redhat.com 2026 5-steps-triage]; [medium.com/@kaige.yang0110].
Situation: Deploying Qwen-72B on a single H100 (80 GB). At BF16, model is 144 GB → doesn't fit. Options: (a) FP8, 72 GB → fits with tight KV budget; (b) AWQ-4, 36 GB → comfortable KV budget, larger batches.
Tension:
Resolution heuristic:
Evidence: [arxiv.org/abs/2411.02355]; [docs.gpustack.ai vLLM quantization].
Situation: 4×A100-80GB available, NVLink. Choice: one big replica (TP=4, max single-request latency optimized) or two replicas (TP=2, more concurrency).
Tension:
Resolution heuristic:
vllm bench serve on representative ISL/OSL.Evidence: [docs.vllm.ai parallelism_scaling]; [docs.jarvislabs.ai scaling-llm-inference-dp-pp-tp].
Situation: Production RAG service. Each request has a 2 KB system prompt + 4–16 KB context + short user query. They wonder whether enable_prefix_caching helps enough to justify the held KV blocks.
Tension:
Resolution heuristic:
Evidence: [docs.vllm.ai/en/stable/design/prefix_caching/]; [bentoml.com/llm/inference-optimization/prefix-caching].
Situation: Single-user, latency-sensitive coding assistant on Llama-3-70B. Decode dominates total latency.
Tension:
Resolution heuristic:
Evidence: [docs.vllm.ai/en/latest/features/speculative_decoding/]; [developers.redhat.com 2025 fly-eagle3-fly].
max_model_len to the model's architectural max "just in case"A Llama-3.1 model supports 128k context. If you serve 4k-prompt workloads but set --max-model-len 131072, vLLM reserves KV-cache slots for the worst case → smaller batches → throughput collapse. Set it to your realistic max (longest_prompt + longest_output + safety margin) [markaicode.com 2026].
APC accelerates prefill only. If outputs are long and prefixes don't repeat, gain ≈ 0 — and you've consumed KV memory for nothing [docs.vllm.ai automatic_prefix_caching]. Measure shared-prefix fraction before enabling for high-pressure workloads.
Tensor parallelism issues an all-reduce after each layer. NVLink (≈900 GB/s bidirectional) handles it; PCIe Gen4 x16 (≈32 GB/s, ~28× slower) does not [spheron.network 2026]. On L40S / consumer cards / cross-NUMA setups, use pipeline parallelism or independent replicas instead [docs.vllm.ai parallelism_scaling].
Academic benchmarks (MMLU, HellaSwag) show GPTQ ≈ AWQ; real coding workloads can show meaningfully larger gaps between methods, and quality at INT3 collapses (~6-point drop) [arxiv.org/abs/2411.02355]. Always test on your eval set.
vLLM needs minimum 2 + N physical CPU cores for N GPUs; under-provisioning CPUs makes the API server, tokenizer, and detokenizer the bottleneck before the GPU is touched [docs.vllm.ai/en/stable/configuration/optimization/]. Scale CPU and use --api-server-count when needed.
--enforce-eager in production "for stability"--enforce-eager skips CUDA graph capture → slower decode. It's a debugging/memory-OOM workaround, not a steady-state production setting. Fix the underlying OOM (KV dtype, max_model_len, max_num_seqs) and re-enable CUDA graphs.
vLLM is the right tool when:
vLLM is the wrong tool when:
guided_decoding (consider SGLang for complex agent state machines).| Engine | Sweet spot | When to pick over vLLM |
|---|---|---|
| vLLM | Production GPU serving, mixed traffic, open weights, vendor-neutral | Default first choice for throughput-oriented LLM serving in 2026 |
| TGI (HuggingFace) | Was the HF default | Officially in maintenance mode; HF themselves now recommend vLLM or SGLang [yottalabs.ai 2026] |
| SGLang | Heavy prefix sharing, agent/RAG state machines, structured generation | ~29% higher throughput than vLLM when requests share context (chatbots, RAG, agents) thanks to RadixAttention prefix tree [n1n.ai 2026] |
| TensorRT-LLM | Single-vendor NVIDIA, max throughput, willing to invest setup | Up to 30–50% higher throughput than vLLM in high-concurrency NVIDIA-only deployments; 1–2 weeks setup; vendor lock-in [n1n.ai 2026] |
| llama.cpp | CPU, edge, Apple Silicon, single-user, GGUF | No GPU available; ≤1 concurrent user; minimal-deps deploy [aimadetools.com 2026] |
| Ollama | Local dev, prototyping, model switching | Developer ergonomics over throughput; 5-minute setup [contracollective.com 2026] |
Common pattern: develop on Ollama → benchmark with vLLM → consider SGLang if prefix-sharing workload → consider TensorRT-LLM only if NVIDIA-locked and engineering budget is large.
A production-ish vLLM serve for Llama-3.1-70B-Instruct-FP8 on 2×H100 with NVLink, RAG-style workload:
vllm serve meta-llama/Llama-3.1-70B-Instruct \
--quantization fp8 \
--tensor-parallel-size 2 \
--max-model-len 8192 \
--max-num-seqs 256 \
--max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.90 \
--enable-prefix-caching \
--enable-chunked-prefill \
--kv-cache-dtype fp8Then triage with the §4 OP-5 5-step workflow on real traffic before tuning further.
© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (references) in skills/agentsop-vllm of agentsope/SkillAlchemy.
Open the folder on GitHubat commit d0f0355
Agentsop Vllm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agentsop Vllm this skillagentsope/SkillAlchemy | 466 | — | ~6.1k | Automated safety check: Pass | MIT | |
| Jetson Inference Mem TuneNVIDIA/skills | 3.5k | 1 repos | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Dstack Prototypingdstackai/dstack | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 925 | — | ~2.8k | Automated safety check: Pass | None | |
| Add Export Formatintel/auto-round | 1.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 |
NVIDIA/skills
Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.
dstackai/dstack
Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
intel/auto-round
Add a new model export format to AutoRound (e.g., autoround, autogptq, autoawq, gguf, llmcompressor).
Orchestra-Research/AI-Research-SKILLs
Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.
agentsope/SkillAlchemy
SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).
agentsope/SkillAlchemy
Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…
agentsope/SkillAlchemy
Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…
agentsope/SkillAlchemy
SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.
agentsope/SkillAlchemy
SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.
agentsope/SkillAlchemy
Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.
Works with
Categories
Decision SOP for serving LLMs with vLLM. An agent skill from agentsope/SkillAlchemy. Agentsop Vllm is an agent skill from agentsope/SkillAlchemy. Decision SOP for serving LLMs with vLLM.
Agentsop Vllm fits situations like: tasks that involve LLM inference and serving.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-vllm -a claude-code`. Or copy the skill folder (skills/agentsop-vllm in agentsope/SkillAlchemy) into .claude/skills/agentsop-vllm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentsope/SkillAlchemy --skill agentsop-vllm -a codex`. Or copy the skill folder (skills/agentsop-vllm in agentsope/SkillAlchemy) into .agents/skills/agentsop-vllm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-vllm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-vllm, .gemini/skills/agentsop-vllm, .github/skills/agentsop-vllm and .opencode/skills/agentsop-vllm in your project.
SKILL.md names no scripts, command-line tools or credentials: Agentsop Vllm is instructions for the agent only.
SKILL.md names 5 domains. As links in the text: arxiv.org, developers.redhat.com, docs.vllm.ai, github.com and yottalabs.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agentsop Vllm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Agentsop Vllm: Jetson Inference Mem Tune (NVIDIA/skills, 3.5k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars), Graphsignal (graphsignal/graphsignal, 257 stars) and LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 466 GitHub stars. The repository holds 46 skills in this directory. The repository was last updated on October 9, 2026.
Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.