Agent skill

Agentsop LLM Engine Selection

by agentsope in agentsope/SkillAlchemy

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

MITAuto-check passedAI & LLM Engineering

Install Agentsop LLM Engine Selection

skills CLI
$ npx skills add agentsope/SkillAlchemy --skill agentsop-llm-engine-selection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentsope/SkillAlchemy agentsop-llm-engine-selection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentsope/SkillAlchemy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agentsop-llm-engine-selection .claude/skills/agentsop-llm-engine-selection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agentsop-llm-engine-selection
GitHub stars
457
Token cost
~6.1k tokens
SKILL.md length
2,587 words
Files
5 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
MIT

At a glance

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

  • Works in 7 steps: 何时激活 (When to activate) → 核心心智模型 (Core Mental Model) → SOP 工作流 (SOP Workflow) → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers 1. 何时激活 (When to activate), 2. 核心心智模型 (Core Mental Model), 3. SOP 工作流 (SOP Workflow) and 4. 操作模型 (Operation Model), plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agentsop LLM Engine Selection is an agent skill from agentsope/SkillAlchemy. Cross-engine decision rubric for self-hosting or recommending an LLM serving stack. Picks among vLLM, SGLang, TensorRT-LLM, TGI, llama.cpp, Ollama, and MLX as a function of (hardware × workload × constraint), not "which is fastest". Activates whenever a coder-agent must choose, defend, or migrate a serving runtime.

Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `intermediate/operation_candidates.json` and `references/R1-source-evidence.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with NVIDIA AI Platform, vLLM, Ollama and SGLang. The repository describes itself as: From thought to skill. From signal to structure. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “which is fastest”
  • “/agentsop-llm-engine-selection”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. 何时激活 (When to activate)
  2. 核心心智模型 (Core Mental Model)
  3. SOP 工作流 (SOP Workflow)
  4. 操作模型 (Operation Model)
  5. 困境决策案例 (Dilemma Cases)
  6. 反模式与边界 (Anti-patterns & Boundaries)
  7. 跨框架对照 (Cross-framework Comparison)

What it can do on your machine

Read from SKILL.md and the folder at commit 6ea799f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agentsop LLM Engine Selection loads about 6.1k tokens when it runs, and up to ~8.9k if it reads all its reference files. Until then it costs about 87 tokens; SKILL.md has 2,587 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~6.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentsope/SkillAlchemy at commit 6ea799f, republished under its MIT licence (© agentsope). 2,587 words, ~6,095 tokens.

Download SKILL.mdSave it as .claude/skills/agentsop-llm-engine-selection/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
agentsop-llm-engine-selection
description
Cross-engine decision rubric for self-hosting or recommending an LLM serving stack. Picks among vLLM, SGLang, TensorRT-LLM, TGI, llama.cpp, Ollama, and MLX as a function of (hardware × workload × constraint), not "which is fastest". Activates whenever a coder-agent must choose, defend, or migrate a serving runtime.
domain
LLM inference engine selection
version
0.1.0
dated
2026-05
sources
https://www.yottalabs.ai/post/best-llm-inference-engines-in-2026-vllm-tensorrt-llm-tgi-and-sglang-compared…

LLM Engine Selection SOP

State-of-the-art warning. This skill is dated May 2026. The inference-engine landscape moves in 3–6 month cycles (TGI exited maintenance into deprecation in late 2025; SGLang's RadixAttention regressed vLLM's lead in 2024; TensorRT-LLM dropped its proprietary-engine requirement in 2025; Ollama added concurrent-request support mid-2025). Re-verify before betting a quarter of eng budget on any choice below.


1. 何时激活 (When to activate)

Activate this skill any time a coder-agent must:

  • Pick a serving stack for a new project (production hosting / batch / edge / dev laptop / multi-tenant SaaS / structured-output service).
  • Defend an existing stack against a "let's switch to X" pressure.
  • Migrate: justify or block a swap (e.g. TGI → vLLM, Ollama → vLLM, vLLM → TensorRT-LLM).
  • Mix: design a multi-tier deployment (e.g. premium tier on TensorRT-LLM, free tier on vLLM-AWQ, dev on Ollama).
  • Audit a recommendation that smells like benchmark-cherry-picking ("X is 5× faster").

Do not activate for:

  • Tuning a single chosen engine — defer to the dedicated skill (vllm, sglang, tensorrt-llm, llama-cpp).
  • Training/fine-tuning runtime selection — different problem class (accelerate, deepspeed, axolotl).
  • Hosted-API procurement (OpenAI / Anthropic / Bedrock) — engine choice doesn't apply.

2. 核心心智模型 (Core Mental Model)

Engine choice is a function of (hardware × workload × constraint), not "which is fastest".

There is no global ranking. Every "X beats Y by N%" headline holds only inside an unstated (hardware, batch size, ISL/OSL, model, quantization, concurrency) tuple. Change any axis and the ranking flips.

2.1 The four-axis decision space
  1. Hardware axis — NVIDIA H100/A100 (NVLink) ≠ NVIDIA L40S/RTX (PCIe-only) ≠ AMD MI300 ≠ Apple Silicon ≠ CPU-only. The interconnect topology matters as much as raw FLOPS [spheron.network 2026]. PCIe-only tensor parallelism collapses; NVLink rescues it.
  2. Workload axis — production multi-user (throughput) vs latency-bound single-stream vs offline batch vs edge single-user vs structured-output service vs multi-LoRA SaaS. Each has a different winner.
  3. Constraint axis — license (Apache vs proprietary), vendor lock-in tolerance, engineering budget (1 day vs 2 weeks setup), commercial-use clauses, on-prem vs cloud, P50 vs P99 SLA.
  4. Maturity axis — the engine's coverage of YOUR model family. A 2025-launched MoE may run on vLLM day-1 but need a 3-month wait for TensorRT-LLM, and may never get a stable GGUF.
2.2 The default in 2026

For GPU-backed, multi-user, open-weights serving, the default is vLLM. It owns the production slot because it is vendor-neutral (NVIDIA/AMD/Intel/TPU/Apple-experimental), supports 200+ architectures including MoE/multimodal, ships an OpenAI-compatible API, and HuggingFace themselves recommend it over their own (now-maintenance-mode) TGI [yottalabs.ai 2026; vllm-project README].

You only reach past vLLM when one of three conditions binds:

  • Workload-binding (heavy prefix sharing or structured generation → SGLang).
  • Hardware-binding (NVIDIA-only + max throughput goal → TensorRT-LLM; CPU/edge/Apple → llama.cpp/MLX/Ollama).
  • Operator-binding (dev laptop, want 5-min setup → Ollama; ≤1 concurrent user → llama.cpp).
2.3 "Fastest" is a category error

Throughput-per-GPU, throughput-per-dollar, P50 TTFT, P99 ITL, and developer-time-to-first-request are five different goals, and the engines optimize for different combinations:

EngineWhat it optimizes for
vLLMThroughput-per-GPU across mixed traffic, model breadth
SGLangThroughput when requests share prefix; structured-gen TPS
TensorRT-LLMPeak throughput on NVIDIA at saturation; per-token cost at scale
TGIWas generic; now mostly a migration source
llama.cppSingle-user TPS on CPU/Apple/edge; minimal-deps install
OllamaDeveloper-time-to-first-request (5 min)
MLXApple Silicon throughput and Apple-native dev UX

Pick by which axis your project is binding on, not by which engine has the most stars.


3. SOP 工作流 (SOP Workflow)

[Step 0] Define the four-axis constraint vector
   ├─ Hardware: GPU vendor, count, interconnect (NVLink? PCIe?), VRAM/GPU
   ├─ Workload: # concurrent users, ISL/OSL distribution, shared-prefix %, structured-out %
   ├─ Constraint: license, vendor-lock tolerance, eng-days budget, P50/P99 SLA
   └─ Model: family (Llama/Qwen/Mixtral/DeepSeek/Mamba/...), size, quantization preference

[Step 1] Eliminate incompatible engines (hard filters)
   ├─ No NVIDIA GPU?            → drop TensorRT-LLM
   ├─ CPU/Apple/edge only?      → drop vLLM (production), TGI, TensorRT-LLM
   ├─ Need OSS-permissive only? → drop TensorRT-LLM (NVIDIA license)
   ├─ Mamba / brand-new arch?   → check vLLM+SGLang coverage; likely drop others
   └─ Multi-LoRA hot-swap?      → vLLM (best), TensorRT-LLM (good), SGLang (good); others drop

[Step 2] Map workload to engine strength
   ├─ Throughput + mixed traffic         → vLLM
   ├─ Shared prefixes (chatbot, agent)   → SGLang (≈29% over vLLM on shared-context [n1n.ai 2026])
   ├─ Structured JSON/regex at scale     → SGLang (RadixAttention + state machines)
   ├─ Peak throughput, NVIDIA-only, big budget → TensorRT-LLM (+30–50% over vLLM [n1n.ai 2026])
   ├─ Single user, dev laptop            → Ollama
   ├─ CPU/Apple/edge/embedded            → llama.cpp / MLX
   └─ Multi-tenant SaaS with LoRA fleet  → vLLM (multi-LoRA mature) or SGLang

[Step 3] Sanity-check hardware topology
   ├─ TP requires NVLink (≈900 GB/s) — not PCIe Gen4 (≈32 GB/s, ~28× slower)
   ├─ PCIe-only box → smaller TP + replicas, or PP, or switch engines
   ├─ Cross-NUMA across the same node → as bad as PCIe; pin to NUMA-local GPUs
   └─ Multi-node → require IB / RoCE, not Ethernet

[Step 4] Benchmark top 2 on YOUR workload
   ├─ Replay representative ISL/OSL distribution (NOT MMLU prompts)
   ├─ Measure P50 + P99 TTFT, P50 + P99 ITL, tokens/s/GPU, $/M-tokens
   ├─ Watch saturation: queue depth, KV occupancy, preemption count
   └─ Decide; document the constraint vector that justified the pick

[Step 5] Plan the escape hatch
   ├─ Note the workload threshold that would force a switch
   ├─ Keep the OpenAI-compatible API layer so swaps are mechanical
   └─ Re-evaluate every 6 months — engines evolve in quarters, not years

4. 操作模型 (Operation Model)

OP-1: Decision matrix lookup by (hardware, workload)
  • Trigger: User asks "what should we serve X on?" and gives hardware + workload.
  • Action: Look up the row in §7 table; eliminate by Step 1 hard filters; pick by Step 2 strength match.
  • Output: A primary engine + a fallback, with the constraint vector that justified the pick.
  • Evidence: §7 cross-framework matrix; [yottalabs.ai 2026]; [aimadetools.com 2026].
OP-2: Detect "wrong-tool-for-the-job" symptoms
  • Trigger: A team is using engine E but reporting one of: low throughput despite tuning, unsupported model architecture, structured-output workarounds, single-user pain on a multi-user engine.
  • Action: Match symptom → engine mismatch:
    • vLLM stuck below expected throughput on PCIe-only L40S → switch to PP or 2× independent TP=2 replicas (engine isn't the problem; topology is).
    • vLLM with heavy guided_decoding overhead on agent traffic → A/B SGLang.
    • TGI deployment, HF themselves now recommend vLLM/SGLang → plan migration.
    • Ollama in "production" with ≥10 concurrent users → graduate to vLLM.
    • llama.cpp on a serving-many-users box → graduate to vLLM.
  • Output: Migration recommendation with cited reason.
  • Evidence: [yottalabs.ai 2026]; [contracollective.com 2026].
OP-3: Quantization-driven engine pick
  • Trigger: Quantization format dictates engine; team has GGUF / AWQ / GPTQ / FP8 weights already.
  • Action:
    • GGUF only → llama.cpp / Ollama / LM Studio. vLLM has experimental GGUF; not production-grade.
    • AWQ / GPTQ → vLLM (first-class), SGLang (good), TensorRT-LLM (supported, needs engine rebuild).
    • FP8 (Hopper/Ada+) → vLLM, SGLang, TensorRT-LLM all support; near-lossless [arxiv 2411.02355].
    • MLX format (Apple) → MLX / mlx-lm; not portable to other engines.
  • Output: Engine shortlist gated by available weights format.
  • Evidence: [docs.vllm.ai/quantization]; [arxiv.org/abs/2411.02355].
OP-4: Structured-output workload routing
  • Trigger: ≥30% of requests need JSON / regex / grammar-constrained output, OR the project is an agent with tool-call state machines.
  • Action: Default to SGLang. RadixAttention + first-class state machines beat vLLM's guided_decoding at scale. If staying on vLLM, expect throughput hit on constrained requests; budget for it.
  • Output: SGLang shortlisted as primary; vLLM as fallback if SGLang doesn't support the model.
  • Evidence: [sglang.ai blog]; [n1n.ai 2026 SGLang +29% on shared-context].
OP-5: Apple Silicon / edge routing
  • Trigger: Target platform is M-series Mac, iPhone/iPad, Jetson, Raspberry Pi, or other ARM/edge.
  • Action:
    • M-series Mac, dev/local: Ollama (best UX) or MLX (best perf, Apple-native).
    • M-series Mac, production server: llama.cpp Metal backend or MLX. vLLM Metal/MPS is experimental as of 2026; do not ship on it [aimadetools.com 2026; contracollective.com 2026].
    • Jetson / ARM Linux: llama.cpp (broadest hardware support) or TensorRT-LLM if Jetson Orin.
    • CPU-only x86 server: llama.cpp with AVX-512; expect single-digit users.
  • Output: Apple/edge-specific engine pick with format requirement (GGUF / MLX).
  • Evidence: [contracollective.com 2026]; [aimadetools.com 2026].
OP-6: Multi-tier deployment design
  • Trigger: Project has both interactive (premium, latency-bound) and batch (free, throughput-bound) tiers.
  • Action: Run heterogeneous engines:
    • Premium / interactive on TensorRT-LLM (if NVIDIA-locked + budget) or vLLM FP8 with prefix caching.
    • Free / batch on vLLM AWQ-4 (high concurrency, lower cost) or SGLang if prefix-sharing.
    • Dev / staging on Ollama for fast model swaps.
    • Edge / on-device client on llama.cpp (GGUF).
  • Output: 2–3 engine deployment plan with a single OpenAI-compatible gateway in front.
  • Evidence: Common production pattern documented in [yottalabs.ai 2026; contracollective.com 2026].
OP-7: Bench-then-pick when the choice is close
  • Trigger: Two engines both pass Step 1 + 2; team disagrees.
  • Action: Run a 1-day benchmark with replayed traffic. Required metrics: P50 + P99 TTFT, P50 + P99 ITL, tokens/s/GPU, $/M-tokens, peak memory, model-load time. Do not benchmark on MMLU-style prompts; use real production ISL/OSL.
  • Output: A decision memo with the constraint vector + measured numbers + the threshold that would flip the decision.
  • Evidence: Standard practice; all cited comparison reports note that workload-specific benchmarks override published numbers.

5. 困境决策案例 (Dilemma Cases)

Dilemma 1: Team wants vLLM, but only has a PCIe-only 4-GPU box

Situation: A team standardized on vLLM but their cluster is 4× L40S on PCIe Gen4 (no NVLink). They configure --tensor-parallel-size 4 for Llama-3-70B-FP8 and see throughput collapse — single-replica TPS is ~30% of what the H100×4 benchmark advertises.

Tension:

  • vLLM is the right engine, but TP=4 over PCIe is the wrong topology. The all-reduce after every layer chokes on the ~32 GB/s PCIe link vs the ~900 GB/s NVLink the benchmarks assumed [spheron.network 2026].
  • Switching engines doesn't help — the same all-reduce penalty hits TensorRT-LLM and SGLang.

Resolution:

  • First: drop to TP=2 × 2 replicas (each replica uses an NVLink-paired or NUMA-local pair, if any), or use pipeline parallelism PP=4, TP=1. Replicas eliminate the cross-GPU all-reduce on the hot path [developers.redhat.com 2026 step-5].
  • Second: if 70B doesn't fit at TP=2 with FP8 on L40S (48 GB/GPU), step down to AWQ-4 or use Llama-3-8B until hardware upgrades.
  • Third: if neither works, swap to TGI (similar PCIe penalty but lower TP-coupling overhead in some configs) or accept a switch to llama.cpp for low-concurrency workloads.
  • Key lesson: the bottleneck was hardware topology, not engine choice. Always check interconnect before blaming the engine.

Evidence: [docs.vllm.ai parallelism_scaling]; [developers.redhat.com 2026]; [spheron.network 2026].

Dilemma 2: Structured JSON output at scale — SGLang vs vLLM

Situation: An agent platform serves Qwen-2.5-32B-Instruct with ~70% of requests demanding strict JSON schema output (tool-calling, function arguments). On vLLM with guided_decoding, throughput drops ~40% vs unconstrained baseline; P99 TTFT regresses.

Tension:

  • vLLM has the broader model + production track record; team already operates it.
  • SGLang's RadixAttention + native state-machine compilation gives ≈29% higher throughput on shared-context workloads and meaningful gains on constrained decoding [n1n.ai 2026]; agent traffic has both properties.
  • Migration cost is real: deployment scripts, observability, ops runbooks.

Resolution:

  • A/B: run SGLang as a sidecar on a fraction of traffic. Measure constrained-output TPS + correctness rate on real JSON schemas (not toy schemas).
  • If SGLang wins ≥25% on YOUR traffic → migrate the structured-output path; keep vLLM for unconstrained traffic, OR consolidate on SGLang if it covers all your models.
  • If SGLang model coverage misses a key model in your fleet → stay on vLLM, optimize guided_decoding (xgrammar backend), and re-evaluate next quarter.
  • Heuristic: structured-output workload share above ~30% justifies SGLang shortlisting; above ~50% it usually wins.

Evidence: [sglang.ai/blog]; [n1n.ai 2026]; [yottalabs.ai 2026].

Show full SKILL.md (1,091 more words)Show less
Dilemma 3: NVIDIA shop with 2-week budget — vLLM or TensorRT-LLM?

Situation: NVIDIA-only cluster (8×H100, NVLink-full), single-tenant, throughput-per-GPU is the dominant cost line. Team has 2 weeks of senior eng time. TensorRT-LLM promises +30–50% throughput [n1n.ai 2026] but requires engine build per (model × precision × max-batch × max-seq) tuple.

Tension:

  • TensorRT-LLM wins on saturated throughput and is the right tool when GPU-hours are the binding cost.
  • It loses on agility: every model update or context-length change requires an engine rebuild (10 min – 2 hours); architecture support lags vLLM by 3–6 months for new releases; vendor lock-in if you ever consider AMD/TPU.
  • vLLM tolerates new models on day-1 and supports any NVIDIA + AMD + Intel hardware.

Resolution:

  • If model is stable for ≥3 months and throughput dominates cost → TensorRT-LLM is correct. Spend the 2 weeks.
  • If model churn is monthly OR you ship new models often OR you want vendor optionality → vLLM. The +30–50% throughput is rarely worth the agility loss.
  • Hybrid: vLLM during model-fit phase; freeze a winning model into a TensorRT-LLM engine for steady-state production.

Evidence: [n1n.ai 2026]; [yottalabs.ai 2026]; [spheron.network 2026].

Dilemma 4: "We want one engine for laptop dev → cloud prod"

Situation: A small team wants the same engine for local dev on M-series Macs AND production on cloud NVIDIA. They are tempted to standardize on Ollama (since it runs everywhere) or vLLM (since it's the prod default).

Tension:

  • Ollama everywhere: great dev UX, but tops out at single-digit concurrent users; vLLM-class throughput is unreachable. Using Ollama in cloud production wastes 10–20× the GPU spend [contracollective.com 2026].
  • vLLM everywhere: production-grade, but Metal/MPS is experimental on Mac; dev velocity is worse than Ollama; single-user latency is no better than llama.cpp.

Resolution:

  • Do not standardize on one engine. Standardize on the API contract (OpenAI-compatible HTTP).
  • Local dev: Ollama (M-series) or vLLM (if devs have NVIDIA workstations). The OpenAI compatibility means application code is identical.
  • Production: vLLM (default) or SGLang/TensorRT-LLM (if Step 2 binds).
  • Edge / on-device: llama.cpp with GGUF.
  • The "one engine" goal is the wrong abstraction; the right abstraction is "one API, multiple runtimes" — same pattern as JVM languages, same pattern as POSIX [yottalabs.ai 2026 common pattern].

Evidence: [contracollective.com 2026]; [yottalabs.ai 2026].


6. 反模式与边界 (Anti-patterns & Boundaries)

Anti-pattern 1: Pick by GitHub stars or community vibes

vLLM has ~50k stars, llama.cpp has ~70k, TensorRT-LLM has ~10k. Stars do not predict fit for YOUR workload. A 70-star project may be the only thing that runs your model on your hardware. Pick by the constraint vector (§3 Step 0), not by popularity.

Anti-pattern 2: Pick by the tutorial you read last week

The "X is 5× faster than Y" blog you skimmed had specific (hardware, model, batch, ISL/OSL) parameters. Most blogs do not state them. Treat every comparison as ungeneralizable until you've replicated it with YOUR parameters.

Anti-pattern 3: Ignore hardware topology

PCIe-only TP. Cross-NUMA all-reduce. Ethernet between nodes for tensor-parallel. Each of these turns a "fastest engine" into a slowest engine. The engine isn't the variable; the interconnect is. Check the topology before blaming the engine [spheron.network 2026].

Anti-pattern 4: Confuse single-user latency with throughput

Ollama and llama.cpp can give you a faster first token on a single request than vLLM's cold path. This proves nothing about serving 50 users. Benchmark at YOUR concurrency, not at concurrency=1.

Anti-pattern 5: Forget the license

TensorRT-LLM ships under NVIDIA's proprietary license. If your shop forbids non-OSS in the inference path (e.g. air-gapped, regulated, or open-core-only orgs), this is a hard filter — no benchmark matters. Same caveat for any Triton-integrated engine that pulls in NVIDIA-proprietary components.

Anti-pattern 6: Standardize on TGI in 2026

HuggingFace themselves moved their internal recommendation to vLLM and SGLang; TGI is in maintenance mode [yottalabs.ai 2026]. Existing TGI deployments are fine if they meet SLA, but new projects in 2026 should not pick TGI.

Anti-pattern 7: Treat Ollama as production-grade

Ollama has added concurrent-request support, but its sweet spot is local dev. A production deployment on Ollama at 10+ QPS wastes 80–95% of GPU capacity vs vLLM [contracollective.com 2026; aimadetools.com 2026]. Use Ollama for dev, swap to vLLM for prod.

Boundaries

This skill is correct when:

  • You are picking, defending, or migrating an inference engine for a new or existing deployment.
  • The four-axis constraint vector (§3 Step 0) is knowable.
  • The user can run a 1-day benchmark on representative traffic.

This skill is not the right tool when:

  • The choice has already been made and the question is how to tune — defer to vllm, sglang, tensorrt-llm, llama-cpp skills.
  • The workload is a hosted-API consumer; engine choice doesn't apply.
  • The choice is training/fine-tuning runtime, not inference.

7. 跨框架对照 (Cross-framework Comparison)

7.1 Vendor matrix
EngineBest forHardwareLicenseOne-line strengthOne-line weakness
vLLMProduction GPU serving, mixed traffic, open weightsNVIDIA / AMD / Intel / TPU / Apple (exp.)Apache 2.0Vendor-neutral, broadest model coverage, HF-blessed defaultNot the fastest in any single niche
SGLangShared-prefix workloads, agents, RAG, structured generationNVIDIA / AMDApache 2.0RadixAttention prefix tree + native state machines; ~29% over vLLM on shared-contextSmaller model coverage; smaller ecosystem
TensorRT-LLMPeak NVIDIA throughput at saturationNVIDIA onlyNVIDIA proprietary+30–50% over vLLM in NVIDIA-only deploymentsVendor lock-in; 1–2 weeks setup; per-config engine rebuild
TGIExisting HF-stack deploymentsNVIDIA / AMDHFOIL (Apache-restricted)Was the HF defaultMaintenance mode; HF now recommends vLLM/SGLang
llama.cppCPU / Apple Silicon / edge / single user / GGUFx86 / ARM / Apple / consumer GPU / CUDA / VulkanMITRuns anywhere; minimal deps; GGUF ecosystemPer-request overhead doesn't scale to many concurrent users
OllamaLocal dev, prototyping, model switchingWraps llama.cppMIT5-minute install; best dev UXConcurrent throughput is poor vs vLLM-class engines
MLX / vllm-mlxApple Silicon nativeM-series onlyMITApple's first-party path; best Mac throughputMac-only; smaller ecosystem
7.2 Headline performance gaps (2025–2026 vintage)
  • vLLM vs HF Transformers: 14–24× higher throughput for Llama on NVIDIA [yottalabs.ai 2026].
  • vLLM vs early TGI: 2.2–3.5× higher throughput at launch [yottalabs.ai 2026].
  • SGLang vs vLLM: ~29% higher throughput when requests share prefix [n1n.ai 2026]; pure-batch parity.
  • TensorRT-LLM vs vLLM: +30–50% in saturated NVIDIA-only high-concurrency [n1n.ai 2026].
  • vLLM vs llama.cpp at many-user concurrency: vLLM wins decisively [aimadetools.com 2026].
  • llama.cpp vs vLLM at concurrency=1 on consumer GPU/CPU: llama.cpp wins on simplicity and often on TPS.

All numbers assume the engine's preferred topology (NVLink for vLLM/TRT-LLM, etc.). Cross-topology comparisons are meaningless.

7.3 The default decision flow
GPU-backed multi-user production?
├─ Yes
│  ├─ Heavy shared prefixes or structured-out >30%?  → SGLang
│  ├─ NVIDIA-locked + max throughput + 2wk budget?   → TensorRT-LLM
│  ├─ Multi-LoRA SaaS / multi-tenant?                → vLLM (or SGLang)
│  └─ Default / mixed traffic / new model            → vLLM
└─ No
   ├─ Dev laptop / 5-min setup?                      → Ollama
   ├─ CPU / Apple / edge / single user?              → llama.cpp (or MLX on Mac)
   ├─ Apple production server?                       → MLX or llama.cpp Metal
   └─ Existing TGI that works?                       → stay; plan migration to vLLM
7.4 Common production pattern in 2026
Local dev          → Ollama
Prod benchmark     → vLLM (default)
If prefix-heavy    → A/B with SGLang
If NVIDIA-only +   → A/B with TensorRT-LLM
   eng budget
Edge / on-device   → llama.cpp + GGUF
Apple-native       → MLX

vLLM occupies the default production slot for open-weights LLM serving in 2026; everything else is a workload-justified deviation.


Cited sources (primary)

  • yottalabs.ai/post/best-llm-inference-engines-in-2026-vllm-tensorrt-llm-tgi-and-sglang-compared
  • explore.n1n.ai/blog/llm-inference-engine-comparison-vllm-tgi-tensorrt-sglang-2026-03-13
  • aimadetools.com/blog/vllm-vs-ollama-vs-llamacpp-vs-tgi/
  • contracollective.com/blog/llama-cpp-vs-mlx-ollama-vllm-apple-silicon-2026
  • developers.redhat.com/articles/2025/09/30/vllm-or-llamacpp-choosing-right-llm-inference-engine-your-use-case
  • spheron.network/blog/vllm-vs-tensorrt-llm-vs-sglang-benchmarks
  • arxiv.org/abs/2309.06180 (PagedAttention)
  • arxiv.org/abs/2411.02355 ("Give Me BF16 or Give Me Death", quantization tradeoffs)
  • github.com/vllm-project/vllm (README)
  • sglang.ai/blog

© agentsope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/agentsop-llm-engine-selection of agentsope/SkillAlchemy.

  • SKILL.md
  • README.md
  • intermediate/operation_candidates.json
  • references/R1-source-evidence.md
  • references/R2-decision-flowchart.md

Open the folder on GitHubat commit 6ea799f

Compare with similar skills

Agentsop LLM Engine Selection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agentsop LLM Engine Selection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agentsop LLM Engine Selection this skillagentsope/SkillAlchemy457—~6.1kAutomated safety check: PassMIT
Jetson Inference Mem TuneNVIDIA/skills3.5k1 repos~2.9kAutomated safety check: PassApache-2.0
Jetson LLM BenchmarkNVIDIA/skills3.5k1 repos~3.1kAutomated safety check: PassApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Debug InferenceNVIDIA/OpenShell15k—~1.9kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

    3.5k GitHub starsUsed in 1 repo~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

    3.5k GitHub starsUsed in 1 repo~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Debug Inference

    NVIDIA/OpenShell

    Official

    Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM.

    15k GitHub stars~1.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed

More from agentsope/SkillAlchemy

All 45 skills in this repo
  • Agentsop Aider

    agentsope/SkillAlchemy

    SOP for terminal-based, git-native AI pair programming with Aider (git work-tree + tree-sitter repo-map + edit-format + human-in-loop REPL).

    457 GitHub stars~3.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Context Scope Discipline

    agentsope/SkillAlchemy

    Coder-agent working-file budget discipline: keep the editable working set (files you /add into writable context) under ~25k tokens, separate "read" from "edit", delegate breadth to a read-only…

    457 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Cost Tiered Models

    agentsope/SkillAlchemy

    Split a multi-call LM workflow by cognitive load, not by accuracy: let one strong model make the few reasoning decisions and a cheap model do the many mechanical executions (Aider architect+editor…

    457 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Crewai

    agentsope/SkillAlchemy

    SOP for building multi-agent systems with CrewAI — role-based collaboration, sequential/hierarchical processes, Flows, memory, delegation.

    457 GitHub stars~4.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Agentsop Dify

    agentsope/SkillAlchemy

    SOP for building LLM applications on Dify — visual workflow + chatflow + agent + RAG knowledge base + plugin marketplace + observability, self-hostable.

    457 GitHub stars~5.4k tokensUpdated 1 mo ago
    Auto-check: notes
  • Agentsop Multiscale Chunking

    agentsope/SkillAlchemy

    Designs multiscale chunking for RAG by embedding small units for retrieval precision and returning larger context for synthesis.

    457 GitHub stars~4.9k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Agentsop LLM Engine Selection

What does Agentsop LLM Engine Selection do?

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack. Agentsop LLM Engine Selection is an agent skill from agentsope/SkillAlchemy. Cross-engine decision rubric for self-hosting or recommending an LLM serving stack.

When should I use Agentsop LLM Engine Selection?

Agentsop LLM Engine Selection fits situations like: tasks that involve LLM inference and serving.

How do I install Agentsop LLM Engine Selection in Claude Code?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-llm-engine-selection -a claude-code`. Or copy the skill folder (skills/agentsop-llm-engine-selection in agentsope/SkillAlchemy) into .claude/skills/agentsop-llm-engine-selection in your project. Claude Code loads it when a task matches its description.

How do I install Agentsop LLM Engine Selection in Codex?

Run `npx skills add agentsope/SkillAlchemy --skill agentsop-llm-engine-selection -a codex`. Or copy the skill folder (skills/agentsop-llm-engine-selection in agentsope/SkillAlchemy) into .agents/skills/agentsop-llm-engine-selection in your project. Codex loads it when a task matches its description.

Can I use Agentsop LLM Engine Selection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentsope/SkillAlchemy --skill agentsop-llm-engine-selection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agentsop-llm-engine-selection, .gemini/skills/agentsop-llm-engine-selection, .github/skills/agentsop-llm-engine-selection and .opencode/skills/agentsop-llm-engine-selection in your project.

What does Agentsop LLM Engine Selection need to run?

SKILL.md names no scripts, command-line tools or credentials: Agentsop LLM Engine Selection is instructions for the agent only.

Does Agentsop LLM Engine Selection access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agentsop LLM Engine Selection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agentsop LLM Engine Selection use?

Agentsop LLM Engine Selection is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agentsop LLM Engine Selection use?

About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.

What are the alternatives to Agentsop LLM Engine Selection?

Skills that share tags, products or a category with Agentsop LLM Engine Selection: Jetson Inference Mem Tune (NVIDIA/skills, 3.5k stars), Jetson LLM Benchmark (NVIDIA/skills, 3.5k stars), Aider Delegate (amElnagdy/delegate-skills, 2.3k stars) and Dstack Prototyping (dstackai/dstack, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agentsop LLM Engine Selection?

agentsope (a GitHub user) maintains it in agentsope/SkillAlchemy, which has 457 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on September 2, 2026.

Source: agentsope/SkillAlchemy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.