Agent skill

TensorRT-LLM Inference

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

MITAuto-check passedAI & LLM Engineering

Install TensorRT-LLM Inference

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill tensorrt-llm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs tensorrt-llm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/12-inference-serving/tensorrt-llm .claude/skills/tensorrt-llm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tensorrt-llm
GitHub stars
13k
Used in
4 other repos
Token cost
~1.3k tokens
SKILL.md length
284 words
Files
4 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

  • Deploying a model for production on NVIDIA GPUs with low latency
  • SKILL.md covers When to use TensorRT-LLM, Quick start, Key features and Common patterns, plus 4 more sections
  • Calls docker, pip and curl
  • Serving a quantized FP8 or INT4 model with in-flight batching

What it does

TensorRT-LLM is NVIDIA's open-source library for fast LLM inference on its own GPUs. The skill covers pulling the Docker image, running inference with the Python LLM and SamplingParams classes, and starting a server with trtllm-serve, which downloads and compiles the model automatically and takes a tensor-parallel size. Its feature list names in-flight batching, a paged KV cache, Flash Attention, CUDA graphs, FP8, INT4 and FP4 quantization, tensor, pipeline and expert parallelism, speculative decoding, LoRA serving and disaggregated serving.

Worked patterns show loading an FP8 quantized model, spreading a very large Llama model across eight GPUs and processing a batch of prompts. The skill recommends vLLM for a simpler Python-first setup or AMD hardware, and llama.cpp for CPU, Apple Silicon or edge deployment. Reference files cover multi-GPU use, optimization and serving. The excerpt is cut off before the benchmarks section.

When your agent uses it

  • Deploying a model for production on NVIDIA GPUs with low latency
  • Serving a quantized FP8 or INT4 model with in-flight batching
  • Splitting one large model across several GPUs or nodes
  • Deciding between TensorRT-LLM, vLLM and llama.cpp for a deployment

Example prompts

  • “Start trtllm-serve for Llama 3 8B with tensor parallelism across four GPUs.”
  • “Load an FP8 quantized model with TensorRT-LLM and run a batch of prompts.”
  • “Write a Docker command to run the TensorRT-LLM image with GPU access.”
  • “Should we use TensorRT-LLM or vLLM for our single-node deployment?”

Requirements

  • NVIDIA GPUs
  • Docker, or a TensorRT-LLM installation with Python

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • pip
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • nvidia.github.io
    • github.com
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

TensorRT-LLM Inference loads about 1.3k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 77 tokens; SKILL.md has 284 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 284 words, ~1,259 tokens.

Download SKILL.mdSave it as .claude/skills/tensorrt-llm/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
tensorrt-llm
description
Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
version
1.0.0
author
Orchestra Research
license
MIT
tags
Inference Serving, TensorRT-LLM, NVIDIA, Inference Optimization, High Throughput, Low Latency, Production, FP8, INT4, In-Flight Batching, Multi-GPU
dependencies
tensorrt-llm, torch

TensorRT-LLM

NVIDIA's open-source library for optimizing LLM inference with state-of-the-art performance on NVIDIA GPUs.

When to use TensorRT-LLM

Use TensorRT-LLM when:

  • Deploying on NVIDIA GPUs (A100, H100, GB200)
  • Need maximum throughput (24,000+ tokens/sec on Llama 3)
  • Require low latency for real-time applications
  • Working with quantized models (FP8, INT4, FP4)
  • Scaling across multiple GPUs or nodes

Use vLLM instead when:

  • Need simpler setup and Python-first API
  • Want PagedAttention without TensorRT compilation
  • Working with AMD GPUs or non-NVIDIA hardware

Use llama.cpp instead when:

  • Deploying on CPU or Apple Silicon
  • Need edge deployment without NVIDIA GPUs
  • Want simpler GGUF quantization format

Quick start

Installation
bash
# Docker (recommended)
docker pull nvidia/tensorrt_llm:latest

# pip install
pip install tensorrt_llm==1.2.0rc3

# Requires CUDA 13.0.0, TensorRT 10.13.2, Python 3.10-3.12
Basic inference
python
from tensorrt_llm import LLM, SamplingParams

# Initialize model
llm = LLM(model="meta-llama/Meta-Llama-3-8B")

# Configure sampling
sampling_params = SamplingParams(
    max_tokens=100,
    temperature=0.7,
    top_p=0.9
)

# Generate
prompts = ["Explain quantum computing"]
outputs = llm.generate(prompts, sampling_params)

for output in outputs:
    print(output.text)
Serving with trtllm-serve
bash
# Start server (automatic model download and compilation)
trtllm-serve meta-llama/Meta-Llama-3-8B \
    --tp_size 4 \              # Tensor parallelism (4 GPUs)
    --max_batch_size 256 \
    --max_num_tokens 4096

# Client request
curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3-8B",
    "messages": [{"role": "user", "content": "Hello!"}],
    "temperature": 0.7,
    "max_tokens": 100
  }'

Key features

Performance optimizations
  • In-flight batching: Dynamic batching during generation
  • Paged KV cache: Efficient memory management
  • Flash Attention: Optimized attention kernels
  • Quantization: FP8, INT4, FP4 for 2-4× faster inference
  • CUDA graphs: Reduced kernel launch overhead
Parallelism
  • Tensor parallelism (TP): Split model across GPUs
  • Pipeline parallelism (PP): Layer-wise distribution
  • Expert parallelism: For Mixture-of-Experts models
  • Multi-node: Scale beyond single machine
Advanced features
  • Speculative decoding: Faster generation with draft models
  • LoRA serving: Efficient multi-adapter deployment
  • Disaggregated serving: Separate prefill and generation

Common patterns

Quantized model (FP8)
python
from tensorrt_llm import LLM

# Load FP8 quantized model (2× faster, 50% memory)
llm = LLM(
    model="meta-llama/Meta-Llama-3-70B",
    dtype="fp8",
    max_num_tokens=8192
)

# Inference same as before
outputs = llm.generate(["Summarize this article..."])
Multi-GPU deployment
python
# Tensor parallelism across 8 GPUs
llm = LLM(
    model="meta-llama/Meta-Llama-3-405B",
    tensor_parallel_size=8,
    dtype="fp8"
)
Batch inference
python
# Process 100 prompts efficiently
prompts = [f"Question {i}: ..." for i in range(100)]

outputs = llm.generate(
    prompts,
    sampling_params=SamplingParams(max_tokens=200)
)

# Automatic in-flight batching for maximum throughput

Performance benchmarks

Meta Llama 3-8B (H100 GPU):

  • Throughput: 24,000 tokens/sec
  • Latency: ~10ms per token
  • vs PyTorch: 100× faster

Llama 3-70B (8× A100 80GB):

  • FP8 quantization: 2× faster than FP16
  • Memory: 50% reduction with FP8

Supported models

  • LLaMA family: Llama 2, Llama 3, CodeLlama
  • GPT family: GPT-2, GPT-J, GPT-NeoX
  • Qwen: Qwen, Qwen2, QwQ
  • DeepSeek: DeepSeek-V2, DeepSeek-V3
  • Mixtral: Mixtral-8x7B, Mixtral-8x22B
  • Vision: LLaVA, Phi-3-vision
  • 100+ models on HuggingFace

References

Resources

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in 12-inference-serving/tensorrt-llm of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/multi-gpu.md
  • references/optimization.md
  • references/serving.md

Open the folder on GitHubat commit 773a529

Used in 4 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 4 other GitHub owners. This page covers the copy in Orchestra-Research/AI-Research-SKILLs, which our catalogue first saw on October 7, 2026.

Compare with similar skills

TensorRT-LLM Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

TensorRT-LLM Inference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
TensorRT-LLM Inference this skillOrchestra-Research/AI-Research-SKILLs13k4 repos~1.3kAutomated safety check: PassMIT
Dstack Presetsdstackai/dstack2.3k—~403Automated safety check: PassMPL-2.0
Dstackdstackai/dstack2.3k—~6.2kAutomated safety check: WarnMPL-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNone

Similar skills

  • Dstack Presets

    dstackai/dstack

    Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

    2.3k GitHub stars~403 tokensUpdated today
    DevOps & CloudAuto-check passed
  • Dstack

    dstackai/dstack

    dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.

    2.3k GitHub stars~6.2k tokensUpdated today
    DevOps & CloudAuto-check: warnings
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Serving Framework Benchmark

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

    925 GitHub stars~7.5k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Questions about TensorRT-LLM Inference

What does TensorRT-LLM Inference do?

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. TensorRT-LLM is NVIDIA's open-source library for fast LLM inference on its own GPUs. The skill covers pulling the Docker image, running inference with the Python LLM and SamplingParams classes, and starting a server with trtllm-serve, which downloads and compiles the model automatically and takes a tensor-parallel size.

When should I use TensorRT-LLM Inference?

TensorRT-LLM Inference fits situations like: deploying a model for production on NVIDIA GPUs with low latency; serving a quantized FP8 or INT4 model with in-flight batching; splitting one large model across several GPUs or nodes; deciding between TensorRT-LLM, vLLM and llama.cpp for a deployment.

How do I install TensorRT-LLM Inference in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill tensorrt-llm -a claude-code`. Or copy the skill folder (12-inference-serving/tensorrt-llm in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/tensorrt-llm in your project. Claude Code loads it when a task matches its description.

How do I install TensorRT-LLM Inference in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill tensorrt-llm -a codex`. Or copy the skill folder (12-inference-serving/tensorrt-llm in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/tensorrt-llm in your project. Codex loads it when a task matches its description.

Can I use TensorRT-LLM Inference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill tensorrt-llm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tensorrt-llm, .gemini/skills/tensorrt-llm, .github/skills/tensorrt-llm and .opencode/skills/tensorrt-llm in your project.

What does TensorRT-LLM Inference need to run?

Going by SKILL.md and its folder, TensorRT-LLM Inference needs the command-line tools its instructions call (docker, pip and curl). Our summary lists: NVIDIA GPUs; Docker, or a TensorRT-LLM installation with Python.

Does TensorRT-LLM Inference access the network?

SKILL.md names 3 domains. As links in the text: nvidia.github.io, github.com and huggingface.co. This is read from the text; nothing was executed.

Is TensorRT-LLM Inference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does TensorRT-LLM Inference use?

TensorRT-LLM Inference is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does TensorRT-LLM Inference use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.

What are the alternatives to TensorRT-LLM Inference?

Skills that share tags, products or a category with TensorRT-LLM Inference: Dstack Presets (dstackai/dstack, 2.3k stars), Dstack (dstackai/dstack, 2.3k stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains TensorRT-LLM Inference?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,374 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.