Agent skill

Tensorrt LLM

by Luciole-Studio in Luciole-Studio/Misaka-Agent

High-throughput LLM inference on NVIDIA GPUs. An agent skill from Luciole-Studio/Misaka-Agent.

MITAuto-check passedAI & LLM Engineering

Install Tensorrt LLM

skills CLI
$ npx skills add Luciole-Studio/Misaka-Agent --skill tensorrt-llm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Luciole-Studio/Misaka-Agent tensorrt-llm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Luciole-Studio/Misaka-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/misaka/core/skills/assets/optional/mlops/tensorrt-llm .claude/skills/tensorrt-llm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tensorrt-llm
GitHub stars
139
Used in
1 other repo
Token cost
~1.3k tokens
SKILL.md length
284 words
Files
4 (incl. references)
Skills in repo
76
Repo updated
First seen
Licence
MIT

At a glance

High-throughput LLM inference on NVIDIA GPUs. An agent skill from Luciole-Studio/Misaka-Agent.

  • Tasks that involve LLM inference and serving
  • SKILL.md covers When to use TensorRT-LLM, Quick start, Key features and Common patterns, plus 4 more sections
  • Calls docker, pip and curl; reaches catalog.ngc.nvidia.com
  • Tasks that involve MLOps

What it does

Tensorrt LLM is an agent skill from Luciole-Studio/Misaka-Agent. High-throughput LLM inference on NVIDIA GPUs.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/multi-gpu.md`, `references/optimization.md` and `references/serving.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving and MLOps. It works with NVIDIA AI Platform and DeepSeek. The repository describes itself as: A multi-agent research system for the humanities and social sciences. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve MLOps

Example prompts

  • “/tensorrt-llm”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit b94464a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • pip
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • catalog.ngc.nvidia.com

    Also links to:

    • nvidia.github.io
    • github.com
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tensorrt LLM loads about 1.3k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 15 tokens; SKILL.md has 284 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~15
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Luciole-Studio/Misaka-Agent at commit b94464a, republished under its MIT licence (© Luciole-Studio). 284 words, ~1,265 tokens.

Download SKILL.mdSave it as .claude/skills/tensorrt-llm/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
tensorrt-llm
description
High-throughput LLM inference on NVIDIA GPUs.
version
1.0.1
author
Orchestra Research
license
MIT
dependencies
tensorrt-llm, torch
platforms
linux, macos

TensorRT-LLM

NVIDIA's open-source library for optimizing LLM inference with high performance on NVIDIA GPUs.

When to use TensorRT-LLM

Use TensorRT-LLM when:

  • Deploying on NVIDIA GPUs (A100, H100, GB200)
  • Need maximum throughput (24,000+ tokens/sec on Llama 3)
  • Require low latency for real-time applications
  • Working with quantized models (FP8, INT4, FP4)
  • Scaling across multiple GPUs or nodes

Use vLLM instead when:

  • Need simpler setup and Python-first API
  • Want PagedAttention without TensorRT compilation
  • Working with AMD GPUs or non-NVIDIA hardware

Use llama.cpp instead when:

  • Deploying on CPU or Apple Silicon
  • Need edge deployment without NVIDIA GPUs
  • Want simpler GGUF quantization format

Quick start

Installation
bash
# Docker (recommended) — images are on NGC (nvcr.io), not Docker Hub.
# Replace x.y.z with the desired version (e.g. 1.2.1). Browse tags on NGC:
# https://catalog.ngc.nvidia.com/orgs/nvidia/teams/tensorrt-llm/containers/release/tags
docker pull nvcr.io/nvidia/tensorrt-llm/release:x.y.z

# pip install (current stable GA)
pip install tensorrt_llm

# Requires CUDA 13.2.1, TensorRT 10.x, Python 3.10-3.12
Basic inference
python
from tensorrt_llm import LLM, SamplingParams

# Initialize model
llm = LLM(model="meta-llama/Meta-Llama-3-8B")

# Configure sampling
sampling_params = SamplingParams(
    max_tokens=100,
    temperature=0.7,
    top_p=0.9
)

# Generate
prompts = ["Explain quantum computing"]
outputs = llm.generate(prompts, sampling_params)

for output in outputs:
    print(output.text)
Serving with trtllm-serve
bash
# Start server (automatic model download and compilation)
trtllm-serve meta-llama/Meta-Llama-3-8B \
    --tp_size 4 \              # Tensor parallelism (4 GPUs)
    --max_batch_size 256 \
    --max_num_tokens 4096

# Client request
curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3-8B",
    "messages": [{"role": "user", "content": "Hello!"}],
    "temperature": 0.7,
    "max_tokens": 100
  }'

Key features

Performance optimizations
  • In-flight batching: Dynamic batching during generation
  • Paged KV cache: Efficient memory management
  • Flash Attention: Optimized attention kernels
  • Quantization: FP8, INT4, FP4 for 2-4× faster inference
  • CUDA graphs: Reduced kernel launch overhead
Parallelism
  • Tensor parallelism (TP): Split model across GPUs
  • Pipeline parallelism (PP): Layer-wise distribution
  • Expert parallelism: For Mixture-of-Experts models
  • Multi-node: Scale beyond single machine
Advanced features
  • Speculative decoding: Faster generation with draft models
  • LoRA serving: Efficient multi-adapter deployment
  • Disaggregated serving: Separate prefill and generation

Common patterns

Quantized model (FP8)
python
from tensorrt_llm import LLM

# Load FP8 quantized model (2× faster, 50% memory)
llm = LLM(
    model="meta-llama/Meta-Llama-3-70B",
    dtype="fp8",
    max_num_tokens=8192
)

# Inference same as before
outputs = llm.generate(["Summarize this article..."])
Multi-GPU deployment
python
# Tensor parallelism across 8 GPUs
llm = LLM(
    model="meta-llama/Meta-Llama-3-405B",
    tensor_parallel_size=8,
    dtype="fp8"
)
Batch inference
python
# Process 100 prompts efficiently
prompts = [f"Question {i}: ..." for i in range(100)]

outputs = llm.generate(
    prompts,
    sampling_params=SamplingParams(max_tokens=200)
)

# Automatic in-flight batching for maximum throughput

Performance benchmarks

Meta Llama 3-8B (H100 GPU):

  • Throughput: 24,000 tokens/sec
  • Latency: ~10ms per token
  • vs PyTorch: 100× faster

Llama 3-70B (8× A100 80GB):

  • FP8 quantization: 2× faster than FP16
  • Memory: 50% reduction with FP8

Supported models

  • LLaMA family: Llama 2, Llama 3, CodeLlama
  • GPT family: GPT-2, GPT-J, GPT-NeoX
  • Qwen: Qwen, Qwen2, QwQ
  • DeepSeek: DeepSeek-V2, DeepSeek-V3
  • Mixtral: Mixtral-8x7B, Mixtral-8x22B
  • Vision: LLaVA, Phi-3-vision
  • 100+ models on HuggingFace

References

Resources

© Luciole-Studio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in misaka/core/skills/assets/optional/mlops/tensorrt-llm of Luciole-Studio/Misaka-Agent.

  • SKILL.md
  • references/multi-gpu.md
  • references/optimization.md
  • references/serving.md

Open the folder on GitHubat commit b94464a

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Luciole-Studio/Misaka-Agent, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Tensorrt LLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tensorrt LLM compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tensorrt LLM this skillLuciole-Studio/Misaka-Agent1391 repos~1.3kAutomated safety check: PassMIT
Serving LLMs On Instinctamd/skills398—~4kAutomated safety check: NotesMIT
SGLang Model Day-0 SupportBBuf/AI-Infra-Auto-Driven-SKILLS911—~2.3kAutomated safety check: PassNone
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~2.8kAutomated safety check: PassNone

Similar skills

  • Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

    398 GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • SGLang Model Day-0 Support

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

    911 GitHub stars~2.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    333 GitHub stars~4.2k tokensUpdated 28 days ago
    AI & LLM EngineeringAuto-check: notes

More from Luciole-Studio/Misaka-Agent

All 76 skills in this repo
  • Kanban Video Orchestrator

    Luciole-Studio/Misaka-Agent

    Plan and run multi-agent video production pipelines. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check: notes
  • Ast Grep

    Luciole-Studio/Misaka-Agent

    AST-aware structural code search and rewrite via ast-grep. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Drug Discovery

    Luciole-Studio/Misaka-Agent

    Drug discovery: ChEMBL search, drug-likeness, interactions. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Fitness Nutrition

    Luciole-Studio/Misaka-Agent

    Workout planning, macros, and body metrics via wger/USDA. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Hyperframes

    Luciole-Studio/Misaka-Agent

    Render MP4/WebM videos from HTML compositions. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check passed
  • Osint Investigation

    Luciole-Studio/Misaka-Agent

    Follow the money via public records and sanctions data. An agent skill from Luciole-Studio/Misaka-Agent.

    139 GitHub starsUsed in 1 repo~2.9k tokens
    Auto-check passed

Questions about Tensorrt LLM

What does Tensorrt LLM do?

High-throughput LLM inference on NVIDIA GPUs. An agent skill from Luciole-Studio/Misaka-Agent. Tensorrt LLM is an agent skill from Luciole-Studio/Misaka-Agent. High-throughput LLM inference on NVIDIA GPUs.

When should I use Tensorrt LLM?

Tensorrt LLM fits situations like: tasks that involve LLM inference and serving; tasks that involve MLOps.

How do I install Tensorrt LLM in Claude Code?

Run `npx skills add Luciole-Studio/Misaka-Agent --skill tensorrt-llm -a claude-code`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/tensorrt-llm in Luciole-Studio/Misaka-Agent) into .claude/skills/tensorrt-llm in your project. Claude Code loads it when a task matches its description.

How do I install Tensorrt LLM in Codex?

Run `npx skills add Luciole-Studio/Misaka-Agent --skill tensorrt-llm -a codex`. Or copy the skill folder (misaka/core/skills/assets/optional/mlops/tensorrt-llm in Luciole-Studio/Misaka-Agent) into .agents/skills/tensorrt-llm in your project. Codex loads it when a task matches its description.

Can I use Tensorrt LLM in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Luciole-Studio/Misaka-Agent --skill tensorrt-llm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tensorrt-llm, .gemini/skills/tensorrt-llm, .github/skills/tensorrt-llm and .opencode/skills/tensorrt-llm in your project.

What does Tensorrt LLM need to run?

Going by SKILL.md and its folder, Tensorrt LLM needs the command-line tools its instructions call (docker, pip and curl). Our summary lists: Python 3; Docker.

Does Tensorrt LLM access the network?

SKILL.md names 4 domains. In commands or code: catalog.ngc.nvidia.com; the agent is likely to contact it when it follows the instructions. As links in the text: nvidia.github.io, github.com and huggingface.co. This is read from the text; nothing was executed.

Is Tensorrt LLM safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tensorrt LLM use?

Tensorrt LLM is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tensorrt LLM use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.

What are the alternatives to Tensorrt LLM?

Skills that share tags, products or a category with Tensorrt LLM: Serving LLMs On Instinct (amd/skills, 398 stars), SGLang Model Day-0 Support (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), Dstack Prototyping (dstackai/dstack, 2.3k stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tensorrt LLM?

Luciole-Studio (a GitHub organization) maintains it in Luciole-Studio/Misaka-Agent, which has 139 GitHub stars. The repository holds 76 skills in this directory. The repository was last updated on October 8, 2026.

Source: Luciole-Studio/Misaka-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.