Agent skill

Tensorrt Optimization

by majiayu000 in majiayu000/claude-skill-registry

NVIDIA TensorRT model optimization and deployment. An agent skill from majiayu000/claude-skill-registry.

MITAuto-check: notesAI & LLM Engineering

Install Tensorrt Optimization

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill tensorrt-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry tensorrt-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-ml/tensorrt-optimization .claude/skills/tensorrt-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tensorrt-optimization
GitHub stars
666
Used in
1 other repo
Token cost
~2.4k tokens
SKILL.md length
189 words
Files
2
Skills in repo
971
Repo updated
First seen
Licence
MIT

At a glance

NVIDIA TensorRT model optimization and deployment. An agent skill from majiayu000/claude-skill-registry.

  • Works in 8 steps: Model Conversion to TensorRT → Precision Configuration → INT8 Calibration → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Overview, Prerequisites, Capabilities and Command Line Tools, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tensorrt Optimization is an agent skill from majiayu000/claude-skill-registry. NVIDIA TensorRT model optimization and deployment. Convert models to TensorRT engines, configure optimization profiles and precision modes, apply INT8 calibration, analyze kernel fusion, generate custom plugins, and profile inference performance.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in AI & LLM Engineering, covering LLM inference and serving and Performance reviews. It works with NVIDIA AI Platform. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve Performance reviews

Example prompts

  • “/tensorrt-optimization”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash(*), Read, Write, Edit, Glob, Grep, WebFetch

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Model Conversion to TensorRT
  2. Precision Configuration
  3. INT8 Calibration
  4. Dynamic Shapes
  5. Inference Execution
  6. Plugin Development
  7. Performance Profiling
  8. Kernel Fusion Analysis

What it can do on your machine

Read from SKILL.md and the folder at commit 000116a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(*)
    • Read
    • Write
    • Edit
    • Glob
    • Grep
    • WebFetch

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, bash, cpp and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tensorrt Optimization loads about 2.4k tokens when it runs. Until then it costs about 67 tokens; SKILL.md has 189 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash(*), Read, Write, Edit, Glob, Grep, WebFetch

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 000116a, republished under its MIT licence (© majiayu000). 189 words, ~2,351 tokens.

Download SKILL.mdSave it as .claude/skills/tensorrt-optimization/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
tensorrt-optimization
description
NVIDIA TensorRT model optimization and deployment. Convert models to TensorRT engines, configure optimization profiles and precision modes, apply INT8 calibration, analyze kernel fusion, generate custom plugins, and profile inference performance.
allowed-tools
Bash(*), Read, Write, Edit, Glob, Grep, WebFetch
metadata.author
babysitter-sdk
metadata.version
1.0.0
metadata.category
ml-inference
metadata.backlog-id
SK-008

tensorrt-optimization

You are tensorrt-optimization - a specialized skill for NVIDIA TensorRT model optimization and deployment. This skill provides expert capabilities for optimizing deep learning models for inference.

Overview

This skill enables AI-powered TensorRT optimization including:

  • Convert models to TensorRT engines
  • Configure optimization profiles and precision modes
  • Apply INT8 calibration and quantization
  • Analyze kernel fusion opportunities
  • Generate custom TensorRT plugins
  • Profile inference latency and throughput
  • Handle dynamic shapes and batch sizes
  • Compare TensorRT vs framework inference

Prerequisites

  • TensorRT 8.5+
  • CUDA Toolkit 11.0+
  • ONNX Runtime (for ONNX models)
  • Python TensorRT package

Capabilities

1. Model Conversion to TensorRT

Convert models from various frameworks:

python
import tensorrt as trt

# Create builder and network
logger = trt.Logger(trt.Logger.WARNING)
builder = trt.Builder(logger)
network = builder.create_network(
    1 << int(trt.NetworkDefinitionCreationFlag.EXPLICIT_BATCH))

# Parse ONNX model
parser = trt.OnnxParser(network, logger)
with open("model.onnx", "rb") as f:
    parser.parse(f.read())

# Configure builder
config = builder.create_builder_config()
config.set_memory_pool_limit(trt.MemoryPoolType.WORKSPACE, 1 << 30)  # 1GB

# Build engine
engine = builder.build_serialized_network(network, config)

# Save engine
with open("model.engine", "wb") as f:
    f.write(engine)
2. Precision Configuration

Configure FP16, INT8, and TF32:

python
# Enable FP16
config.set_flag(trt.BuilderFlag.FP16)

# Enable INT8 (requires calibration)
config.set_flag(trt.BuilderFlag.INT8)

# Enable TF32 (Ampere+)
config.clear_flag(trt.BuilderFlag.TF32)  # Disable if needed

# Enable sparse tensor cores
config.set_flag(trt.BuilderFlag.SPARSE_WEIGHTS)

# Prefer precision per layer
config.set_flag(trt.BuilderFlag.PREFER_PRECISION_CONSTRAINTS)

# Force strict types
config.set_flag(trt.BuilderFlag.STRICT_TYPES)
3. INT8 Calibration
python
class Calibrator(trt.IInt8EntropyCalibrator2):
    def __init__(self, data_loader, cache_file):
        super().__init__()
        self.data_loader = iter(data_loader)
        self.cache_file = cache_file
        self.batch_size = data_loader.batch_size
        self.device_input = cuda.mem_alloc(
            self.batch_size * 3 * 224 * 224 * 4)

    def get_batch_size(self):
        return self.batch_size

    def get_batch(self, names):
        try:
            batch = next(self.data_loader)
            cuda.memcpy_htod(self.device_input, batch.numpy())
            return [int(self.device_input)]
        except StopIteration:
            return None

    def read_calibration_cache(self):
        if os.path.exists(self.cache_file):
            with open(self.cache_file, "rb") as f:
                return f.read()
        return None

    def write_calibration_cache(self, cache):
        with open(self.cache_file, "wb") as f:
            f.write(cache)

# Use calibrator
calibrator = Calibrator(calibration_loader, "calibration.cache")
config.int8_calibrator = calibrator
config.set_flag(trt.BuilderFlag.INT8)
4. Dynamic Shapes

Handle variable input sizes:

python
# Create optimization profile
profile = builder.create_optimization_profile()

# Define shape ranges [min, optimal, max]
profile.set_shape("input",
    min=(1, 3, 224, 224),     # Minimum shape
    opt=(8, 3, 224, 224),     # Optimal shape
    max=(32, 3, 224, 224))    # Maximum shape

config.add_optimization_profile(profile)

# Multiple profiles for different scenarios
profile_small = builder.create_optimization_profile()
profile_small.set_shape("input", (1, 3, 224, 224), (4, 3, 224, 224), (8, 3, 224, 224))
config.add_optimization_profile(profile_small)

profile_large = builder.create_optimization_profile()
profile_large.set_shape("input", (16, 3, 224, 224), (32, 3, 224, 224), (64, 3, 224, 224))
config.add_optimization_profile(profile_large)
5. Inference Execution
python
# Load engine
runtime = trt.Runtime(logger)
with open("model.engine", "rb") as f:
    engine = runtime.deserialize_cuda_engine(f.read())

# Create execution context
context = engine.create_execution_context()

# Set input shape for dynamic shapes
context.set_input_shape("input", (batch_size, 3, 224, 224))

# Allocate buffers
inputs = []
outputs = []
bindings = []

for i in range(engine.num_io_tensors):
    name = engine.get_tensor_name(i)
    dtype = trt.nptype(engine.get_tensor_dtype(name))
    shape = context.get_tensor_shape(name)
    size = trt.volume(shape)

    buffer = cuda.mem_alloc(size * dtype.itemsize)
    bindings.append(int(buffer))

    if engine.get_tensor_mode(name) == trt.TensorIOMode.INPUT:
        inputs.append(buffer)
    else:
        outputs.append(buffer)

# Execute inference
cuda.memcpy_htod(inputs[0], input_data)
context.execute_v2(bindings)
cuda.memcpy_dtoh(output_data, outputs[0])
6. Plugin Development

Create custom operations:

cpp
// Plugin class
class CustomPlugin : public nvinfer1::IPluginV2DynamicExt {
public:
    int getNbOutputs() const noexcept override { return 1; }

    nvinfer1::DimsExprs getOutputDimensions(
        int outputIndex,
        const nvinfer1::DimsExprs* inputs,
        int nbInputs,
        nvinfer1::IExprBuilder& exprBuilder) noexcept override {
        return inputs[0];  // Same shape as input
    }

    int enqueue(
        const nvinfer1::PluginTensorDesc* inputDesc,
        const nvinfer1::PluginTensorDesc* outputDesc,
        const void* const* inputs,
        void* const* outputs,
        void* workspace,
        cudaStream_t stream) noexcept override {
        // Launch custom CUDA kernel
        customKernel<<<blocks, threads, 0, stream>>>(
            inputs[0], outputs[0], inputDesc[0].dims);
        return 0;
    }
};

// Register plugin
REGISTER_TENSORRT_PLUGIN(CustomPluginCreator);
7. Performance Profiling
python
# Enable profiling
config.profiling_verbosity = trt.ProfilingVerbosity.DETAILED

# Use timing cache for faster builds
timing_cache_file = "timing.cache"
if os.path.exists(timing_cache_file):
    with open(timing_cache_file, "rb") as f:
        cache = config.create_timing_cache(f.read())
else:
    cache = config.create_timing_cache(b"")
config.set_timing_cache(cache, ignore_mismatch=False)

# Profile inference
profiler = trt.Profiler()
context.profiler = profiler

# Benchmark
import time
warmup = 10
iterations = 100

for _ in range(warmup):
    context.execute_v2(bindings)
cuda.Context.synchronize()

start = time.perf_counter()
for _ in range(iterations):
    context.execute_v2(bindings)
cuda.Context.synchronize()
end = time.perf_counter()

latency = (end - start) / iterations * 1000
throughput = batch_size * iterations / (end - start)
print(f"Latency: {latency:.2f} ms, Throughput: {throughput:.2f} samples/s")
8. Kernel Fusion Analysis
bash
# Use trtexec for analysis
trtexec --onnx=model.onnx \
    --fp16 \
    --workspace=4096 \
    --verbose \
    --dumpLayerInfo \
    --exportLayerInfo=layers.json

# Profile with Nsight Systems
nsys profile -o trt_profile \
    trtexec --loadEngine=model.engine --iterations=100

# View layer timing
trtexec --loadEngine=model.engine \
    --dumpProfile \
    --separateProfileRun

Command Line Tools

bash
# Convert ONNX to TensorRT
trtexec --onnx=model.onnx --saveEngine=model.engine

# With FP16
trtexec --onnx=model.onnx --fp16 --saveEngine=model_fp16.engine

# With INT8 calibration
trtexec --onnx=model.onnx --int8 \
    --calib=calibration.cache --saveEngine=model_int8.engine

# Dynamic shapes
trtexec --onnx=model.onnx \
    --minShapes=input:1x3x224x224 \
    --optShapes=input:8x3x224x224 \
    --maxShapes=input:32x3x224x224 \
    --saveEngine=model_dynamic.engine

# Benchmark existing engine
trtexec --loadEngine=model.engine \
    --iterations=1000 \
    --warmUp=500 \
    --duration=10

Process Integration

This skill integrates with the following processes:

  • ml-inference-optimization.js - ML inference optimization
  • tensor-core-programming.js - Tensor core usage

Output Format

json
{
  "operation": "build-engine",
  "status": "success",
  "input_model": "model.onnx",
  "output_engine": "model.engine",
  "configuration": {
    "precision": ["FP16", "INT8"],
    "workspace_mb": 1024,
    "dynamic_shapes": true
  },
  "optimization": {
    "layer_fusions": 23,
    "reformats_eliminated": 8,
    "tactics_selected": 156
  },
  "performance": {
    "build_time_s": 45.2,
    "engine_size_mb": 28.5,
    "estimated_latency_ms": 1.2
  }
}

Dependencies

  • TensorRT 8.5+
  • CUDA Toolkit 11.0+
  • ONNX Runtime (optional)
  • Python tensorrt package

Constraints

  • INT8 requires representative calibration data
  • Dynamic shapes increase build time
  • Custom plugins need careful memory management
  • Engine files are GPU-architecture specific

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ai-ml/tensorrt-optimization of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 000116a

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Tensorrt Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tensorrt Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tensorrt Optimization this skillmajiayu000/claude-skill-registry6661 repos~2.4kAutomated safety check: NotesMIT
Astreawarpfront/hipfire652—~2.6kAutomated safety check: PassCustom licence
Nemotron Nano3NVIDIA-NeMo/Nemotron2.1k—~1.9kAutomated safety check: PassApache-2.0
Nemotron Super3NVIDIA-NeMo/Nemotron2.1k—~2.4kAutomated safety check: PassApache-2.0
Add Vlm Modelintel/auto-round1.6k—~2.4kAutomated safety check: PassApache-2.0
SGLang Structured ServingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.9kAutomated safety check: PassMIT

Similar skills

  • Astrea

    warpfront/hipfire

    A skill your agent uses for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform…

    652 GitHub stars~2.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Nemotron Nano3

    NVIDIA-NeMo/Nemotron

    Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment.

    2.1k GitHub stars~1.9k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Nemotron Super3

    NVIDIA-NeMo/Nemotron

    Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes.

    2.1k GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Add Vlm Model

    intel/auto-round

    Official

    Add support for a new Vision-Language Model (VLM) to AutoRound, including multimodal block handler, calibration dataset template, and special model handling.

    1.6k GitHub stars~2.4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • SGLang Structured Serving

    Orchestra-Research/AI-Research-SKILLs

    Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads.

    13k GitHub starsUsed in 3 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Nemotron Ultra

    NVIDIA-NeMo/Nemotron

    Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference.

    2.1k GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from majiayu000/claude-skill-registry

All 971 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed
  • Open Notebook

    majiayu000/claude-skill-registry

    Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis.

    666 GitHub starsUsed in 4 repos~2.4k tokens
    Auto-check passed

Questions about Tensorrt Optimization

What does Tensorrt Optimization do?

NVIDIA TensorRT model optimization and deployment. An agent skill from majiayu000/claude-skill-registry. Tensorrt Optimization is an agent skill from majiayu000/claude-skill-registry. NVIDIA TensorRT model optimization and deployment.

When should I use Tensorrt Optimization?

Tensorrt Optimization fits situations like: tasks that involve LLM inference and serving; tasks that involve Performance reviews.

How do I install Tensorrt Optimization in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill tensorrt-optimization -a claude-code`. Or copy the skill folder (skills/ai-ml/tensorrt-optimization in majiayu000/claude-skill-registry) into .claude/skills/tensorrt-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Tensorrt Optimization in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill tensorrt-optimization -a codex`. Or copy the skill folder (skills/ai-ml/tensorrt-optimization in majiayu000/claude-skill-registry) into .agents/skills/tensorrt-optimization in your project. Codex loads it when a task matches its description.

Can I use Tensorrt Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill tensorrt-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tensorrt-optimization, .gemini/skills/tensorrt-optimization, .github/skills/tensorrt-optimization and .opencode/skills/tensorrt-optimization in your project.

What does Tensorrt Optimization need to run?

SKILL.md names no scripts, command-line tools or credentials: Tensorrt Optimization is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(*), Read, Write, Edit, Glob, Grep, WebFetch.

Does Tensorrt Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tensorrt Optimization safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tensorrt Optimization use?

Tensorrt Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tensorrt Optimization use?

About 2.4k tokens (SKILL.md is roughly 9.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tensorrt Optimization?

Skills that share tags, products or a category with Tensorrt Optimization: Astrea (warpfront/hipfire, 652 stars), Nemotron Nano3 (NVIDIA-NeMo/Nemotron, 2.1k stars), Nemotron Super3 (NVIDIA-NeMo/Nemotron, 2.1k stars) and Add Vlm Model (intel/auto-round, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tensorrt Optimization?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 971 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.