Official agent skill

Jetson Speculative Decoding

by NVIDIA in NVIDIA/skills

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Jetson Speculative Decoding

skills CLI
$ npx skills add NVIDIA/skills --skill jetson-speculative-decoding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills jetson-speculative-decoding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/jetson-speculative-decoding .claude/skills/jetson-speculative-decoding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
jetson-speculative-decoding
GitHub stars
3.5k
Used in
1 other repo
Token cost
~1.2k tokens
SKILL.md length
572 words
Files
5
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

  • Works in 3 steps: Run jetson-llm-benchmark (vLLM path) at… → Acceptance: target ≥30% improvement in… → If improvement is <10%, or…
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Purpose, When to use, When NOT to use and Prerequisites, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Jetson Speculative Decoding is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `BENCHMARK.md`, `evals/evals.json` and `skill-card.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving and GPU and accelerator computing. It works with NVIDIA AI Platform and vLLM. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM inference and serving
  • Tasks that involve GPU and accelerator computing

Example prompts

  • “/jetson-speculative-decoding”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Run jetson-llm-benchmark (vLLM path) at --concurrency 1 before and after enabling speculation.
  2. Acceptance: target ≥30% improvement in throughput_tok_s and ≥20% drop in tpot_ms_p50 at concurrency 1.
  3. If improvement is <10%, or throughput_tok_s regresses at concurrency 8, disable speculation. The draft model is costing more than it…

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • jetson-ai-lab.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Jetson Speculative Decoding loads about 1.2k tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 572 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 572 words, ~1,205 tokens.

Download SKILL.mdSave it as .claude/skills/jetson-speculative-decoding/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
jetson-speculative-decoding
description
Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.
version
0.0.1
license
Apache-2.0
metadata.author
Jetson Team
metadata.tags
jetson, llm, speculative-decoding
metadata.languages
markdown
metadata.data-classification
public

Jetson Speculative Decoding (vLLM)

Speculative decoding lets a small "draft" model propose tokens that the target model verifies in a single forward pass, reducing per-token latency. On Jetson, the win/loss is dominated by VRAM headroom, not by the draft quality. This skill encodes the parts an LLM won't already know.

Purpose

Tune an existing Jetson vLLM deployment for faster token generation by appending the right --speculative-config and validating whether it improves single-stream decode speed.

When to use

  • TPOT/ITL is the bottleneck (TTFT is fine, output is just slow).
  • Workload is single-stream or low-concurrency (≤2). Speculation usually loses at high concurrency.
  • Jetson family is Thor or AGX Orin. Do not suggest EAGLE-3 on Orin Nano/NX — there is rarely enough VRAM headroom to host both target and draft, and you'll OOM at startup.

When NOT to use

  • High-concurrency serving (≥8): batched decode usually beats speculation; the draft model just steals VRAM.
  • Models without a published EAGLE-3 head — do not train one ad-hoc as a "fix".
  • After applying jetson-inference-mem-tune flags that already pushed --gpu-memory-utilization near the ceiling. Free at least ~2 GB first.

Prerequisites

  • A working vLLM server recipe from jetson-llm-serve.
  • Enough memory headroom for the draft model or EAGLE-3 head in addition to the target model.
  • A benchmark baseline from jetson-llm-benchmark before enabling speculation.
  • A target model with a compatible EAGLE-3 head, or a small same-family draft model for the fallback path.

Instructions

Append --speculative-config to the vllm serve command shown in jetson-llm-serve.

EAGLE-3 (preferred when a head is published for the target model):

bash
--speculative-config '{
  "method": "eagle3",
  "model": "<eagle3-head-repo-id>",
  "num_speculative_tokens": 5,
  "draft_tensor_parallel_size": 1
}'

Draft-model (fallback — pair a small same-family model):

bash
--speculative-config '{
  "method": "draft_model",
  "model": "<small-draft-model-repo-id>",
  "num_speculative_tokens": 4,
  "draft_tensor_parallel_size": 1
}'
Jetson-specific tuning rules
  • num_speculative_tokens: start at 5 on Thor, 3 on AGX Orin. Higher values pay off only if the draft acceptance rate is >0.6.
  • Always pair with the same vLLM runtime path used by jetson-llm-serve: upstream vLLM 0.20+ (vllm/vllm-openai:latest) or validated native vLLM 0.20+ on Thor, upstream vLLM 0.20+ on Orin JetPack 7.2 / L4T r39+, or the NVIDIA-AI-IOT vLLM image on older Orin. Do not use an Orin NVIDIA-AI-IOT vLLM image on Thor. Older runtimes may lack EAGLE-3 or the current --speculative-config shape.
  • Drop --gpu-memory-utilization by ~0.05 vs the non-speculative baseline to give the draft model headroom.
Show full SKILL.md (215 more words)Show less

How to verify it actually helped

  1. Run jetson-llm-benchmark (vLLM path) at --concurrency 1 before and after enabling speculation.
  2. Acceptance: target ≥30% improvement in throughput_tok_s and ≥20% drop in tpot_ms_p50 at concurrency 1.
  3. If improvement is <10%, or throughput_tok_s regresses at concurrency 8, disable speculation. The draft model is costing more than it returns.

Limitations

  • Speculative decoding improves decode-heavy workloads; it does not reduce TTFT-dominated latency.
  • High concurrency can erase the benefit because continuous batching already keeps the GPU busy.
  • Orin Nano/NX usually lack enough memory headroom for both target and draft models.
  • Acceptance rate and draft overhead are model-specific, so benchmark before and after instead of assuming a speedup.

Error handling

  • If vLLM rejects --speculative-config, verify that Thor and Orin JetPack 7.2 / L4T r39+ are using vLLM 0.20+ and that older Orin is using a JetPack-matched NVIDIA-AI-IOT vLLM image; then switch back to the non-speculative serving command if the runtime still rejects it.
  • If startup OOMs, lower --gpu-memory-utilization, use a smaller draft, or disable speculation and hand off to jetson-inference-mem-tune.
  • If benchmark throughput regresses, remove --speculative-config; a bad draft path is worse than no speculation.

Hand off to

  • jetson-llm-benchmark to quantify the change.
  • jetson-inference-mem-tune if startup OOMs after enabling speculation.

Source

vLLM speculative decoding docs and the Jetson AI Lab GenAI tutorial.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/jetson-speculative-decoding of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 0e0d506

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NVIDIA/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Jetson Speculative Decoding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Jetson Speculative Decoding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Jetson Speculative Decoding this skillNVIDIA/skills3.5k1 repos~1.2kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS900—~2.8kAutomated safety check: PassNone
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Ascend Model Adapter for vLLMvllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    900 GitHub stars~2.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ascend Model Adapter for vLLM

    vllm-project/vllm-ascend

    Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

    2.9k GitHub stars~2.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • TensorRT-LLM Inference

    Orchestra-Research/AI-Research-SKILLs

    Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

    13k GitHub starsUsed in 5 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Jetson Speculative Decoding

What does Jetson Speculative Decoding do?

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck. Jetson Speculative Decoding is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

When should I use Jetson Speculative Decoding?

Jetson Speculative Decoding fits situations like: tasks that involve LLM inference and serving; tasks that involve GPU and accelerator computing.

How do I install Jetson Speculative Decoding in Claude Code?

Run `npx skills add NVIDIA/skills --skill jetson-speculative-decoding -a claude-code`. Or copy the skill folder (skills/jetson-speculative-decoding in NVIDIA/skills) into .claude/skills/jetson-speculative-decoding in your project. Claude Code loads it when a task matches its description.

How do I install Jetson Speculative Decoding in Codex?

Run `npx skills add NVIDIA/skills --skill jetson-speculative-decoding -a codex`. Or copy the skill folder (skills/jetson-speculative-decoding in NVIDIA/skills) into .agents/skills/jetson-speculative-decoding in your project. Codex loads it when a task matches its description.

Can I use Jetson Speculative Decoding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill jetson-speculative-decoding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/jetson-speculative-decoding, .gemini/skills/jetson-speculative-decoding, .github/skills/jetson-speculative-decoding and .opencode/skills/jetson-speculative-decoding in your project.

What does Jetson Speculative Decoding need to run?

SKILL.md names no scripts, command-line tools or credentials: Jetson Speculative Decoding is instructions for the agent only.

Does Jetson Speculative Decoding access the network?

SKILL.md names 1 domain. As links in the text: jetson-ai-lab.com. This is read from the text; nothing was executed.

Is Jetson Speculative Decoding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Jetson Speculative Decoding use?

Jetson Speculative Decoding is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Jetson Speculative Decoding use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Jetson Speculative Decoding?

Skills that share tags, products or a category with Jetson Speculative Decoding: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 900 stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Dstack Prototyping (dstackai/dstack, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Jetson Speculative Decoding?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.