Official agent skill

Foundationpose Setup

by NVIDIA in NVIDIA/skills

Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Foundationpose Setup

skills CLI
$ npx skills add NVIDIA/skills --skill foundationpose-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills foundationpose-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/foundationpose-setup .claude/skills/foundationpose-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
foundationpose-setup
GitHub stars
3.5k
Token cost
~1.5k tokens
SKILL.md length
671 words
Files
9 (incl. references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines.

  • Works in 5 steps: Preflight before installing. Check… → Install into the product's Python 3.12… → Verify checkpoint access. SAM3 is gated… → …
  • SAM3/TAO dependency conflicts
  • SKILL.md covers Purpose, Requirements, Instructions and Troubleshooting, plus 2 more sections
  • Calls python and uv

What it does

Foundationpose Setup is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines. Use for SAM3/TAO dependency conflicts, CUDA library failures, and depth-engine shape or precision decisions.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `BENCHMARK.md`, `agents/openai.yaml` and `evals/config.yml`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • SAM3/TAO dependency conflicts
  • CUDA library failures
  • Depth-engine shape
  • Precision decisions

Example prompts

  • “/foundationpose-setup”

Requirements

  • Python 3
  • Docker

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Preflight before installing. Check glibc, GLIBCXX, driver, nvcc, tools, and Docker GPU
  2. Install into the product's Python 3.12 venv. Follow
  3. Verify checkpoint access. SAM3 is gated at Hugging Face; an existing authorized token
  4. Prepare the depth engine. Read engine construction. Use the
  5. Set runtime paths and verify. From the product root

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Foundationpose Setup loads about 1.5k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 671 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 671 words, ~1,521 tokens.

Download SKILL.mdSave it as .claude/skills/foundationpose-setup/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
foundationpose-setup
description
Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines. Use for SAM3/TAO dependency conflicts, CUDA library failures, and depth-engine shape or precision decisions.
license
Apache-2.0
metadata.author
zwdoescode <zhengwang@nvidia.com>
metadata.version
0.1.0

FoundationPose perception pipeline setup

Purpose

Prepare the FoundationPose perception pipeline for depth, segmentation, and pose inference. Environment installation and engine construction belong here; dataset adaptation, inference, and pose evaluation belong to foundationpose-pipeline when that skill is installed.

Requirements

Work from the product checkout, not the installed skill directory. Find the user's checkout by checking for pyproject.toml (project foundationpose-perception-pipeline), tools/build_tao_engine.py, and config/defaults.yaml. If absent and setup was requested, clone the product URL above into the user's workspace and enter it. For advice-only requests, use the supplied diagnostics without cloning or installing anything. Commands below use paths relative to the product root; references/ links are relative to this skill.

Read the checkout's README.md Requirements and Install sections for the matching revision. The supported stack requires Linux x86_64, glibc >= 2.38, GLIBCXX_3.4.31, NVIDIA driver >= 580, a CUDA toolkit >= 12.8 with nvcc, Python 3.12, uv, Git, Docker with GPU access, wget, and unzip. Start GPU sizing at 24 GB and measure the densest scene; 32 GB was tested. Budget depth-cache disk as roughly width * height * 4 * 3 bytes per scene, plus predictions and models.

Keep sibling directories for sam3/, foundation-pose-inference-library/, and models/ beside the product checkout. models/ contains the deployable ONNX and engine, not FoundationStereo source.

Instructions

  1. Preflight before installing. Check glibc, GLIBCXX, driver, nvcc, tools, and Docker GPU access. An Ubuntu 22.04 host with glibc 2.35 cannot load the shipped FoundationPose library; report the unsupported runtime and stop setup there. Do not replace system libc or try to solve this with LD_LIBRARY_PATH. See installation.

  2. Install into the product's Python 3.12 venv. Follow installation for uv, SAM3, the FoundationPose build, and TAO Deploy. On a fresh venv use uv sync --extra foundationpose; on an existing venv use uv sync --inexact --extra foundationpose to preserve out-of-band packages.

  3. Verify checkpoint access. SAM3 is gated at Hugging Face; an existing authorized token or usable cached checkpoint is sufficient. Request user action only if access is missing. FoundationPose and the documented FoundationStereo export are public Hugging Face downloads; credential hunting is not the first response to a network failure.

  4. Prepare the depth engine. Read engine construction. Use the user's ONNX location or the sibling models/ directory. Adapt the dataset before measuring the engine shape; use tools/bop_adapt/adapt.py --config <profile> --src <source> as described in the checkout's README Dataset adaptation section. Only registered adapters are supported. Build with --shape-from-scene on an adapted scene and FP32 unless the user requests a precision experiment. Set overrides.depth.engine in config/<profile>.yaml.

  5. Set runtime paths and verify. From the product root:

    bash
    source .venv/bin/activate
    export FOUNDATIONPOSE_ROOT="$(realpath ../foundation-pose-inference-library)"
    PIPELINE_SITE="$(realpath .venv/lib/python3.12/site-packages)"
    export LD_LIBRARY_PATH="${PIPELINE_SITE}/tensorrt_libs:${PIPELINE_SITE}/nvidia/cu13/lib:${LD_LIBRARY_PATH:-}"
    python tools/verify_sam3.py
    python tools/verify_foundationpose.py
    python tools/verify_foundationstereo.py --config <profile> --engine <engine-path>
    python test/check_engine_depth_smoke.py --config <profile> --engine <engine-path>

    The first three verify components; the last also needs an adapted dataset. Expect backend=tao, normalization=imagenet, the intended fixed shape, and no cropping N rows warning. An unloaded or unavailable engine is an incomplete verification, not a pass.

Show full SKILL.md (215 more words)Show less

Troubleshooting

SymptomAction
GLIBC_2.38 not foundUse a supported OS/runtime; a venv or library search path cannot upgrade host libc.
libcudart.so.13 missingCheck the product venv runtime wheels and absolute library paths before retrying pose.
SAM3 breaks after syncUse --inexact; confirm numpy 1.26.x and reinstall the sibling SAM3 package if pruned.
pycuda build cannot find cuda.hCheck the CUDA toolkit, nvcc on PATH, or CUDA_ROOT.
TAO import or dependency conflictUse TAO Deploy 7.1.0 with --no-deps; sync declared dependencies with --inexact.
Engine sidecar mismatch or croppingRebuild for this GPU, TensorRT version, precision, and adapted scene shape.

Examples

  • "Install the FoundationPose perception pipeline on this Ubuntu 24.04 GPU machine."
  • "The pipeline cannot load libcudart.so.13 after I moved the checkout."
  • "Build the TAO depth engine for my adapted T-LESS scenes."

Limitations and completion

Engines are machine-specific and must not be committed. With no dataset, download the ONNX and report shape-dependent engine construction and scene validation as pending; do not invent a rig shape. The pipeline's Apache license does not cover separately downloaded model weights; retain their upstream terms and SAM3's access requirements.

Report which preflight, install, checkpoint, engine, and verification steps actually passed, the checkout and engine paths, versions used, and remaining blockers. Do not equate installation or a smoke check with measured pose accuracy.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/foundationpose-setup of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • agents/openai.yaml
  • evals/config.yml
  • evals/evals.json
  • references/engine.md
  • references/installation.md
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 0e0d506

Compare with similar skills

Foundationpose Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Foundationpose Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Foundationpose Setup this skillNVIDIA/skills3.5k—~1.5kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS900—~2.8kAutomated safety check: PassNone
Llama CppOrchestra-Research/AI-Research-SKILLs13k4 repos~1.5kAutomated safety check: PassMIT
Vllm Deploy Simplevllm-project/vllm-skills103—~1.6kAutomated safety check: PassApache-2.0
Quark Env Preflightamd/Quark181—~1.4kAutomated safety check: PassMIT

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    900 GitHub stars~2.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Llama Cpp

    Orchestra-Research/AI-Research-SKILLs

    Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware.

    13k GitHub starsUsed in 4 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Simple

    vllm-project/vllm-skills

    Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Foundationpose Setup

What does Foundationpose Setup do?

Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines. Foundationpose Setup is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines.

When should I use Foundationpose Setup?

Foundationpose Setup fits situations like: SAM3/TAO dependency conflicts; CUDA library failures; depth-engine shape; precision decisions.

How do I install Foundationpose Setup in Claude Code?

Run `npx skills add NVIDIA/skills --skill foundationpose-setup -a claude-code`. Or copy the skill folder (skills/foundationpose-setup in NVIDIA/skills) into .claude/skills/foundationpose-setup in your project. Claude Code loads it when a task matches its description.

How do I install Foundationpose Setup in Codex?

Run `npx skills add NVIDIA/skills --skill foundationpose-setup -a codex`. Or copy the skill folder (skills/foundationpose-setup in NVIDIA/skills) into .agents/skills/foundationpose-setup in your project. Codex loads it when a task matches its description.

Can I use Foundationpose Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill foundationpose-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/foundationpose-setup, .gemini/skills/foundationpose-setup, .github/skills/foundationpose-setup and .opencode/skills/foundationpose-setup in your project.

What does Foundationpose Setup need to run?

Going by SKILL.md and its folder, Foundationpose Setup needs the command-line tools its instructions call (python and uv). Our summary lists: Python 3; Docker.

Does Foundationpose Setup access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Foundationpose Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Foundationpose Setup use?

Foundationpose Setup is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Foundationpose Setup use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Foundationpose Setup?

Skills that share tags, products or a category with Foundationpose Setup: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 900 stars), Llama Cpp (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Vllm Deploy Simple (vllm-project/vllm-skills, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Foundationpose Setup?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.