LLM Torch Profiler Analysis
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
CenterPose for keypoint / pose estimation. An agent skill from NVIDIA/skills.
$ npx skills add NVIDIA/skills --skill tao-train-centerpose -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-train-centerpose --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-train-centerpose .claude/skills/tao-train-centerpose && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-train-centerpose" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-centerpose into .claude/skills/tao-train-centerpose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-centerpose", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-train-centerposeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-train-centerpose -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-train-centerpose --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-train-centerpose .agents/skills/tao-train-centerpose && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-train-centerpose" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-centerpose into .agents/skills/tao-train-centerpose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-centerpose", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-train-centerpose -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-train-centerpose --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-train-centerpose .cursor/skills/tao-train-centerpose && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-train-centerpose" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-centerpose into .cursor/skills/tao-train-centerpose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-centerpose", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-train-centerpose--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-train-centerpose -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-train-centerpose --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-train-centerpose .gemini/skills/tao-train-centerpose && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-train-centerpose" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-centerpose into .gemini/skills/tao-train-centerpose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-centerpose", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-train-centerposeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-train-centerpose -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-train-centerpose .github/skills/tao-train-centerpose && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-train-centerpose" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-centerpose into .github/skills/tao-train-centerpose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-centerpose", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-train-centerpose -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-train-centerpose --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-train-centerpose .opencode/skills/tao-train-centerpose && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-train-centerpose" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-centerpose into .opencode/skills/tao-train-centerpose/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-centerpose", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-train-centerposeCenterPose for keypoint / pose estimation. An agent skill from NVIDIA/skills.
Tao Train Centerpose is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. CenterPose for keypoint / pose estimation. Detects object centers and regresses keypoint locations for 6-DoF object pose estimation. Use when training, evaluating, exporting, or running inference for a TAO CenterPose model. Trigger phrases include "train CenterPose", "6-DoF object pose", "keypoint estimation", "object pose regression".
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit.
It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires docker + nvidia-container-toolkit.
From compatibility in the SKILL.md frontmatter.
Tao Train Centerpose loads about 2.9k tokens when it runs, and up to ~9.1k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 1,105 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,105 words, ~2,858 tokens.
.claude/skills/tao-train-centerpose/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
CenterPose for keypoint / pose estimation. Detects object centers and regresses keypoint locations. Used for 6-DoF object pose estimation.
Set model.backbone.pretrained_backbone_path.
For TAO Deploy TensorRT actions (gen_trt_engine, TensorRT evaluate, and TensorRT inference), use the deploy spec templates packaged in this skill's references/ folder with the spec_template_deploy_*.yaml prefix.
Generated TAO Core schemas are packaged in schemas/<action>.schema.json, with schemas/manifest.json listing available actions. Each generated schema also emits references/spec_template_<action>.yaml from the schema top-level default field. AutoML enablement is declared at the model layer in references/skill_info.yaml via automl_enabled. Runnable AutoML for an action requires schemas/<action>.schema.json and references/spec_template_<action>.yaml to exist and parse. Use the packaged selected-action schema for automl_default_parameters, automl_disabled_parameters, defaults, min/max bounds, enums, option weights, math conditions, dependencies, and popular parameters. Do not expect ~/tao-core at runtime; maintainers regenerate schemas/templates before packaging the skill bank.
This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skill_info.yaml and resolve the run override from either an explicit automl_policy value or the user's workflow request. Use automl_policy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automl_policy: off for this run only. When automl_policy: on, automl_enabled: true, and both schemas/train.schema.json and references/spec_template_train.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skill_dir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automl_policy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.
Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.
val_3DIoU, val_2DMPEtest_3DIoU, test_2DMPEtest_3DIoU with maximize direction for
the required evaluation-backed baseline, every recommendation through
eval_fn, best-model selection, and final evaluation. val_3DIoU is emitted
by training but not by the evaluate action, so it is only suitable for an
explicitly accepted training-proxy run without an impact baseline.| Action | Spec Key | Source | Files | List? |
|---|---|---|---|---|
| evaluate | dataset.test_data | eval_dataset | test.tar.gz | No |
| gen_trt_engine | gen_trt_engine.tensorrt.calibration.cal_image_dir | calibration_dataset | train.tar.gz | Yes |
| inference | dataset.inference_data | inference_dataset | val.tar.gz | No |
| train | dataset.train_data | train_datasets | train.tar.gz | No |
| train | dataset.val_data | eval_dataset | val.tar.gz | No |
Data source overrides are mandatory for every action — the agent MUST construct data source paths from the Per-Action Dataset Requirements table above and include them in spec_overrides.
TRAIN_DIR = "/path/to/extracted/train"
VAL_DIR = "/path/to/extracted/val"
TEST_DIR = "/path/to/extracted/test"
INFER_DIR = VAL_DIR
CAL_IMAGE_DIRS = ["/path/to/extracted/train/<sequence_or_image_dir>"]train (mandatory data sources):
{
"train.num_epochs": 30,
"train.checkpoint_interval": 10,
"train.validation_interval": 10,
"train.num_gpus": 1,
"dataset.category": "bike",
"dataset.batch_size": 4,
"dataset.train_data": TRAIN_DIR,
"dataset.val_data": VAL_DIR,
}evaluate (mandatory data sources):
{
"dataset.category": "bike",
"dataset.test_data": TEST_DIR,
}inference (mandatory data sources):
{
"dataset.category": "bike",
"dataset.inference_data": INFER_DIR,
}gen_trt_engine (mandatory data sources):
{
"gen_trt_engine.tensorrt.calibration.cal_image_dir": CAL_IMAGE_DIRS,
}Optional. Val and test datasets are provided as separate tarballs. Training
writes val_3DIoU/val_2DMPE KPIs, while evaluate writes
test_3DIoU/test_2DMPE; do not configure an evaluate-backed AutoML callback
to extract the training-prefixed names.
Launch method: Lightning-managed (single python process, Lightning spawns workers).
| Spec Key | Description | Default |
|---|---|---|
train.num_gpus | Number of GPUs | 1 |
train.gpu_ids | GPU device indices | [0] |
auto (Lightning picks the best strategy automatically)num_nodes or distributed_strategy config — single-node onlysync_batchnormMinimum 1 GPU(s), recommended 2 GPU(s). 16GB+ VRAM per GPU. CenterPose is moderately memory-intensive depending on input resolution and number of keypoints.
num_joints mismatch: Ensure dataset.num_joints matches the keypoint count in your annotations.
Extract S3 tarballs for local Docker: The starter-kit S3 data is packaged as
train.tar.gz, val.tar.gz, and test.tar.gz, but the CenterPose TAO actions
consume extracted folders. Extract each archive and set dataset.train_data,
dataset.val_data, dataset.test_data, and dataset.inference_data to the
extracted split directories.
Checkpoint handoff: CenterPose training writes concrete checkpoints such as
model_epoch_000_step_00008.pth and a centerpose_model_latest.pth symlink.
Use the SDK/model checkpoint resolver or the exact epoch/step checkpoint for
evaluate, inference, export, and resume. Use the symlink only when the user
explicitly asks for latest.
TAO Deploy postprocessor compatibility: Use the deploy image resolved from
the skill's pinned deploy image or the selected platform. A successful gen_trt_engine run does
not prove deploy evaluate or inference works; inspect those action exit codes
and logs separately, especially for CenterPose postprocessor errors such as
TypeError: only 0-dimensional arrays can be converted to Python scalars.
Model-specific inference mappings belong in this MD file, not in config.json. Generated runners should read this section and apply the mappings with SDK helpers before create_job(). This mirrors the old microservices infer_params.py flow.
Inference mappings from TAO Core centerpose.config.json:
| Action | Spec Field | Inference Function | Meaning |
|---|---|---|---|
| evaluate | encryption_key | key | encryption key |
| evaluate | evaluate.checkpoint | parent_model | model file inferred from the parent job results folder |
| evaluate | evaluate.trt_engine | parent_model | model file inferred from the parent job results folder |
| evaluate | results_dir | output_dir | current job results directory |
| export | encryption_key | key | encryption key |
| export | export.checkpoint | parent_model | model file inferred from the parent job results folder |
| export | export.onnx_file | create_onnx_file | output ONNX path |
| export | results_dir | output_dir | current job results directory |
| gen_trt_engine | encryption_key | key | encryption key |
| gen_trt_engine | gen_trt_engine.onnx_file | parent_model | model file inferred from the parent job results folder |
| gen_trt_engine | gen_trt_engine.tensorrt.calibration.cal_cache_file | create_cal_cache | calibration cache path |
| gen_trt_engine | gen_trt_engine.trt_engine | create_engine_file | output TensorRT engine path |
| gen_trt_engine | results_dir | output_dir | current job results directory |
| inference | encryption_key | key | encryption key |
| inference | inference.checkpoint | parent_model | model file inferred from the parent job results folder |
| inference | inference.trt_engine | parent_model | model file inferred from the parent job results folder |
| inference | results_dir | output_dir | current job results directory |
| train | encryption_key | key | encryption key |
| train | model.backbone.pretrained_backbone_path | ptm_if_no_resume_model | PTM when no resume checkpoint exists |
| train | results_dir | output_dir | current job results directory |
| train | train.resume_training_checkpoint_path | resume_model | model file inferred from the current job results folder |
For parent_model or parent_model_folder, pass the upstream train/export/AutoML child job id as parent_job_id. The SDK lists the parent result folder, filters checkpoint artifacts, and returns the selected model file or folder. Do not add these mappings back to config.json and do not patch generated runner scripts to guess checkpoint paths.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 22 other files (references) in skills/tao-train-centerpose of NVIDIA/skills.
Open the folder on GitHubat commit dfdd080
Tao Train Centerpose next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tao Train Centerpose this skillNVIDIA/skills | 3.5k | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | |
| LLM Torch Profiler Analysissgl-project/sglang | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | |
| Skill InspectorNVIDIA/SkillSpector | 20k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Megatron-LM Container and Dependency SetupNVIDIA/Megatron-LM | 18k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | |
| Embeddings via 9Routerdecolua/9router | 30k | — | ~604 | Automated safety check: Pass | MIT | |
| Megatron-LM Base Image BumpNVIDIA/Megatron-LM | 18k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 |
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
NVIDIA/SkillSpector
Decides whether an agent skill is safe to install by combining a SkillSpector static scan with the agent's own source review, ending in APPROVE, CAUTION or REJECT.
NVIDIA/Megatron-LM
Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.
decolua/9router
Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.
NVIDIA/Megatron-LM
Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.
NVIDIA/NemoClaw
Remove bracketed NemoClaw tags from GitHub issue and PR titles.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
CenterPose for keypoint / pose estimation. An agent skill from NVIDIA/skills. Tao Train Centerpose is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. CenterPose for keypoint / pose estimation.
Tao Train Centerpose fits situations like: running inference for a TAO CenterPose model; phrases include train CenterPose; 6-DoF object pose; keypoint estimation.
Run `npx skills add NVIDIA/skills --skill tao-train-centerpose -a claude-code`. Or copy the skill folder (skills/tao-train-centerpose in NVIDIA/skills) into .claude/skills/tao-train-centerpose in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-train-centerpose -a codex`. Or copy the skill folder (skills/tao-train-centerpose in NVIDIA/skills) into .agents/skills/tao-train-centerpose in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-train-centerpose -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-train-centerpose, .gemini/skills/tao-train-centerpose, .github/skills/tao-train-centerpose and .opencode/skills/tao-train-centerpose in your project.
SKILL.md names no scripts, command-line tools or credentials: Tao Train Centerpose is instructions for the agent only. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Tao Train Centerpose is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tao Train Centerpose: LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Skill Inspector (NVIDIA/SkillSpector, 20k stars), Megatron-LM Container and Dependency Setup (NVIDIA/Megatron-LM, 18k stars) and Embeddings via 9Router (decolua/9router, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.