A skill your agent uses when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS…
Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines
Install Synthetic Data
$ npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install microsoft/physical-ai-toolchain synthetic-data --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/synthetic-data .claude/skills/synthetic-data && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "synthetic-data" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/synthetic-data into .claude/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/synthetic-dataType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install microsoft/physical-ai-toolchain synthetic-data --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.github/skills/synthetic-data .agents/skills/synthetic-data && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "synthetic-data" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/synthetic-data into .agents/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install microsoft/physical-ai-toolchain synthetic-data --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.github/skills/synthetic-data .cursor/skills/synthetic-data && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "synthetic-data" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/synthetic-data into .cursor/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/microsoft/physical-ai-toolchain.git --path .github/skills/synthetic-data--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install microsoft/physical-ai-toolchain synthetic-data --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.github/skills/synthetic-data .gemini/skills/synthetic-data && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "synthetic-data" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/synthetic-data into .gemini/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install microsoft/physical-ai-toolchain synthetic-dataInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .github/skills && cp -r skills-src/.github/skills/synthetic-data .github/skills/synthetic-data && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "synthetic-data" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/synthetic-data into .github/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install microsoft/physical-ai-toolchain synthetic-data --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.github/skills/synthetic-data .opencode/skills/synthetic-data && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "synthetic-data" agent skill from https://github.com/microsoft/physical-ai-toolchain/tree/main/.github/skills/synthetic-data into .opencode/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
- Skill name
synthetic-data- GitHub stars
- 122
- Token cost
- ~469 tokens
- SKILL.md length
- 165 words
- Files
- 1
- Skills in repo
- 7
- Repo updated
- First seen
- Licence
- MIT
At a glance
Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines
- Tasks that involve Test data and fixtures
- SKILL.md covers Overview, Pipeline Stages, Workflow Submission and Key Files, plus 2 more sections
- Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
What it does
Synthetic Data is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization. Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines
Its SKILL.md is about 470 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Testing & QA, covering Test data and fixtures. It works with NVIDIA AI Platform. The licence is MIT.
When your agent uses it
- Tasks that involve Test data and fixtures
Example prompts
- “/synthetic-data”
What it can do on your machine
Read from SKILL.md and the folder at commit 5d38197. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Network
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Synthetic Data loads about 469 tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 165 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
SKILL.md
The full file from microsoft/physical-ai-toolchain at commit 5d38197, republished under its MIT licence (© microsoft). 165 words, ~469 tokens.
.claude/skills/synthetic-data/SKILL.md (or your agent's skills folder).- name
- synthetic-data
- description
- Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines
Synthetic Data Skill
Generate photorealistic training data using NVIDIA Cosmos world foundation models — Cosmos Transfer, Cosmos Predict, and Cosmos Reason.
Overview
The Synthetic Data domain provides SDG pipelines that transform simulation-rendered frames into photorealistic training data, predict future environment states, and curate output for training quality.
Pipeline Stages
| Stage | Model | Purpose |
|---|---|---|
| Transfer | Cosmos Transfer 2.5 | Convert Isaac Sim renders to photorealistic images |
| Predict | Cosmos Predict 2.5 | Generate future frame sequences from observations |
| Reason | Cosmos Reason 2 | Assess data quality and filter training samples |
Workflow Submission
SDG workflows can be submitted via OSMO or AzureML:
- OSMO workflows:
synthetic-data/workflows/osmo/ - AzureML jobs:
synthetic-data/workflows/azureml/
Key Files
| File | Purpose |
|---|---|
synthetic-data/README.md | Domain overview and directory structure |
synthetic-data/workflows/osmo/sdg-pipeline.yaml | End-to-end OSMO SDG pipeline |
synthetic-data/cosmos/configs/README.md | Model configuration reference |
synthetic-data/specifications/synthetic-data.specification.md | SDG pipeline specification |
synthetic-data/specifications/cosmos-integration.specification.md | Cosmos model integration specification |
Environment Requirements
All NVIDIA Cosmos containers require:
| Variable | Value |
|---|---|
ACCEPT_EULA | Y |
PRIVACY_CONSENT | Y |
NVIDIA_DRIVER_CAPABILITIES | all |
GPU Requirements
Each Cosmos model stage requires a minimum of 1x A100 (40 GB) GPU. H100 (80 GB) recommended for production workloads.
© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Files
Just SKILL.md in .github/skills/synthetic-data of microsoft/physical-ai-toolchain.
Open the folder on GitHubat commit 5d38197
Compare with similar skills
Synthetic Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Synthetic Data this skillmicrosoft/physical-ai-toolchain | 122 | — | ~469 | Automated safety check: Pass | MIT | |
| Physical AI Infrastructure Setup And Resilient ScalingNVIDIA/skills | 3.5k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | |
| Jest Testing PatternsChrisWiles/claude-code-showcase | 6.1k | 7 repos | ~1.5k | Automated safety check: Pass | None | |
| Java SDK E2E Test with Replay Snapshotgithub/copilot-sdk | 11k | — | ~1.8k | Automated safety check: Pass | MIT | |
| source-mssql E2E Test Harnessairbytehq/airbyte | 22k | — | ~4.4k | Automated safety check: Pass | Custom licence | |
| RTK Filter TDD in Rustrtk-ai/rtk | 83k | — | ~1.9k | Automated safety check: Notes | Apache-2.0 |
Similar skills
- Official
Jest Testing Patterns
ChrisWiles/claude-code-showcase
Jest patterns for React Native style tests: TDD discipline, mock factory functions, module and GraphQL hook mocking, custom render helpers and anti-patterns to avoid.
- Official
Java SDK E2E Test with Replay Snapshot
github/copilot-sdk
Creates a Java SDK end-to-end test for the Copilot SDK that runs against a recorded YAML snapshot through a replay proxy, so CI needs no real authentication.
- Official
source-mssql E2E Test Harness
airbytehq/airbyte
Stands up a throwaway local SQL Server 2022 backend, applies SQL fixtures and runs Airbyte spec, check, discover and read against source-mssql images.
RTK Filter TDD in Rust
rtk-ai/rtk
Enforces red-green-refactor for new RTK output filters in Rust, using real captured fixtures, snapshot tests with insta and token-savings assertions.
OpenLogi Device Fixture Contribution
AprilNEA/OpenLogi
Guides recording, privacy review and offline verification of OpenLogi device fixtures with the fixture contribute and verify commands, without treating replay as proof of hardware behavior.
More from microsoft/physical-ai-toolchain
- Official
Osmo Lerobot Training
microsoft/physical-ai-toolchain
Submit, monitor, analyze, and evaluate LeRobot imitation learning training jobs on OSMO with Azure ML MLflow integration and inference evaluation - Brought to you by microsoft/physical-ai-toolchain
- Official
Azureml K3s Compute Target Setup
microsoft/physical-ai-toolchain
Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.
- Official
Environment Deployment
microsoft/physical-ai-toolchain
Generate, transfer, and consume environment-specific Azure, AKS, OSMO, ACR, and Azure ML deployment bundles.
- Official
Fleet Deployment
microsoft/physical-ai-toolchain
Deploy trained robot policies to edge fleets via FluxCD GitOps, image automation, and deployment gating
- Official
Fleet Intelligence
microsoft/physical-ai-toolchain
Monitor robot fleet telemetry via Azure IoT Operations, drift detection, Grafana dashboards, and Fabric analytics
- Official
Infrastructure
microsoft/physical-ai-toolchain
Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology
Related
Works with
Categories
Questions about Synthetic Data
What does Synthetic Data do?
Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines. Synthetic Data is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization.
When should I use Synthetic Data?
Synthetic Data fits situations like: tasks that involve Test data and fixtures.
How do I install Synthetic Data in Claude Code?
Run `npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a claude-code`. Or copy the skill folder (.github/skills/synthetic-data in microsoft/physical-ai-toolchain) into .claude/skills/synthetic-data in your project. Claude Code loads it when a task matches its description.
How do I install Synthetic Data in Codex?
Run `npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a codex`. Or copy the skill folder (.github/skills/synthetic-data in microsoft/physical-ai-toolchain) into .agents/skills/synthetic-data in your project. Codex loads it when a task matches its description.
Can I use Synthetic Data in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/synthetic-data, .gemini/skills/synthetic-data, .github/skills/synthetic-data and .opencode/skills/synthetic-data in your project.
What does Synthetic Data need to run?
SKILL.md names no scripts, command-line tools or credentials: Synthetic Data is instructions for the agent only.
Does Synthetic Data access the network?
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Is Synthetic Data safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
What licence does Synthetic Data use?
Synthetic Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Synthetic Data use?
About 469 tokens (SKILL.md is roughly 1.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
What are the alternatives to Synthetic Data?
Skills that share tags, products or a category with Synthetic Data: Physical AI Infrastructure Setup And Resilient Scaling (NVIDIA/skills, 3.5k stars), Jest Testing Patterns (ChrisWiles/claude-code-showcase, 6.1k stars), Java SDK E2E Test with Replay Snapshot (github/copilot-sdk, 11k stars) and source-mssql E2E Test Harness (airbytehq/airbyte, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Synthetic Data?
microsoft (a GitHub organization, an official publisher) maintains it in microsoft/physical-ai-toolchain, which has 122 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 6, 2026.
Source: microsoft/physical-ai-toolchain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.
