Official agent skill

Synthetic Data

by microsoft in microsoft/physical-ai-toolchain

Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines

OfficialMITAuto-check passedTesting & QA

Install Synthetic Data

skills CLI
$ npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/physical-ai-toolchain synthetic-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/physical-ai-toolchain.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/synthetic-data .claude/skills/synthetic-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
synthetic-data
GitHub stars
122
Token cost
~469 tokens
SKILL.md length
165 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines

  • Tasks that involve Test data and fixtures
  • SKILL.md covers Overview, Pipeline Stages, Workflow Submission and Key Files, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Synthetic Data is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization. Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines

Its SKILL.md is about 470 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test data and fixtures. It works with NVIDIA AI Platform. The licence is MIT.

When your agent uses it

  • Tasks that involve Test data and fixtures

Example prompts

  • “/synthetic-data”

What it can do on your machine

Read from SKILL.md and the folder at commit 5d38197. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Synthetic Data loads about 469 tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 165 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~469

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/physical-ai-toolchain at commit 5d38197, republished under its MIT licence (© microsoft). 165 words, ~469 tokens.

Download SKILL.mdSave it as .claude/skills/synthetic-data/SKILL.md (or your agent's skills folder).
name
synthetic-data
description
Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines

Synthetic Data Skill

Generate photorealistic training data using NVIDIA Cosmos world foundation models — Cosmos Transfer, Cosmos Predict, and Cosmos Reason.

Overview

The Synthetic Data domain provides SDG pipelines that transform simulation-rendered frames into photorealistic training data, predict future environment states, and curate output for training quality.

Pipeline Stages

StageModelPurpose
TransferCosmos Transfer 2.5Convert Isaac Sim renders to photorealistic images
PredictCosmos Predict 2.5Generate future frame sequences from observations
ReasonCosmos Reason 2Assess data quality and filter training samples

Workflow Submission

SDG workflows can be submitted via OSMO or AzureML:

  • OSMO workflows: synthetic-data/workflows/osmo/
  • AzureML jobs: synthetic-data/workflows/azureml/

Key Files

FilePurpose
synthetic-data/README.mdDomain overview and directory structure
synthetic-data/workflows/osmo/sdg-pipeline.yamlEnd-to-end OSMO SDG pipeline
synthetic-data/cosmos/configs/README.mdModel configuration reference
synthetic-data/specifications/synthetic-data.specification.mdSDG pipeline specification
synthetic-data/specifications/cosmos-integration.specification.mdCosmos model integration specification

Environment Requirements

All NVIDIA Cosmos containers require:

VariableValue
ACCEPT_EULAY
PRIVACY_CONSENTY
NVIDIA_DRIVER_CAPABILITIESall

GPU Requirements

Each Cosmos model stage requires a minimum of 1x A100 (40 GB) GPU. H100 (80 GB) recommended for production workloads.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/synthetic-data of microsoft/physical-ai-toolchain.

Open the folder on GitHubat commit 5d38197

Compare with similar skills

Synthetic Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Synthetic Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Synthetic Data this skillmicrosoft/physical-ai-toolchain122—~469Automated safety check: PassMIT
Physical AI Infrastructure Setup And Resilient ScalingNVIDIA/skills3.5k—~2.8kAutomated safety check: NotesApache-2.0
Jest Testing PatternsChrisWiles/claude-code-showcase6.1k7 repos~1.5kAutomated safety check: PassNone
Java SDK E2E Test with Replay Snapshotgithub/copilot-sdk11k—~1.8kAutomated safety check: PassMIT
source-mssql E2E Test Harnessairbytehq/airbyte22k—~4.4kAutomated safety check: PassCustom licence
RTK Filter TDD in Rustrtk-ai/rtk83k—~1.9kAutomated safety check: NotesApache-2.0

Similar skills

  • A skill your agent uses when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS…

    3.5k GitHub stars~2.8k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Jest Testing Patterns

    ChrisWiles/claude-code-showcase

    Jest patterns for React Native style tests: TDD discipline, mock factory functions, module and GraphQL hook mocking, custom render helpers and anti-patterns to avoid.

    6.1k GitHub starsUsed in 7 repos~1.5k tokens
    Testing & QAAuto-check passed
  • Official

    Creates a Java SDK end-to-end test for the Copilot SDK that runs against a recorded YAML snapshot through a replay proxy, so CI needs no real authentication.

    11k GitHub stars~1.8k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Official

    Stands up a throwaway local SQL Server 2022 backend, applies SQL fixtures and runs Airbyte spec, check, discover and read against source-mssql images.

    22k GitHub stars~4.4k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Enforces red-green-refactor for new RTK output filters in Rust, using real captured fixtures, snapshot tests with insta and token-savings assertions.

    83k GitHub stars~1.9k tokensUpdated yesterday
    Testing & QAAuto-check: notes
  • Guides recording, privacy review and offline verification of OpenLogi device fixtures with the fixture contribute and verify commands, without treating replay as proof of hardware behavior.

    23k GitHub stars~1.2k tokensUpdated 4 days ago
    Testing & QAAuto-check passed

More from microsoft/physical-ai-toolchain

  • Osmo Lerobot Training

    microsoft/physical-ai-toolchain

    Official

    Submit, monitor, analyze, and evaluate LeRobot imitation learning training jobs on OSMO with Azure ML MLflow integration and inference evaluation - Brought to you by microsoft/physical-ai-toolchain

    122 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check: notes
  • Azureml K3s Compute Target Setup

    microsoft/physical-ai-toolchain

    Official

    Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.

    122 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check: notes
  • Environment Deployment

    microsoft/physical-ai-toolchain

    Official

    Generate, transfer, and consume environment-specific Azure, AKS, OSMO, ACR, and Azure ML deployment bundles.

    122 GitHub stars~5.8k tokensUpdated yesterday
    Auto-check passed
  • Fleet Deployment

    microsoft/physical-ai-toolchain

    Official

    Deploy trained robot policies to edge fleets via FluxCD GitOps, image automation, and deployment gating

    122 GitHub stars~518 tokensUpdated yesterday
    Auto-check passed
  • Fleet Intelligence

    microsoft/physical-ai-toolchain

    Official

    Monitor robot fleet telemetry via Azure IoT Operations, drift detection, Grafana dashboards, and Fabric analytics

    122 GitHub stars~598 tokensUpdated yesterday
    Auto-check passed
  • Infrastructure

    microsoft/physical-ai-toolchain

    Official

    Deploy and manage Azure infrastructure for the Physical AI Toolchain including Terraform IaC, Kubernetes setup, GPU configuration, and network topology

    122 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Synthetic Data

What does Synthetic Data do?

Generate synthetic training data using NVIDIA Cosmos world foundation models for SDG pipelines. Synthetic Data is an agent skill from microsoft/physical-ai-toolchain, published by the product's own GitHub organization.

When should I use Synthetic Data?

Synthetic Data fits situations like: tasks that involve Test data and fixtures.

How do I install Synthetic Data in Claude Code?

Run `npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a claude-code`. Or copy the skill folder (.github/skills/synthetic-data in microsoft/physical-ai-toolchain) into .claude/skills/synthetic-data in your project. Claude Code loads it when a task matches its description.

How do I install Synthetic Data in Codex?

Run `npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a codex`. Or copy the skill folder (.github/skills/synthetic-data in microsoft/physical-ai-toolchain) into .agents/skills/synthetic-data in your project. Codex loads it when a task matches its description.

Can I use Synthetic Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/physical-ai-toolchain --skill synthetic-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/synthetic-data, .gemini/skills/synthetic-data, .github/skills/synthetic-data and .opencode/skills/synthetic-data in your project.

What does Synthetic Data need to run?

SKILL.md names no scripts, command-line tools or credentials: Synthetic Data is instructions for the agent only.

Does Synthetic Data access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Synthetic Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Synthetic Data use?

Synthetic Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Synthetic Data use?

About 469 tokens (SKILL.md is roughly 1.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Synthetic Data?

Skills that share tags, products or a category with Synthetic Data: Physical AI Infrastructure Setup And Resilient Scaling (NVIDIA/skills, 3.5k stars), Jest Testing Patterns (ChrisWiles/claude-code-showcase, 6.1k stars), Java SDK E2E Test with Replay Snapshot (github/copilot-sdk, 11k stars) and source-mssql E2E Test Harness (airbytehq/airbyte, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Synthetic Data?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/physical-ai-toolchain, which has 122 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 6, 2026.

Source: microsoft/physical-ai-toolchain on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.