Official agent skill

Data Designer

by NVIDIA in NVIDIA/skills

A skill your agent uses when the user wants to create a dataset, generate synthetic data, or build a data generation pipeline.

OfficialApache-2.0Auto-check: warningsTesting & QA

Install Data Designer

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add NVIDIA/skills --skill data-designer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills data-designer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-designer .claude/skills/data-designer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-designer
GitHub stars
3.5k
Token cost
~1.2k tokens
SKILL.md length
440 words
Files
11 (incl. scripts, references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user wants to create a dataset, generate synthetic data, or build a data generation pipeline.

  • The user wants to create a dataset
  • Runs Python scripts from its folder
  • Generate synthetic data
  • Build a data generation pipeline

What it does

Data Designer is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use when the user wants to create a dataset, generate synthetic data, or build a data generation pipeline.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts and reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/person-sampling.md`).

It sits in Testing & QA, covering Test data and fixtures. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • The user wants to create a dataset
  • Generate synthetic data
  • Build a data generation pipeline

Example prompts

  • “/data-designer”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Designer loads about 1.2k tokens when it runs, and up to ~2.5k if it reads all its reference files. Until then it costs about 30 tokens; SKILL.md has 440 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningContains instruction-override wording (e.g. “without asking the user”)SKILL.md:45
    it themselves. Do not install anything without the user's permission.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 440 words, ~1,178 tokens.

Download SKILL.mdSave it as .claude/skills/data-designer/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
data-designer
description
Use when the user wants to create a dataset, generate synthetic data, or build a data generation pipeline.
argument-hint
describe the dataset you want to generate
license
Apache-2.0
metadata.owner
DataDesigner

Before You Start

Do not explore the workspace first. The workflow's Learn step gives you everything you need.

Goal

Build a synthetic dataset using the Data Designer library that matches this description:

$ARGUMENTS

Workflow

Use Autopilot mode if the user implies they don't want to answer questions — e.g., they say something like "be opinionated", "you decide", "make reasonable assumptions", "just build it", "surprise me", etc. Otherwise, use Interactive mode (default).

Read only the workflow file that matches the selected mode, then follow it:

  • Interactive → read workflows/interactive.md
  • Autopilot → read workflows/autopilot.md

Rules

  • Keep all columns in the output by default. The only exceptions for dropping a column are: (1) the user explicitly asks, or (2) it is a helper column that exists solely to derive other columns (e.g., a sampled person object used to extract name, city, etc.). When in doubt, keep the column.
  • Do not suggest or ask about seed datasets. Only use one when the user explicitly provides seed data or asks to build from existing records. When using a seed, read references/seed-datasets.md.
  • When the dataset requires person data (names, demographics, addresses), read references/person-sampling.md.
  • If a dataset script that matches the dataset description already exists, ask the user whether to edit it or create a new one.

Usage Tips and Common Pitfalls

  • Sampler and validation columns need both a type and params. E.g., sampler_type="category" with params=dd.CategorySamplerParams(...).
  • Jinja2 templates in prompt, system_prompt, and expr fields: reference columns with {{ column_name }}, nested fields with {{ column_name.field }}.
  • SamplerColumnConfig: Takes params, not sampler_params.
  • LLM judge score access: LLMJudgeColumnConfig produces a nested dict where each score name maps to {reasoning: str, score: int}. To get the numeric score, use the .score attribute. For example, for a judge column named quality with a score named correctness, use {{ quality.correctness.score }}. Using {{ quality.correctness }} returns the full dict, not the numeric score.
Show full SKILL.md (140 more words)Show less

Troubleshooting

  • data-designer CLI not found: Tell the user that data-designer is not installed in this environment (requires Python >= 3.10). Ask if they would like you to create a virtual environment and install it, or if they prefer to do it themselves. Do not install anything without the user's permission.
  • Network errors during preview: A sandbox environment may be blocking outbound requests. Ask the user for permission to retry the command with the sandbox disabled. Only as a last resort, if retrying outside the sandbox also fails, tell the user to run the command themselves.

Output Template

Write a Python file to the current directory with a load_config_builder() function returning a DataDesignerConfigBuilder. Name the file descriptively (e.g., customer_reviews.py). Use PEP 723 inline metadata for dependencies.

python
# /// script
# dependencies = [
#   "data-designer", # always required
#   "pydantic", # only if this script imports from pydantic
#   # add additional dependencies here
# ]
# ///
import data_designer.config as dd
from pydantic import BaseModel, Field


# Use Pydantic models when the output needs to conform to a specific schema
class MyStructuredOutput(BaseModel):
    field_one: str = Field(description="...")
    field_two: int = Field(description="...")


# Use custom generators when built-in column types aren't enough
@dd.custom_column_generator(
    required_columns=["col_a"],
    side_effect_columns=["extra_col"],
)
def generator_function(row: dict) -> dict:
    # add custom logic here that depends on "col_a" and update row in place
    row["name_in_custom_column_config"] = "custom value"
    row["extra_col"] = "extra value"
    return row


def load_config_builder() -> dd.DataDesignerConfigBuilder:
    config_builder = dd.DataDesignerConfigBuilder()

    # Seed dataset (only if the user explicitly mentions a seed dataset path)
    # config_builder.with_seed_dataset(dd.LocalFileSeedSource(path="path/to/seed.parquet"))

    # config_builder.add_column(...)
    # config_builder.add_processor(...)

    return config_builder

Only include Pydantic models, custom generators, seed datasets, and extra dependencies when the task requires them.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts, references) in skills/data-designer of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • references/person-sampling.md
  • references/preview-review.md
  • references/seed-datasets.md
  • scripts/get_person_object_schema.py
  • skill-card.md
  • skill.oms.sig
  • workflows/autopilot.md
  • workflows/interactive.md

Open the folder on GitHubat commit 0e0d506

Compare with similar skills

Data Designer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Designer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Designer this skillNVIDIA/skills3.5k—~1.2kAutomated safety check: WarnApache-2.0
Jest Testing PatternsChrisWiles/claude-code-showcase6.1k7 repos~1.5kAutomated safety check: PassNone
Java SDK E2E Test with Replay Snapshotgithub/copilot-sdk11k—~1.8kAutomated safety check: PassMIT
source-mssql E2E Test Harnessairbytehq/airbyte22k—~4.4kAutomated safety check: PassCustom licence
RTK Filter TDD in Rustrtk-ai/rtk83k—~1.9kAutomated safety check: NotesApache-2.0
OpenLogi Device Fixture ContributionAprilNEA/OpenLogi23k—~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • Jest Testing Patterns

    ChrisWiles/claude-code-showcase

    Jest patterns for React Native style tests: TDD discipline, mock factory functions, module and GraphQL hook mocking, custom render helpers and anti-patterns to avoid.

    6.1k GitHub starsUsed in 7 repos~1.5k tokens
    Testing & QAAuto-check passed
  • Official

    Creates a Java SDK end-to-end test for the Copilot SDK that runs against a recorded YAML snapshot through a replay proxy, so CI needs no real authentication.

    11k GitHub stars~1.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Official

    Stands up a throwaway local SQL Server 2022 backend, applies SQL fixtures and runs Airbyte spec, check, discover and read against source-mssql images.

    22k GitHub stars~4.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Enforces red-green-refactor for new RTK output filters in Rust, using real captured fixtures, snapshot tests with insta and token-savings assertions.

    83k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check: notes
  • Guides recording, privacy review and offline verification of OpenLogi device fixtures with the fixture contribute and verify commands, without treating replay as proof of hardware behavior.

    23k GitHub stars~1.2k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Official

    Refreshes stored golden values from a GitHub Actions run, reports signed percentage changes per model, and writes a summary ready for a pull request description.

    18k GitHub stars~2.8k tokensUpdated today
    Testing & QAAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Categories

Questions about Data Designer

What does Data Designer do?

A skill your agent uses when the user wants to create a dataset, generate synthetic data, or build a data generation pipeline. Data Designer is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use when the user wants to create a dataset, generate synthetic data, or build a data generation pipeline.

When should I use Data Designer?

Data Designer fits situations like: the user wants to create a dataset; generate synthetic data; build a data generation pipeline.

How do I install Data Designer in Claude Code?

Run `npx skills add NVIDIA/skills --skill data-designer -a claude-code`. Or copy the skill folder (skills/data-designer in NVIDIA/skills) into .claude/skills/data-designer in your project. Claude Code loads it when a task matches its description.

How do I install Data Designer in Codex?

Run `npx skills add NVIDIA/skills --skill data-designer -a codex`. Or copy the skill folder (skills/data-designer in NVIDIA/skills) into .agents/skills/data-designer in your project. Codex loads it when a task matches its description.

Can I use Data Designer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill data-designer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-designer, .gemini/skills/data-designer, .github/skills/data-designer and .opencode/skills/data-designer in your project.

What does Data Designer need to run?

Going by SKILL.md and its folder, Data Designer needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Data Designer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Designer safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): contains instruction-override wording (e.g. “without asking the user”). Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Designer use?

Data Designer is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Designer use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Data Designer?

Skills that share tags, products or a category with Data Designer: Jest Testing Patterns (ChrisWiles/claude-code-showcase, 6.1k stars), Java SDK E2E Test with Replay Snapshot (github/copilot-sdk, 11k stars), source-mssql E2E Test Harness (airbytehq/airbyte, 22k stars) and RTK Filter TDD in Rust (rtk-ai/rtk, 83k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Designer?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.