Agent skill

Dummy Dataset Generator

by phuryn in phuryn/pm-skills

Generates realistic test datasets with custom columns, row counts and business constraints, output as CSV, JSON, SQL inserts or a runnable Python script.

MITAuto-check passedTesting & QA

Install Dummy Dataset Generator

skills CLI
$ npx skills add phuryn/pm-skills --skill dummy-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install phuryn/pm-skills dummy-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/phuryn/pm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pm-execution/skills/dummy-dataset .claude/skills/dummy-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dummy-dataset
GitHub stars
27k
Token cost
~983 tokens
SKILL.md length
331 words
Files
1
Skills in repo
67
Repo updated
First seen
Licence
MIT

At a glance

Generates realistic test datasets with custom columns, row counts and business constraints, output as CSV, JSON, SQL inserts or a runnable Python script.

  • Works in 8 steps: Identify dataset type - Understand the… → Define column specifications - Names,… → Determine row count - How many sample… → …
  • Creating a mock dataset for a demo or development environment
  • SKILL.md covers Step-by-Step Process, Template: Python Script Output, Example Dataset Specification and Output Deliverables, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The agent works through a fixed process: identify the data domain, define column names, types and value ranges, decide the row count, pick an output format, apply realistic-looking patterns, apply any business constraints such as skewed rating distributions or category-dependent values, then generate the data or a script that generates it and validate the result. Parameters named in the skill are the product or system name, the dataset type, row count, which columns to include, the output format and any extra constraints, with a default of 100 rows when not specified.

A worked example builds a customer feedback dataset with columns like feedback_id, customer_name, email, feedback_date, rating, category and free-text feedback, using constraints such as a skewed star-rating distribution, bug reports only carrying low ratings, and realistic email domains. Deliverables include either a ready-to-execute Python script or a direct data file, plus documentation of how the data was generated and quick-start instructions for using it.

When your agent uses it

  • Creating a mock dataset for a demo or development environment
  • Generating sample rows with realistic constraints for a new database table
  • Producing test data in CSV, JSON or SQL insert format

Example prompts

  • “Generate 200 rows of dummy customer feedback data as CSV.”
  • “Create a mock transactions dataset with realistic amounts and give me SQL inserts.”
  • “Build a Python script that generates sample user profiles with valid emails.”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Identify dataset type - Understand the data domain
  2. Define column specifications - Names, data types, and value ranges
  3. Determine row count - How many sample records needed
  4. Select output format - CSV, JSON, SQL INSERT, or Python script
  5. Apply realistic patterns - Ensure data looks authentic and valid
  6. Add business constraints - Respect business logic and relationships
  7. Generate or script data - Create executable output
  8. Validate output - Ensure data quality and completeness

What it can do on your machine

Read from SKILL.md and the folder at commit 8607e3b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dummy Dataset Generator loads about 983 tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 331 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~983

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from phuryn/pm-skills at commit 8607e3b, republished under its MIT licence (© phuryn). 331 words, ~983 tokens.

Download SKILL.mdSave it as .claude/skills/dummy-dataset/SKILL.md (or your agent's skills folder).
name
dummy-dataset
description
Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script). Use when creating test data, building mock datasets, or generating sample data for development and demos.

Dummy Dataset Generation

Generate realistic dummy datasets for testing with customizable columns, constraints, and output formats (CSV, JSON, SQL, Python script). Creates executable scripts or direct data files for immediate use.

Use when: Creating test data, generating sample datasets, building realistic mock data for development, or populating test environments.

Arguments:

  • $PRODUCT: The product or system name
  • $DATASET_TYPE: Type of data (e.g., customer feedback, transactions, user profiles)
  • $ROWS: Number of rows to generate (default: 100)
  • $COLUMNS: Specific columns or fields to include
  • $FORMAT: Output format (CSV, JSON, SQL, Python script)
  • $CONSTRAINTS: Additional constraints or business rules

Step-by-Step Process

  1. Identify dataset type - Understand the data domain
  2. Define column specifications - Names, data types, and value ranges
  3. Determine row count - How many sample records needed
  4. Select output format - CSV, JSON, SQL INSERT, or Python script
  5. Apply realistic patterns - Ensure data looks authentic and valid
  6. Add business constraints - Respect business logic and relationships
  7. Generate or script data - Create executable output
  8. Validate output - Ensure data quality and completeness

Template: Python Script Output

python
import csv
import json
from datetime import datetime, timedelta
import random

# Configuration
ROWS = $ROWS
FILENAME = "$DATASET_TYPE.csv"

# Column definitions with realistic value generators
columns = {
    "id": "auto-increment",
    "name": "first_last_name",
    "email": "email",
    "created_at": "timestamp",
    # Add more columns...
}

def generate_dataset():
    """Generate realistic dummy dataset"""
    data = []
    for i in range(1, ROWS + 1):
        record = {
            "id": f"U{i:06d}",
            # Generate values based on column definitions
        }
        data.append(record)
    return data

def save_as_csv(data, filename):
    """Save dataset as CSV"""
    with open(filename, 'w', newline='') as f:
        writer = csv.DictWriter(f, fieldnames=data[0].keys())
        writer.writeheader()
        writer.writerows(data)

if __name__ == "__main__":
    dataset = generate_dataset()
    save_as_csv(dataset, FILENAME)
    print(f"Generated {len(dataset)} records in {FILENAME}")

Example Dataset Specification

Dataset Type: Customer Feedback

Columns:

  • feedback_id (auto-increment, U001, U002...)
  • customer_name (realistic names)
  • email (valid email format)
  • feedback_date (dates last 90 days)
  • rating (1-5 stars)
  • category (Bug, Feature Request, Complaint, Praise)
  • text (realistic feedback)
  • product (electronics, clothing, home)

Constraints:

  • Ratings skewed: 40% 5-star, 30% 4-star, 20% 3-star, 10% 1-2 star
  • Bug category only with ratings 1-3
  • Feature requests only with ratings 3-5
  • Email domains realistic (gmail, yahoo, company.com)

Output Deliverables

  • Ready-to-execute Python script OR direct data file
  • CSV file with proper headers and formatting
  • JSON file with valid structure and types
  • SQL INSERT statements for database population
  • Data validation and constraint compliance
  • Realistic, business-appropriate values
  • Documentation of data generation logic
  • Quick-start instructions for using the dataset

Output Formats

CSV: Flat tabular format, easy to import into spreadsheets and databases

JSON: Nested structure, ideal for APIs and NoSQL databases

SQL: INSERT statements, directly executable on relational databases

Python Script: Executable generator for custom or large datasets

© phuryn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in pm-execution/skills/dummy-dataset of phuryn/pm-skills.

Open the folder on GitHubat commit 8607e3b

Compare with similar skills

Dummy Dataset Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dummy Dataset Generator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dummy Dataset Generator this skillphuryn/pm-skills27k—~983Automated safety check: PassMIT
Phy Test Data FactoryLeoYeAI/openclaw-master-skills2.2k—~5.5kAutomated safety check: PassApache-2.0
Django Filter Benchmarksaleor/saleor23k—~2.3kAutomated safety check: PassBSD-3-Clause
Dummy Datasetkillvxk/pm-skills-zh167—~595Automated safety check: PassMIT
Update Megatron Golden ValuesNVIDIA/Megatron-LM18k—~2.8kAutomated safety check: PassCustom licence
Forge Project Inityaomindong1996/forge-admin125—~1.9kAutomated safety check: NotesApache-2.0

Similar skills

  • Phy Test Data Factory

    LeoYeAI/openclaw-master-skills

    Schema-driven test data factory generator. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~5.5k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Benchmarks Django ORM filters in Saleor by generating bulk data, extracting the SQL and running EXPLAIN ANALYZE to check index usage.

    23k GitHub stars~2.3k tokensUpdated today
    DatabasesAuto-check passed
  • Dummy Dataset

    killvxk/pm-skills-zh

    生成用于测试的逼真虚拟数据集,支持自定义列、约束条件及输出格式(CSV、JSON、SQL、Python 脚本)。适用于创建测试数据、构建模拟数据集,或为开发和演示生成示例数据。

    167 GitHub stars~595 tokensUpdated 6 mo ago
    DatabasesAuto-check passed
  • Official

    Refreshes stored golden values from a GitHub Actions run, reports signed percentage changes per model, and writes a summary ready for a pull request description.

    18k GitHub stars~2.8k tokensUpdated today
    Testing & QAAuto-check passed
  • Forge Project Init

    yaomindong1996/forge-admin

    Bootstrap a new project from this Forge template end to end, or reset a Forge database into a clean template system.

    125 GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check: notes
  • Agent-Core Python Testing

    openJiuwen-ai/agent-core

    Pytest patterns for the agent-core codebase: a red-green-refactor workflow, conftest fixtures, custom marks, monkeypatch and patch mocking, and async tests.

    446 GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed

More from phuryn/pm-skills

All 67 skills in this repo
  • Reviews a diff by anchoring on agreements between two sides of a boundary, forcing a concrete violating execution, and refuting each finding before reporting it.

    27k GitHub stars~3.6k tokensUpdated 24 days ago
    Auto-check passed
  • A/B Test Analysis

    phuryn/pm-skills

    Validates an experiment's setup, works out lift, p-value and confidence interval from A/B test data, and recommends whether to ship, extend or stop.

    27k GitHub stars~893 tokensUpdated 24 days ago
    Auto-check passed
  • Team OKR Brainstorm

    phuryn/pm-skills

    Drafts three alternative sets of team OKRs, each with an inspiring objective and measurable key results, tied to the company strategy you provide.

    27k GitHub stars~1.1k tokensUpdated 24 days ago
    Auto-check passed
  • Analyzes uploaded cohort data to compute retention curves and feature adoption trends, builds heatmaps and charts, and suggests qualitative follow-up research.

    27k GitHub stars~1.3k tokensUpdated 24 days ago
    Auto-check passed
  • Grammar and Flow Checker

    phuryn/pm-skills

    Reviews a draft for grammar, logic and flow problems and returns located, prioritized fix suggestions without rewriting the whole text.

    27k GitHub stars~2.4k tokensUpdated 24 days ago
    Auto-check passed
  • Builds a structured customer interview script with opening, warm-up, jobs-to-be-done exploration and wrap-up sections, following Mom Test rules against leading questions.

    27k GitHub stars~1.2k tokensUpdated 24 days ago
    Auto-check passed

Works with

Categories

Questions about Dummy Dataset Generator

What does Dummy Dataset Generator do?

Generates realistic test datasets with custom columns, row counts and business constraints, output as CSV, JSON, SQL inserts or a runnable Python script. The agent works through a fixed process: identify the data domain, define column names, types and value ranges, decide the row count, pick an output format, apply realistic-looking patterns, apply any business constraints such as skewed rating distributions or category-dependent values, then generate the data or a script that generates it and validate the result. Parameters named in the skill are the product or system name, the dataset type, row count, which columns to include, the output format and any extra constraints, with a default of 100 rows when not specified.

When should I use Dummy Dataset Generator?

Dummy Dataset Generator fits situations like: creating a mock dataset for a demo or development environment; generating sample rows with realistic constraints for a new database table; producing test data in CSV, JSON or SQL insert format.

How do I install Dummy Dataset Generator in Claude Code?

Run `npx skills add phuryn/pm-skills --skill dummy-dataset -a claude-code`. Or copy the skill folder (pm-execution/skills/dummy-dataset in phuryn/pm-skills) into .claude/skills/dummy-dataset in your project. Claude Code loads it when a task matches its description.

How do I install Dummy Dataset Generator in Codex?

Run `npx skills add phuryn/pm-skills --skill dummy-dataset -a codex`. Or copy the skill folder (pm-execution/skills/dummy-dataset in phuryn/pm-skills) into .agents/skills/dummy-dataset in your project. Codex loads it when a task matches its description.

Can I use Dummy Dataset Generator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add phuryn/pm-skills --skill dummy-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dummy-dataset, .gemini/skills/dummy-dataset, .github/skills/dummy-dataset and .opencode/skills/dummy-dataset in your project.

What does Dummy Dataset Generator need to run?

SKILL.md names no scripts, command-line tools or credentials: Dummy Dataset Generator is instructions for the agent only.

Does Dummy Dataset Generator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dummy Dataset Generator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dummy Dataset Generator use?

Dummy Dataset Generator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dummy Dataset Generator use?

About 983 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dummy Dataset Generator?

Skills that share tags, products or a category with Dummy Dataset Generator: Phy Test Data Factory (LeoYeAI/openclaw-master-skills, 2.2k stars), Django Filter Benchmark (saleor/saleor, 23k stars), Dummy Dataset (killvxk/pm-skills-zh, 167 stars) and Update Megatron Golden Values (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dummy Dataset Generator?

phuryn (a GitHub user) maintains it in phuryn/pm-skills, which has 26,845 GitHub stars. The repository holds 67 skills in this directory. The repository was last updated on September 14, 2026.

Source: phuryn/pm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.