Agent skill

Dummy Dataset

by killvxk in killvxk/pm-skills-zh

生成用于测试的逼真虚拟数据集,支持自定义列、约束条件及输出格式(CSV、JSON、SQL、Python 脚本)。适用于创建测试数据、构建模拟数据集,或为开发和演示生成示例数据。

MITAuto-check passedDatabases

Install Dummy Dataset

skills CLI
$ npx skills add killvxk/pm-skills-zh --skill dummy-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install killvxk/pm-skills-zh dummy-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/killvxk/pm-skills-zh.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pm-execution/skills/dummy-dataset .claude/skills/dummy-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dummy-dataset
GitHub stars
167
Token cost
~595 tokens
SKILL.md length
113 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

生成用于测试的逼真虚拟数据集,支持自定义列、约束条件及输出格式(CSV、JSON、SQL、Python 脚本)。适用于创建测试数据、构建模拟数据集,或为开发和演示生成示例数据。

  • Works in 8 steps: 确定数据集类型 - 理解数据领域 → 定义列规格 - 名称、数据类型和取值范围 → 确定行数 - 需要多少条样本记录 → …
  • Tasks that involve Test data and fixtures
  • SKILL.md covers Step-by-Step Process(分步流程), Template: Python Script…, Example Dataset… and Output Deliverables(输出交付物), plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Dummy Dataset is an agent skill from killvxk/pm-skills-zh. 生成用于测试的逼真虚拟数据集,支持自定义列、约束条件及输出格式(CSV、JSON、SQL、Python 脚本)。适用于创建测试数据、构建模拟数据集,或为开发和演示生成示例数据。

Its SKILL.md is about 600 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering Test data and fixtures, SQL and CSV and tabular files. It works with Python and SQL. The repository describes itself as: PM Skills Marketplace 简体中文版 - 65个产品经理技能和36个工作流,翻译自 https://github.com/phuryn/pm-skills/. The licence is MIT.

When your agent uses it

  • Tasks that involve Test data and fixtures
  • Tasks that involve SQL
  • Tasks that involve CSV and tabular files

Example prompts

  • “/dummy-dataset”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. 确定数据集类型 - 理解数据领域
  2. 定义列规格 - 名称、数据类型和取值范围
  3. 确定行数 - 需要多少条样本记录
  4. 选择输出格式 - CSV、JSON、SQL INSERT 或 Python 脚本
  5. 应用真实规律 - 确保数据看起来真实有效
  6. 添加业务约束 - 遵守业务逻辑和关联关系
  7. 生成或脚本化数据 - 创建可执行的输出
  8. 验证输出 - 确保数据质量和完整性

What it can do on your machine

Read from SKILL.md and the folder at commit 5179784. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dummy Dataset loads about 595 tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 113 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~26
When it runs · the whole SKILL.md, loaded when a task matches
~595

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from killvxk/pm-skills-zh at commit 5179784, republished under its MIT licence (© killvxk). 113 words, ~595 tokens.

Download SKILL.mdSave it as .claude/skills/dummy-dataset/SKILL.md (or your agent's skills folder).
name
dummy-dataset
description
生成用于测试的逼真虚拟数据集,支持自定义列、约束条件及输出格式(CSV、JSON、SQL、Python 脚本)。适用于创建测试数据、构建模拟数据集,或为开发和演示生成示例数据。

虚拟数据集生成

生成用于测试的逼真虚拟数据集,支持自定义列、约束条件及输出格式(CSV、JSON、SQL、Python 脚本)。生成可直接执行的脚本或数据文件,即开即用。

适用场景: 创建测试数据、生成示例数据集、为开发构建逼真的模拟数据,或填充测试环境。

参数:

  • $PRODUCT:产品或系统名称
  • $DATASET_TYPE:数据类型(如客户反馈、交易记录、用户画像)
  • $ROWS:生成的行数(默认:100)
  • $COLUMNS:需要包含的具体列或字段
  • $FORMAT:输出格式(CSV、JSON、SQL、Python 脚本)
  • $CONSTRAINTS:附加约束条件或业务规则

Step-by-Step Process(分步流程)

  1. 确定数据集类型 - 理解数据领域
  2. 定义列规格 - 名称、数据类型和取值范围
  3. 确定行数 - 需要多少条样本记录
  4. 选择输出格式 - CSV、JSON、SQL INSERT 或 Python 脚本
  5. 应用真实规律 - 确保数据看起来真实有效
  6. 添加业务约束 - 遵守业务逻辑和关联关系
  7. 生成或脚本化数据 - 创建可执行的输出
  8. 验证输出 - 确保数据质量和完整性

Template: Python Script Output(Python 脚本输出模板)

python
import csv
import json
from datetime import datetime, timedelta
import random

# 配置
ROWS = $ROWS
FILENAME = "$DATASET_TYPE.csv"

# 列定义及逼真值生成器
columns = {
    "id": "auto-increment",
    "name": "first_last_name",
    "email": "email",
    "created_at": "timestamp",
    # 添加更多列...
}

def generate_dataset():
    """生成逼真的虚拟数据集"""
    data = []
    for i in range(1, ROWS + 1):
        record = {
            "id": f"U{i:06d}",
            # 根据列定义生成值
        }
        data.append(record)
    return data

def save_as_csv(data, filename):
    """将数据集保存为 CSV 格式"""
    with open(filename, 'w', newline='') as f:
        writer = csv.DictWriter(f, fieldnames=data[0].keys())
        writer.writeheader()
        writer.writerows(data)

if __name__ == "__main__":
    dataset = generate_dataset()
    save_as_csv(dataset, FILENAME)
    print(f"已在 {FILENAME} 中生成 {len(dataset)} 条记录")

Example Dataset Specification(数据集规格示例)

数据集类型: 客户反馈

列:

  • feedback_id(自增,U001、U002……)
  • customer_name(真实姓名)
  • email(有效邮箱格式)
  • feedback_date(过去 90 天内的日期)
  • rating(1-5 星)
  • category(缺陷、功能请求、投诉、好评)
  • text(真实的反馈内容)
  • product(电子产品、服装、家居)

约束条件:

  • 评分分布偏斜:40% 五星,30% 四星,20% 三星,10% 一二星
  • 缺陷类别仅出现在 1-3 星评分中
  • 功能请求仅出现在 3-5 星评分中
  • 邮箱域名真实(gmail、yahoo、company.com)

Output Deliverables(输出交付物)

  • 可直接执行的 Python 脚本,或直接的数据文件
  • 格式正确、带表头的 CSV 文件
  • 结构有效、类型正确的 JSON 文件
  • 可在数据库中直接执行的 SQL INSERT 语句
  • 数据验证和约束条件合规
  • 真实、符合业务实际的数据值
  • 数据生成逻辑说明文档
  • 快速上手使用指南

Output Formats(输出格式)

CSV: 平面表格格式,易于导入电子表格和数据库

JSON: 嵌套结构,适用于 API 和 NoSQL 数据库

SQL: INSERT 语句,可直接在关系型数据库上执行

Python 脚本: 可执行的生成器,适用于自定义或大型数据集

© killvxk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in pm-execution/skills/dummy-dataset of killvxk/pm-skills-zh.

Open the folder on GitHubat commit 5179784

Compare with similar skills

Dummy Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dummy Dataset compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dummy Dataset this skillkillvxk/pm-skills-zh167—~595Automated safety check: PassMIT
Chdb SQLvemetric/vemetric3941 repos~1.2kAutomated safety check: PassApache-2.0
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Django Filter Benchmarksaleor/saleor23k—~2.3kAutomated safety check: PassBSD-3-Clause
Rust SQL Testshencangsheng/easydb_app590—~1.2kAutomated safety check: PassMIT
Jeecg Onlformjeecgboot/skills237—~7.7kAutomated safety check: PassApache-2.0

Similar skills

  • Chdb SQL

    vemetric/vemetric

    A skill your agent uses when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse…

    394 GitHub starsUsed in 1 repo~1.2k tokens
    DatabasesAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Benchmarks Django ORM filters in Saleor by generating bulk data, extracting the SQL and running EXPLAIN ANALYZE to check index usage.

    23k GitHub stars~2.3k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Rust SQL Test

    shencangsheng/easydb_app

    Enforces unit test requirements for the Rust data-processing modules in src-tauri/src/sql/ (generator.rs, parse.rs) and src-tauri/src/reader/ (excel.rs and other readers).

    590 GitHub stars~1.2k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Jeecg Onlform

    jeecgboot/skills

    JeecgBoot Online表单(cgform)全生命周期管理——通过API自动创建/编辑数据库表和表单配置, 支持单表、主子表、树表,26种控件类型,以及JS/Java/SQL增强、权限配置、数据CRUD、积木报表集成。

    237 GitHub stars~7.7k tokensUpdated 20 days ago
    DatabasesAuto-check passed
  • Duckdb En

    aAAaqwq/AGI-Super-Team

    DuckDB CLI specialist for SQL analysis, data processing and file conversion.

    105 GitHub starsUsed in 1 repo~1.6k tokens
    DatabasesAuto-check passed

More from killvxk/pm-skills-zh

All 11 skills in this repo
  • Ab Test Analysis

    killvxk/pm-skills-zh

    分析 A/B 测试结果,涵盖统计显著性检验、样本量验证、置信区间计算及上线/延长/停止的决策建议。适用于评估实验结果、判断测试是否达到显著性、解读分流测试数据,或决定是否上线某个实验组。

    167 GitHub stars~480 tokensUpdated 6 mo ago
    Auto-check passed
  • Brainstorm Okrs

    killvxk/pm-skills-zh

    集思广益制定团队级 OKR(目标与关键成果),对齐公司目标——定性目标与可量化关键成果。适用于制定季度 OKR、将团队目标与公司战略对齐、起草目标,或学习如何编写有效的 OKR。

    167 GitHub stars~533 tokensUpdated 6 mo ago
    Auto-check passed
  • Cohort Analysis

    killvxk/pm-skills-zh

    对用户参与度数据执行同期群分析——留存曲线、功能采用趋势及分层洞察。适用于按同期群分析用户留存、研究功能随时间的采用情况、调查流失规律,或识别参与度趋势。

    167 GitHub stars~562 tokensUpdated 6 mo ago
    Auto-check passed
  • Interview Script

    killvxk/pm-skills-zh

    创建结构化的用户访谈脚本,包含 JTBD(用户待办任务)探测问题、破冰环节、核心探索及收尾部分。遵循"老妈测试"原则——不引导、不推销,聚焦过去的实际行为。适用于准备用户访谈、创建访谈指南或规划探索性研究时使用。

    167 GitHub stars~547 tokensUpdated 6 mo ago
    Auto-check passed
  • Metrics Dashboard

    killvxk/pm-skills-zh

    定义并设计产品指标看板,包含核心指标、数据来源、可视化类型和告警阈值。适用于创建指标看板、定义 KPI(关键绩效指标)、搭建产品分析体系或制定数据监控计划时使用。

    167 GitHub stars~741 tokensUpdated 6 mo ago
    Auto-check passed
  • Pre Mortem

    killvxk/pm-skills-zh

    对 PRD(产品需求文档)或发布计划进行事前剖析(Pre-mortem)。将风险分类为老虎(真实问题)、纸老虎(被夸大的担忧)和大象(未被说出的隐忧),并按发布阻断、快速跟进或持续跟踪进行分级。适用于发布准备、对产品计划进行压力测试,或识别可能出错的地方。

    167 GitHub stars~475 tokensUpdated 6 mo ago
    Auto-check passed

Works with

Questions about Dummy Dataset

What does Dummy Dataset do?

生成用于测试的逼真虚拟数据集,支持自定义列、约束条件及输出格式(CSV、JSON、SQL、Python 脚本)。适用于创建测试数据、构建模拟数据集,或为开发和演示生成示例数据。. Dummy Dataset is an agent skill from killvxk/pm-skills-zh.

When should I use Dummy Dataset?

Dummy Dataset fits situations like: tasks that involve Test data and fixtures; tasks that involve SQL; tasks that involve CSV and tabular files.

How do I install Dummy Dataset in Claude Code?

Run `npx skills add killvxk/pm-skills-zh --skill dummy-dataset -a claude-code`. Or copy the skill folder (pm-execution/skills/dummy-dataset in killvxk/pm-skills-zh) into .claude/skills/dummy-dataset in your project. Claude Code loads it when a task matches its description.

How do I install Dummy Dataset in Codex?

Run `npx skills add killvxk/pm-skills-zh --skill dummy-dataset -a codex`. Or copy the skill folder (pm-execution/skills/dummy-dataset in killvxk/pm-skills-zh) into .agents/skills/dummy-dataset in your project. Codex loads it when a task matches its description.

Can I use Dummy Dataset in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add killvxk/pm-skills-zh --skill dummy-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dummy-dataset, .gemini/skills/dummy-dataset, .github/skills/dummy-dataset and .opencode/skills/dummy-dataset in your project.

What does Dummy Dataset need to run?

SKILL.md names no scripts, command-line tools or credentials: Dummy Dataset is instructions for the agent only. Our summary lists: Python 3.

Does Dummy Dataset access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dummy Dataset safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dummy Dataset use?

Dummy Dataset is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dummy Dataset use?

About 595 tokens (SKILL.md is roughly 2.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dummy Dataset?

Skills that share tags, products or a category with Dummy Dataset: Chdb SQL (vemetric/vemetric, 394 stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Django Filter Benchmark (saleor/saleor, 23k stars) and Rust SQL Test (shencangsheng/easydb_app, 590 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dummy Dataset?

killvxk (a GitHub user) maintains it in killvxk/pm-skills-zh, which has 167 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on March 16, 2026.

Source: killvxk/pm-skills-zh on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.