Agent skill

Generating Database Seed Data

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments.

MITAuto-check passedTesting & QA

Install Generating Database Seed Data

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill generating-database-seed-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace generating-database-seed-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/generating-database-seed-data .claude/skills/generating-database-seed-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
generating-database-seed-data
GitHub stars
2.8k
Token cost
~1.9k tokens
SKILL.md length
868 words
Files
4 (incl. scripts, references, assets)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments.

  • Works in 10 steps: Analyze the database schema to catalog… → Determine the seeding order by… → Map each column to an appropriate Faker… → …
  • Working with databases
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
  • With phrases like database

What it does

Generating Database Seed Data is an agent skill from jeremylongshore/tons-of-skills-marketplace. Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts, reference files and assets (for example `assets/README.md`, `references/README.md` and `scripts/README.md`). Compatibility notes: Designed for Claude Code

It sits in Testing & QA, covering Test data and fixtures. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Working with databases
  • With phrases like database

Example prompts

  • “database”
  • “schema”
  • “/generating-database-seed-data”

Requirements

  • Python 3
  • Node.js
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Grep, Glob, Bash(cmd:*)

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. Analyze the database schema to catalog all tables, columns, data types, constraints, and foreign key relationships. Build a dependency…
  2. Determine the seeding order by topologically sorting the dependency graph. Tables with no foreign keys are seeded first (users…
  3. Map each column to an appropriate Faker generator based on column name and data type
  4. Generate foreign key values by referencing previously inserted parent records. Store parent IDs in arrays during generation and randomly…
  5. Handle unique constraints by tracking generated values in a Set and regenerating on collision. For email addresses, append a counter or…
  6. Respect CHECK constraints and ENUM types by reading the allowed values from the schema and restricting random selection to valid options…
  7. Generate the seed script in the appropriate format
  8. Make seed scripts idempotent: wrap in a transaction, truncate target tables in reverse dependency order before inserting, or use upsert…
  9. Add configurable volume control: accept a scale factor parameter that multiplies base counts (scale=1: 100 users, scale=10: 1000 users)…
  10. Validate the generated seed data by running it against an empty database, then checking: all foreign key references resolve, unique…

What it can do on your machine

Read from SKILL.md and the folder at commit 80f86df. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Grep
    • Glob
    • Bash(cmd:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • fakerjs.dev
    • faker.readthedocs.io
    • prisma.io
    • knexjs.org
    • postgresql.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Generating Database Seed Data loads about 1.9k tokens when it runs, and up to ~1.9k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 868 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit 80f86df, republished under its MIT licence (© jeremylongshore). 868 words, ~1,897 tokens.

Download SKILL.mdSave it as .claude/skills/generating-database-seed-data/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
generating-database-seed-data
description
Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. it uses faker libraries to create realistic data, maintains relational integrity, and allows configurable data volumes. u... Use when working with databases or data models. Trigger with phrases like 'database', 'query', or 'schema'.
allowed-tools
Read, Write, Edit, Grep, Glob, Bash(cmd:*)
compatibility
Designed for Claude Code
version
1.22.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT
tags
database, testing, database-seed

Data Seeder Generator

Overview

Generate realistic database seed scripts that populate development and testing environments with representative data. This skill creates seed data that respects foreign key relationships, unique constraints, check constraints, and data type validations using Faker libraries (faker.js, Faker for Python, or raw SQL with random functions).

Prerequisites

  • Database schema definition (SQL DDL, ORM models, or Prisma schema) to understand table structures
  • Target database connection for schema introspection (optional, can work from DDL files)
  • Faker library available: @faker-js/faker (Node.js), faker (Python), or Bogus (.NET)
  • Knowledge of referential integrity constraints (foreign keys, cascades)
  • Target data volume per table (e.g., 100 users, 1000 orders, 5000 line items)

Instructions

  1. Analyze the database schema to catalog all tables, columns, data types, constraints, and foreign key relationships. Build a dependency graph where parent tables (referenced by foreign keys) must be seeded before child tables.

  2. Determine the seeding order by topologically sorting the dependency graph. Tables with no foreign keys are seeded first (users, categories, products), then tables referencing them (orders, reviews), then junction tables and deeply nested tables last.

  3. Map each column to an appropriate Faker generator based on column name and data type:

    • first_name, last_name -> faker.person.firstName(), faker.person.lastName()
    • email -> faker.internet.email() with unique enforcement
    • phone -> faker.phone.number()
    • address, city, state, zip -> faker.location.*
    • created_at, updated_at -> faker.date.between({ from: '2023-01-01', to: '2024-12-31' })
    • price, amount -> faker.commerce.price({ min: 1, max: 999 })
    • description, bio -> faker.lorem.paragraph()
    • status -> Random selection from CHECK constraint values or enum values
    • uuid -> faker.string.uuid()
  4. Generate foreign key values by referencing previously inserted parent records. Store parent IDs in arrays during generation and randomly select from them for child records. Ensure every parent has at least one child (if the relationship is expected) and distribute children realistically (e.g., Zipf distribution where some users have many orders, most have few).

  5. Handle unique constraints by tracking generated values in a Set and regenerating on collision. For email addresses, append a counter or use faker.internet.email({ firstName, lastName }) with unique names.

  6. Respect CHECK constraints and ENUM types by reading the allowed values from the schema and restricting random selection to valid options. For range constraints (CHECK (age >= 18 AND age <= 120)), configure Faker to generate within the valid range.

  7. Generate the seed script in the appropriate format:

    • Raw SQL: INSERT INTO users (name, email, ...) VALUES ('John Doe', 'john@example.com', ...); with proper escaping
    • TypeORM/Prisma: TypeScript seed file using prisma.user.createMany() or repository.save()
    • Django: Python fixtures in JSON format or management command
    • Knex: JavaScript seed file using knex('users').insert([...])
  8. Make seed scripts idempotent: wrap in a transaction, truncate target tables in reverse dependency order before inserting, or use upsert operations (ON CONFLICT DO NOTHING).

  9. Add configurable volume control: accept a scale factor parameter that multiplies base counts (scale=1: 100 users, scale=10: 1000 users). Maintain consistent ratios between related tables (1 user : 5 orders : 15 line items).

  10. Validate the generated seed data by running it against an empty database, then checking: all foreign key references resolve, unique constraints hold, check constraints pass, and row counts match expectations.

Output

  • Seed script files in SQL, TypeScript, Python, or JavaScript format
  • Faker configuration mapping columns to appropriate generators
  • Dependency order listing the correct table insertion sequence
  • Validation queries to verify seed data integrity after insertion
  • Volume configuration with scale factor and per-table row counts
Show full SKILL.md (326 more words)Show less

Error Handling

ErrorCauseSolution
Foreign key constraint violation during seedingChild records reference parent IDs that do not existVerify seeding order follows dependency graph; ensure parent seed completes before child seed starts
Unique constraint violationFaker generated duplicate values for unique columnsTrack generated values in a Set; use faker.helpers.unique() wrapper; append sequential suffix for high-volume unique fields
CHECK constraint violationGenerated value outside allowed range or not in enum listRead CHECK constraints from schema; configure Faker min/max ranges; restrict enum selection to valid values
Seed script too slow for large volumesIndividual INSERT statements instead of batch operationsUse batch inserts (INSERT INTO ... VALUES (...), (...), (...)); use COPY command for PostgreSQL; disable indexes during bulk insert
Unrealistic data distributionAll records have uniform random valuesUse weighted random selection for status fields; apply Zipf distribution for popularity-based relationships; generate time-series data with realistic patterns

Examples

Seeding an e-commerce database with 10,000 orders: Generate 500 users, 200 products across 15 categories, 10,000 orders (distributed over 12 months with higher volume in November-December), and 35,000 line items. Each order has 1-5 line items, prices follow a realistic distribution ($5-$500 with most under $50), and order statuses follow a funnel pattern (70% delivered, 15% shipped, 10% processing, 5% cancelled).

Creating test data for a multi-tenant SaaS application: Generate 5 tenants, each with 20-100 users, organization settings, and tenant-specific data. Tenant isolation is maintained in seed data by assigning all records to a specific tenant_id. One "demo" tenant has curated showcase data with meaningful names and descriptions.

Populating a social media prototype: Generate 1,000 users with profile photos (sample image URLs from picsum.photos), 5,000 posts with timestamps following a realistic posting pattern (more activity on weekdays, peak at noon), 15,000 comments with reply threading (30% of comments are replies to other comments), and 50,000 likes distributed by post popularity.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references, assets) in skills/.curated/generating-database-seed-data of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • assets/README.md
  • references/README.md
  • scripts/README.md

Open the folder on GitHubat commit 80f86df

Compare with similar skills

Generating Database Seed Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Generating Database Seed Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Generating Database Seed Data this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.9kAutomated safety check: PassMIT
Fs Fixtureprivatenumber/fs-fixture100—~1.2kAutomated safety check: PassMIT
Dev Tenant APInightscout/nocturne139—~1.4kAutomated safety check: PassNone
Rsibench Data Factoryevolvent-ai/RSIBench-Data171—~640Automated safety check: NotesNone
Eval Designagentscope-ai/OpenJudge870—~2.8kAutomated safety check: WarnApache-2.0
Data GenerationRed-Hat-AI-Innovation-Team/sdg_hub164—~381Automated safety check: PassApache-2.0

Similar skills

  • Fs Fixture

    privatenumber/fs-fixture

    Create disposable file system test fixtures from objects, templates, or empty directories with automatic cleanup.

    100 GitHub stars~1.2k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Dev Tenant API

    nightscout/nocturne

    Interact with Nocturne's local dev-only API: seed a loginable tenant preloaded with realistic sample data, obtain a browser session (loginLink) or bearer token headlessly, export/re-seed the dev…

    139 GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Rsibench Data Factory

    evolvent-ai/RSIBench-Data

    Use inside RSIBench-Data when testing whether an automation agent can improve a target model on a configured benchmark through synthetic Tinker SFT data, Tinker sampling, and E2B-based Harbor…

    171 GitHub stars~640 tokensUpdated 1 mo ago
    Testing & QAAuto-check: notes
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    870 GitHub stars~2.8k tokensUpdated 28 days ago
    Testing & QAAuto-check: warnings
  • Data Generation

    Red-Hat-AI-Innovation-Team/sdg_hub

    A skill your agent uses when the user wants to run synthetic data generation via scripts — detect environment, execute a flow, and present results.

    164 GitHub stars~381 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Official

    Use this skill before answering or editing whenever an MSTest v1/v2 project is being upgraded or repaired for v3.

    5.6k GitHub starsUsed in 1 repo~5.5k tokens
    Testing & QAAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Generating Database Seed Data

What does Generating Database Seed Data do?

Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments. Generating Database Seed Data is an agent skill from jeremylongshore/tons-of-skills-marketplace. Process this skill enables AI assistant to generate realistic test data and database seed scripts for development and testing environments.

When should I use Generating Database Seed Data?

Generating Database Seed Data fits situations like: working with databases; with phrases like database.

How do I install Generating Database Seed Data in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill generating-database-seed-data -a claude-code`. Or copy the skill folder (skills/.curated/generating-database-seed-data in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/generating-database-seed-data in your project. Claude Code loads it when a task matches its description.

How do I install Generating Database Seed Data in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill generating-database-seed-data -a codex`. Or copy the skill folder (skills/.curated/generating-database-seed-data in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/generating-database-seed-data in your project. Codex loads it when a task matches its description.

Can I use Generating Database Seed Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill generating-database-seed-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/generating-database-seed-data, .gemini/skills/generating-database-seed-data, .github/skills/generating-database-seed-data and .opencode/skills/generating-database-seed-data in your project.

What does Generating Database Seed Data need to run?

SKILL.md names no scripts, command-line tools or credentials: Generating Database Seed Data is instructions for the agent only. Our summary lists: Python 3; Node.js. Its frontmatter pre-approves these tools: Read, Write, Edit, Grep, Glob, Bash(cmd:*). Compatibility (from SKILL.md): Designed for Claude Code.

Does Generating Database Seed Data access the network?

SKILL.md names 5 domains. As links in the text: fakerjs.dev, faker.readthedocs.io, prisma.io, knexjs.org and postgresql.org. This is read from the text; nothing was executed.

Is Generating Database Seed Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Generating Database Seed Data use?

Generating Database Seed Data is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Generating Database Seed Data use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16 tokens, read only when the agent opens those files.

What are the alternatives to Generating Database Seed Data?

Skills that share tags, products or a category with Generating Database Seed Data: Fs Fixture (privatenumber/fs-fixture, 100 stars), Dev Tenant API (nightscout/nocturne, 139 stars), Rsibench Data Factory (evolvent-ai/RSIBench-Data, 171 stars) and Eval Design (agentscope-ai/OpenJudge, 870 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Generating Database Seed Data?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,825 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 9, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.