Agent skill

Database Seeding

by seb1n in seb1n/awesome-ai-agent-skills

Populate databases with realistic, reproducible test data for development, testing, and staging environments.

MITAuto-check passedTesting & QA

Install Database Seeding

skills CLI
$ npx skills add seb1n/awesome-ai-agent-skills --skill database-seeding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install seb1n/awesome-ai-agent-skills database-seeding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/database/database-seeding .claude/skills/database-seeding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
database-seeding
GitHub stars
206
Token cost
~2.8k tokens
SKILL.md length
784 words
Files
1
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Populate databases with realistic, reproducible test data for development, testing, and staging environments.

  • Works in 6 steps: Analyze the target schema: Inspect the… → Design the seed data strategy: Choose… → Generate seed scripts: Write seed… → …
  • The user requests database seeding
  • SKILL.md covers Workflow, Supported Technologies, Usage and Examples, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Database Seeding is an agent skill from seb1n/awesome-ai-agent-skills. Populate databases with realistic, reproducible test data for development, testing, and staging environments. Use when the user requests database seeding or provides relevant inputs for this workflow.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Test data and fixtures. The repository describes itself as: 103 ready-to-use AI agent skills for Claude Code, OpenAI Codex, Gemini CLI, Cursor, GitHub Copilot, Windsurf, and other Agent Skills-compatible tools. Complete SKILL.md… The licence is MIT.

When your agent uses it

  • The user requests database seeding
  • Provides relevant inputs for this workflow

Example prompts

  • “/database-seeding”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Analyze the target schema: Inspect the database schema to identify all tables, their columns, data types, constraints (NOT NULL, UNIQUE…
  2. Design the seed data strategy: Choose the appropriate approach based on the use case. Use deterministic data with fixed seeds for…
  3. Generate seed scripts: Write seed scripts in the project's language (Python, JavaScript, SQL, etc.) that create data matching all schema…
  4. Ensure idempotency: Design scripts to be safely re-runnable. Use INSERT ON CONFLICT DO NOTHING, UPSERT patterns, or truncate-then-insert…
  5. Support environment-specific profiles: Create different seed profiles — a small dataset (10-50 records per table) for local development, a…
  6. Execute and verify: Run the seed script against the target database, verify row counts match expectations, and confirm relational…

What it can do on your machine

Read from SKILL.md and the folder at commit 75865a5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Database Seeding loads about 2.8k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 784 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from seb1n/awesome-ai-agent-skills at commit 75865a5, republished under its MIT licence (© seb1n). 784 words, ~2,829 tokens.

Download SKILL.mdSave it as .claude/skills/database-seeding/SKILL.md (or your agent's skills folder).
name
database-seeding
description
Populate databases with realistic, reproducible test data for development, testing, and staging environments. Use when the user requests database seeding or provides relevant inputs for this workflow.
license
MIT
metadata.author
AI Agent Skills Community
metadata.version
1.0.0

Database Seeding

This skill enables an AI agent to generate and insert realistic test data into databases for development, testing, and staging environments. The agent creates idempotent seed scripts using deterministic generators or faker libraries, handles relational data with proper foreign key ordering, supports environment-specific seed profiles (minimal dev data vs. large-scale load testing), and ensures seeds can be run repeatedly without duplicating data.

Workflow

  1. Analyze the target schema: Inspect the database schema to identify all tables, their columns, data types, constraints (NOT NULL, UNIQUE, CHECK, foreign keys), and relationships. Determine the correct insertion order to satisfy foreign key dependencies — parent tables must be seeded before child tables.

  2. Design the seed data strategy: Choose the appropriate approach based on the use case. Use deterministic data with fixed seeds for reproducible test suites. Use faker-based generation for realistic-looking development data. Use anonymized production snapshots for staging environments that need realistic data distributions. Define the volume of data for each table.

  3. Generate seed scripts: Write seed scripts in the project's language (Python, JavaScript, SQL, etc.) that create data matching all schema constraints. Use the Faker library or equivalent for realistic names, emails, addresses, and dates. Handle unique constraints by generating unique values or using sequence-based patterns. Wrap inserts in transactions for atomicity.

  4. Ensure idempotency: Design scripts to be safely re-runnable. Use INSERT ON CONFLICT DO NOTHING, UPSERT patterns, or truncate-then-insert strategies. Check for existing data before inserting to avoid duplicates or constraint violations on repeated runs.

  5. Support environment-specific profiles: Create different seed profiles — a small dataset (10-50 records per table) for local development, a medium dataset (1,000-10,000 records) for integration testing, and a large dataset (100K+ records) for performance testing. Control the profile via environment variables or command-line arguments.

  6. Execute and verify: Run the seed script against the target database, verify row counts match expectations, and confirm relational integrity by checking that all foreign keys reference existing rows. Log the seeding results with counts per table.

Supported Technologies

  • Python: Faker, Factory Boy, SQLAlchemy, psycopg2
  • JavaScript/TypeScript: @faker-js/faker, Prisma seed, Knex seed files, TypeORM
  • SQL: Raw INSERT statements, COPY FROM CSV
  • Ruby: FactoryBot, Faker gem, Rails db:seed
  • Frameworks: Django fixtures, Laravel seeders, Rails seeds.rb

Usage

Provide the database schema (or point to your migration files) and specify the target environment and desired data volume. The agent will generate a complete seed script that respects all constraints and relationships. You can request specific data characteristics (e.g., "include users from multiple time zones" or "create orders spanning the last 12 months").

Examples

Example 1: Python Seed Script Using Faker

Request: Seed a PostgreSQL database with users, products, and orders for development.

python
"""seed.py — Seed development database with realistic test data."""
import random
from datetime import datetime, timedelta
from faker import Faker
import psycopg2

fake = Faker()
Faker.seed(42)  # Deterministic output for reproducibility
random.seed(42)

DB_CONFIG = {
    "host": "localhost",
    "port": 5432,
    "dbname": "dev_db",
    "user": "dev_user",
    "password": "dev_password",
}

NUM_USERS = 50
NUM_PRODUCTS = 30
NUM_ORDERS = 100


def seed():
    conn = psycopg2.connect(**DB_CONFIG)
    cur = conn.cursor()

    # Seed users
    user_ids = []
    for _ in range(NUM_USERS):
        cur.execute(
            """INSERT INTO users (email, password_hash, full_name, created_at)
               VALUES (%s, %s, %s, %s)
               ON CONFLICT (email) DO NOTHING
               RETURNING id""",
            (
                fake.unique.email(),
                fake.sha256(),
                fake.name(),
                fake.date_time_between(start_date="-2y", end_date="now"),
            ),
        )
        row = cur.fetchone()
        if row:
            user_ids.append(row[0])

    # Seed products
    product_ids = []
    for i in range(NUM_PRODUCTS):
        cur.execute(
            """INSERT INTO products (name, description, price, stock_quantity, sku)
               VALUES (%s, %s, %s, %s, %s)
               ON CONFLICT (sku) DO NOTHING
               RETURNING id""",
            (
                fake.catch_phrase(),
                fake.paragraph(nb_sentences=3),
                round(random.uniform(9.99, 499.99), 2),
                random.randint(0, 500),
                f"SKU-{i+1:05d}",
            ),
        )
        row = cur.fetchone()
        if row:
            product_ids.append(row[0])

    # Seed orders with order items
    statuses = ["pending", "confirmed", "shipped", "delivered"]
    for _ in range(NUM_ORDERS):
        user_id = random.choice(user_ids)
        status = random.choice(statuses)
        items = random.sample(product_ids, k=random.randint(1, 5))
        total = 0.0

        cur.execute(
            """INSERT INTO orders (user_id, status, total_amount, shipping_address, ordered_at)
               VALUES (%s, %s, 0, %s, %s) RETURNING id""",
            (user_id, status, fake.address(), fake.date_time_between("-1y", "now")),
        )
        order_id = cur.fetchone()[0]

        for pid in items:
            qty = random.randint(1, 4)
            price = round(random.uniform(9.99, 499.99), 2)
            total += qty * price
            cur.execute(
                """INSERT INTO order_items (order_id, product_id, quantity, unit_price)
                   VALUES (%s, %s, %s, %s)""",
                (order_id, pid, qty, price),
            )

        cur.execute(
            "UPDATE orders SET total_amount = %s WHERE id = %s", (round(total, 2), order_id)
        )

    conn.commit()
    cur.close()
    conn.close()
    print(f"Seeded {len(user_ids)} users, {len(product_ids)} products, {NUM_ORDERS} orders.")


if __name__ == "__main__":
    seed()
Example 2: SQL Seed File with Realistic Test Data

Request: Create a plain SQL seed file for a small development dataset.

sql
-- seed.sql — Idempotent seed data for local development
-- Run with: psql -U dev_user -d dev_db -f seed.sql

BEGIN;

-- Users
INSERT INTO users (id, email, password_hash, full_name, created_at) VALUES
  (1, 'alice@example.com',  'hash_alice',  'Alice Johnson',  '2024-03-15 09:00:00'),
  (2, 'bob@example.com',    'hash_bob',    'Bob Martinez',   '2024-05-20 14:30:00'),
  (3, 'carol@example.com',  'hash_carol',  'Carol Chen',     '2024-07-01 11:15:00'),
  (4, 'dave@example.com',   'hash_dave',   'Dave Okafor',    '2024-09-10 08:45:00'),
  (5, 'eve@example.com',    'hash_eve',    'Eve Andersson',  '2024-11-28 16:00:00')
ON CONFLICT (id) DO NOTHING;

-- Products
INSERT INTO products (id, name, description, price, stock_quantity, sku) VALUES
  (1, 'Wireless Keyboard',   'Bluetooth mechanical keyboard',  79.99,  150, 'SKU-00001'),
  (2, 'USB-C Hub',           '7-in-1 USB-C docking station',   49.99,  300, 'SKU-00002'),
  (3, 'Noise-Cancelling Headphones', 'Over-ear ANC headphones', 199.99, 75, 'SKU-00003'),
  (4, '4K Monitor',          '27-inch IPS 4K display',         399.99,  40, 'SKU-00004'),
  (5, 'Laptop Stand',        'Adjustable aluminum stand',       34.99, 200, 'SKU-00005')
ON CONFLICT (id) DO NOTHING;

-- Orders
INSERT INTO orders (id, user_id, status, total_amount, shipping_address, ordered_at) VALUES
  (1, 1, 'delivered',  129.98, '123 Oak St, Portland, OR 97201',   '2024-12-01 10:00:00'),
  (2, 2, 'shipped',    199.99, '456 Elm Ave, Austin, TX 78701',    '2025-01-05 14:20:00'),
  (3, 3, 'confirmed',  484.98, '789 Pine Rd, Seattle, WA 98101',   '2025-01-10 09:30:00'),
  (4, 1, 'pending',     49.99, '123 Oak St, Portland, OR 97201',   '2025-01-12 16:45:00')
ON CONFLICT (id) DO NOTHING;

-- Order items
INSERT INTO order_items (id, order_id, product_id, quantity, unit_price) VALUES
  (1, 1, 1, 1, 79.99),
  (2, 1, 2, 1, 49.99),
  (3, 2, 3, 1, 199.99),
  (4, 3, 4, 1, 399.99),
  (5, 3, 5, 1, 34.99),
  (6, 4, 2, 1, 49.99)
ON CONFLICT (id) DO NOTHING;

-- Reset sequences to avoid conflicts with future inserts
SELECT setval('users_id_seq',    (SELECT MAX(id) FROM users));
SELECT setval('products_id_seq', (SELECT MAX(id) FROM products));
SELECT setval('orders_id_seq',   (SELECT MAX(id) FROM orders));
SELECT setval('order_items_id_seq', (SELECT MAX(id) FROM order_items));

COMMIT;
Show full SKILL.md (322 more words)Show less

Best Practices

  • Set a fixed random seed (e.g., Faker.seed(42)) to produce deterministic data that makes test results reproducible and diffs in seed output meaningful.
  • Always respect foreign key ordering — insert parent rows before child rows. Map out the dependency graph before writing the script to avoid constraint violations.
  • Use ON CONFLICT DO NOTHING or UPSERT patterns to make seed scripts idempotent. Running the seed twice should produce the same result, not duplicate data or throw errors.
  • Separate seed profiles by environment — a small, fast seed for local development, a larger seed for CI integration tests, and an even larger one for load testing. Control via environment variables.
  • Never seed production databases with test data. Use environment checks (e.g., assert os.environ["ENV"] != "production") at the top of seed scripts as a safety guard.
  • Reset auto-increment sequences after seeding with explicit IDs to prevent primary key collisions when the application inserts new rows.

Edge Cases

  • Unique constraint collisions with faker: Faker does not guarantee uniqueness across large datasets. Use fake.unique.email() or append a counter to generated values to avoid duplicates. Reset the unique tracker between test runs with fake.unique.clear().
  • Circular foreign keys: When table A references table B and table B references table A, insert rows into both tables with nullable FK columns first, then update the FK values in a second pass.
  • Large seed datasets and performance: For seeding more than 10,000 rows, use bulk insert methods (COPY in PostgreSQL, LOAD DATA INFILE in MySQL) rather than individual INSERT statements. Disable indexes and constraints during bulk load, then re-enable them afterward.
  • Time-dependent test data: If your application logic depends on dates (e.g., "orders from the last 30 days"), generate dates relative to the current date rather than hardcoded dates that will become stale.
  • Seeding binary or file data: For tables that reference uploaded files or images, seed with placeholder file paths or small base64-encoded test images rather than attempting to generate full binary content.

© seb1n, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in database/database-seeding of seb1n/awesome-ai-agent-skills.

Open the folder on GitHubat commit 75865a5

Compare with similar skills

Database Seeding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Database Seeding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Database Seeding this skillseb1n/awesome-ai-agent-skills206—~2.8kAutomated safety check: PassMIT
Fs Fixtureprivatenumber/fs-fixture100—~1.2kAutomated safety check: PassMIT
Dev Tenant APInightscout/nocturne139—~1.4kAutomated safety check: PassNone
Rsibench Data Factoryevolvent-ai/RSIBench-Data169—~640Automated safety check: NotesNone
Eval Designagentscope-ai/OpenJudge868—~2.8kAutomated safety check: WarnApache-2.0
Data GenerationRed-Hat-AI-Innovation-Team/sdg_hub164—~381Automated safety check: PassApache-2.0

Similar skills

  • Fs Fixture

    privatenumber/fs-fixture

    Create disposable file system test fixtures from objects, templates, or empty directories with automatic cleanup.

    100 GitHub stars~1.2k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Dev Tenant API

    nightscout/nocturne

    Interact with Nocturne's local dev-only API: seed a loginable tenant preloaded with realistic sample data, obtain a browser session (loginLink) or bearer token headlessly, export/re-seed the dev…

    139 GitHub stars~1.4k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Rsibench Data Factory

    evolvent-ai/RSIBench-Data

    Use inside RSIBench-Data when testing whether an automation agent can improve a target model on a configured benchmark through synthetic Tinker SFT data, Tinker sampling, and E2B-based Harbor…

    169 GitHub stars~640 tokensUpdated 1 mo ago
    Testing & QAAuto-check: notes
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    868 GitHub stars~2.8k tokensUpdated 27 days ago
    Testing & QAAuto-check: warnings
  • Data Generation

    Red-Hat-AI-Innovation-Team/sdg_hub

    A skill your agent uses when the user wants to run synthetic data generation via scripts — detect environment, execute a flow, and present results.

    164 GitHub stars~381 tokensUpdated today
    Testing & QAAuto-check passed
  • App Control

    ad-repo/nullplayer

    Launch, configure, drive, screenshot and measure the running NullPlayer app.

    157 GitHub stars~9k tokensUpdated today
    Testing & QAAuto-check passed

More from seb1n/awesome-ai-agent-skills

All 91 skills in this repo
  • Agent Red Teaming

    seb1n/awesome-ai-agent-skills

    Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings.

    206 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Eu AI Act Readiness

    seb1n/awesome-ai-agent-skills

    Build a preliminary, evidence-based EU AI Act readiness assessment across AI-system inventory, territorial scope, operator roles, prohibited-practice screening, risk classification, transparency…

    206 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Human In The Loop

    seb1n/awesome-ai-agent-skills

    Design and verify auditable human oversight, approval gates, escalation paths, and safe state transitions for AI agent workflows.

    206 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • MCP Server Building

    seb1n/awesome-ai-agent-skills

    Design, implement, harden, and verify Model Context Protocol (MCP) servers with precise tool contracts, least-privilege authorization, safe transports, structured errors, and interoperability tests.

    206 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Skill Supply Chain Audit

    seb1n/awesome-ai-agent-skills

    Audit agent skills, plugins, prompts, manifests, scripts, dependencies, and bundled assets for provenance, prompt-injection, permission, execution, exfiltration, persistence, and update risk.

    206 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spreadsheet Analysis

    seb1n/awesome-ai-agent-skills

    Inspect, profile, clean, reconcile, analyze, visualize, and verify spreadsheet data while preserving formulas, formatting, types, and source files.

    206 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Database Seeding

What does Database Seeding do?

Populate databases with realistic, reproducible test data for development, testing, and staging environments. Database Seeding is an agent skill from seb1n/awesome-ai-agent-skills. Populate databases with realistic, reproducible test data for development, testing, and staging environments.

When should I use Database Seeding?

Database Seeding fits situations like: the user requests database seeding; provides relevant inputs for this workflow.

How do I install Database Seeding in Claude Code?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill database-seeding -a claude-code`. Or copy the skill folder (database/database-seeding in seb1n/awesome-ai-agent-skills) into .claude/skills/database-seeding in your project. Claude Code loads it when a task matches its description.

How do I install Database Seeding in Codex?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill database-seeding -a codex`. Or copy the skill folder (database/database-seeding in seb1n/awesome-ai-agent-skills) into .agents/skills/database-seeding in your project. Codex loads it when a task matches its description.

Can I use Database Seeding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add seb1n/awesome-ai-agent-skills --skill database-seeding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/database-seeding, .gemini/skills/database-seeding, .github/skills/database-seeding and .opencode/skills/database-seeding in your project.

What does Database Seeding need to run?

SKILL.md names no scripts, command-line tools or credentials: Database Seeding is instructions for the agent only. Our summary lists: Python 3.

Does Database Seeding access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Database Seeding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Database Seeding use?

Database Seeding is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Database Seeding use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Database Seeding?

Skills that share tags, products or a category with Database Seeding: Fs Fixture (privatenumber/fs-fixture, 100 stars), Dev Tenant API (nightscout/nocturne, 139 stars), Rsibench Data Factory (evolvent-ai/RSIBench-Data, 169 stars) and Eval Design (agentscope-ai/OpenJudge, 868 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Database Seeding?

seb1n (a GitHub user) maintains it in seb1n/awesome-ai-agent-skills, which has 206 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on August 9, 2026.

Source: seb1n/awesome-ai-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.