Agent skill

Test Data Management

by petrkindlmann in petrkindlmann/qa-skills

Create and manage test data with factory patterns, fixture strategies, data anonymization, and synthetic data generation.

MITAuto-check passedTesting & QA

Install Test Data Management

skills CLI
$ npx skills add petrkindlmann/qa-skills --skill test-data-management -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install petrkindlmann/qa-skills test-data-management --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-data-management .claude/skills/test-data-management && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-data-management
GitHub stars
165
Token cost
~4.4k tokens
SKILL.md length
2,158 words
Files
3 (incl. references)
Skills in repo
45
Repo updated
First seen
Licence
MIT

At a glance

Create and manage test data with factory patterns, fixture strategies, data anonymization, and synthetic data generation.

  • Works in 5 steps: Each Test Owns Its Data → Factories Over Fixtures for Dynamic Data → Anonymize Production Data Before Use → …
  • Data anonymization. Not for: migration/integrity testing of the DB itself — use database-testing
  • SKILL.md covers Quick Route, Discovery Questions, Core Principles and Factory Patterns, plus 5 more sections
  • Calls psql, pytest and supabase

What it does

Test Data Management is an agent skill from petrkindlmann/qa-skills. Create and manage test data with factory patterns, fixture strategies, data anonymization, and synthetic data generation. Covers Fishery (TypeScript), FactoryBot (Ruby), Factory Boy (Python), database seeding, cleanup strategies, and GDPR-compliant data handling. Use when: "test data," "fixtures," "factories," "seed data," "synthetic data," "test database," "data anonymization." Not for: migration/integrity testing of the DB itself — use database-testing; environment provisioning and database branching strategy —…

Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/factories.md` and `references/seeding-and-synthetic.md`).

It sits in Testing & QA, covering Test data and fixtures. It works with Python, Ruby and TypeScript. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.

When your agent uses it

  • Data anonymization. Not for: migration/integrity testing of the DB itself — use database-testing
  • Environment provisioning and database branching strategy — use test-environments

Example prompts

  • “test data,”
  • “fixtures,”
  • “factories,”
  • “/test-data-management”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Each Test Owns Its Data
  2. Factories Over Fixtures for Dynamic Data
  3. Anonymize Production Data Before Use
  4. Deterministic Data Enables Reproducible Tests
  5. Minimize Data, Maximize Signal

What it can do on your machine

Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • psql
    • pytest
    • supabase
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use supabase and npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Data Management loads about 4.4k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 159 tokens; SKILL.md has 2,158 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~159
When it runs · the whole SKILL.md, loaded when a task matches
~4.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,158 words, ~4,425 tokens.

Download SKILL.mdSave it as .claude/skills/test-data-management/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
test-data-management
description
Create and manage test data with factory patterns, fixture strategies, data anonymization, and synthetic data generation. Covers Fishery (TypeScript), FactoryBot (Ruby), Factory Boy (Python), database seeding, cleanup strategies, and GDPR-compliant data handling. Use when: "test data," "fixtures," "factories," "seed data," "synthetic data," "test database," "data anonymization." Not for: migration/integrity testing of the DB itself — use database-testing; environment provisioning and database branching strategy — use test-environments. Related: test-environments, database-testing, api-testing, unit-testing.
license
MIT
metadata.author
kindlmann
metadata.version
2.0
metadata.category
infrastructure
<objective>
Create, maintain, and clean up test data that is deterministic, isolated, realistic, and safe. Good test data is the foundation of reliable tests -- without it, tests are either flaky (shared mutable state), unrealistic (hardcoded nonsense values), or dangerous (production PII in test environments). This skill delivers factories, fixtures, idempotent seeds, anonymization pipelines, and cleanup strategies that survive parallel execution.
</objective>

Quick Route

SituationGo to
Need fresh entity data with per-test overridesFactory Patterns → references/factories.md
Mocking an API response or golden fileFixture Strategies → references/factories.md
Copying production data anywhere non-prodData Anonymization
Populating a test DB / reference data idempotentlyDatabase Seeding → references/seeding-and-synthetic.md
Cleaning up after tests / parallel isolationCleanup Strategies → references/seeding-and-synthetic.md
Generating edge cases and boundary valuesSynthetic Data → references/seeding-and-synthetic.md

Discovery Questions

Before designing a test data strategy, understand the current state. Check .agents/qa-project-context.md first -- if it exists, use it as the foundation and skip questions already answered there.

Current Data Practices
  • How is test data created today? (manually, scripts, copy of production, none)
  • Do tests share data or does each test create its own?
  • How is test data cleaned up? (truncate, rollback, manual, never)
  • Are there seed scripts? Are they idempotent?
Privacy and Compliance
  • Does the product handle PII? (names, emails, addresses, phone numbers, SSNs)
  • Are there GDPR, HIPAA, PCI-DSS, or other data protection requirements?
  • Is production data ever used in test environments?
Scale and Complexity
  • How large are the test datasets? (dozens of records, thousands, millions)
  • How complex are the data relationships? (simple CRUD, deep nested hierarchies, polymorphic)
  • Are there cross-service data dependencies? (microservices sharing data)

Core Principles

1. Each Test Owns Its Data

Tests that rely on pre-existing shared data are fragile. When Test A modifies shared data, Test B breaks. Every test should create exactly the data it needs, verify against that data, and clean up after itself. This enables parallel execution and eliminates ordering dependencies.

2. Factories Over Fixtures for Dynamic Data

Static fixtures (JSON/YAML files) are appropriate for reference data that does not change (country codes, currency lists). For entity data that tests create and manipulate (users, orders, products), use factory functions that generate fresh instances with sensible defaults and allow per-test overrides.

3. Anonymize Production Data Before Use

Production databases contain the most realistic data, but they also contain real user information. Never copy production data to test environments without anonymization. Replace PII with synthetic equivalents while preserving data distributions and relationships.

4. Deterministic Data Enables Reproducible Tests

Tests should produce the same results regardless of when or where they run. Avoid Math.random(), Date.now(), or auto-increment IDs in assertions. Use seeded random generators (faker.seed(n)), fixed timestamps, factory sequences, and -- when an ID must be a UUID you assert on -- a seeded faker.string.uuid() so it stays stable across runs.

5. Minimize Data, Maximize Signal

Create only the data each test needs. A test for user search does not need a complete user profile with billing address, payment method, and order history. Over-specified test data obscures the intent of the test and increases maintenance burden.


Factory Patterns

Factories are functions that produce test data with sensible defaults, allowing individual tests to override only what matters for their scenario. Fishery (2.4.0) is the default for TypeScript, FactoryBot (6.6.0) for Ruby, factory-boy (3.3.3) for Python; all pair with faker (v10.4.0) for realistic field values.

See references/factories.md for the full Fishery (with associations and deterministic UUIDs), FactoryBot (User + Product out_of_stock/discounted traits), Factory Boy (class Params + Trait), and Playwright fixture implementations. The shape every factory follows:

  • Defaults + overrides — Factory.define produces sensible defaults; tests pass overrides for the one field they care about (userFactory.build({ role: 'admin' })).
  • Sequences for unique fields — sequence (Fishery), sequence(:email) (FactoryBot), factory.Sequence (Factory Boy) — never hardcode IDs or emails.
  • Traits for variants — name common states (:admin, :inactive, out_of_stock, discounted) instead of spawning a fixture file per combination.
  • Associations — one factory builds another (an order builds its user), keeping referential structure without manual wiring.
When to Use Factories vs Fixtures
ScenarioFactoriesStatic Fixtures
Entity data that tests create/modifyYesNo
Reference data (countries, currencies, configs)NoYes
Data with many variations per testYesNo -- file explosion
Data with complex relationshipsYes -- associationsNo -- hard to maintain
API response mocksNoYes -- JSON fixtures
Snapshot/golden file comparisonsNoYes

Decision rule: If the data has a lifecycle (created, modified, deleted during tests), use a factory. If the data is read-only reference material, use a fixture file.


Fixture Strategies

Three fixture shapes, all in references/factories.md:

  • Static fixtures (JSON/YAML) — best for API response mocks (page.route + route.fulfill), config data, and golden file comparisons.
  • Dynamic fixtures (Playwright) — test.extend creates data via API before the test and deletes it after await use(...). The standard per-test setup/teardown.
  • Fixture composition — combine factory-built data (userFactory.build(), orderFactory.buildList(3)) inside a single test.extend that seeds and cleans up in one step.

Data Anonymization

When production data is needed for realistic testing, anonymize it before use.

PII Masking Rules
Data TypeAnonymization MethodExample
EmailFaker email with original domain patternjane.doe@acme.com -> user-7291@test.example.com
Full nameFaker nameJane Doe -> Alice Johnson
Phone numberFaker phone, preserve format+1-555-123-4567 -> +1-555-987-6543
AddressFaker address, preserve country/region123 Main St, NYC -> 456 Oak Ave, NYC
SSN/National IDTest pattern123-45-6789 -> 000-00-0001
Credit cardTest card numbers4111-... -> 4242-4242-4242-4242
Date of birthShift by fixed offset1990-03-15 -> 1987-07-22

The anonymization pipeline -- seeded Faker for determinism, an in-memory lookup table, parent-records-first ordering, and a wrapping transaction -- is in references/seeding-and-synthetic.md (Anonymization with Faker.js, Referential Integrity During Anonymization). Anonymizing a user's email must also update that email everywhere it is referenced (orders, comments, audit logs); process parents first, children second, using the same lookup, all inside one transaction.

GDPR Compliance Checklist
  • No real PII exists in any non-production environment
  • Anonymization is irreversible (no lookup table mapping back to originals is stored)
  • Anonymization preserves data distributions (age ranges, geographic spread) for realistic testing
  • Anonymized data cannot be re-identified through combination of quasi-identifiers
  • Data retention policies apply to test environments (auto-delete after N days)
  • The anonymization pipeline runs automatically, not manually (eliminates human error)

Database Seeding

Idempotent Seed Scripts

Seed scripts must be safe to run multiple times without duplicating data. Use upsert -- INSERT ... ON CONFLICT (natural_key) DO UPDATE SET ... -- keyed on a stable natural key, not the primary key. A DELETE-then-INSERT "reset" is not idempotent: it breaks foreign keys and reassigns serial IDs. See references/seeding-and-synthetic.md (Idempotent Seed Scripts) for the full INSERT ... ON CONFLICT (code) DO UPDATE countries/currencies example and the reasoning.

Database Branching (DB-as-a-Service)

If your prod DB lives on Neon, Supabase, or PlanetScale, branching can give a PR its own database instead of seeding from scratch -- but the providers differ on whether the branch carries data:

  • Neon Branching — copy-on-write Postgres branches in seconds, with data; ideal for ephemeral preview envs. The strongest "PR gets a real DB copy" story.
  • Supabase Branching — supabase branches create pr-123 clones schema and (optionally, from a backup) data; preview env points at the branch URL.
  • PlanetScale Branching — MySQL branches are schema-only by default (no data), so you still seed the branch. Note: PlanetScale removed its free Hobby tier (April 2024); MySQL now starts at ~$39/mo, Postgres ~$5/mo.

Pair with the Preview Environments pattern in test-environments.

Avoid: Snaplet (hosted) — shut down 31 Aug 2024; the team joined Supabase. @snaplet/seed lives on as supabase-community/seed (community-maintained, last meaningful release v0.98.0, July 2024, no feature work since). For new projects, prefer the DB-branching providers above plus factory-generated seeds.

Per-Test vs Per-Suite Data
StrategyWhen to UseProsCons
Per-test setup/teardownTests that modify dataFull isolation, parallel-safeSlower, more setup code
Per-suite seedRead-only reference dataFast, simpleCannot be modified by tests
Per-worker seedPlaywright parallel workersBalances speed and isolationRequires worker-scoped fixtures
Global seedEnvironment bootstrapRuns once, sets up baselineMust be idempotent, shared state risk

For the worker-scoped fixture (test.extend with { scope: 'worker' }) that powers per-worker seeding, see references/factories.md (Worker-Scoped Seeding).

Show full SKILL.md (855 more words)Show less
Cleanup Strategies
StrategyWhen to useSpeed
Transaction rollbackUnit/integration tests with direct DB accessFastest
Truncation (TRUNCATE ... CASCADE)Resetting tables between suitesMedium
API-based cleanupE2E tests with no direct DB accessSlowest

Transaction rollback cannot clean up E2E tests -- the app opens its own DB connections, so a test-side transaction can't undo the app's writes; use API-based cleanup (delete in reverse creation order) there. All three implementations are in references/seeding-and-synthetic.md (Cleanup Strategies).


Synthetic Data Generation

Factories should make it easy to generate edge cases and boundary values without hand-writing them per test. The reusable arrays and helpers -- edgeCaseStrings (empty, whitespace, very long, XSS, SQL injection, null/control chars, RTL override), edgeCaseDates, and boundaryValues(min, max) driving a test.each -- are in references/seeding-and-synthetic.md (Synthetic Data Generation).


Anti-Patterns

Shared Mutable Test Data

Multiple tests reading and writing the same database rows. Test A creates a user, Test B modifies it, Test C asserts on the original state and fails. Fix by having each test create its own data through factories.

Production Data Without Anonymization

Copying the production database to staging for "realistic testing." This violates GDPR, risks data breaches in less-secured environments, and creates compliance liability. Always anonymize before use, or generate synthetic data that matches production distributions.

Non-Deterministic Data

Using Math.random() or Date.now() in test data creation without seeding. Tests pass on Monday and fail on Tuesday because the random name generated happens to exceed a field length limit. Use seeded Faker instances and fixed timestamps.

No Cleanup Strategy

Tests that create data and never clean it up. The test database grows until it affects performance, or stale data causes false positives in other tests. Every data creation must have a corresponding cleanup.

Fixture File Explosion

Creating a separate JSON fixture file for every test variation. Instead of user-admin.json, user-inactive.json, user-admin-inactive.json, use a factory with traits. Fixtures should be reserved for static reference data and API response mocks.

Over-Specified Test Data

Creating a complete user object with 30 fields when the test only cares about role. This obscures intent and makes tests brittle. Factories with sensible defaults solve this: override only what the test cares about.

Hard-Coded IDs

Using userId: '1' in tests. This couples tests to database state and breaks when running in parallel (ID collision) or against a database with existing data. Use factory sequences or seeded UUIDs (see Core Principle 4).


Verification

Prove the data layer is deterministic, isolated, and PII-free, smallest check first:

  1. Seeds are idempotent — run the seed twice back-to-back and diff the row counts: psql -c "SELECT count(*) FROM countries" && <seed> && psql -c "SELECT count(*) FROM countries" returns the same number both times and exits 0. A growing count means a missing ON CONFLICT.
  2. No shared mutable state — run the suite under parallelism and randomized order: npx playwright test --workers=4 (or pytest -n auto -p randomly) stays green. A failure that only appears here is an ordering or shared-data dependency.
  3. Determinism holds — run the same data-generating test twice; with faker.seed(n) set, generated names/IDs/UUIDs match across runs. If they drift, an unseeded Faker call or Date.now()/crypto.randomUUID() leaked in.
  4. No real PII — grep -rE '@(gmail|outlook|yahoo)\.com|[0-9]{3}-[0-9]{2}-[0-9]{4}' tests/ fixtures/ returns nothing (real-looking emails and SSNs). Anything it finds is an anonymization gap.
  5. Cleanup returns to baseline — snapshot row counts before the suite, run it, snapshot again: the test DB is back to baseline with no orphaned records.

Done When

  • Every entity type the suite creates has a factory or fixture (no inline ad-hoc object literals in tests for shared entities -- grep the test dir for hand-built fixtures and confirm none remain).
  • Test data is isolated per test -- the suite passes with parallelism on (--workers=N / pytest -n auto) and under randomized order (--shuffle / -p randomly), proving no shared mutable state or ordering dependency.
  • Seed scripts are idempotent -- running the seed twice in a row produces the same row count and exits 0; the CI job runs them with no manual intervention.
  • No real PII used in test fixtures -- all sensitive data anonymized or synthetic (grep for production domains / real-looking SSNs returns nothing).
  • Data cleanup verified -- row counts in the test DB return to baseline after the suite (no orphaned records accumulate across runs).

Reference Files (in references/)

  • factories.md — Full Fishery (associations, deterministic UUIDs), FactoryBot (User + Product traits), Factory Boy (Params/Trait), static/dynamic/composed Playwright fixtures, and the worker-scoped seeding fixture.
  • seeding-and-synthetic.md — Idempotent ON CONFLICT seed script, Faker.js anonymization + referential-integrity pipeline, cleanup strategies (rollback / truncate / API), and synthetic edge-case + boundary-value generators.
  • unit-testing -- Unit tests are the primary consumer of factory-generated data; this skill provides the data layer.
  • api-testing -- API tests use both factories (for request bodies) and fixtures (for mocked responses).
  • playwright-automation -- E2E tests need test data seeded via API or fixtures before browser interaction.
  • test-reliability -- Deterministic test data eliminates a major source of test flakiness.
  • test-environments -- Owns environment provisioning and database-branching strategy (Neon, Supabase, PlanetScale) for preview envs; this skill owns the data that fills them.
  • database-testing -- Migration testing, data-integrity assertions, and Testcontainers for the database layer specifically -- go there to test the DB, come here to populate it.
  • ci-cd-integration -- Database seeding and cleanup must be integrated into CI pipeline stages.

© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/test-data-management of petrkindlmann/qa-skills.

  • SKILL.md
  • references/factories.md
  • references/seeding-and-synthetic.md

Open the folder on GitHubat commit b3bb61b

Compare with similar skills

Test Data Management next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Data Management compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Data Management this skillpetrkindlmann/qa-skills165—~4.4kAutomated safety check: PassMIT
Supercovsupercorp-ai/supercov1501 repos~415Automated safety check: PassMIT
Supercov Securitysupercorp-ai/supercov1501 repos~236Automated safety check: PassMIT
Testing Strategiesancoleman/ai-design-components526—~3.8kAutomated safety check: PassMIT
Testing Patternssoftspark/ai-toolkit179—~1.6kAutomated safety check: PassApache-2.0
Phy Test Data FactoryLeoYeAI/openclaw-master-skills2.2k—~5.5kAutomated safety check: PassApache-2.0

Similar skills

  • Supercov

    supercorp-ai/supercov

    Measures test coverage and code quality in a repository with the supercov CLI, and turns what it finds into small, focused tests or fixes.

    150 GitHub starsUsed in 1 repo~415 tokens
    Testing & QAAuto-check passed
  • Supercov Security

    supercorp-ai/supercov

    Scans a repository's source for security vulnerabilities with the supercov CLI, pointing to the line of each finding and mapping it to CWE classes.

    150 GitHub starsUsed in 1 repo~236 tokens
    Testing & QAAuto-check passed
  • Testing Strategies

    ancoleman/ai-design-components

    Strategic guidance for choosing and implementing testing approaches across the test pyramid.

    526 GitHub stars~3.8k tokensUpdated 10 mo ago
    Testing & QAAuto-check passed
  • Testing Patterns

    softspark/ai-toolkit

    Testing strategy: pyramid, AAA, mocks/fakes/stubs, flaky tests, coverage.

    179 GitHub stars~1.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Phy Test Data Factory

    LeoYeAI/openclaw-master-skills

    Schema-driven test data factory generator. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~5.5k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • Audit AI Code

    jxnl/personal-monorepo-template

    Audit, de-slop, parameterize, modularize, or safely clean up AI-generated or AI-shaped backend/general code.

    562 GitHub stars~1.7k tokensUpdated 3 mo ago
    DevelopmentAuto-check passed

More from petrkindlmann/qa-skills

All 45 skills in this repo
  • Accessibility Testing

    petrkindlmann/qa-skills

    Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).

    165 GitHub stars~4.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Agentic Browser Testing

    petrkindlmann/qa-skills

    Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.

    165 GitHub stars~4.5k tokensUpdated 4 mo ago
    Auto-check passed
  • AI Test Generation

    petrkindlmann/qa-skills

    Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.

    165 GitHub stars~4.8k tokensUpdated 4 mo ago
    Auto-check passed
  • API Testing

    petrkindlmann/qa-skills

    Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.

    165 GitHub stars~2.7k tokensUpdated 4 mo ago
    Auto-check passed
  • CI CD Integration

    petrkindlmann/qa-skills

    Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.

    165 GitHub stars~4.8k tokensUpdated 4 mo ago
    Auto-check passed
  • Compliance Testing

    petrkindlmann/qa-skills

    Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…

    165 GitHub stars~4.6k tokensUpdated 4 mo ago
    Auto-check passed

Categories

Questions about Test Data Management

What does Test Data Management do?

Create and manage test data with factory patterns, fixture strategies, data anonymization, and synthetic data generation. Test Data Management is an agent skill from petrkindlmann/qa-skills. Create and manage test data with factory patterns, fixture strategies, data anonymization, and synthetic data generation.

When should I use Test Data Management?

Test Data Management fits situations like: data anonymization. Not for: migration/integrity testing of the DB itself — use database-testing; environment provisioning and database branching strategy — use test-environments.

How do I install Test Data Management in Claude Code?

Run `npx skills add petrkindlmann/qa-skills --skill test-data-management -a claude-code`. Or copy the skill folder (skills/test-data-management in petrkindlmann/qa-skills) into .claude/skills/test-data-management in your project. Claude Code loads it when a task matches its description.

How do I install Test Data Management in Codex?

Run `npx skills add petrkindlmann/qa-skills --skill test-data-management -a codex`. Or copy the skill folder (skills/test-data-management in petrkindlmann/qa-skills) into .agents/skills/test-data-management in your project. Codex loads it when a task matches its description.

Can I use Test Data Management in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill test-data-management -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-data-management, .gemini/skills/test-data-management, .github/skills/test-data-management and .opencode/skills/test-data-management in your project.

What does Test Data Management need to run?

Going by SKILL.md and its folder, Test Data Management needs the command-line tools its instructions call (psql, pytest, supabase and npx).

Does Test Data Management access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Data Management safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Data Management use?

Test Data Management is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Data Management use?

About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.

What are the alternatives to Test Data Management?

Skills that share tags, products or a category with Test Data Management: Supercov (supercorp-ai/supercov, 150 stars), Supercov Security (supercorp-ai/supercov, 150 stars), Testing Strategies (ancoleman/ai-design-components, 526 stars) and Testing Patterns (softspark/ai-toolkit, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Data Management?

petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 165 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.

Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.