Agent skill

Data Model Extractor

by EmeaAppGbb in EmeaAppGbb/spec2cloud

Extract database schemas, data models, and entity relationships from code.

MITAuto-check passedDatabases

Install Data Model Extractor

skills CLI
$ npx skills add EmeaAppGbb/spec2cloud --skill data-model-extractor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EmeaAppGbb/spec2cloud data-model-extractor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EmeaAppGbb/spec2cloud.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/data-model-extractor .claude/skills/data-model-extractor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-model-extractor
GitHub stars
100
Token cost
~2.5k tokens
SKILL.md length
876 words
Files
1
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Extract database schemas, data models, and entity relationships from code.

  • Works in 7 steps: Identify Data Layer Technology → Extract Entity Definitions → Extract Relationships → …
  • Tasks that involve Database schema design
  • SKILL.md covers Role, Inputs, Process and Output Format, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Model Extractor is an agent skill from EmeaAppGbb/spec2cloud. Extract database schemas, data models, and entity relationships from code. Produce Mermaid ERD diagrams. Pure extraction from ORM models, migration files, and schema definitions — no assessment of normalization, no suggestions for schema changes, no judgment on data modeling decisions.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering Database schema design, Diagrams and ORMs and data access. It works with Mermaid, Prisma, Mongoose and SQL. The licence is MIT.

When your agent uses it

  • Tasks that involve Database schema design
  • Tasks that involve Diagrams
  • Tasks that involve ORMs and data access

Example prompts

  • “/data-model-extractor”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Identify Data Layer Technology
  2. Extract Entity Definitions
  3. Extract Relationships
  4. Extract Indexes and Constraints
  5. Parse Migration History
  6. Identify Seed Data and Fixtures
  7. Generate ERD Diagrams

What it can do on your machine

Read from SKILL.md and the folder at commit 8e76618. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are mermaid and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Model Extractor loads about 2.5k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 876 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EmeaAppGbb/spec2cloud at commit 8e76618, republished under its MIT licence (© EmeaAppGbb). 876 words, ~2,525 tokens.

Download SKILL.mdSave it as .claude/skills/data-model-extractor/SKILL.md (or your agent's skills folder).
name
data-model-extractor
description
Extract database schemas, data models, and entity relationships from code. Produce Mermaid ERD diagrams. Pure extraction from ORM models, migration files, and schema definitions — no assessment of normalization, no suggestions for schema changes, no judgment on data modeling decisions.

Data Model Extractor

Role

You are the Data Model Extractor — a factual agent that reads ORM models, migration files, raw SQL schemas, and configuration to produce a complete picture of the project's data layer. You extract entities, fields, types, relationships, indexes, and constraints.

You are a surveyor, not a consultant. You measure and record what exists. You NEVER assess normalization level, suggest schema changes, flag "missing" indexes, or comment on data modeling quality. If the schema has a users table with 47 columns and no indexes, you document exactly that.

Inputs

  • The project source tree
  • Output from codebase-scanner (specs/docs/technology/stack.md) if available — to know which ORM/database technologies to look for
  • Output from architecture-mapper if available — to understand which services own which data

Process

Step 1 — Identify Data Layer Technology

Determine how the project defines its data model:

TechnologyFiles to Scan
Prismaschema.prisma, prisma/schema.prisma
TypeORM*.entity.ts, files with @Entity() decorator
Sequelize*.model.ts, files with sequelize.define(), migration files
Drizzleschema.ts files with pgTable/mysqlTable definitions
MikroORM*.entity.ts with @Entity() from @mikro-orm/core
Mongoose*.model.ts, files with mongoose.Schema / mongoose.model
SQLAlchemymodels.py, files with class X(Base) or class X(db.Model)
Django ORMmodels.py with class X(models.Model)
Entity Framework*.cs with DbContext, entity classes, *.Designer.cs
Hibernate/JPA*.java with @Entity, @Table annotations
GORM*.go with gorm.Model embedding
ActiveRecord*.rb with class X < ApplicationRecord, db/schema.rb
Dieselschema.rs, *.rs with diesel macros
Raw SQL*.sql files in migrations/, db/, sql/ directories
KnexMigration files in migrations/, knex config

A project may use multiple approaches. Document all of them.

Step 2 — Extract Entity Definitions

For each entity/model/table found, extract:

  1. Entity name: The class name, table name, or collection name
  2. Table/collection name: The actual database table or collection name (may differ from the entity name due to naming conventions)
  3. Source file: Where the entity is defined
  4. Fields/columns: For each field:
    • Name
    • Type (as declared — e.g., VARCHAR(255), String, text, INTEGER)
    • Nullable (yes/no)
    • Default value (if declared)
    • Primary key (yes/no)
    • Auto-generated (auto-increment, UUID generation, etc.)
    • Unique constraint (yes/no)
  5. Timestamps: Are createdAt/updatedAt/deletedAt fields present? Is soft delete used?
Step 3 — Extract Relationships

Identify all entity relationships:

Relationship TypeWhat to Look For
One-to-One@OneToOne, hasOne, ForeignKey with unique constraint, belongs_to + unique
One-to-Many@OneToMany/@ManyToOne, hasMany/belongsTo, foreign key columns
Many-to-Many@ManyToMany, join tables, has_and_belongs_to_many, intermediate models
Self-referentialEntity referencing itself (e.g., parentId on same table)
Polymorphictype + id columns referencing different tables

For each relationship, record:

  • Source entity and field
  • Target entity and field
  • Relationship type
  • Cascade behavior (if declared): CASCADE, SET NULL, RESTRICT, NO ACTION
  • Whether it's bidirectional or unidirectional
Step 4 — Extract Indexes and Constraints

Document all declared indexes and constraints:

  1. Primary keys: Single-column or composite
  2. Unique constraints: Single-column or composite
  3. Foreign keys: With referential actions
  4. Check constraints: Validation rules at the database level
  5. Indexes: Name, columns, type (btree, hash, gin, gist), unique/non-unique, partial index conditions
  6. Full-text indexes: If declared
Step 5 — Parse Migration History

If migration files exist, extract the schema evolution:

  1. List all migrations in chronological order.
  2. For each migration, record:
    • Migration name/identifier
    • Timestamp or sequence number
    • Operations performed (create table, add column, add index, etc.)
    • Whether it's been applied (check migration status table if accessible)
  3. Note if the current model definitions match the latest migration state. If they diverge, document both — models represent "intended" state, migrations represent "applied" state.
Show full SKILL.md (326 more words)Show less
Step 6 — Identify Seed Data and Fixtures

Look for data seeding mechanisms:

  • Seed files (seeds/, fixtures/, data/)
  • Factory definitions (for testing)
  • Initial data migrations
  • Enum/lookup table population

Document what seed data exists and which entities it populates.

Step 7 — Generate ERD Diagrams

Produce Mermaid Entity-Relationship Diagrams:

mermaid
erDiagram
    USERS {
        uuid id PK
        varchar email UK
        varchar password_hash
        varchar name
        timestamp created_at
        timestamp updated_at
    }
    POSTS {
        uuid id PK
        uuid author_id FK
        varchar title
        text content
        varchar status
        timestamp published_at
        timestamp created_at
    }
    COMMENTS {
        uuid id PK
        uuid post_id FK
        uuid user_id FK
        text content
        timestamp created_at
    }
    USERS ||--o{ POSTS : "authors"
    USERS ||--o{ COMMENTS : "writes"
    POSTS ||--o{ COMMENTS : "has"

If the schema is large (>15 entities), produce multiple diagrams grouped by domain/feature area.

Output Format

Produce specs/docs/architecture/data-models.md:

markdown
# Data Models — [Project Name]

_Extracted on [date]. Documents the data layer as defined in code._

## Data Layer Technology

| Aspect | Value |
|--------|-------|
| ORM/ODM | Prisma 5.8.0 |
| Database | PostgreSQL (inferred from provider config) |
| Migration tool | Prisma Migrate |
| Migration count | 23 |
| Schema source | prisma/schema.prisma |

## Entity-Relationship Diagram

[Mermaid ERD — full or grouped by domain]

## Entity Catalog

### Entity: User

| Field | Type | Nullable | Default | Key | Notes |
|-------|------|----------|---------|-----|-------|
| id | UUID | No | uuid_generate_v4() | PK | Auto-generated |
| email | VARCHAR(255) | No | — | UK | — |
| passwordHash | VARCHAR(255) | No | — | — | — |
| name | VARCHAR(100) | Yes | NULL | — | — |
| role | ENUM(admin,user) | No | 'user' | — | — |
| createdAt | TIMESTAMP | No | now() | — | Auto-set |
| updatedAt | TIMESTAMP | No | now() | — | Auto-updated |

**Relationships:**
- User → Posts (one-to-many via Post.authorId)
- User → Comments (one-to-many via Comment.userId)

**Indexes:**
- `users_email_key` — UNIQUE on (email)
- `users_role_idx` — BTREE on (role)

[Repeat for each entity]

## Relationship Summary

| Source | Target | Type | FK Column | Cascade |
|--------|--------|------|-----------|---------|
| Post | User | Many-to-One | authorId | SET NULL |
| Comment | Post | Many-to-One | postId | CASCADE |
| Comment | User | Many-to-One | userId | SET NULL |

## Migration History

| # | Name | Date | Operations |
|---|------|------|-----------|
| 1 | 20240101_init | 2024-01-01 | Create users, posts tables |
| 2 | 20240115_add_comments | 2024-01-15 | Create comments table |
| ... | ... | ... | ... |

## Model vs Migration Discrepancies

[List any differences between ORM models and migration state, or "None found"]

## Seed Data

[Document any seed/fixture files found]

Rules

  1. Document actual schema. Extract what the code defines, not what you think the schema should be.
  2. Models AND migrations. If both exist, document both. If they tell different stories, document the discrepancy factually.
  3. No normalization opinions. Do not assess whether the schema is properly normalized. Do not say "denormalized", "over-normalized", or suggest restructuring.
  4. No missing index commentary. Do not flag queries that "would benefit from an index". Document indexes that exist, period.
  5. No schema suggestions. Banned words: "should", "consider", "missing", "recommend", "improve", "optimize". You document, not advise.
  6. Type fidelity. Report types as declared in the ORM/schema, not translated to a normalized form. If Prisma says String, write String. If SQL says VARCHAR(255), write VARCHAR(255).
  7. Complete entity catalog. Missing an entity that exists in the code is a failure. Scan every model file, every migration, every schema definition.
  8. Relationship accuracy. Every foreign key and relationship decorator must appear in the output. Verify by checking both sides of bidirectional relationships.

Mandatory Completion Checklist

The orchestrator MUST verify ALL of the following before marking data-model-extractor as complete:

  • specs/docs/architecture/data-models.md exists with a Mermaid ERD diagram
  • Every ORM model, migration file, and schema definition in the codebase is covered
  • All entities have their fields, types, and constraints documented
  • All relationships (foreign keys, one-to-many, many-to-many) are documented with both sides verified
  • Index definitions are cataloged where present in schema files

BLOCKING: If any item is unchecked, the skill has NOT completed successfully. The orchestrator must loop back and complete the missing items before advancing to the next extraction step.

© EmeaAppGbb, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/data-model-extractor of EmeaAppGbb/spec2cloud.

Open the folder on GitHubat commit 8e76618

Compare with similar skills

Data Model Extractor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Model Extractor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Model Extractor this skillEmeaAppGbb/spec2cloud100—~2.5kAutomated safety check: PassMIT
Drizzle Erdhiroppy/mf-dashboard418—~627Automated safety check: PassMIT
Database FundamentalsDanielPodolsky/ownyourcode2901 repos~1.6kAutomated safety check: PassMIT
DB SculptorEliasOulkadi/shokunin114—~3.1kAutomated safety check: NotesMIT
Prisma Migrationopenathleteorg/openathlete100—~575Automated safety check: PassAGPL-3.0
Drizzle Orm Expertdavila7/claude-code-templates32k3 repos~2.6kAutomated safety check: PassMIT

Similar skills

  • Drizzle Erd

    hiroppy/mf-dashboard

    A skill your agent uses when needing to visualize database schema, generate ERD diagrams from Drizzle ORM schemas, or understand table relationships

    418 GitHub stars~627 tokensUpdated today
    DatabasesAuto-check passed
  • Database Fundamentals

    DanielPodolsky/ownyourcode

    Reviews schema design, SQL queries, ORM patterns. An agent skill from DanielPodolsky/ownyourcode.

    290 GitHub starsUsed in 1 repo~1.6k tokens
    DatabasesAuto-check passed
  • DB Sculptor

    EliasOulkadi/shokunin

    Design database schemas with Prisma/Drizzle, PostgreSQL index strategy (B-tree, GIN, GiST, BRIN, Hash), query optimization (EXPLAIN ANALYZE), migration safety (expand/contract, zero-downtime), and…

    114 GitHub stars~3.1k tokensUpdated 2 days ago
    DatabasesAuto-check: notes
  • Prisma Migration

    openathleteorg/openathlete

    Change the OpenAthlete database schema safely with Prisma, covering the migration SQL, production safety, account deletion coverage and tests.

    100 GitHub stars~575 tokensUpdated today
    DatabasesAuto-check passed
  • Drizzle Orm Expert

    davila7/claude-code-templates

    Expert in Drizzle ORM for TypeScript — schema design, relational queries, migrations, and serverless database integration.

    32k GitHub starsUsed in 3 repos~2.6k tokens
    DatabasesAuto-check passed
  • Data Model Creation

    TencentCloudBase/CloudBase-AI-Toolkit

    [Deprecated] Optional advanced tool for complex data modeling.

    1.1k GitHub starsUsed in 1 repo~1.8k tokens
    DatabasesAuto-check passed

More from EmeaAppGbb/spec2cloud

All 39 skills in this repo
  • Azure Deployment

    EmeaAppGbb/spec2cloud

    Provision Azure infrastructure, deploy to Azure Container Apps, and verify via smoke tests.

    100 GitHub stars~1.8k tokensUpdated 5 mo ago
    Auto-check passed
  • Contract Generation

    EmeaAppGbb/spec2cloud

    Generate API contracts, shared TypeScript types, and infrastructure resource definitions from Gherkin scenarios and test files.

    100 GitHub stars~1.6k tokensUpdated 5 mo ago
    Auto-check passed
  • Ddd Modeling

    EmeaAppGbb/spec2cloud

    Create Domain-Driven Design proposals from product specs or brownfield extraction outputs.

    100 GitHub stars~2.4k tokensUpdated 5 mo ago
    Auto-check passed
  • Implementation

    EmeaAppGbb/spec2cloud

    Write application code to make failing tests pass using contract-driven, slice-based architecture.

    100 GitHub stars~2.8k tokensUpdated 5 mo ago
    Auto-check passed
  • Spec Refinement

    EmeaAppGbb/spec2cloud

    Review PRDs and FRDs through product and technical lenses. An agent skill from EmeaAppGbb/spec2cloud.

    100 GitHub stars~2.2k tokensUpdated 5 mo ago
    Auto-check passed
  • State Management

    EmeaAppGbb/spec2cloud

    Read, write, and maintain .spec2cloud/state.json across phases and increments.

    100 GitHub stars~1.5k tokensUpdated 5 mo ago
    Auto-check passed

Categories

Questions about Data Model Extractor

What does Data Model Extractor do?

Extract database schemas, data models, and entity relationships from code. Data Model Extractor is an agent skill from EmeaAppGbb/spec2cloud. Extract database schemas, data models, and entity relationships from code.

When should I use Data Model Extractor?

Data Model Extractor fits situations like: tasks that involve Database schema design; tasks that involve Diagrams; tasks that involve ORMs and data access.

How do I install Data Model Extractor in Claude Code?

Run `npx skills add EmeaAppGbb/spec2cloud --skill data-model-extractor -a claude-code`. Or copy the skill folder (.github/skills/data-model-extractor in EmeaAppGbb/spec2cloud) into .claude/skills/data-model-extractor in your project. Claude Code loads it when a task matches its description.

How do I install Data Model Extractor in Codex?

Run `npx skills add EmeaAppGbb/spec2cloud --skill data-model-extractor -a codex`. Or copy the skill folder (.github/skills/data-model-extractor in EmeaAppGbb/spec2cloud) into .agents/skills/data-model-extractor in your project. Codex loads it when a task matches its description.

Can I use Data Model Extractor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EmeaAppGbb/spec2cloud --skill data-model-extractor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-model-extractor, .gemini/skills/data-model-extractor, .github/skills/data-model-extractor and .opencode/skills/data-model-extractor in your project.

What does Data Model Extractor need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Model Extractor is instructions for the agent only.

Does Data Model Extractor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Model Extractor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Model Extractor use?

Data Model Extractor is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Model Extractor use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Model Extractor?

Skills that share tags, products or a category with Data Model Extractor: Drizzle Erd (hiroppy/mf-dashboard, 418 stars), Database Fundamentals (DanielPodolsky/ownyourcode, 290 stars), DB Sculptor (EliasOulkadi/shokunin, 114 stars) and Prisma Migration (openathleteorg/openathlete, 100 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Model Extractor?

EmeaAppGbb (a GitHub organization) maintains it in EmeaAppGbb/spec2cloud, which has 100 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on April 16, 2026.

Source: EmeaAppGbb/spec2cloud on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.