Agent skill

Datajunction Semantic Model

by DataJunction in DataJunction/dj

Activate this skill for DataJunction (DJ) semantic modeling decisions — choosing the right node shape (fact, dimension, transform, metric, cube), turning a draft SQL query into well-designed nodes…

MITAuto-check passedDatabases

Install Datajunction Semantic Model

skills CLI
$ npx skills add DataJunction/dj --skill datajunction-semantic-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install DataJunction/dj datajunction-semantic-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/DataJunction/dj.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/datajunction/skills/datajunction-semantic-model .claude/skills/datajunction-semantic-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
datajunction-semantic-model
GitHub stars
161
Token cost
~4.3k tokens
SKILL.md length
1,717 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Activate this skill for DataJunction (DJ) semantic modeling decisions — choosing the right node shape (fact, dimension, transform, metric, cube), turning a draft SQL query into well-designed nodes…

  • Tasks that involve SQL
  • SKILL.md covers General principles, The dimensional model, Metrics and Query → DJ decomposition…
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Datajunction Semantic Model is an agent skill from DataJunction/dj. Activate this skill for DataJunction (DJ) semantic modeling decisions — choosing the right node shape (fact, dimension, transform, metric, cube), turning a draft SQL query into well-designed nodes, and the cross-cutting conventions (ownership, naming, namespace organization). Format-agnostic modeling guidance. Keywords: - semantic modeling - decompose query, model query, query to nodes - how should I model this metric, what shape should this node be - design a cube, what belongs in this cube - ratio metric…

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering SQL. It works with SQL. The licence is MIT.

When your agent uses it

  • Tasks that involve SQL

Example prompts

  • “/datajunction-semantic-model”

What it can do on your machine

Read from SKILL.md and the folder at commit 904745e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Datajunction Semantic Model loads about 4.3k tokens when it runs. Until then it costs about 186 tokens; SKILL.md has 1,717 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~186
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from DataJunction/dj at commit 904745e, republished under its MIT licence (© DataJunction). 1,717 words, ~4,291 tokens.

Download SKILL.mdSave it as .claude/skills/datajunction-semantic-model/SKILL.md (or your agent's skills folder).
name
datajunction-semantic-model
description
Activate this skill for DataJunction (DJ) semantic modeling decisions — choosing the right node shape (fact, dimension, transform, metric, cube), turning a draft SQL query into well-designed nodes, and the cross-cutting conventions (ownership, naming, namespace organization). Format-agnostic modeling guidance. Keywords: - semantic modeling - decompose query, model query, query to nodes - how should I model this metric, what shape should this node be - design a cube, what belongs in this cube - ratio metric, derived metric, base metric - composable metrics - metric query constraints - node ownership, metric ownership - metric naming, namespace organization - grain, fact vs dimension - dimension link, not JOIN
user-invocable
false

DataJunction Semantic Modeling

Use this skill when designing how something gets expressed as DJ nodes — independent of whether you'll write the result as YAML files in a repo or POST it to the DJ API.

For DJ vocabulary (what node types exist, how dimension links work mechanically), see the datajunction skill. This skill assumes that vocabulary and focuses on the modeling decisions.


General principles

Reuse over reinvent

Before authoring a new node, check whether an authoritative one already exists. Use the datajunction-query skill's MCP tools (search_nodes, get_node_details) to find candidates. Every additional transform / dim / metric fragments the catalog; every reuse strengthens it.

When you can't find a fit:

  • Build on the closest existing parent rather than creating a sibling source/transform that duplicates upstream logic.
  • Refuse to build on non-git-backed namespaces. Nodes without a source repo can change or disappear without review — depending on them silently couples your work to someone's UI session. Stick to git-backed namespaces (see datajunction-api's "Checking if a Namespace Is Repo-Backed" section for how to verify).
Every node has owners

Not just metrics — every DJ node should declare owners. The reasons apply uniformly across all node types:

  • No one to contact when the node breaks or produces unexpected values
  • Undefined responsibility for keeping definitions accurate
  • Difficult to coordinate changes or deprecations
  • Team knowledge scattered across individuals

Best practices:

  • ✅ Prefer team emails over individuals — teams outlast individual contributors
  • ✅ Use Google groups or distribution lists for teams
  • ✅ Include both data team AND business stakeholder teams where applicable
  • ❌ Never leave owners field empty or omit it
✅ GOOD — team ownership:
  data-platform-team@company.com
  finance-analytics@company.com

⚠️ ACCEPTABLE — individual ownership (less sustainable):
  alice@company.com

❌ BAD — no owners (governance nightmare!)

This applies to sources, transforms, dimensions, metrics, and cubes alike.

Naming conventions

Use fully qualified names with namespace:

namespace.node_name

Examples:

  • ✅ finance.total_revenue
  • ✅ common.dimensions.users
  • ✅ clean.user_events
  • ❌ revenue (missing namespace)

Names should be readable and business-meaningful — what a stakeholder would call this thing, not the column transformations behind it. total_revenue over sum_amount_usd. avg_session_duration_secs over avg_session_dur.

Namespace organization

Namespaces are organized by business area:

Common conventions:

  • common.dimensions.* — shared dimensions (users, dates, regions)
  • finance.* — financial metrics & facts
  • growth.* — user engagement & activation
  • product.* — product usage & features
  • source.* — raw source tables
Model named, reusable entities — not anonymous SQL

A general spirit that runs through every layer:

  • Don't bury an aggregate inside a derived metric's query when it should be its own named base metric (see the Metrics section below for the reusability rule).
  • Don't bury a join condition inside a query when it should be a dimension link (see "Joins → dimension links" below).

The semantic layer has two goals working together:

  • Express the business in business concepts. Stakeholders should see signups, revenue, active customers — not a chain of tables and SQL transformations. Every node should have a name and meaning a non-engineer would recognize.
  • Make every meaningful thing findable, queryable, and reusable. Anonymous SQL — inline aggregates, hardcoded joins — works against both goals: it hides business concepts behind technical detail and forecloses reuse.

The dimensional model

DJ uses normalized star schema modeling. The decisions you make here shape everything downstream.

Facts vs dimensions vs transforms
NodeUsed forExamples
SourcePhysical table in the warehousewarehouse.finance.transactions_table
DimensionAn entity with attributes you'll slice byusers, products, dates, regions, geo_country
TransformCleaned/derived fact data — the aggregable rowsclean_transactions, daily_user_activity, enriched_orders
MetricOne aggregation expression over a transformtotal_revenue, num_orders, avg_session_duration
CubeA curated set of metrics + dimensions for downstream consumersrevenue_dashboard, weekly_orders_report

When to author a transform vs use a source directly:

  • Use the source directly when its row shape is the semantic entity you want to aggregate.
  • Author a transform when the source needs cleaning (status IN ('complete', 'completed') → 'completed'), joining, or filtering before it represents the entity meaningfully.
  • The transform is where you attach dimension links — that determines what slices are available to every metric built on it.
Grain is the most important decision in a fact

What does one row in your fact transform represent?

  • "One row per order" → grain is order_id
  • "One row per order per snapshot day" → grain is (order_id, snapshot_date)
  • "One row per customer per month" → grain is (customer_id, month)

Grains don't mix in one fact. If you find yourself with two different grains in one transform, you have two facts trying to live in one node — split them.

The DJ idiom: every JOIN that connects a fact to a dimension belongs in a dimension_links: declaration on the fact's transform — not inside a metric's query.

Why: dimension links make the join optional. Consumers slice by that dim only when they ask for it. A hardcoded JOIN in a metric forces every query that touches the fact to pay the join cost even when nobody's slicing by that dim. It also fragments behavior — metrics on the same fact would each carry their own copy of the same JOIN, with the inevitable drift.

Same applies to dim-to-dim joins (e.g., users → countries → regions) — express them as dimension links on the dimension node, creating a dimensional graph DJ can traverse automatically.

See the datajunction skill's "Dimension Links" section for the mechanics; this is about when to use them — which is essentially always.


Metrics

A metric is one aggregation expression over a single source, transform, or dimension node.

Metric query constraints

Metrics cannot contain WHERE clauses

  • Use CASE WHEN for conditional aggregation instead
sql
-- ❌ NOT ALLOWED — WHERE clause in metric
SELECT SUM(amount_usd)
FROM finance.transactions
WHERE status = 'completed'

-- ✅ CORRECT — use CASE WHEN
SELECT SUM(
  CASE WHEN status = 'completed' THEN amount_usd ELSE 0 END
) FROM finance.transactions

A WHERE permanently constrains scope; CASE WHEN lets the metric coexist with a broader-population sibling on the same fact. Reflect the scope in the metric name (active_signups, not signups).

Metrics select a single expression from a single node

  • Cannot join multiple nodes in a metric query
  • Define joins via dimension links on the source/transform node instead
Metrics can reference other metrics (composition)

Build derived metrics by referencing other metrics in the query. Excellent for ratios, rates, and complex calculations.

sql
-- Create base metrics first
SELECT COUNT(*) FROM finance.transactions              -- metric: finance.transaction_count
SELECT SUM(amount_usd) FROM finance.transactions       -- metric: finance.total_revenue

-- Then compose them into a derived metric
SELECT finance.total_revenue / finance.transaction_count  -- metric: finance.avg_transaction_value
-- DJ handles divide-by-zero automatically; NULLIF() is optional safety
The reusability rule: every aggregate is its own named metric

This is the highest-leverage discipline in DJ metric modeling. Apply it even when "the user only cares about the final ratio."

❌ Less reusable — anonymous SQL blob:

sql
SELECT SUM(amount_usd)
       / NULLIF(SUM(CASE WHEN status = 'completed' THEN 1 ELSE 0 END), 0)
FROM finance.transactions

✅ Reusable — three named, discoverable, composable metrics:

sql
-- metric: total_revenue
SELECT SUM(amount_usd) FROM finance.transactions

-- metric: num_completed_orders
SELECT SUM(CASE WHEN status = 'completed' THEN 1 ELSE 0 END)
FROM finance.transactions

-- metric: avg_order_value (derived)
SELECT total_revenue / NULLIF(num_completed_orders, 0)

Why: the bottom shape gives three named metrics anyone can find, query, and reuse. The top shape gives one — the base aggregates are anonymous SQL blobs nobody else can reference. Different teams asking "what's our total revenue?" can't find it; they'd have to know to read inside avg_order_value's query.

Show full SKILL.md (682 more words)Show less
Metric patterns

Common shapes a metric can take. Each shows the query: expression — the actual file/API format lives in datajunction-repo or datajunction-api.

Base metrics (simple aggregations):

sql
-- COUNT
SELECT COUNT(transaction_id) FROM finance.transactions

-- COUNT DISTINCT
SELECT COUNT(DISTINCT customer_id) FROM finance.transactions

-- APPROX_COUNT_DISTINCT (HyperLogLog — for large datasets)
SELECT APPROX_COUNT_DISTINCT(profile_id) FROM finance.transactions

-- SUM
SELECT SUM(amount_usd) FROM finance.transactions

-- AVG
SELECT AVG(amount_usd) FROM finance.transactions

-- Conditional aggregation
SELECT SUM(
  CASE
    WHEN status = 'completed' AND refund_flag = false
    THEN amount_usd
    ELSE 0
  END
) FROM finance.transactions

Derived metrics (composed from base metrics):

sql
-- ratio
SELECT finance.total_revenue / NULLIF(finance.transaction_count, 0)

-- rate as percentage
SELECT finance.clicks * 100.0 / NULLIF(finance.impressions, 0)

-- revenue per thousand impressions (RPM)
SELECT finance.total_revenue / NULLIF(finance.impressions, 0) * 1000

Statistical metrics:

sql
SELECT VAR_POP(amount_usd) FROM finance.transactions
SELECT STDDEV_POP(amount_usd) FROM finance.transactions
SELECT PERCENTILE_APPROX(amount_usd, 0.5)  FROM finance.transactions  -- median
SELECT PERCENTILE_APPROX(amount_usd, 0.95) FROM finance.transactions  -- p95

Rolling window metrics (set required_dimensions to include the ORDER BY dim):

sql
-- trailing 7-day sum
SELECT SUM(finance.daily_revenue) OVER (
  ORDER BY common.dimensions.date.dateint
  ROWS BETWEEN 6 PRECEDING AND CURRENT ROW
)

Period-over-period metrics:

sql
-- week-over-week % change
SELECT
  (finance.weekly_revenue - LAG(finance.weekly_revenue, 1) OVER (
    ORDER BY common.dimensions.date.week_code
  )) * 100.0 /
  LAG(finance.weekly_revenue, 1) OVER (
    ORDER BY common.dimensions.date.week_code
  )

Same shape for MoM (month_code), QoQ (quarter_code), YoY (year).

Metric metadata quick reference
FieldRequiredValid ValuesNotes
name✅ Yesnamespace.metric_nameFully qualified name
query✅ YesSQL SELECT expressionSingle aggregation from single node
description❌ OptionalStringRecommended for clarity
direction❌ Optionalhigher_is_better / lower_is_better / neutralIndicates performance direction
unit❌ Optionaldollar / unitless / ⚠️ NOT countServer rejects count — use unitless
mode❌ Optionaldraft / publishedDefault: published
fixed_grain❌ OptionalList of dimension namesGrain the aggregate is computed at; omit for query grain, [] for global
required_dimensions❌ OptionalList of dimension namesFor time-based / windowed metrics
owners❌ Optional but strongly recommendedList of email addressesPrefer team emails
Key patterns summary
  • DJ handles divide-by-zero automatically — NULLIF() is optional safety
  • Use CASE WHEN instead of WHERE clauses for filtering
  • Window functions enable rolling windows and period-over-period comparisons
  • required_dimensions should include the dimension used in window ORDER BY clauses
  • Derived metrics can reference other metrics for ratios and calculations
  • Always specify owners — use team emails for sustainability

Query → DJ decomposition workflow

When a user arrives with an existing SQL query and wants to express it as DJ nodes, don't generate node definitions on the first pass. Propose a structured decomposition, get critique, iterate on the shape, then have the user create the nodes (in YAML via datajunction-repo, or via API via datajunction-api).

The workflow

1. Parse the query mechanically. From the user's SQL, extract:

  • Aggregates (SUM, COUNT, COUNT DISTINCT, MIN, MAX, AVG, APPROX_COUNT_DISTINCT, percentile/window aggregates) → each is a candidate base metric.
  • FROM / JOIN tables → candidate parent nodes (transforms or dimensions).
  • GROUP BY columns, JOIN-bound columns, top-level WHERE columns → candidate dimension link targets.
  • Top-level SELECT expressions that combine the aggregates (ratios, sums-of-sums, computed expressions) → candidate derived metrics.

2. Resolve parents against existing nodes first. For each candidate parent table or dimension, use the datajunction-query skill's MCP tools (search_nodes, get_node_details) to check whether DJ already has an authoritative node for it. Prefer building on existing nodes — every additional transform fragments the catalog.

3. Treat every JOIN as a candidate dimension link, not a baked-in join. See "Joins → dimension links" above. The DJ idiom is to express joins as dimension_links: on the fact's transform, not hardcoded in the metric query.

4. Apply the reusability rule. Even if the user's query has aggregates only inside a ratio expression, decompose them: one named base metric per aggregate, then a derived metric for the ratio.

5. Handle WHERE clauses correctly. A WHERE in the user's query is one of two things:

  • A dimensional filter the consumer should apply at query time → keep it out of the metric definition entirely.
  • Part of the metric's identity → bake it into the aggregate as CASE WHEN, not as a WHERE on the metric's query.

6. Name things meaningfully. Propose readable, business-meaningful names that match what a stakeholder would call the metric or entity.

7. Check for duplicates before producing nodes. For each proposed metric / transform / dimension name, check whether something with the same name (or doing the same thing under a different name) already exists in the catalog. Reuse rather than recreate; rename when the names collide but the semantics differ.

8. Propose, don't produce. Present the decomposition as a structured list (parents, base metrics, derived metrics, dim links) and ask the user to critique the shape. Only after they confirm, hand off to datajunction-repo (for YAML) or datajunction-api (for curl) to produce the actual node definitions.

Worked example: ratio decomposition

User's query:

sql
SELECT
  region_id,
  SUM(amount_usd)
  / SUM(CASE WHEN status = 'completed' THEN 1 ELSE 0 END) AS avg_order_value
FROM finance.transactions
GROUP BY region_id

Decomposition:

Parent transform:
  Reuse: finance.transactions  (existing)

Base metrics (one per aggregate):
  1. total_revenue          = SUM(amount_usd)
  2. num_completed_orders   = SUM(CASE WHEN status = 'completed' THEN 1 ELSE 0 END)

Derived metric:
  1. avg_order_value = total_revenue / NULLIF(num_completed_orders, 0)

Dimension links:
  - region_id → common.dimensions.region  (existing shared dim)

Why this shape is worth the extra metric nodes: downstream tools re-aggregate the materialized output. Three independent metrics roll up correctly across any dimensional slice — sum the numerator, sum the denominator, then divide. A single pre-divided ratio doesn't compose: when consumers re-aggregate, the numerator and denominator can drift apart across slices in subtle ways (especially when NULL rows are dropped from one but not the other).

© DataJunction, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/datajunction/skills/datajunction-semantic-model of DataJunction/dj.

Open the folder on GitHubat commit 904745e

Compare with similar skills

Datajunction Semantic Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Datajunction Semantic Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Datajunction Semantic Model this skillDataJunction/dj161—~4.3kAutomated safety check: PassMIT
Evolving The Data ModelTriliumNext/Trilium38k—~2.1kAutomated safety check: PassAGPL-3.0
Orchardcore Data MigrationOrchardCMS/OrchardCore8.2k—~1.7kAutomated safety check: PassBSD-3-Clause
SQL Optimization Patternsynulihao/AgentSkillOS61711 repos~3.3kAutomated safety check: PassNone
SQL PortabilityHL7/sql-on-fhir150—~512Automated safety check: PassCustom licence
StmoSAP/project-foxhound1802 repos~1.8kAutomated safety check: PassGPL-3.0

Similar skills

  • Evolving The Data Model

    TriliumNext/Trilium

    A skill your agent uses when adding a DB migration or a new column/field to a Becca entity in Trilium ("add a migration", "new column on notes/attributes", "ALTER TABLE", "add a field to…

    38k GitHub stars~2.1k tokensUpdated today
    DatabasesAuto-check passed
  • Orchardcore Data Migration

    OrchardCMS/OrchardCore

    Creates and updates OrchardCore data migrations (DataMigration classes with CreateAsync/UpdateFromX).

    8.2k GitHub stars~1.7k tokensUpdated today
    DatabasesAuto-check passed
  • SQL Optimization Patterns

    ynulihao/AgentSkillOS

    Master SQL query optimization, indexing strategies, and EXPLAIN analysis to dramatically improve database performance and eliminate slow queries.

    617 GitHub starsUsed in 11 repos~3.3k tokens
    DatabasesAuto-check passed
  • SQL Portability

    HL7/sql-on-fhir

    Analyse whether a SQL query is portable across database implementations using sqlglot transpilation.

    150 GitHub stars~512 tokensUpdated yesterday
    DatabasesAuto-check passed
  • Stmo

    SAP/project-foxhound

    Official

    Manage Redash queries and dashboards on Mozilla's STMO (sql.telemetry.mozilla.org) using stmo-cli.

    180 GitHub starsUsed in 2 repos~1.8k tokens
    DatabasesAuto-check passed
  • DB Migrations

    kurealnum/dotfiles

    A skill your agent uses when generating or regenerating Drizzle migration files, changing database schema tables or columns, resolving migration sequence conflicts after rebase, reviewing migration…

    290 GitHub stars~820 tokensUpdated 5 mo ago
    DatabasesAuto-check passed

More from DataJunction/dj

  • Datajunction

    DataJunction/dj

    Activate this skill whenever working with DataJunction (DJ) semantic layer.

    161 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Datajunction API

    DataJunction/dj

    Activate this skill when authoring DataJunction (DJ) nodes via the REST API directly (curl, HTTP clients) — typically for exploration, ad-hoc prototyping, or namespaces that aren't repo-backed.

    161 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Datajunction Query

    DataJunction/dj

    Activate this skill for querying DataJunction (DJ) — finding nodes, generating SQL, fetching metric data, exploring lineage, visualizing results — via the DJ UI, MCP tools, or REST/GraphQL APIs.

    161 GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Categories

Questions about Datajunction Semantic Model

What does Datajunction Semantic Model do?

Activate this skill for DataJunction (DJ) semantic modeling decisions — choosing the right node shape (fact, dimension, transform, metric, cube), turning a draft SQL query into well-designed nodes…. Datajunction Semantic Model is an agent skill from DataJunction/dj. Activate this skill for DataJunction (DJ) semantic modeling decisions — choosing the right node shape (fact, dimension, transform, metric, cube), turning a draft SQL query into well-designed nodes, and the cross-cutting conventions (ownership, naming, namespace organization).

When should I use Datajunction Semantic Model?

Datajunction Semantic Model fits situations like: tasks that involve SQL.

How do I install Datajunction Semantic Model in Claude Code?

Run `npx skills add DataJunction/dj --skill datajunction-semantic-model -a claude-code`. Or copy the skill folder (plugins/datajunction/skills/datajunction-semantic-model in DataJunction/dj) into .claude/skills/datajunction-semantic-model in your project. Claude Code loads it when a task matches its description.

How do I install Datajunction Semantic Model in Codex?

Run `npx skills add DataJunction/dj --skill datajunction-semantic-model -a codex`. Or copy the skill folder (plugins/datajunction/skills/datajunction-semantic-model in DataJunction/dj) into .agents/skills/datajunction-semantic-model in your project. Codex loads it when a task matches its description.

Can I use Datajunction Semantic Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add DataJunction/dj --skill datajunction-semantic-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datajunction-semantic-model, .gemini/skills/datajunction-semantic-model, .github/skills/datajunction-semantic-model and .opencode/skills/datajunction-semantic-model in your project.

What does Datajunction Semantic Model need to run?

SKILL.md names no scripts, command-line tools or credentials: Datajunction Semantic Model is instructions for the agent only.

Does Datajunction Semantic Model access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Datajunction Semantic Model safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Datajunction Semantic Model use?

Datajunction Semantic Model is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Datajunction Semantic Model use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Datajunction Semantic Model?

Skills that share tags, products or a category with Datajunction Semantic Model: Evolving The Data Model (TriliumNext/Trilium, 38k stars), Orchardcore Data Migration (OrchardCMS/OrchardCore, 8.2k stars), SQL Optimization Patterns (ynulihao/AgentSkillOS, 617 stars) and SQL Portability (HL7/sql-on-fhir, 150 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Datajunction Semantic Model?

DataJunction (a GitHub organization) maintains it in DataJunction/dj, which has 161 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 7, 2026.

Source: DataJunction/dj on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.