Agent skill

Malloy Define

by malloydata in malloydata/publisher

Propose scope, a source plan and field definitions for a Malloy model.

MITAuto-check passedData & Analytics

Install Malloy Define

skills CLI
$ npx skills add malloydata/publisher --skill malloy-define -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install malloydata/publisher malloy-define --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/malloy-define .claude/skills/malloy-define && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
malloy-define
GitHub stars
116
Token cost
~3.4k tokens
SKILL.md length
1,731 words
Files
1
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Propose scope, a source plan and field definitions for a Malloy model.

  • Works in 4 steps: A table summary with rows, columns, role… → 2-3 analytical focuses, each with the… → Tables to skip, with reasons:… → …
  • Data & Analytics work in your project
  • SKILL.md covers Propose the analytical scope, Propose a source plan, Propose definitions and Data-driven proposals, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Malloy Define is an agent skill from malloydata/publisher. Propose scope, a source plan and field definitions for a Malloy model. Which tables and questions, which sources at what grain, then renames, dimensions and measures, each backed by querying data.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.

When your agent uses it

  • Data & Analytics work in your project

Example prompts

  • “/malloy-define”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. A table summary with rows, columns, role and key relationships. Roles: Fact (events you measure: orders, sessions), Dimension (entities…
  2. 2-3 analytical focuses, each with the tables it covers (for example "Order Analysis: revenue, trends, product performance"), and one…
  3. Tables to skip, with reasons: operational tables, staging copies of a table you model, and pre-aggregated summaries (compute fresh in…
  4. Options the user can answer in one word. If your host has an ask_user tool, put the options in that card and nowhere else, with the…

What it can do on your machine

Read from SKILL.md and the folder at commit 39a546f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Malloy Define loads about 3.4k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 1,731 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from malloydata/publisher at commit 39a546f, republished under its MIT licence (© malloydata). 1,731 words, ~3,413 tokens.

Download SKILL.mdSave it as .claude/skills/malloy-define/SKILL.md (or your agent's skills folder).
name
malloy-define
description
Propose scope, a source plan and field definitions for a Malloy model. Which tables and questions, which sources at what grain, then renames, dimensions and measures, each backed by querying data.
<!--
Copyright (c) Credible Data Inc.
SPDX-License-Identifier: MIT
-->

Propose sources and definitions

This skill covers three consecutive activities when building or extending a Malloy semantic model:

Tool names are written bare here - get_context, execute_query, search_malloy_docs. The exact prefixed name depends on the host surface; match each against the tools you actually have.

  • Propose scope: which tables and which questions the model is for.
  • Propose sources: the architectural blueprint (which sources, what grain).
  • Propose definitions: the specific fields per base source (renames, dimensions, measures).

Both happen in conversation. Propose, let the user confirm or adjust, then carry the confirmed plan forward into the actual .malloy model. There is no separate plan-file store: keep the source plan and field proposals in the conversation, and write the model itself when the user has confirmed. See your modeling workflow for the broader picture.

Read the existing model first so you propose against what is really there. Use get_context with a plain-English description to inspect the current sources and fields and find the most relevant existing sources. Confirm the scope (below) before proposing the source plan.

Propose the analytical scope

When: after you have read the package's sources and fields and looked at the data distributions. Goal: present what you found, recommend one analytical focus, and let the user pick. Query the data with execute_query for row counts and data-quality problems, and record the proposal and the user's decision in your modeling workflow's modeling-notes.md.

Scope is "which questions", not just "which tables". A model aimed at recommendation looks different from one aimed at catalog analysis over the same tables. A single A/B/C question about table inclusion is the source-plan question asked too early.

Present four things:

  1. A table summary with rows, columns, role and key relationships. Roles: Fact (events you measure: orders, sessions), Dimension (entities you slice by: customers, products), Bridge (many-to-many links: order_items, tags), Operational (ETL, staging, audit; not analytical).
  2. 2-3 analytical focuses, each with the tables it covers (for example "Order Analysis: revenue, trends, product performance"), and one recommendation with reasons. Say what kind of model this is and what that rules out. A dataset of events supports trends over time. A dataset of entities with cumulative counters cannot do period-over-period analysis however well it is built ("release dates are 98% null, so this model cannot do calendar analysis"). The user should learn what the model will never answer before choosing.
  3. Tables to skip, with reasons: operational tables, staging copies of a table you model, and pre-aggregated summaries (compute fresh in Malloy; a pre-aggregated source ignores your filters, see skill:malloy-gotchas-modeling).
  4. Options the user can answer in one word. If your host has an ask_user tool, put the options in that card and nowhere else, with the evidence and your recommendation as prose. Otherwise give lettered one-line options (label, tables, key questions), mark the recommended one, and invite a mix ("A plus suppliers").

Be opinionated: recommend one option clearly. With 20+ tables, group them by domain and focus on the most relevant cluster. Show evidence (row counts, relationship density) and flag data-quality problems you saw ("orders has about 3 percent duplicate rows on order_id").

After the user confirms, restate the confirmed scope (connection and schema, tables in scope with row count and role, analytical focus in one line, deferred tables with reasons) and record it in modeling-notes.md. Then continue with the source plan.

Propose a source plan

Goal: Propose the full source architecture for the tables in scope.

Base sources

One base source per table in scope. For each, specify:

SourceTableGrainPrimary KeyRole
orderssales.ordersone row per orderorder_idFact, transactions
customerssales.customersone row per customercustomer_idDimension, who
productssales.productsone row per productproduct_idDimension, what
Computed sources

Computed sources are created from queries, not physical tables. Propose them when:

  1. Grain mismatch: the analytical scope requires a grain that no physical table provides (e.g., customer-level metrics from an order-grain table).
  2. Repeated aggregation patterns: the same group-by plus aggregate pattern would be used in multiple places.
  3. Cross-entity aggregations: inspecting the model and querying the data shows that an aggregate rolled up to a different entity would be reused.

For each computed source, explain:

SourceSource QueryGrainRationale
user_order_factsorders grouped by customer_idone row per customerNeed customer-level order metrics (LTV, order count, recency) for customer health analysis
Dependencies

Show which sources depend on which:

customers (physical) ← user_order_facts (derived, sources from orders)
orders (physical) → user_order_facts (derived)
products (physical): independent
Deferred sources

List sources considered but not included, with reasoning:

  • order_items: bridge table, defer until line-item analysis is needed.
  • monthly_product_facts: derived, defer until product trend analysis is requested.
User interaction

The user will:

  • Confirm the source plan as-is.
  • Add missing sources (physical or derived).
  • Remove unnecessary sources.
  • Validate grain assignments.
  • Defer sources to later iterations.

Once the source plan is confirmed, carry it forward into the definitions step below. Keep the confirmed map in the conversation rather than persisting it to a separate file.

Propose definitions

Goal: Propose specific fields per base source with data evidence, working from the confirmed source plan.

For each base source

Present a table of proposed fields.

Renames (schema cleanup):

Raw ColumnProposed NameReason
Order Dateorder_dateWhitespace in column name
Typeorder_typeReserved word
numberitem_numberReserved word

Dimensions:

FieldLogicData EvidencePriority
order_statusstatus column5 distinct values: pending, processing, shipped, delivered, cancelledmust-have
order_monthsubmitted_at.monthTime trendingmust-have
order_sizetotal buckets (data-driven)Distribution: min $5, p25 $35, median $85, p75 $150, p95 $450, max $2,400. Proposed breaks at p25/p75: <$35, $35-$150, >$150nice-to-have
is_returnedreturned_at is not null8% of orders have non-null returned_atnice-to-have

Data-driven tiers: For bucketed dimensions like order_size, always derive boundaries from the actual data distribution (percentiles, natural breaks, clustering). Query min, max, p25, p50, p75, p95 and propose boundaries based on the distribution. Malloy has no percentile function; use the two-stage nearest-rank query in skill:malloy-discover § Example Queries. Show the evidence so the user can confirm or adjust. Never use arbitrary hardcoded thresholds unless the user explicitly provides them.

Measures:

FieldLogicData EvidencePriority
order_countcount()Basic metricmust-have
revenuesum(total)Total column includes tax. Range: $5 - $2,400must-have
avg_order_valuerevenue / nullif(order_count, 0)Derived from abovemust-have
return_ratereturned_count / nullif(order_count, 0)8% overall return ratenice-to-have
Show full SKILL.md (703 more words)Show less
For each computed source

Show the source query and additional fields.

user_order_facts, derived from orders grouped by customer_id:

Aggregated FieldLogic
total_orderscount()
total_revenuesum(total_price)
first_order_datemin(submitted_at)
last_order_datemax(submitted_at)

Additional dimensions on top:

FieldLogicEvidence
days_since_last_orderdays(last_order_date::timestamp to now)Recency metric (cast: a date column will not measure against a timestamp)
is_repeat_buyertotal_orders > 162% of customers are repeat
buyer_frequencytotal_orders bucketsDistribution: 1 (38%), 2-4 (35%), 5-19 (22%), 20+ (5%)
Business logic questions

Flag decisions the agent can't make from data alone. Be specific and data-grounded:

Q1: Your orders table has both created_at and submitted_at. 87% of rows have them within 1 minute, but 13% differ by 1-3 days. Which should be the canonical order date?

Q2: I'm proposing order_size tiers based on the data distribution: small (<$35, below p25), medium ($35-$150, p25-p75), large (>$150, above p75). Do these data-driven breaks work for you, or do you have specific business thresholds?

Q3: The status column has 5 values. Should "cancelled" orders be excluded from revenue calculations, or included with a separate measure?

Priority ranking

Group proposals into:

  • Must-have: core metrics that every analyst needs (counts, sums, primary dimensions).
  • Nice-to-have: useful but not critical (bucketed dimensions, rates).
  • Value-add: new insights the data supports but may not be asked for yet (computed sources, complex measures).
User interaction

The user will:

  • Confirm business logic decisions.
  • Adjust thresholds and bucket boundaries.
  • Add missing fields.
  • Remove fields they don't need.
  • Change priorities.

Once the definitions are confirmed, write them into the .malloy model (see your modeling workflow). Use #(doc) annotations to document sources and fields, and given: parameters to declare runtime-filterable dimensions where appropriate. #(filter) is deprecated. Never add a #(filter) annotation: every use, including required, implicit, and date/number ranges, has a given: form. See skill:malloy-model § Parameterizing sources with given:. Keep the confirmed definitions in the conversation; there is no separate plan-file store.

Data-driven proposals

Every recommendation must be backed by a query result. Do not propose based on column names or schema structure alone. Always run execute_query to check the actual data before presenting. To learn what sources and fields exist, ground yourself with get_context: it returns the model's sources, views, and fields, so there is no separate schema-search step.

Proposal TypeWhat to query first
Dimension (bucketed)Distribution: min, p25, median, p75, p95, max (two-stage query in skill:malloy-discover § Example Queries; Malloy has no percentile function). Propose boundaries from natural breaks, not arbitrary values.
Dimension (categorical)Distinct values and frequencies. Show the actual categories and their counts.
Measure (sum/avg)Sample values: min, max, avg. Verify the column contains what you think (e.g., is total gross or net?).
Measure (rate/ratio)Query both numerator and denominator. Verify they make sense together.
Denormalized field vs joinCompare the pre-computed column against the joined aggregate. Report match rate. Recommend whichever is more reliable.
Computed sourceRun the proposed group-by plus aggregation. Verify the grain collapses as expected and the result is useful.
Date field selectionQuery all candidate date columns. Show % of rows where they differ and by how much.
Column renameVerify the column has data worth exposing (not 100% NULL).

Example, denormalized vs joined:

"Your customers table has an order_count column. I compared it against count() from the orders table:

  • 94% of customers match exactly
  • 6% have stale counts (the denormalized value is lower than the actual count)
  • The max discrepancy is 12 orders

I'd recommend using the joined count from orders rather than the denormalized order_count. Want to keep the denormalized column as internal, or drop it?"

Tips

  • Show data, not assumptions: every proposed dimension or measure should have evidence (distinct values, distributions, ranges).
  • Use execute_query to verify any data questions before presenting to the user.
  • Don't over-propose: 5-8 dimensions and 4-6 measures per base source is usually enough to start.
  • Rank everything: users appreciate knowing what's essential vs. optional.
  • Business logic questions must be specific. "What date should I use?" is bad. "Your table has created_at and submitted_at that differ by 1-3 days in 13% of rows, which is canonical?" is good.

Output

A confirmed source architecture and a confirmed set of field definitions (renames, dimensions, measures, business decisions), held in the conversation and ready to write into the .malloy model via your modeling workflow.

© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/malloy-define of malloydata/publisher.

Open the folder on GitHubat commit 39a546f

Compare with similar skills

Malloy Define next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Malloy Define compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Malloy Define this skillmalloydata/publisher116—~3.4kAutomated safety check: PassMIT
Exploratory Data Analysisspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: PassMIT
MatplotlibzLanqing/codex-claude-academic-skills4.6k18 repos~2.9kAutomated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.6k17 repos~3.9kAutomated safety check: PassBSD-3-Clause
Chart Visualizationbytedance/deer-flow83k2 repos~840Automated safety check: PassMIT
TimesFM Forecastinggoogle-research/timesfm34k—~4.7kAutomated safety check: PassApache-2.0

Similar skills

  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Matplotlib

    zLanqing/codex-claude-academic-skills

    Low-level plotting library for full customization. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 18 repos~2.9k tokens
    Data & AnalyticsAuto-check passed
  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Chart Visualization

    bytedance/deer-flow

    Picks a suitable chart type from 26 options for your data, maps the data to that chart's parameters and generates a chart image through a JavaScript script.

    83k GitHub starsUsed in 2 repos~840 tokens
    Data & AnalyticsAuto-check passed
  • TimesFM Forecasting

    google-research/timesfm

    Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.

    34k GitHub stars~4.7k tokensUpdated 8 days ago
    Data & AnalyticsAuto-check passed
  • Sandbox Bench

    vercel/next.js

    Official

    Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…

    143k GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from malloydata/publisher

All 29 skills in this repo
  • Eval Answer

    malloydata/publisher

    Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.

    116 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Fix Scan Finding

    malloydata/publisher

    Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…

    116 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Eval Import

    malloydata/publisher

    Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.

    116 GitHub stars~5.9k tokensUpdated today
    Auto-check passed
  • Eval Loop

    malloydata/publisher

    Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.

    116 GitHub stars~7.8k tokensUpdated today
    Auto-check passed
  • Eval Improve

    malloydata/publisher

    Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.

    116 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Eval Judge

    malloydata/publisher

    Decide whether ONE answer matches its golden, and say whether you believe the golden.

    116 GitHub stars~3.4k tokensUpdated today
    Auto-check passed

Questions about Malloy Define

What does Malloy Define do?

Propose scope, a source plan and field definitions for a Malloy model. Malloy Define is an agent skill from malloydata/publisher. Propose scope, a source plan and field definitions for a Malloy model.

When should I use Malloy Define?

Malloy Define fits situations like: data & Analytics work in your project.

How do I install Malloy Define in Claude Code?

Run `npx skills add malloydata/publisher --skill malloy-define -a claude-code`. Or copy the skill folder (skills/malloy-define in malloydata/publisher) into .claude/skills/malloy-define in your project. Claude Code loads it when a task matches its description.

How do I install Malloy Define in Codex?

Run `npx skills add malloydata/publisher --skill malloy-define -a codex`. Or copy the skill folder (skills/malloy-define in malloydata/publisher) into .agents/skills/malloy-define in your project. Codex loads it when a task matches its description.

Can I use Malloy Define in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill malloy-define -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/malloy-define, .gemini/skills/malloy-define, .github/skills/malloy-define and .opencode/skills/malloy-define in your project.

What does Malloy Define need to run?

SKILL.md names no scripts, command-line tools or credentials: Malloy Define is instructions for the agent only.

Does Malloy Define access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Malloy Define safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Malloy Define use?

Malloy Define is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Malloy Define use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Malloy Define?

Skills that share tags, products or a category with Malloy Define: Exploratory Data Analysis (spacering-net/codeg, 3.8k stars), Matplotlib (zLanqing/codex-claude-academic-skills, 4.6k stars), Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars) and Chart Visualization (bytedance/deer-flow, 83k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Malloy Define?

malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.

Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.