Exploratory Data Analysis
spacering-net/codeg
Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.
Propose scope, a source plan and field definitions for a Malloy model.
$ npx skills add malloydata/publisher --skill malloy-define -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install malloydata/publisher malloy-define --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/malloy-define .claude/skills/malloy-define && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "malloy-define" agent skill from https://github.com/malloydata/publisher/tree/main/skills/malloy-define into .claude/skills/malloy-define/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "malloy-define", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/malloydata/publisher/tree/main/skills/malloy-defineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add malloydata/publisher --skill malloy-define -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install malloydata/publisher malloy-define --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/malloy-define .agents/skills/malloy-define && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "malloy-define" agent skill from https://github.com/malloydata/publisher/tree/main/skills/malloy-define into .agents/skills/malloy-define/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "malloy-define", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill malloy-define -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install malloydata/publisher malloy-define --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/malloy-define .cursor/skills/malloy-define && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "malloy-define" agent skill from https://github.com/malloydata/publisher/tree/main/skills/malloy-define into .cursor/skills/malloy-define/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "malloy-define", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/malloydata/publisher.git --path skills/malloy-define--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add malloydata/publisher --skill malloy-define -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install malloydata/publisher malloy-define --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/malloy-define .gemini/skills/malloy-define && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "malloy-define" agent skill from https://github.com/malloydata/publisher/tree/main/skills/malloy-define into .gemini/skills/malloy-define/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "malloy-define", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install malloydata/publisher malloy-defineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add malloydata/publisher --skill malloy-define -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/malloy-define .github/skills/malloy-define && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "malloy-define" agent skill from https://github.com/malloydata/publisher/tree/main/skills/malloy-define into .github/skills/malloy-define/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "malloy-define", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add malloydata/publisher --skill malloy-define -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install malloydata/publisher malloy-define --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/malloydata/publisher.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/malloy-define .opencode/skills/malloy-define && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "malloy-define" agent skill from https://github.com/malloydata/publisher/tree/main/skills/malloy-define into .opencode/skills/malloy-define/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "malloy-define", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
malloy-definePropose scope, a source plan and field definitions for a Malloy model.
Malloy Define is an agent skill from malloydata/publisher. Propose scope, a source plan and field definitions for a Malloy model. Which tables and questions, which sources at what grain, then renames, dimensions and measures, each backed by querying data.
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics. The repository describes itself as: Publisher is the open-source analytics engine for Malloy. It lets you define data models once — and use them everywhere. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 39a546f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Malloy Define loads about 3.4k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 1,731 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from malloydata/publisher at commit 39a546f, republished under its MIT licence (© malloydata). 1,731 words, ~3,413 tokens.
.claude/skills/malloy-define/SKILL.md (or your agent's skills folder).<!--
Copyright (c) Credible Data Inc.
SPDX-License-Identifier: MIT
-->
This skill covers three consecutive activities when building or extending a Malloy semantic model:
Tool names are written bare here -
get_context,execute_query,search_malloy_docs. The exact prefixed name depends on the host surface; match each against the tools you actually have.
Both happen in conversation. Propose, let the user confirm or adjust, then carry the confirmed plan forward into the actual .malloy model. There is no separate plan-file store: keep the source plan and field proposals in the conversation, and write the model itself when the user has confirmed. See your modeling workflow for the broader picture.
Read the existing model first so you propose against what is really there. Use get_context with a plain-English description to inspect the current sources and fields and find the most relevant existing sources. Confirm the scope (below) before proposing the source plan.
When: after you have read the package's sources and fields and looked at the data distributions. Goal: present what you found, recommend one analytical focus, and let the user pick. Query the data with execute_query for row counts and data-quality problems, and record the proposal and the user's decision in your modeling workflow's modeling-notes.md.
Scope is "which questions", not just "which tables". A model aimed at recommendation looks different from one aimed at catalog analysis over the same tables. A single A/B/C question about table inclusion is the source-plan question asked too early.
Present four things:
skill:malloy-gotchas-modeling).ask_user tool, put the options in that card and nowhere else, with the evidence and your recommendation as prose. Otherwise give lettered one-line options (label, tables, key questions), mark the recommended one, and invite a mix ("A plus suppliers").Be opinionated: recommend one option clearly. With 20+ tables, group them by domain and focus on the most relevant cluster. Show evidence (row counts, relationship density) and flag data-quality problems you saw ("orders has about 3 percent duplicate rows on order_id").
After the user confirms, restate the confirmed scope (connection and schema, tables in scope with row count and role, analytical focus in one line, deferred tables with reasons) and record it in modeling-notes.md. Then continue with the source plan.
Goal: Propose the full source architecture for the tables in scope.
One base source per table in scope. For each, specify:
| Source | Table | Grain | Primary Key | Role |
|---|---|---|---|---|
| orders | sales.orders | one row per order | order_id | Fact, transactions |
| customers | sales.customers | one row per customer | customer_id | Dimension, who |
| products | sales.products | one row per product | product_id | Dimension, what |
Computed sources are created from queries, not physical tables. Propose them when:
For each computed source, explain:
| Source | Source Query | Grain | Rationale |
|---|---|---|---|
| user_order_facts | orders grouped by customer_id | one row per customer | Need customer-level order metrics (LTV, order count, recency) for customer health analysis |
Show which sources depend on which:
customers (physical) ← user_order_facts (derived, sources from orders)
orders (physical) → user_order_facts (derived)
products (physical): independentList sources considered but not included, with reasoning:
The user will:
Once the source plan is confirmed, carry it forward into the definitions step below. Keep the confirmed map in the conversation rather than persisting it to a separate file.
Goal: Propose specific fields per base source with data evidence, working from the confirmed source plan.
Present a table of proposed fields.
Renames (schema cleanup):
| Raw Column | Proposed Name | Reason |
|---|---|---|
Order Date | order_date | Whitespace in column name |
Type | order_type | Reserved word |
number | item_number | Reserved word |
Dimensions:
| Field | Logic | Data Evidence | Priority |
|---|---|---|---|
| order_status | status column | 5 distinct values: pending, processing, shipped, delivered, cancelled | must-have |
| order_month | submitted_at.month | Time trending | must-have |
| order_size | total buckets (data-driven) | Distribution: min $5, p25 $35, median $85, p75 $150, p95 $450, max $2,400. Proposed breaks at p25/p75: <$35, $35-$150, >$150 | nice-to-have |
| is_returned | returned_at is not null | 8% of orders have non-null returned_at | nice-to-have |
Data-driven tiers: For bucketed dimensions like order_size, always derive boundaries from the actual data distribution (percentiles, natural breaks, clustering). Query min, max, p25, p50, p75, p95 and propose boundaries based on the distribution. Malloy has no percentile function; use the two-stage nearest-rank query in skill:malloy-discover § Example Queries. Show the evidence so the user can confirm or adjust. Never use arbitrary hardcoded thresholds unless the user explicitly provides them.
Measures:
| Field | Logic | Data Evidence | Priority |
|---|---|---|---|
| order_count | count() | Basic metric | must-have |
| revenue | sum(total) | Total column includes tax. Range: $5 - $2,400 | must-have |
| avg_order_value | revenue / nullif(order_count, 0) | Derived from above | must-have |
| return_rate | returned_count / nullif(order_count, 0) | 8% overall return rate | nice-to-have |
Show the source query and additional fields.
user_order_facts, derived from orders grouped by customer_id:
| Aggregated Field | Logic |
|---|---|
| total_orders | count() |
| total_revenue | sum(total_price) |
| first_order_date | min(submitted_at) |
| last_order_date | max(submitted_at) |
Additional dimensions on top:
| Field | Logic | Evidence |
|---|---|---|
| days_since_last_order | days(last_order_date::timestamp to now) | Recency metric (cast: a date column will not measure against a timestamp) |
| is_repeat_buyer | total_orders > 1 | 62% of customers are repeat |
| buyer_frequency | total_orders buckets | Distribution: 1 (38%), 2-4 (35%), 5-19 (22%), 20+ (5%) |
Flag decisions the agent can't make from data alone. Be specific and data-grounded:
Q1: Your
orderstable has bothcreated_atandsubmitted_at. 87% of rows have them within 1 minute, but 13% differ by 1-3 days. Which should be the canonical order date?Q2: I'm proposing
order_sizetiers based on the data distribution: small (<$35, below p25), medium ($35-$150, p25-p75), large (>$150, above p75). Do these data-driven breaks work for you, or do you have specific business thresholds?Q3: The
statuscolumn has 5 values. Should "cancelled" orders be excluded from revenue calculations, or included with a separate measure?
Group proposals into:
The user will:
Once the definitions are confirmed, write them into the .malloy model (see your modeling workflow). Use #(doc) annotations to document sources and fields, and given: parameters to declare runtime-filterable dimensions where appropriate. #(filter) is deprecated. Never add a #(filter) annotation: every use, including required, implicit, and date/number ranges, has a given: form. See skill:malloy-model § Parameterizing sources with given:. Keep the confirmed definitions in the conversation; there is no separate plan-file store.
Every recommendation must be backed by a query result. Do not propose based on column names or schema structure alone. Always run execute_query to check the actual data before presenting. To learn what sources and fields exist, ground yourself with get_context: it returns the model's sources, views, and fields, so there is no separate schema-search step.
| Proposal Type | What to query first |
|---|---|
| Dimension (bucketed) | Distribution: min, p25, median, p75, p95, max (two-stage query in skill:malloy-discover § Example Queries; Malloy has no percentile function). Propose boundaries from natural breaks, not arbitrary values. |
| Dimension (categorical) | Distinct values and frequencies. Show the actual categories and their counts. |
| Measure (sum/avg) | Sample values: min, max, avg. Verify the column contains what you think (e.g., is total gross or net?). |
| Measure (rate/ratio) | Query both numerator and denominator. Verify they make sense together. |
| Denormalized field vs join | Compare the pre-computed column against the joined aggregate. Report match rate. Recommend whichever is more reliable. |
| Computed source | Run the proposed group-by plus aggregation. Verify the grain collapses as expected and the result is useful. |
| Date field selection | Query all candidate date columns. Show % of rows where they differ and by how much. |
| Column rename | Verify the column has data worth exposing (not 100% NULL). |
Example, denormalized vs joined:
"Your
customerstable has anorder_countcolumn. I compared it againstcount()from theorderstable:
- 94% of customers match exactly
- 6% have stale counts (the denormalized value is lower than the actual count)
- The max discrepancy is 12 orders
I'd recommend using the joined count from
ordersrather than the denormalizedorder_count. Want to keep the denormalized column as internal, or drop it?"
execute_query to verify any data questions before presenting to the user.created_at and submitted_at that differ by 1-3 days in 13% of rows, which is canonical?" is good.A confirmed source architecture and a confirmed set of field definitions (renames, dimensions, measures, business decisions), held in the conversation and ready to write into the .malloy model via your modeling workflow.
© malloydata, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/malloy-define of malloydata/publisher.
Open the folder on GitHubat commit 39a546f
Malloy Define next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Malloy Define this skillmalloydata/publisher | 116 | — | ~3.4k | Automated safety check: Pass | MIT | |
| Exploratory Data Analysisspacering-net/codeg | 3.8k | 15 repos | ~3.6k | Automated safety check: Pass | MIT | |
| MatplotlibzLanqing/codex-claude-academic-skills | 4.6k | 18 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Scikit LearnzLanqing/codex-claude-academic-skills | 4.6k | 17 repos | ~3.9k | Automated safety check: Pass | BSD-3-Clause | |
| Chart Visualizationbytedance/deer-flow | 83k | 2 repos | ~840 | Automated safety check: Pass | MIT | |
| TimesFM Forecastinggoogle-research/timesfm | 34k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 |
spacering-net/codeg
Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.
zLanqing/codex-claude-academic-skills
Low-level plotting library for full customization. An agent skill from zLanqing/codex-claude-academic-skills.
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
bytedance/deer-flow
Picks a suitable chart type from 26 options for your data, maps the data to that chart's parameters and generates a chart image through a JavaScript script.
google-research/timesfm
Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.
vercel/next.js
Benchmark React or Next.js changes on Vercel Sandbox VMs with paired A/B statistics: react PR/commit vs base, or Next.js PR/commit vs base, measured end-to-end through the bench/render-pipeline app…
malloydata/publisher
Score one analytical answer against a verified golden, and score which of the entities the golden depends on retrieval delivered to the answerer.
malloydata/publisher
Fix a CRITICAL Trivy finding that is failing CI in this repo (a vulnerability, misconfiguration, or secret from security-scan.yml or image-scan.yml), or add, review, or retire an entry in…
malloydata/publisher
Turn a list of questions into an eval set, whatever shape it arrived in: a JSONL a customer sent, a CSV, a spreadsheet export, a markdown doc, an email thread, or a pull from production logs.
malloydata/publisher
Conduct a local Publisher evaluation loop in five steps: scrape/run, eval, diagnose, improve, checkpoint.
malloydata/publisher
Make the smallest safe Malloy model edit that closes a diagnosed model-owned gap, with a probe receipt for every factual claim.
malloydata/publisher
Decide whether ONE answer matches its golden, and say whether you believe the golden.
Categories
Propose scope, a source plan and field definitions for a Malloy model. Malloy Define is an agent skill from malloydata/publisher. Propose scope, a source plan and field definitions for a Malloy model.
Malloy Define fits situations like: data & Analytics work in your project.
Run `npx skills add malloydata/publisher --skill malloy-define -a claude-code`. Or copy the skill folder (skills/malloy-define in malloydata/publisher) into .claude/skills/malloy-define in your project. Claude Code loads it when a task matches its description.
Run `npx skills add malloydata/publisher --skill malloy-define -a codex`. Or copy the skill folder (skills/malloy-define in malloydata/publisher) into .agents/skills/malloy-define in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add malloydata/publisher --skill malloy-define -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/malloy-define, .gemini/skills/malloy-define, .github/skills/malloy-define and .opencode/skills/malloy-define in your project.
SKILL.md names no scripts, command-line tools or credentials: Malloy Define is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Malloy Define is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Malloy Define: Exploratory Data Analysis (spacering-net/codeg, 3.8k stars), Matplotlib (zLanqing/codex-claude-academic-skills, 4.6k stars), Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars) and Chart Visualization (bytedance/deer-flow, 83k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
malloydata (a GitHub organization) maintains it in malloydata/publisher, which has 116 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on October 8, 2026.
Source: malloydata/publisher on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.