Official agent skill

Core

by microsoft in microsoft/data-formulator

The analyst's built-in capabilities: data-inspection tools and the always-available actions (visualize and askuser).

OfficialMITAuto-check passedData & Analytics

Install Core

skills CLI
$ npx skills add microsoft/data-formulator --skill core -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/data-formulator core --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/data-formulator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/py-src/data_formulator/analyst/skills/core .claude/skills/core && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
core
GitHub stars
18k
Token cost
~5.4k tokens
SKILL.md length
3,003 words
Files
4
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

The analyst's built-in capabilities: data-inspection tools and the always-available actions (visualize and askuser).

  • Data & Analytics work in your project
  • SKILL.md covers Tools (for data gathering), Actions, Choosing what to do and Chart Creation Guide
  • Runs Python scripts from its folder

What it does

Core is an agent skill from microsoft/data-formulator, published by the product's own GitHub organization. The analyst's built-in capabilities: data-inspection tools and the always-available actions (visualize and askuser).

Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `__init__.py`, `skill.py` and `tools.json`).

It sits in Data & Analytics. It works with Python. The repository describes itself as: 🪄 Data Formulator is an interactive AI-powered data analysis system makes it easy to connect, explore and visualize data. The licence is MIT.

When your agent uses it

  • Data & Analytics work in your project

Example prompts

  • “/core”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 5477f0e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Core loads about 5.4k tokens when it runs. Until then it costs about 31 tokens; SKILL.md has 3,003 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~31
When it runs · the whole SKILL.md, loaded when a task matches
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/data-formulator at commit 5477f0e, republished under its MIT licence (© microsoft). 3,003 words, ~5,443 tokens.

Download SKILL.mdSave it as .claude/skills/core/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
core
description
The analyst's built-in capabilities: data-inspection tools and the always-available actions (visualize and ask_user).
when_to_use
Always loaded by default — this is the agent's baseline.
always_on
true
tools
execute_python_script, inspect_source_data
actions
visualize, ask_user

Core capabilities

This describes the built-in inspection tools you use to gather data and the always-available actions you take on it. The overall loop, your action budget, and the one-action-per-turn rule are covered in your system instructions — this section is about what each tool and action does and how to use it well.

Tools (for data gathering)

  • execute_python_script(code) — run a general-purpose Python script to inspect data, compute stats, transform tables, or verify assumptions. Its stdout is returned to you (use print()); the script is for your analysis and its output is never shown to the user. pandas, numpy, duckdb, sklearn, scipy are available. Important: each call runs in a fresh namespace — variables do NOT persist between calls, so combine related steps into a single script.
  • inspect_source_data(table_names) — get schema, stats, and sample rows for source tables (cheaper than execute_python_script for basic inspection).
  • load_skill(name) — load a skill's instructions into context so you can use the action it unlocks (see the Skills section of your system instructions).

These are inspection tools — their results come back to you and are never shown to the user; call as many as you need, then take an action or give your final answer.

You analyse data that is already in the workspace. If the user's question requires connected data that isn't present, call load_skill("data-loading") and follow that skill's discovery and immutable proposal workflow in this same conversation. Do not hand off to the standalone Data Loading agent.

The initial context already includes sample rows and statistics for each table. If the data is straightforward, go straight to the action without calling tools. Tool results are returned to you before you act.

Actions

Call an action as a tool call when you want to act on the data. Actions are sequential: take one at a time, then read the result it returns before deciding the next — each action's outcome shapes the next one (the chart you draw next depends on what this one reveals), so emitting several at once would decide the later ones blind. After each result you choose what to do — take another action, or stop. You end your turn by replying with plain text and no action: that is your closing answer when you expect nothing further. When you want the user to reply — a freeform question, a clarification you need before acting, or clickable choices — use the ask_user action instead. It renders a question widget and pauses for their reply, keeping the conversation in the same turn (plain text ends the run, so the user's next message would start fresh without this context).

Match the response to what the user asked for. Two different cases:

  • A direct answer — they asked a question, so answer it. Length follows the question: one line when that settles it, more when it genuinely takes more.
  • A finding after you acted — they asked for the work, not a write-up, so this is unsolicited. The artifact already shows what it shows; add only what they'd miss by looking at it, and default to short.

Open with the point rather than announcing one is coming, and don't close by restating what you just said. Never narrate what you're about to do or recap a chart's axes; let the artifact speak for itself. When an action pauses for the user, give enough context to explain what you found and what their choices mean.

visualize — chart a transform

Run code that produces a DataFrame and render it as a chart. You then observe the result and decide your next move.

  • display_instruction — ≤12 words; the question/hypothesis the chart investigates (don't recap x/y/color — those are visible). Wrap a column in **…** if it anchors the question.
  • title — a concise, neutral analytical heading naming the subject, measure, and analytical lens, such as “Year-over-year price change peaks.” Prefer a stable description of the view over a takeaway claim or narrated trend. Do not mention the chart type, imply causality, or editorialize. This field is required; put interpretation in the closing response instead.
  • subtitle — concise supporting context not already clear from the title or axes. Use one phrase of at most 16 words to provide contextual details. Do not restate the measure or analytical lens named in the title.
  • code — Python producing a DataFrame assigned to output_variable.
  • output_variable — snake_case name the code assigns.
  • chart — {chart_type, encodings:{x,y,…}, config:{}} (chart_type from the chart type reference).
  • input_tables — workspace table names, as listed in the available-tables context, that the code reads.
  • field_metadata — field → semantic annotation. Include units, index baselines, intrinsic domains, and ordinal order when supported by the data; never invent a unit. Distinguish percentages from percentage points and identifiers from quantities.
  • field_display_names — field → concise human-readable label for axes, legends, and table headers. Expand technical names, preserve established domain abbreviations, include units when useful, and use the user's language.

Silently classify the analytical intent before choosing a chart: comparison, trend, distribution, relationship, composition, deviation, ranking, uncertainty, or spatial pattern. Choose encodings and chart type from that intent and the data shape. Set ordering deliberately: chronological for time, semantic order for ordinal fields, and measure order for rankings. Avoid line charts or legends with excessive series, labels that collide, and color that does not encode additional information; aggregate, bin, facet, or limit categories when needed without hiding material data.

ask_user — ask the user and pause for their reply (pauses the run)

Ask the user something and pause for their input. Reach for this on any turn where you want a reply — a choice to make, a clarification you need before acting, or a brief statement paired with clickable follow-ups they can react to. Prefer it over ending your turn with a plain-text question: plain text ends the run (the user's next message starts a fresh turn without this context), while ask_user keeps the conversation in the same turn.

  • questions — 1–3 items, each something the user acts on: a choice (single_choice with options) or an open question they type an answer to (free_text). Put your reasoning, rationale, and context in your reply text — not here. Never add a questions item that only states a rationale or explanation with nothing for the user to answer or click.
  • each question: text (wrap a column in **…**), responseType (single_choice when you offer options, else free_text — the user types their own open-ended answer, not a slot for your exposition), required (true when the run depends on the answer, false for an optional follow-up), and options (plain-text choices, at most 3 — just the most likely answers; the user can always type a freeform reply, so don't enumerate every case).

This is terminal: the run pauses after it and resumes when the user replies.

Choosing what to do

Match the response depth to the user's request. Create charts that materially contribute to the answer, and stop when the answer is sufficient.

  • For conceptual or informational questions, answer directly when a chart would not improve the answer.
  • For specific analytical questions, create the view or views needed to answer them clearly.
  • For diagnostic or exploratory questions, follow relevant findings across multiple views when doing so adds meaningful insight.
  • If essential intent is unclear, use ask_user rather than guessing.
  • Missing data (needs tables not in the workspace): load_skill("data-loading"), discover the source, and propose immutable loading options inline.
  • Report / write-up request (e.g. "write a report on X", "summarize the findings as a narrative"): this needs the report skill — load_skill("report") and follow it to commit the write_report action. Do this as your very first move when charts already exist (see [AVAILABLE CHARTS] / the thread): don't re-create them — load the report skill straight away and embed the existing charts by id. Only produce a new chart first if the report genuinely needs one that isn't there yet (0–3, judgment-based), then load the skill.

Follow explicit requests about scope, depth, and format. Never repeat a visualization already in the trajectory or in another thread.

Chart Creation Guide

The following reference material applies when you call the visualize tool.

A. Code Execution Rules

About the execution environment:

  • You can use BOTH DuckDB SQL and pandas operations in the same script
  • The script will run in the workspace data directory (all data files are in the current directory)
  • Each table in [CONTEXT] has a file path (e.g., student_exam.parquet, sales.csv). Use EXACTLY that path to load data:
    • .parquet: pd.read_parquet('file.parquet') or DuckDB read_parquet('file.parquet')
    • .csv: pd.read_csv('file.csv') or DuckDB read_csv_auto('file.csv')
    • .json: pd.read_json('file.json')
    • .xlsx/.xls: pd.read_excel('file.xlsx')
    • .txt: pd.read_csv('file.txt', sep='\t')
  • IMPORTANT: Use the exact filename from the context — do NOT change the file extension or assume all files are parquet.
  • Allowed libraries: pandas, numpy, duckdb, math, datetime, json, statistics, collections, re, sklearn, scipy, random, itertools, functools, operator, time
  • Not allowed: matplotlib, plotly, seaborn, requests, subprocess, os, sys, io, or any other library not listed above.
  • File system access (open, write) and network access are also forbidden.

When to use DuckDB vs pandas:

  • Prefer plain pandas for most tasks — it's simpler and more readable.
  • Only use DuckDB when the dataset is very large and you need efficient SQL aggregations, filtering, joins, or window functions.
  • You can combine both: DuckDB for initial loading/filtering on large files, then pandas for complex operations.

Code structure: standalone script (no function wrapper), imports at top. CRITICAL: The final result DataFrame MUST be assigned to the exact variable name you specified in "output_variable" — the system uses this name to extract the result. For example, if your output_variable is sales_by_region, the script must contain sales_by_region = ....

DuckDB notes:

  • Escape single quotes with '' (not ')
  • No Unicode escapes (\u0400); use character ranges directly: [а-яА-Я]
  • Cast date columns explicitly: CAST(col AS DATE), CAST(col AS TIMESTAMP)
  • For complex datetime operations, load data first then use pandas datetime functions
  • Critical identifier quoting rule:
    • If a table/column name contains non-ASCII characters (e.g., Chinese, Japanese, Korean, Cyrillic, etc.), spaces, or punctuation, you MUST wrap it in double quotes, e.g. SELECT "金额" FROM "客户表".
    • Never output placeholder identifiers like your_table_name, your_column, your_condition.

Datetime handling:

  • date columns contain date-only values (YYYY-MM-DD). datetime columns contain date+time (ISO 8601).
  • time columns contain time-only values (HH:mm:ss). duration columns are time intervals.
  • Year → number. Year-month / year-month-day → string ("2020-01" / "2020-01-01").
  • Hour alone → number. Hour:min or h:m:s → string. Never return raw datetime objects.
Show full SKILL.md (1,321 more words)Show less
B. Chart Type Reference

The chart_type value in the visualize action MUST be one of the names listed below (exact spelling, including capitalization). When a row lists multiple names, pick whichever fits the "when to use" hint best.

Choosing a chart — prefer simple, escalate when it fits. Reach for the Everyday set first: it answers most questions and is the safest, most legible choice. But when the data or question genuinely fits a Specialized type (a distribution's shape, a cumulative curve, a rank race, a before→after, a geographic pattern…), prefer it — a well-matched specialized chart is more insightful than forcing a generic one. Don't pick a specialized type for novelty; use it because its "when to use" condition is met.

Everyday — reach for these first

chart_typeencodingsconfigwhen to use
Scatter Plotx, y, color, size, facetopacity (0.1–1.0)Relationships between two quantitative fields
Regressionx, y, color, size, facetregressionMethod ("linear","log","exp","pow","quad","poly"), polyOrder (2–10)Trend line over scatter; one line per color group
Bar Chart / Stacked Bar Chart / Lollipop Chart / Waterfall Chartx, y, color, facet—Bar: categorical comparison (auto-stacks when color is set). Stacked Bar: explicit stacked totals, color = the stack. Lollipop: cleaner for ranked lists / sparse categories. Waterfall: cumulative gain/loss, each bar starts where the previous ended
Grouped Bar Chartx, y, group, facet—Side-by-side bars across a second categorical dimension
Line Chartx, y, color, strokeDash, facetinterpolate ("linear","monotone","step")Trends over an ordered (usually temporal) x-axis
Area Chartx, y, color, facet—Magnitude over ordered x; auto-stacks when color is set
Histogram / Density Plotx, color, facet—Distribution of one quantitative field. Histogram: discrete bins, auto-binned. Density Plot: smooth KDE curve
Boxplotx, y, color, facet—Distribution summary (median/quartiles/outliers) by category
Pie Chartsize, color, facetinnerRadius (0–100; 0=pie, >0=donut)Part-of-whole with ≤7 categories. Wedge value goes on size, not theta
Heatmapx, y, color, facetcolorScheme — sequential ("viridis","blues","reds","oranges","greens") or diverging ("blueorange","redblue")Matrix / 2D density; color encodes the quantitative cell value

Specialized — use when the data/question fits the "when to use"

chart_typeencodingsconfigwhen to use
Connected Scatter Plotx, y, order, color, facet—Two quantitative fields traced in sequence — needs an order field (e.g. time) so points are joined in order, not by x
Ranged Dot Plotx, y, color, facet—Min–max range or two-point comparison per category
Violin Plotx, y, color, facet—Distribution SHAPE (KDE silhouette) by category; better than a boxplot when data is multimodal. x = category, y = value
Strip Plotx, y, color, size, facet—Every individual point by category (jittered); good for small/medium n where raw values matter, not just a summary
ECDF Plotx, color, facet—Cumulative distribution of one quantitative field. Pass the RAW field on x (do NOT pre-compute the CDF); color for per-group curves
Bump Chartx, y, color, facet—How RANKINGS change over ordered x; y = rank, color = entity (long-form: one row per entity × x)
Slope Chartx, y, color, facet—Change between exactly TWO points (before → after) per entity; x = the two labels, y = value, color = entity
Streamgraphx, y, color, facet—Several series' magnitude over ordered x, stacked around a center baseline (color = series) — theme/volume shifts over time
Range Area Chartx, y, y2, color, facet—A shaded band between a lower (y) and upper (y2) bound over ordered x — e.g. min–max or a confidence interval
Rose Chartx, y, color, facet—Cyclical/categorical magnitude as angular wedges (polar bars); x = category/angle, y = value
Pyramid Chartx, y, color, facet—Back-to-back bars split by a binary group (e.g. population by age × sex); y = category, x = value, color = the two-sided group
Radar Chartx, y, color, facet—Multi-metric profile/comparison; x = metric name, y = value, color = entity (long-form data)
Bar Tablex, y, color, facet—Ranked horizontal table with inline bars; one row per category. y = category, x = value
KPI Cardmetric, value, goal—"Big number" dashboard tile(s); one row per tile. value must be pre-aggregated; goal is optional
Candlestick Chartx, open, high, low, close, facet—OHLC financial data
Maplongitude, latitude, color, sizeprojection ("mercator","equalEarth","naturalEarth1","orthographic","albersUsa"), projectionCenter ([lon,lat])Geographic POINTS/bubbles by lon/lat (use projection "albersUsa" for a US-only map)
Choroplethid, color, facetregion ("world","usa",…)Filled REGIONS shaded by value; id = the region key (country/state name or code), color = the quantitative value

Critical chart rules:

  • Scatter Plot: use config opacity (0.1–1.0) for dense data instead of encoding opacity.
  • Regression: trend line is automatic — do NOT compute regression coefficients/predictions in Python. Use color to get separate trend lines per group.
  • Bar Chart: x=categorical, y=quantitative (vertical bars). Swap x↔y for horizontal bars. Same-x rows are auto-stacked when color is set.
  • Grouped Bar Chart: use the group channel (not color) for side-by-side bars.
  • Histogram: do NOT pre-bin in Python — pass the raw quantitative field on x and the chart bins automatically. Pre-aggregating gives wrong bin widths.
  • Line Chart: use strokeDash to differentiate line styles (e.g. actual vs forecast).
  • Pie Chart: use the size channel (not theta) for wedge values. Avoid when >7–8 categories.
  • Radar Chart: data must be long-form — one row per (entity, metric, value). If your data is wide-form (one column per metric), melt it first in the Python step.
  • Heatmap: pick colorScheme by the meaning of the values. Use a sequential scheme (viridis/blues/reds/oranges/greens) for single-direction magnitudes (counts, rates, prices, scores — higher is simply more). Use a diverging scheme (blueorange/redblue) ONLY when the values have a meaningful center to read away from (e.g. profit/loss around 0, change vs. a baseline, temperature around freezing).
  • Bar Table: y is the category column to rank; x is the quantitative value driving bar length. Don't sort in Python — the template sorts.
  • KPI Card: channels are metric, value, goal (not x/y). One DataFrame row = one tile. The value column must already contain the final number to display (aggregate upstream in the Python step).
  • Candlestick Chart: requires open, high, low, close columns.
  • Connected Scatter Plot: provide an order field (usually time) so points are joined in sequence, not by x-order.
  • ECDF Plot: pass the RAW quantitative field on x — the chart computes the cumulative curve; do NOT pre-compute it in Python.
  • Range Area Chart: y is the lower bound and y2 the upper bound of the band.
  • Bump / Slope Chart: long-form data — one row per (entity, x); color is the entity. Slope's x has exactly two categories (before/after).
  • Violin Plot: like Boxplot but shows the full distribution shape; x = category, y = value.
  • Map / Choropleth: Map plots points via longitude / latitude (set projection "albersUsa" for the US); Choropleth fills regions — put the region key on id and the value on color, not x / y.
  • facet: available for nearly all chart types; use a low-cardinality categorical field.
  • All fields in encodings must also appear in output_fields. Typically use 2–3 channels (x, y, color/size).
C. Semantic Type Reference

Choose the most specific type that fits. Only annotate fields used in chart encodings.

CategoryTypes
TemporalDateTime, Date, Time, Timestamp, Year, Quarter, Month, Week, Day, Hour, YearMonth, YearQuarter, YearWeek, Decade, Duration
Monetary measuresAmount, Price
Physical measuresQuantity, Temperature
ProportionPercentage
Signed/divergingProfit, PercentageChange, Sentiment, Correlation
Generic measuresCount, Number
Discrete numericRank, Score
IdentifierID
GeographicLatitude, Longitude, Country, State, City, Region, Address, ZipCode
Entity namesCategory, Name
Coded categoricalStatus, Boolean, Direction
Binned rangesRange
FallbackUnknown

Key guidelines:

  • Use Amount for summed monetary totals, Price for per-unit prices, Profit for values that can be negative.
  • Use Temperature (not Quantity) for temperature — it has special diverging behavior.
  • Use Year (not Number) for columns like "year" with values 2020, 2021.
D. Statistical Analysis Guide
  • Regression: use chart_type "Regression" — the trend line is automatic, do NOT compute regression values in Python code. Configure method via {"regressionMethod": "linear"} (options: "linear", "log", "exp", "pow", "quad", "poly"; for poly add {"polyOrder": 3}).
  • Forecasting: compute predicted future values in Python. Use Line Chart with strokeDash to distinguish actual vs forecast, and color for series grouping.
  • Clustering: compute cluster assignments in Python. Output [x, y, cluster_id]. Use Scatter Plot with color → cluster_id.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in py-src/data_formulator/analyst/skills/core of microsoft/data-formulator.

  • SKILL.md
  • __init__.py
  • skill.py
  • tools.json

Open the folder on GitHubat commit 5477f0e

Compare with similar skills

Core next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Core compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Core this skillmicrosoft/data-formulator18k—~5.4kAutomated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.6k17 repos~3.9kAutomated safety check: PassBSD-3-Clause
TimesFM Forecastinggoogle-research/timesfm34k—~4.7kAutomated safety check: PassApache-2.0
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
StatsmodelszLanqing/codex-claude-academic-skills4.6k16 repos~4.9kAutomated safety check: PassBSD-3-Clause
Scientific Figure MakingChenLiu-1996/figures4papers8.1k—~557Automated safety check: PassCustom licence

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • TimesFM Forecasting

    google-research/timesfm

    Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.

    34k GitHub stars~4.7k tokensUpdated 8 days ago
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 16 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • Scientific Figure Making

    ChenLiu-1996/figures4papers

    Covers publication-ready matplotlib figures for academic papers, slides, and reports—bars, trends, scatter, heatmaps, and multi-panel layouts—with this…

    8.1k GitHub stars~557 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Diagnose and fix ModuleNotFoundError in Nuitka standalone binaries caused by missing implicit imports.

    15k GitHub stars~519 tokensUpdated today
    Data & AnalyticsAuto-check passed

More from microsoft/data-formulator

  • Error Handling

    microsoft/data-formulator

    Official

    统一错误处理系统。在添加 API 端点、修改错误处理、添加前端 API 调用、编写错误相关测试时使用. An agent skill from microsoft/data-formulator.

    18k GitHub stars~3.8k tokensUpdated yesterday
    Auto-check passed
  • Language Injection

    microsoft/data-formulator

    Official

    LLM Agent 多语言注入规范。在修改 Agent 提示词、添加新的 Agent 端点、处理用户可见的后端消息(messagecode)时使用。

    18k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Path Safety

    microsoft/data-formulator

    Official

    服务端路径安全与文件访问编码规范。在编写文件下载路由、Agent 工具(文件读取/目录列出)、数据连接器/Loader、Workspace 路径操作、沙箱配置时使用。

    18k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Data Loading

    microsoft/data-formulator

    Official

    Discover connected data sources, add new data connectors through a user-confirmed form, inspect table metadata, and run bounded read-only probes when the current workspace data is insufficient.

    18k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Report

    microsoft/data-formulator

    Official

    Turn an exploration (threads, findings, charts) into a single Markdown report — note, blog post, executive summary, KPI dashboard, slide brief, or multi-section analytical report, with embedded…

    18k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Core

What does Core do?

The analyst's built-in capabilities: data-inspection tools and the always-available actions (visualize and askuser). Core is an agent skill from microsoft/data-formulator, published by the product's own GitHub organization. The analyst's built-in capabilities: data-inspection tools and the always-available actions (visualize and askuser).

When should I use Core?

Core fits situations like: data & Analytics work in your project.

How do I install Core in Claude Code?

Run `npx skills add microsoft/data-formulator --skill core -a claude-code`. Or copy the skill folder (py-src/data_formulator/analyst/skills/core in microsoft/data-formulator) into .claude/skills/core in your project. Claude Code loads it when a task matches its description.

How do I install Core in Codex?

Run `npx skills add microsoft/data-formulator --skill core -a codex`. Or copy the skill folder (py-src/data_formulator/analyst/skills/core in microsoft/data-formulator) into .agents/skills/core in your project. Codex loads it when a task matches its description.

Can I use Core in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/data-formulator --skill core -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/core, .gemini/skills/core, .github/skills/core and .opencode/skills/core in your project.

What does Core need to run?

Going by SKILL.md and its folder, Core needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Core access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Core safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Core use?

Core is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Core use?

About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Core?

Skills that share tags, products or a category with Core: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), TimesFM Forecasting (google-research/timesfm, 34k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars) and Statsmodels (zLanqing/codex-claude-academic-skills, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Core?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/data-formulator, which has 17,538 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 5, 2026.

Source: microsoft/data-formulator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.