Official agent skill

Data Loading

by microsoft in microsoft/data-formulator

Discover connected data sources, add new data connectors through a user-confirmed form, inspect table metadata, and run bounded read-only probes when the current workspace data is insufficient.

OfficialMITAuto-check passed

Install Data Loading

skills CLI
$ npx skills add microsoft/data-formulator --skill data-loading -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/data-formulator data-loading --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/data-formulator.git skills-src && mkdir -p .claude/skills && cp -r skills-src/py-src/data_formulator/analyst/skills/data-loading .claude/skills/data-loading && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-loading
GitHub stars
18k
Token cost
~1.4k tokens
SKILL.md length
770 words
Files
4
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Discover connected data sources, add new data connectors through a user-confirmed form, inspect table metadata, and run bounded read-only probes when the current workspace data is insufficient.

  • Works in 4 steps: Call list_connectors first because… → Once the source type is known, call… → **When the requested source type is… → …
  • SKILL.md covers Adding a connector, Discovery sequence, Proposing loading options and Grounding rules
  • Runs Python scripts from its folder

What it does

Data Loading is an agent skill from microsoft/data-formulator, published by the product's own GitHub organization. Discover connected data sources, add new data connectors through a user-confirmed form, inspect table metadata, and run bounded read-only probes when the current workspace data is insufficient.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `__init__.py`, `skill.py` and `tools.json`).

The repository describes itself as: 🪄 Data Formulator is an interactive AI-powered data analysis system makes it easy to connect, explore and visualize data. The licence is MIT.

Example prompts

  • “/data-loading”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Call list_connectors first because available built-ins and plugins vary by
  2. Once the source type is known, call describe_connector when field or auth
  3. **When the requested source type is known and available, you MUST call
  4. The form is only a proposal. The user reviews it and clicks Connect; the

What it can do on your machine

Read from SKILL.md and the folder at commit 5477f0e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Loading loads about 1.4k tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 770 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/data-formulator at commit 5477f0e, republished under its MIT licence (© microsoft). 770 words, ~1,412 tokens.

Download SKILL.mdSave it as .claude/skills/data-loading/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
data-loading
description
Discover connected data sources, add new data connectors through a user-confirmed form, inspect table metadata, and run bounded read-only probes when the current workspace data is insufficient.
when_to_use
The user's question needs data that is not already available as a workspace input, the user asks what connected data is available, or the user wants to…
always_on
false
tools
list_data, find_data, describe_data, probe_data, list_connectors, describe_connector
actions
propose_data_operation, propose_connection

Skill: Data discovery

The workspace tables listed in your context are the data already loaded into the system, and the only data that can be read directly. Everything these tools return is not loaded yet — it lives in a connected source and only becomes usable after the user selects a loading option and the server materializes it.

Use these tools to determine whether connected sources contain data needed for the user's goal. They are read-only: discovering, describing, or probing a source does not add anything to the workspace analysis inputs.

Adding a connector

When the user wants to connect a new source, do not merely ask them to navigate to settings and do not attempt to connect on their behalf.

  1. Call list_connectors first because available built-ins and plugins vary by deployment. For a broad request such as "help me connect", summarize the concrete available types and ask which one they use.
  2. Once the source type is known, call describe_connector when field or auth details are useful.
  3. When the requested source type is known and available, you MUST call propose_connection in this same turn. Do not stop with text such as "I'll open the form", "you'll need to provide", or a list of required fields. Only the action opens the form. Include one or two helpful sentences alongside the action call explaining what the user should review or supply; this text appears above the chat while the form opens on the canvas. Pass prefilled values the user already supplied, including values parsed from a connection string or config snippet. Never invent missing values.
  4. The form is only a proposal. The user reviews it and clicks Connect; the action must never connect automatically.

Prefilled values may include credentials the user deliberately supplied. Do not repeat those values in prose or subsequent tool output. They are transient form seeds and are removed from persisted UI state.

Discovery sequence

  1. Use find_data when the user names a business concept or table. Use list_data when you need to browse available sources or hierarchy.
  2. Use describe_data before relying on columns, types, row counts, or filter values. Pass the exact source_id and table_key returned by discovery.
  3. Use probe_data only when metadata is insufficient to choose a useful bounded result. Probes are limited, read-only, and may be approximate.
  4. First reconcile discoveries with every table in [PRIMARY TABLE(S)], [OTHER AVAILABLE TABLES], or [AVAILABLE TABLES]. If the needed data is already loaded, use or explain that workspace table instead of proposing it.
  5. When there are genuinely missing useful alternatives, call propose_data_operation with one to three complete immutable plans. This pauses for the user's choice; it does not load data yet.
Show full SKILL.md (322 more words)Show less

Proposing loading options

Write your answer as message text alongside the call — that prose is what the user reads, so it carries the whole answer. Do not put it in an action field, and do not leave the call bare. Say what you went looking for, what you actually found, and what each option would give them — enough that they can choose without opening a single preview. Two to four sentences; more when the options differ in ways that matter (grain, coverage, freshness, joins needed), fewer when the choice is obvious. Name real tables and columns you saw during discovery, and say plainly when an option is a compromise or when you'd pick one yourself. Write it as you'd say it to a colleague, not as a schema summary.

  • Each option is a complete alternative: a concise action label (2–6 words) and one or more tables. The labels are buttons, not sentences — the reasoning belongs in your message text. The application displays table previews separately, so don't list columns as a substitute for explaining.
  • Use only source IDs, table keys, columns, and values grounded by discovery.
  • For a whole table, omit query. Use the optional raw-row query only when the request needs filters, projection, ordering, or an intentional limit. It uses the same filters / columns / order_by / limit vocabulary as probe_data, without aggregation.
  • Do not invent operation IDs, plan IDs, or hashes. The server creates them.
  • Never propose an exact connector query already represented by a workspace table. The server also enforces this using persisted load provenance.

Grounding rules

  • Never invent source IDs, table keys, columns, or category values.
  • Prefer cached catalog discovery before a live probe.
  • Treat probe rows as evidence for planning, not as analysis input data.
  • Keep queries structured and bounded. Do not generate source-specific SQL.
  • If a source is unavailable or permissions changed, report the tool result and ask the user for the needed connection or choose another source.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in py-src/data_formulator/analyst/skills/data-loading of microsoft/data-formulator.

  • SKILL.md
  • __init__.py
  • skill.py
  • tools.json

Open the folder on GitHubat commit 5477f0e

Compare with similar skills

Data Loading next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Loading compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Loading this skillmicrosoft/data-formulator18k—~1.4kAutomated safety check: PassMIT
ConnectComposioHQ/awesome-claude-skills77k3 repos~987Automated safety check: PassNone
MongoDB Source Connector E2E Harnessairbytehq/airbyte22k—~1.9kAutomated safety check: PassCustom licence
Connecting To Data Sourceaws/agent-toolkit-for-aws2.8k1 repos~2.2kAutomated safety check: PassApache-2.0
Source Mapsthedaviddias/Front-End-Checklist74k—~445Automated safety check: PassMIT
Loading Indicatorsthedaviddias/Front-End-Checklist74k—~434Automated safety check: PassMIT

Similar skills

  • Connect

    ComposioHQ/awesome-claude-skills

    Connect Claude to any app. An agent skill from ComposioHQ/awesome-claude-skills.

    77k GitHub starsUsed in 3 repos~987 tokens
    Productivity & AutomationAuto-check passed
  • Official

    Starts a throwaway MongoDB 7.0 replica set and runs the Airbyte spec, check, discover and read commands against source-mongodb-v2 images for local end-to-end testing.

    22k GitHub stars~1.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Connecting To Data Source

    aws/agent-toolkit-for-aws

    Official

    Create and troubleshoot AWS Glue connections to JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS), Redshift, Snowflake, and BigQuery.

    2.8k GitHub starsUsed in 1 repo~2.2k tokens
    DatabasesAuto-check passed
  • Source Maps

    thedaviddias/Front-End-Checklist

    A skill your agent uses when auditing slow page loads, heavy assets, or rendering delays related to Provide source maps for production debugging.

    74k GitHub stars~445 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Loading Indicators

    thedaviddias/Front-End-Checklist

    A skill your agent uses when auditing slow page loads, heavy assets, or rendering delays related to Show loading indicators.

    74k GitHub stars~434 tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Font Loading

    thedaviddias/Front-End-Checklist

    A skill your agent uses when auditing slow page loads, heavy assets, or rendering delays related to Optimize web font loading.

    74k GitHub stars~400 tokensUpdated yesterday
    Frontend & DesignAuto-check passed

More from microsoft/data-formulator

  • Error Handling

    microsoft/data-formulator

    Official

    统一错误处理系统。在添加 API 端点、修改错误处理、添加前端 API 调用、编写错误相关测试时使用. An agent skill from microsoft/data-formulator.

    18k GitHub stars~3.8k tokensUpdated yesterday
    Auto-check passed
  • Language Injection

    microsoft/data-formulator

    Official

    LLM Agent 多语言注入规范。在修改 Agent 提示词、添加新的 Agent 端点、处理用户可见的后端消息(messagecode)时使用。

    18k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Path Safety

    microsoft/data-formulator

    Official

    服务端路径安全与文件访问编码规范。在编写文件下载路由、Agent 工具(文件读取/目录列出)、数据连接器/Loader、Workspace 路径操作、沙箱配置时使用。

    18k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Report

    microsoft/data-formulator

    Official

    Turn an exploration (threads, findings, charts) into a single Markdown report — note, blog post, executive summary, KPI dashboard, slide brief, or multi-section analytical report, with embedded…

    18k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Core

    microsoft/data-formulator

    Official

    The analyst's built-in capabilities: data-inspection tools and the always-available actions (visualize and askuser).

    18k GitHub stars~5.4k tokensUpdated yesterday
    Auto-check passed

Questions about Data Loading

What does Data Loading do?

Discover connected data sources, add new data connectors through a user-confirmed form, inspect table metadata, and run bounded read-only probes when the current workspace data is insufficient. Data Loading is an agent skill from microsoft/data-formulator, published by the product's own GitHub organization. Discover connected data sources, add new data connectors through a user-confirmed form, inspect table metadata, and run bounded read-only probes when the current workspace data is insufficient.

How do I install Data Loading in Claude Code?

Run `npx skills add microsoft/data-formulator --skill data-loading -a claude-code`. Or copy the skill folder (py-src/data_formulator/analyst/skills/data-loading in microsoft/data-formulator) into .claude/skills/data-loading in your project. Claude Code loads it when a task matches its description.

How do I install Data Loading in Codex?

Run `npx skills add microsoft/data-formulator --skill data-loading -a codex`. Or copy the skill folder (py-src/data_formulator/analyst/skills/data-loading in microsoft/data-formulator) into .agents/skills/data-loading in your project. Codex loads it when a task matches its description.

Can I use Data Loading in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/data-formulator --skill data-loading -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-loading, .gemini/skills/data-loading, .github/skills/data-loading and .opencode/skills/data-loading in your project.

What does Data Loading need to run?

Going by SKILL.md and its folder, Data Loading needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Data Loading access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Loading safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Loading use?

Data Loading is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Loading use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Loading?

Skills that share tags, products or a category with Data Loading: Connect (ComposioHQ/awesome-claude-skills, 77k stars), MongoDB Source Connector E2E Harness (airbytehq/airbyte, 22k stars), Connecting To Data Source (aws/agent-toolkit-for-aws, 2.8k stars) and Source Maps (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Loading?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/data-formulator, which has 17,538 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 5, 2026.

Source: microsoft/data-formulator on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.