Excel and CSV Data Analysis
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
Owns data understanding before any model is designed. An agent skill from probabl-ai/skills.
$ npx skills add probabl-ai/skills --skill explore-ml-data -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install probabl-ai/skills explore-ml-data --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/explore-ml-data .claude/skills/explore-ml-data && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "explore-ml-data" agent skill from https://github.com/probabl-ai/skills/tree/main/skills/explore-ml-data into .claude/skills/explore-ml-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explore-ml-data", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/probabl-ai/skills/tree/main/skills/explore-ml-dataType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add probabl-ai/skills --skill explore-ml-data -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install probabl-ai/skills explore-ml-data --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/explore-ml-data .agents/skills/explore-ml-data && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "explore-ml-data" agent skill from https://github.com/probabl-ai/skills/tree/main/skills/explore-ml-data into .agents/skills/explore-ml-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explore-ml-data", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add probabl-ai/skills --skill explore-ml-data -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install probabl-ai/skills explore-ml-data --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/explore-ml-data .cursor/skills/explore-ml-data && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "explore-ml-data" agent skill from https://github.com/probabl-ai/skills/tree/main/skills/explore-ml-data into .cursor/skills/explore-ml-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explore-ml-data", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/probabl-ai/skills.git --path skills/explore-ml-data--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add probabl-ai/skills --skill explore-ml-data -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install probabl-ai/skills explore-ml-data --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/explore-ml-data .gemini/skills/explore-ml-data && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "explore-ml-data" agent skill from https://github.com/probabl-ai/skills/tree/main/skills/explore-ml-data into .gemini/skills/explore-ml-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explore-ml-data", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install probabl-ai/skills explore-ml-dataInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add probabl-ai/skills --skill explore-ml-data -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/explore-ml-data .github/skills/explore-ml-data && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "explore-ml-data" agent skill from https://github.com/probabl-ai/skills/tree/main/skills/explore-ml-data into .github/skills/explore-ml-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explore-ml-data", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add probabl-ai/skills --skill explore-ml-data -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install probabl-ai/skills explore-ml-data --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/probabl-ai/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/explore-ml-data .opencode/skills/explore-ml-data && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "explore-ml-data" agent skill from https://github.com/probabl-ai/skills/tree/main/skills/explore-ml-data into .opencode/skills/explore-ml-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "explore-ml-data", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
explore-ml-dataOwns data understanding before any model is designed. An agent skill from probabl-ai/skills.
Explore ML Data is an agent skill from probabl-ai/skills. Owns data understanding before any model is designed. Place dataanalysis/dataanalysis.py, run cells run, write dataanalysis.md and JOURNAL § Data understanding. Never design the model, edit src/<pkg/, or modify raw data files. TRIGGER when the user asks to explore, profile, or understand the data; triage sent them here (status.dataanalysis missing); a data source changed; they want to refresh a recorded EDA; or leakage or research arises on a recorded EDA. STOP when status.setup.pending is non-empty (load…
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including reference files (for example `evals/evals.json`, `references/cell_anatomy.md` and `references/extra_analyses.md`).
It sits in Data & Analytics, covering Data analysis. It works with Python. The repository describes itself as: Tabular Data Science Skills for guardrailing AI Agents. The licence is BSD-3-Clause.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5edc7a4. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pythongituvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git and uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Explore ML Data loads about 5.7k tokens when it runs, and up to ~7.8k if it reads all its reference files. Until then it costs about 239 tokens; SKILL.md has 2,742 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from probabl-ai/skills at commit 5edc7a4, republished under its BSD-3-Clause licence (© probabl-ai). 2,742 words, ~5,661 tokens.
.claude/skills/explore-ml-data/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.One project-level exploratory data analysis: a notebook the user
can open, HTML reports, a short data_analysis.md that embeds
them, and a JOURNAL index row.
Details: setup-workspace references/human_facing_prose.md.
Notebook markdown, data_analysis.md, JOURNAL text, and #
comments describe this dataset — not the skills framework, the
CLI, or the command that produced an output. Questions, replies,
and the close narrative use the same data-science language — not
skill ids, G-* names, or the wrapper CLI.
<!-- results-embed: … --> is a site marker. Authoring hints stay
in this skill. style is ruff only.
| Path | Audience |
|---|---|
| raw data (anywhere) | User-owned, read-only |
data_analysis/data_analysis.py | Human notebook — TableReport + ML-gap cells |
data_analysis/data_analysis_<slug>.html | Human — TableReport page, one per family |
data_analysis/*.png | Human — figures for implications, never glance |
data_analysis/<slug>.html | Human — Plotly (or other) HTML, iframe in implications |
data_analysis/data_analysis.md | Human + later modelling — TableReport iframes, implications |
scratch/data_analysis/<slug>.json | Agent — TableReport.json() per family; gitignored |
scratch/data_analysis/extras.json | Agent — tables[], target, leakage, png and html paths |
| JOURNAL § Data understanding | Index: status, 2–4 line summary, link |
TableReport owns dtypes, missingness, univariate distributions, cardinality, and top pairwise associations. Extra cells cover duplicates, target distribution, feature-vs-target, and leakage candidates. Do not duplicate TableReport in extra cells.
Details: references/cell_anatomy.md. Extra recipes:
references/extra_analyses.md.
| You came here for… | → next |
|---|---|
| First EDA (triage or free-text) | → write md; then the five-option continuation board |
| Continuation pick | → pre-defined option, query, automatic exploration, or describe a plot; no end-turn yet |
| Close this stage | → convert / site / git end-turn / triage-ml-task if installed |
| Methodology concern while EDA is done | → skip G-DATA-ANALYSIS; Keep exploring § Automatic exploration (named concern skips the canned survey) |
| Changed data source or "also plot X" | → overwrite data_analysis/data_analysis.*, refresh JOURNAL |
python -m skore_skills status. If
status.setup.pending is non-empty and
status.skills.setup-ml-project is true, load
setup-ml-project and stop. Do not start this skill. When it
returns, continue. Do not load it again on this turn. If that
skill is not installed, name the pending pieces in one line
and stop. Do not invent git init, scaffold, or env init.
If status.setup.env or status.setup.workspace is
declined, stop in one line. A declined git or editable
is not asked again; continue. Do not define PROJECT_ROOT by
hand or write a root JOURNAL.md while workspace setup is
pending.measure
board. Cleaning belongs in build-ml-pipeline.data_analysis/. Raw load may point
anywhere.skipped — <date> and stop.
Skip is valid only when data_analysis/data_analysis.md is
absent (status.data_analysis missing or skipped). If
status is present, do not write skipped. Say the written
analysis stays, and offer to run exploration again (overwrites
data_analysis/data_analysis.*) or keep it. Do not overwrite
until the user accepts the re-run. A named methodology concern
while EDA is done still skips this gate. Do not run site build on skip. The unfitted snapshot build in
build-ml-pipeline still runs before Evaluate.add-python-package for
ipython (env route agent). Decline → skip path. Do not
pixi add / fabricate output.data_analysis/data_analysis.py.
status.policy.tabular; else choose-python-library (recommend
pandas) then add-python-package for that lib and skrub,
matplotlib, and seaborn. No silent default. If G-TABULAR,
target, or families are unanswered, stop after the asks —
no default-path notebook, even as a “Deliverable A assuming
pandas.” Do not install sklearn / skore / pytest unless the
user picked an extra that needs them.<TARGET>=None, <TASK>=none;
TableReport + duplicates only. Do not persist a policy key.templates/family.py
only (TableReport + duplicates) — no leakage / target /
bivariate cells on a family that does not hold the target.
Do not persist a joined modeling table.family_a). Inventing
families means writing them before the answer, not proposing
those slugs in the ask. One file → skip this ask.api get this turn for symbols used (cache hits count).
TableReport.json() keys drift — .get(...).data_analysis/data_analysis.py. Repeat the
TableReport cell per confirmed family (or per file if the
user picked that). Re-run overwrites in place. Default
notebook = those templates only (plus the matching target
snippet). Do not add extra histograms, sns.heatmap /
association matrices, unique-ratio (nunique()/n),
column-dicts, or report.json() cells. The duplicate cell
prints the duplicate count only, not a uniqueness percentage
and not nunique()/n. Leakage is the
template table, not a comment. Default figures: seaborn
displot for the target only inside
templates/target_regression.py /
target_classification.py (describe + one target figure).
Copy that snippet into the live notebook; do not invent a
second target histogram next to TableReport. One
faceted relplot → bivariate_grid.png, last expression
g. Do not import matplotlib.pyplot on the default path.data_analysis.md only.data_analysis/. Ignore specific raw patterns
via setup-git if the user asks (default: don't).- [ ] Setup: pending empty | load setup-ml-project and stop
- [ ] Detect: status.data_analysis present|skipped|missing
- [ ] G-DATA-ANALYSIS: run | skip when the analysis file is absent (skip → JOURNAL only, STOP). present → keep or re-run; never write skipped
- [ ] G-TABULAR + add frame lib + skrub + matplotlib + seaborn
- [ ] Target: inferred | AskUserQuestion | none
- [ ] Families: one file | AskUserQuestion grouping
- [ ] IPython available or add-python-package
- [ ] Load plot-ml-figure if installed; place
data_analysis/data_analysis.py from the template (edit to
the live path); cells run
- [ ] scratch/data_analysis/facts.py → <slug>.json per family
+ extras.json
- [ ] Author data_analysis.md + JOURNAL
- [ ] Preview `site build` if `policy.site` (skip on G-DATA-ANALYSIS skip)
- [ ] AskUserQuestion five options, including Close (skip if user
already closed the turn)Tick, then run the matching step. Re-emit the checklist with evidence. End of turn only after Close.
After G-DATA-ANALYSIS, G-TABULAR, target, families, and IPython
are resolved, emit 1–3 natural sentences immediately before the
first notebook write / cells run. Say that this is local
computation: it profiles the confirmed full table family or
families, runs duplicate / target / bivariate / leakage analyses
that apply, and writes data_analysis/ HTML / figures plus
scratch/data_analysis/ JSON facts. Name the data scope and the
report paths; do not dump Pre-flight as the explanation.
Describe cost from facts, not guesses. TableReport and requested plots scale with table size and number of families; unless a measured duration is already available, say timing depends on those inputs and do not invent minutes. Emit this preview once, not before every cell command; refresh it only when a newly selected extra materially changes the work.
Automatic exploration / a methodology discussion is LLM research, not model fitting or testing: say that it will reason over the recorded EDA, may write a scratch research note, and will stop at a measurement-choice board before local analysis. If a mandatory gate is pending, preview the possible work but do not write or execute the notebook.
Resolve <TARGET> / <TASK> (classification | regression
| none) first. Resolve families (stop condition above).
Copy templates/data_analysis.py if it fits, then edit
— do not paste unused branches. The first family is the one
that contains <TARGET> when a target exists; substitute
<pkg>, <LOAD_RAW_DATA> (in-memory concat of that family's
shards; optional _source_file), <slug> (Python
identifier). Each further family → templates/family.py (<OTHER_SLUG>,
<LOAD_OTHER>). Append
templates/target_regression.py or
templates/target_classification.py (set TARGET / TASK
in the first-family load cell). No target → neither snippet,
no TARGET / TASK lines. Datetime columns on a family →
templates/datetime.py for that family (drop the relplot
cell if there is no numeric target). Two families that share
column names → templates/drift.py with <OTHER_FRAME>.
Disjoint schemas → no drift on the default pass. Do not
append join-coverage cells here. Generated notebook must
not contain if TARGET, if TASK, OTHER = None, empty
datetime loops, or “skip this cell”. Load plot-ml-figure
if installed before writing figure cells (including this
first write). Markdown is about this analysis.
python -m skore_skills style after the write.
python -m skore_skills cells run data_analysis/data_analysis.py — writes HTML and PNGs. A
useless TableReport repr in the digest is expected.
Copy templates/facts.py → scratch/data_analysis/facts.py
with the same families and target; run it; read
scratch/data_analysis/<slug>.json (each family) and
extras.json.
Write data_analysis/data_analysis.md from
templates/data_analysis.md: glance (one iframe per family
and nothing else), modelling implications (include
feature-engineering candidates), open questions. Reports
and figures are embedded, not linked;  or an
HTML iframe (<iframe src="<slug>.html" …>) sits beside the
implication it supports, never in the glance. Glance stays
TableReport-only. Every extras["pngs"] and extras["htmls"]
path is embedded, each with a sentence citing numbers from
those JSON files (or a notebook summary table). Do not save a
figure that earns no such sentence.
Ground claims in both JSON files and the HTML. Do not invent
columns.
JOURNAL § Data understanding table: Status done — <date>,
short summary (shape, target balance/skew, one or two findings
that shape modelling), Report
[data_analysis/data_analysis.md](../data_analysis/data_analysis.md).
Skip path: Status row only. Do not convert or git end-turn on
skip.
Continuation board — unless the user already closed
the turn (“EDA is done”, “close the turn”): if policy.site
is true and export-ml-site is installed, run
python -m skore_skills site build first. Skip in one line
otherwise. If site build errors with mkdocs-material is required, load add-python-package for mkdocs-material
(agent) and build once more. Do not pixi add / uv add.
If that skill is missing, or the retry still fails, name the
error in one line. Name a build error; do not fail the gate.
Link
data_analysis/data_analysis.md plus report.html and
html/data_analysis.html when the build ran. Do not
notebook convert or git end-turn on this preview. Then
AskUserQuestion one pick. None is recommended or
preselected. After the first md, always ask (including
when triage sent you here). Do not say “extra-analyses” or
“standard extra analysis” on this board.
| Label | Subtitle |
|---|---|
| Choose additional pre-defined option | Name only items that apply: interactions / pairplot, PCA, hypothesis tests, subgroup, time-series, text or geo, join keys / coverage (2+ families) |
| Provide a query to extend the exploration | Describe an analysis to add to the notebook (table, test, or plot) |
| Automatic exploration related to the data and problem | Do in-depth research related to the problem and data that we are exploring |
| Describe a plot | You name a chart and I add cells for it |
| Close | End this stage. Do not add another analysis. |
Close → End of turn (User-facing close). Do not rewrite
data_analysis.md on Close. The other four picks go to
Keep exploring and do not ask this board again for the
same pick. Duplicate / target / leakage stay in Modelling
implications, not only Open questions.
If the original prompt already named extras (e.g. PCA), include those cells in step 1 and do not re-ask that extra.
Import failures → add-python-package, do not work around.
Refresh (already done, user asks for more plots): edit the
.py, re-run steps 2–5, then step 6. Do not re-ask
G-DATA-ANALYSIS.
Methodology concern while status.data_analysis is present
(leakage / “research this”): skip G-DATA-ANALYSIS; do not
overwrite the notebook; go to Keep exploring § Automatic
exploration with that named concern (skip the canned
extra-analysis survey).
Always load plot-ml-figure if installed before writing or
rewriting figure cells. The template is not a license to skip
the plotting worker. Missing skill → one-line skip and still
follow that tree (seaborn statistical, pandas simple chart,
matplotlib last). Never plt.close in notebook cells: save PNG
then leave the figure/grid as the cell output.
No convert, no git end-turn. Do not run a second site build until the md is rewritten. Do not say “extra-analyses”
or “standard extra analysis” anywhere this turn (chat,
checklists, or the board). The file
references/extra_analyses.md may be named as a path only.
The continuation board was already asked. Handle the picked
label. Do not ask keep-versus-close, and do not repeat the
board for this pick.
Pre-defined option — the recipe file
references/extra_analyses.md (path only; do not say
“extra-analyses” in chat). Its own allow_multiple board,
all unchecked. Load
plot-ml-figure if installed before figure cells.
Query — wait for the user’s analysis request. Append
cells (load plot-ml-figure if a figure). Not the canned
research survey. Then the refresh step.
Automatic exploration — load research-ml-practice
if installed with stage data_analysis and the canned
survey concern below. Missing skill → one-line skip and
re-ask the five-option board without another site build.
Do not ask intake. Pass JOURNAL,
data_analysis.md (implications + open questions), and
scratch/data_analysis/extras.json as context. The
worker abstracts the problem class (no dataset proper
name) before searching. Canned question:
Given the kind of problem in JOURNAL (domain, task, constraints) and the kinds of structure already seen in EDA (not the dataset’s proper name), what extra measurements on a raw table like this are still worth doing?
Read scratch/research/survey-<slug>.md. Summarize in
chat; do not dump the note. If tools did not run, two
sentences on the named concern (for leakage:
provenance / scoring-time availability) plus the
measure board — do not claim a scratch file was read.
Do not say to drop a raw column. AskUserQuestion
allow_multiple (unchecked) on only sourced
measure extras that are not already in the notebook.
Map onto extra_analyses when a recipe exists; else a
custom cell. declare / evaluate / confirm stay off
this board → Open questions as advice, not findings. Do
not copy Open questions onto the board unless the survey
note listed them with a source.
A user-named methodology concern (leakage / “research
this”) skips the canned survey: pass that concern for
depth, then the same measure board. Summarize
as above if tools did not run; do not say to drop a raw
column.
Describe a plot — load plot-ml-figure if installed;
append cells.
Picks that change the .py: style, cells run, refresh
facts, rewrite data_analysis.md from JSON/PNGs/HTML
(implications from results). Then preview site build if
policy.site and re-ask the five-option board (run path
step 6). Do not invent domain checklists.
Called from triage-ml-task (explore-the-data intent, or
explore-first on a modeling request while data_analysis is
missing) and user free-text.
Calls: add-python-package, api get, choose-python-library /
stack for G-TABULAR, research-ml-practice if installed when the
user wants extra-analysis research or a named methodology
concern, plot-ml-figure if installed before any figure
cells (default notebook, extras, free-text, research-measure),
style after data_analysis.py.
Need a package? Load add-python-package if installed; else name
it and stop. Do not env add here.
Run this block only after Close (or when the user already
closed the turn). Keep exploring never reaches here. Do not
rewrite data_analysis.md in this block.
The user-facing message is a short story plus links. It is not Pre-flight, not a dump of markdown, and not JOURNAL table cells alone.
data_analysis.md.site build ran or is about to, link the site and not
the markdown:
[report.html](<workspace>/report.html) and
html/data_analysis.html. Otherwise
[data_analysis/data_analysis.md](data_analysis/data_analysis.md).
No Skore locator on this stage.This skill owns the close. Keep exploring stays a 1–2 sentence
summary (optional md / site link); it never reaches convert /
git end-turn.
If policy.notebooks is true, export-ml-notebook is installed,
run python -m skore_skills notebook convert data_analysis/data_analysis.py, with --html when policy.site
is also true. Skip in one line otherwise. If convert fails
because ipywidgets is missing, load add-python-package for
it (agent) and convert again. Missing jupytext / nbclient /
nbconvert → one-line skip naming add-python-package;
do not fail the turn, do not pixi add.
Then, if policy.site is true, export-ml-site is installed, run
python -m skore_skills site build after durable files are on
disk. Skip in one line otherwise. If site build errors with
mkdocs-material is required, load add-python-package for
mkdocs-material (agent) and build once more. Do not
pixi add / uv add. If that skill is missing, or the retry
still fails, name the error in one line. Name a build error;
do not fail the data-analysis turn. Name report.html and
html/data_analysis.html in the User-facing close when the
build ran. Do not also send the user to the markdown.
python -m skore_skills git end-turn --stage data_analysis. If
JSON action is invoke, load persist-ml-git only if
status.skills.persist-ml-git is true and stop; that skill
returns to triage. If persist is missing, name the pending
staged paths and stop. Otherwise load triage-ml-task only if
status.skills.triage-ml-task is true; else stop. No git commit.
© probabl-ai, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (references) in skills/explore-ml-data of probabl-ai/skills.
Open the folder on GitHubat commit 5edc7a4
Explore ML Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Explore ML Data this skillprobabl-ai/skills | 135 | — | ~5.7k | Automated safety check: Pass | BSD-3-Clause | |
| Excel and CSV Data Analysisbytedance/deer-flow | 83k | 4 repos | ~2.2k | Automated safety check: Pass | MIT | |
| Python Executorcortega26/chile-hub | 113 | 2 repos | ~1.5k | Automated safety check: Pass | MIT | |
| MatlabzLanqing/codex-claude-academic-skills | 4.6k | 9 repos | ~2.3k | Automated safety check: Notes | GPL-3.0 | |
| Raccoon DataanalysisSenseTime-Copilot/raccoon-dataanalysis-skill | 137 | — | ~1.9k | Automated safety check: Pass | None | |
| Meridian MMM Model Buildinggoogle/meridian | 1.6k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 |
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
cortega26/chile-hub
Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).
zLanqing/codex-claude-academic-skills
MATLAB and GNU Octave numerical computing for matrix operations, data analysis, visualization, and scientific computing.
SenseTime-Copilot/raccoon-dataanalysis-skill
Raccoon (小浣熊) Data Analysis - Remote code interpreter and data visualization service powered by SenseTime.
google/meridian
Takes a user through building a Meridian marketing mix model, from loading CSV data and mapping columns to running EDA, fitting and saving the model.
Jeffallan/claude-skills
Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.
probabl-ai/skills
Add a Python dependency through the project env manager, or ask the user to install it when env.managed is false.
probabl-ai/skills
Declare the pipeline from data source to predictor as a skrub DataOps graph.
probabl-ai/skills
Evaluate one learner with skore.evaluate. An agent skill from probabl-ai/skills.
probabl-ai/skills
Detect an existing ML workspace or scaffold a fresh one via python -m skoreskills scaffold --package <pkg.
probabl-ai/skills
Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/.
probabl-ai/skills
Convert a jupytext percent %% Python file into an executed .ipynb with cell outputs.
Works with
Categories
Owns data understanding before any model is designed. An agent skill from probabl-ai/skills. Explore ML Data is an agent skill from probabl-ai/skills. Owns data understanding before any model is designed.
Explore ML Data fits situations like: the user asks to explore; understand the data; triage sent them here (status.dataanalysis missing); A data source changed.
Run `npx skills add probabl-ai/skills --skill explore-ml-data -a claude-code`. Or copy the skill folder (skills/explore-ml-data in probabl-ai/skills) into .claude/skills/explore-ml-data in your project. Claude Code loads it when a task matches its description.
Run `npx skills add probabl-ai/skills --skill explore-ml-data -a codex`. Or copy the skill folder (skills/explore-ml-data in probabl-ai/skills) into .agents/skills/explore-ml-data in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add probabl-ai/skills --skill explore-ml-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/explore-ml-data, .gemini/skills/explore-ml-data, .github/skills/explore-ml-data and .opencode/skills/explore-ml-data in your project.
Going by SKILL.md and its folder, Explore ML Data needs Python for the scripts in its folder and the command-line tools its instructions call (python, git and uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Explore ML Data is published under the BSD-3-Clause licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Explore ML Data: Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars), Python Executor (cortega26/chile-hub, 113 stars), Matlab (zLanqing/codex-claude-academic-skills, 4.6k stars) and Raccoon Dataanalysis (SenseTime-Copilot/raccoon-dataanalysis-skill, 137 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
probabl-ai (a GitHub organization) maintains it in probabl-ai/skills, which has 135 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 6, 2026.
Source: probabl-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.