Scikit Learn
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
Audit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API.
$ npx skills add apache/datafusion-python --skill audit-skill-md -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install apache/datafusion-python audit-skill-md --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.ai/skills/audit-skill-md .claude/skills/audit-skill-md && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "audit-skill-md" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/audit-skill-md into .claude/skills/audit-skill-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-skill-md", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/apache/datafusion-python/tree/main/.ai/skills/audit-skill-mdType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add apache/datafusion-python --skill audit-skill-md -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install apache/datafusion-python audit-skill-md --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.ai/skills/audit-skill-md .agents/skills/audit-skill-md && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "audit-skill-md" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/audit-skill-md into .agents/skills/audit-skill-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-skill-md", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add apache/datafusion-python --skill audit-skill-md -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install apache/datafusion-python audit-skill-md --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.ai/skills/audit-skill-md .cursor/skills/audit-skill-md && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "audit-skill-md" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/audit-skill-md into .cursor/skills/audit-skill-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-skill-md", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/apache/datafusion-python.git --path .ai/skills/audit-skill-md--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add apache/datafusion-python --skill audit-skill-md -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install apache/datafusion-python audit-skill-md --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.ai/skills/audit-skill-md .gemini/skills/audit-skill-md && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "audit-skill-md" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/audit-skill-md into .gemini/skills/audit-skill-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-skill-md", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install apache/datafusion-python audit-skill-mdInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add apache/datafusion-python --skill audit-skill-md -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .github/skills && cp -r skills-src/.ai/skills/audit-skill-md .github/skills/audit-skill-md && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "audit-skill-md" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/audit-skill-md into .github/skills/audit-skill-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-skill-md", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add apache/datafusion-python --skill audit-skill-md -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install apache/datafusion-python audit-skill-md --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.ai/skills/audit-skill-md .opencode/skills/audit-skill-md && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "audit-skill-md" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/audit-skill-md into .opencode/skills/audit-skill-md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audit-skill-md", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
audit-skill-mdAudit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API.
Audit Skill Md is an agent skill from apache/datafusion-python. Audit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API. Find new APIs that should be documented, stale mentions of removed/renamed APIs, examples that drifted from current idiomatic style, and places that need a "requires datafusion-python NN or newer" note. Run after upstream syncs and before each release.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics. It works with Python. The repository describes itself as: Apache DataFusion Python Bindings. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6c5d9ff. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitpythonpytestFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
apache.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Audit Skill Md loads about 3.3k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 1,401 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from apache/datafusion-python at commit 6c5d9ff, republished under its Apache-2.0 licence (© apache). 1,401 words, ~3,253 tokens.
.claude/skills/audit-skill-md/SKILL.md (or your agent's skills folder).<!---
Licensed to the Apache Software Foundation (ASF) under one
or more contributor license agreements. See the NOTICE file
distributed with this work for additional information
regarding copyright ownership. The ASF licenses this file
to you under the Apache License, Version 2.0 (the
"License"); you may not use this file except in compliance
with the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing,
software distributed under the License is distributed on an
"AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
KIND, either express or implied. See the License for the
specific language governing permissions and limitations
under the License.
-->
skills/datafusion_python/SKILL.mdYou are auditing the user-facing skill at
skills/datafusion_python/SKILL.md
against the current state of the Python API. The skill is the source of truth
for how AI coding assistants are taught to write datafusion-python code, so
it must match what the project actually ships. This skill identifies gaps
caused by upstream syncs, refactors, or renames, and (if asked) applies the
edits directly to SKILL.md.
The skill is most usefully run after the check-upstream step of an
upstream sync (see dev/release/upstream-sync.md) — once any new APIs are
exposed, this skill makes sure they get documented.
The user-facing SKILL.md documents these public surfaces. This list is not
exhaustive — if a new top-level area is added (e.g., a new Catalog API
exposed at the package root), include it.
| Surface | Module | Sections in SKILL.md |
|---|---|---|
SessionContext | python/datafusion/context.py | "Data Loading" |
DataFrame | python/datafusion/dataframe.py | "DataFrame Operations Quick Reference", "Executing and Collecting Results", "Idiomatic Patterns" |
Expr | python/datafusion/expr.py | "Expression Building", "Common Pitfalls" |
functions | python/datafusion/functions/__init__.py | "Available Functions (Categorized)", scattered uses throughout |
functions.spark | python/datafusion/functions/spark.py | "Available Functions (Categorized)" → "Spark-Compatible Functions" subsection |
Top-level helpers (col, lit, WindowFrame, ...) | python/datafusion/__init__.py | "Import Conventions", "Core Abstractions" |
The user may specify a scope via $ARGUMENTS to limit the audit. If no scope
is given or all is specified, audit every area.
| Scope | Audit target |
|---|---|
session-context | SessionContext methods and the "Data Loading" section |
dataframe | DataFrame methods and the operations / executing / patterns sections |
expr | Expr methods/operators and the "Expression Building" section |
functions | functions/__init__.py __all__ and the "Available Functions (Categorized)" section |
spark-functions | functions/spark.py __all__, the "Spark-Compatible Functions" subsection, and the divergent-semantics table |
patterns | "Idiomatic Patterns" section — confirm patterns still match recommended style |
pitfalls | "Common Pitfalls" — confirm each pitfall still reproduces, drop ones fixed upstream |
version-notes | Cross-check version annotations (see below) |
all | Everything above |
Before producing the report:
skills/datafusion_python/SKILL.md — the document being audited.__all__ list (where defined) plus class and def symbols not prefixed
with _.Cargo.toml (root) for the current datafusion-python version — read
the version field under [workspace.package] (format NN.0.0). The
major version always matches the upstream datafusion crate, so a
single datafusion-python version expresses both.
python/datafusion/__init__.py's __version__ is the same value
exposed at runtime.git log --oneline -- python/datafusion/dataframe.py | head -20Walk through each scoped area and flag four kinds of issues.
For each public symbol in the module's __all__ (or each public class
method), check whether it appears anywhere in SKILL.md. A symbol is
"covered" if it shows up in:
Decide whether each missing symbol deserves an entry. Not every public
symbol belongs in SKILL.md — the skill is curated for the patterns users
hit daily, not exhaustive API reference. Use these heuristics:
When you flag a missing symbol, include a one-line proposed insertion point (which section / which table row) so a reviewer can decide quickly.
For each function name, method name, or import shown in SKILL.md, verify it
still exists in the current API:
python/datafusion/functions/__init__.py's __all__.python/datafusion/functions/spark.py's
__all__. Also confirm the divergent-semantics table still matches the
current spark vs. main signatures.from datafusion import ...) should succeed against the current
__init__.py.A quick way to check imports without running them:
python -c "from datafusion import SessionContext, col, lit; from datafusion import functions as F; print('ok')"For each stale mention, propose either:
The skill teaches a Pythonic style: prefer plain strings to col(...) when a
column reference is all you need; prefer raw Python values to lit(...)
where auto-wrapping applies. Recent refactors (see the make-pythonic
skill) keep moving more functions toward accepting native types.
For each code example in SKILL.md, check:
lit(value) where a raw value would work? Comparison RHS,
arithmetic with a column, etc. all auto-wrap. (Reserve lit() for the
cases listed in pitfall #2.)col("name") where a plain string would work? select(...),
aggregate([keys], ...), sort(...), sort_by(...) all accept plain
name strings.functions.py calls match the current pythonic signature for that
function? If make-pythonic recently changed a signature (e.g.,
repeat(string, n: Expr | int)), the example should pass 3 rather than
lit(3).For drift, propose the updated snippet. If the change is purely stylistic and the older form still works, mark the suggestion as non-blocking.
When an API depends on a specific version, the skill should say so — otherwise an agent referencing the skill in an older project will write code that fails at import or at runtime.
datafusion-python shares its major version number with the upstream
datafusion crate (e.g., datafusion-python 53.x tracks upstream
datafusion 53). Always express version requirements in terms of
datafusion-python only — there is no need to call out upstream and
package versions separately.
Add a version note when:
DataFrame method that didn't exist before 53).Format for version notes (inline, italicized):
*Requires datafusion-python 53 or newer.*For each missing/stale version note, propose the exact line and where it belongs.
If the user supplies a previous version or commit SHA where the audit was last run, diff against it:
# Public-API-relevant changes since SHA <prev>
git log --oneline <prev>..HEAD -- python/datafusion/
# Whose signatures actually moved
git diff <prev>..HEAD -- python/datafusion/functions.py | grep '^[+-]def 'If no prior audit point is given, fall back to "since the last upstream
sync" by inspecting commits that touch Cargo.toml's datafusion pin:
git log --oneline -- Cargo.toml | grep -i datafusion | head -5Produce a report grouped by scope. Each finding is one bullet with a proposed action, so a maintainer can review the list quickly and apply edits in order.
## SKILL.md Audit (scope: <scope>)
Audited against:
- skills/datafusion_python/SKILL.md @ <git SHA / "working tree">
- datafusion-python <version>
### New APIs to cover
- `DataFrame.foo()` — added in datafusion-python 53. Insert in "DataFrame Operations Quick Reference" under <subsection>.
Proposed snippet:
```python
df.foo(...)df.filter(col("a") > lit(10)) — drop lit(10), auto-wrap applies. (non-blocking)df.aggregate([col("region")], ...) — pass "region" as a plain string per "Projection" guidance.DataFrame.foo() block needs Requires datafusion-python 53 or newer.SessionContext data-loading section — all entries match current API.
If asked to apply the changes, edit `skills/datafusion_python/SKILL.md`
directly with `Edit` tool calls, one finding at a time, and re-run the
relevant doctest sanity check at the end:
```bash
pytest --doctest-modules python/datafusion -qfunctions.py symbol. The
"Available Functions (Categorized)" list is curated by category, not
exhaustive. Adding a single new aggregate to the aggregate list is
enough — the user follows the pointer to the API reference for the rest./check-upstream first to expose any missing upstream APIs into the
Python layer. Without that, this skill cannot recommend documenting
something that is not yet exposed./make-pythonic before this skill if a Pythonic-signature pass is
planned for a release — that way this skill can update examples to the
final signature in one shot rather than churning them twice.dev/release/upstream-sync.md)
is therefore: /check-upstream → /make-pythonic (optional) →
/audit-skill-md.© apache, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .ai/skills/audit-skill-md of apache/datafusion-python.
Open the folder on GitHubat commit 6c5d9ff
Audit Skill Md next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Audit Skill Md this skillapache/datafusion-python | 606 | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| Scikit LearnzLanqing/codex-claude-academic-skills | 4.6k | 17 repos | ~3.9k | Automated safety check: Pass | BSD-3-Clause | |
| TimesFM Forecastinggoogle-research/timesfm | 34k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | |
| Excel and CSV Data Analysisbytedance/deer-flow | 83k | 4 repos | ~2.2k | Automated safety check: Pass | MIT | |
| StatsmodelszLanqing/codex-claude-academic-skills | 4.6k | 16 repos | ~4.9k | Automated safety check: Pass | BSD-3-Clause | |
| Scientific Figure MakingChenLiu-1996/figures4papers | 8.1k | — | ~557 | Automated safety check: Pass | Custom licence |
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
google-research/timesfm
Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
zLanqing/codex-claude-academic-skills
Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.
ChenLiu-1996/figures4papers
Covers publication-ready matplotlib figures for academic papers, slides, and reports—bars, trends, scatter, heatmaps, and multi-panel layouts—with this…
Nuitka/Nuitka
Diagnose and fix ModuleNotFoundError in Nuitka standalone binaries caused by missing implicit imports.
apache/datafusion-python
TRIGGER — read before adding, changing, or reviewing any datafusion capsule getter, any FFI export that asks for a TaskContextProvider or an extension codec, or any code that calls…
apache/datafusion-python
Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project.
apache/datafusion-python
A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.
apache/datafusion-python
Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping.
Works with
Categories
Audit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API. Audit Skill Md is an agent skill from apache/datafusion-python.md against the current public Python API.
Audit Skill Md fits situations like: data & Analytics work in your project.
Run `npx skills add apache/datafusion-python --skill audit-skill-md -a claude-code`. Or copy the skill folder (.ai/skills/audit-skill-md in apache/datafusion-python) into .claude/skills/audit-skill-md in your project. Claude Code loads it when a task matches its description.
Run `npx skills add apache/datafusion-python --skill audit-skill-md -a codex`. Or copy the skill folder (.ai/skills/audit-skill-md in apache/datafusion-python) into .agents/skills/audit-skill-md in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add apache/datafusion-python --skill audit-skill-md -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-skill-md, .gemini/skills/audit-skill-md, .github/skills/audit-skill-md and .opencode/skills/audit-skill-md in your project.
Going by SKILL.md and its folder, Audit Skill Md needs the command-line tools its instructions call (git, python and pytest). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: apache.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Audit Skill Md is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Audit Skill Md: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), TimesFM Forecasting (google-research/timesfm, 34k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars) and Statsmodels (zLanqing/codex-claude-academic-skills, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
apache (a GitHub organization) maintains it in apache/datafusion-python, which has 606 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 6, 2026.
Source: apache/datafusion-python on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.