Agent skill

Audit Skill Md

by apache in apache/datafusion-python

Audit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API.

Apache-2.0Auto-check passedData & Analytics

Install Audit Skill Md

skills CLI
$ npx skills add apache/datafusion-python --skill audit-skill-md -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install apache/datafusion-python audit-skill-md --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.ai/skills/audit-skill-md .claude/skills/audit-skill-md && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit-skill-md
GitHub stars
606
Token cost
~3.3k tokens
SKILL.md length
1,401 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API.

  • Works in 4 steps: New APIs not mentioned → Stale mentions → Examples that drifted from idiomatic style → …
  • Data & Analytics work in your project
  • SKILL.md covers What the skill covers, Scope argument, Inputs to read and What to look for, plus 4 more sections
  • Calls git, python and pytest

What it does

Audit Skill Md is an agent skill from apache/datafusion-python. Audit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API. Find new APIs that should be documented, stale mentions of removed/renamed APIs, examples that drifted from current idiomatic style, and places that need a "requires datafusion-python NN or newer" note. Run after upstream syncs and before each release.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics. It works with Python. The repository describes itself as: Apache DataFusion Python Bindings. The licence is Apache-2.0.

When your agent uses it

  • Data & Analytics work in your project

Example prompts

  • “requires datafusion-python NN or newer”
  • “/audit-skill-md”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. New APIs not mentioned
  2. Stale mentions
  3. Examples that drifted from idiomatic style
  4. Missing or stale version notes

What it can do on your machine

Read from SKILL.md and the folder at commit 6c5d9ff. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • python
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • apache.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit Skill Md loads about 3.3k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 1,401 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from apache/datafusion-python at commit 6c5d9ff, republished under its Apache-2.0 licence (© apache). 1,401 words, ~3,253 tokens.

Download SKILL.mdSave it as .claude/skills/audit-skill-md/SKILL.md (or your agent's skills folder).
name
audit-skill-md
description
Audit the user-facing skill at skills/datafusion_python/SKILL.md against the current public Python API. Find new APIs that should be documented, stale mentions of removed/renamed APIs, examples that drifted from current idiomatic style, and places that need a "requires datafusion-python NN or newer" note. Run after upstream syncs and before each release.
argument-hint
[scope] (e.g., "session-context", "dataframe", "expr", "functions", "patterns", "pitfalls", "version-notes", "all")
<!---
  Licensed to the Apache Software Foundation (ASF) under one
  or more contributor license agreements.  See the NOTICE file
  distributed with this work for additional information
  regarding copyright ownership.  The ASF licenses this file
  to you under the Apache License, Version 2.0 (the
  "License"); you may not use this file except in compliance
  with the License.  You may obtain a copy of the License at

    http://www.apache.org/licenses/LICENSE-2.0

  Unless required by applicable law or agreed to in writing,
  software distributed under the License is distributed on an
  "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
  KIND, either express or implied.  See the License for the
  specific language governing permissions and limitations
  under the License.
-->

Audit skills/datafusion_python/SKILL.md

You are auditing the user-facing skill at skills/datafusion_python/SKILL.md against the current state of the Python API. The skill is the source of truth for how AI coding assistants are taught to write datafusion-python code, so it must match what the project actually ships. This skill identifies gaps caused by upstream syncs, refactors, or renames, and (if asked) applies the edits directly to SKILL.md.

The skill is most usefully run after the check-upstream step of an upstream sync (see dev/release/upstream-sync.md) — once any new APIs are exposed, this skill makes sure they get documented.

What the skill covers

The user-facing SKILL.md documents these public surfaces. This list is not exhaustive — if a new top-level area is added (e.g., a new Catalog API exposed at the package root), include it.

SurfaceModuleSections in SKILL.md
SessionContextpython/datafusion/context.py"Data Loading"
DataFramepython/datafusion/dataframe.py"DataFrame Operations Quick Reference", "Executing and Collecting Results", "Idiomatic Patterns"
Exprpython/datafusion/expr.py"Expression Building", "Common Pitfalls"
functionspython/datafusion/functions/__init__.py"Available Functions (Categorized)", scattered uses throughout
functions.sparkpython/datafusion/functions/spark.py"Available Functions (Categorized)" → "Spark-Compatible Functions" subsection
Top-level helpers (col, lit, WindowFrame, ...)python/datafusion/__init__.py"Import Conventions", "Core Abstractions"

Scope argument

The user may specify a scope via $ARGUMENTS to limit the audit. If no scope is given or all is specified, audit every area.

ScopeAudit target
session-contextSessionContext methods and the "Data Loading" section
dataframeDataFrame methods and the operations / executing / patterns sections
exprExpr methods/operators and the "Expression Building" section
functionsfunctions/__init__.py __all__ and the "Available Functions (Categorized)" section
spark-functionsfunctions/spark.py __all__, the "Spark-Compatible Functions" subsection, and the divergent-semantics table
patterns"Idiomatic Patterns" section — confirm patterns still match recommended style
pitfalls"Common Pitfalls" — confirm each pitfall still reproduces, drop ones fixed upstream
version-notesCross-check version annotations (see below)
allEverything above

Inputs to read

Before producing the report:

  1. skills/datafusion_python/SKILL.md — the document being audited.
  2. The relevant Python module(s) for the chosen scope. Public surface is the __all__ list (where defined) plus class and def symbols not prefixed with _.
  3. Cargo.toml (root) for the current datafusion-python version — read the version field under [workspace.package] (format NN.0.0). The major version always matches the upstream datafusion crate, so a single datafusion-python version expresses both. python/datafusion/__init__.py's __version__ is the same value exposed at runtime.
  4. Recent commits touching the relevant module(s) for context on what changed since the last sync:
    bash
    git log --oneline -- python/datafusion/dataframe.py | head -20

What to look for

Walk through each scoped area and flag four kinds of issues.

1. New APIs not mentioned

For each public symbol in the module's __all__ (or each public class method), check whether it appears anywhere in SKILL.md. A symbol is "covered" if it shows up in:

  • A code block (the strongest signal — it's demonstrated).
  • The "Available Functions (Categorized)" list.
  • The SQL-to-DataFrame Reference table.

Decide whether each missing symbol deserves an entry. Not every public symbol belongs in SKILL.md — the skill is curated for the patterns users hit daily, not exhaustive API reference. Use these heuristics:

  • Add it if it replaces or supersedes something already in the skill (e.g., a new operation that is the idiomatic alternative to a documented workaround).
  • Add it if it fits a category already present (a new aggregate function goes in the aggregate list; a new join type goes in the joining section).
  • Add it if it changes how a documented pattern should be written.
  • Skip it if it is genuinely niche / advanced / experimental.
  • Skip it if it is internal plumbing exposed for FFI but not user-facing.

When you flag a missing symbol, include a one-line proposed insertion point (which section / which table row) so a reviewer can decide quickly.

2. Stale mentions

For each function name, method name, or import shown in SKILL.md, verify it still exists in the current API:

  • Function names mentioned in prose or in the categorized list should appear in python/datafusion/functions/__init__.py's __all__.
  • Spark function names mentioned in the "Spark-Compatible Functions" subsection should appear in python/datafusion/functions/spark.py's __all__. Also confirm the divergent-semantics table still matches the current spark vs. main signatures.
  • Method calls in code blocks should resolve against the current class.
  • Imports (from datafusion import ...) should succeed against the current __init__.py.

A quick way to check imports without running them:

bash
python -c "from datafusion import SessionContext, col, lit; from datafusion import functions as F; print('ok')"

For each stale mention, propose either:

  • a rename to the current name, or
  • removal if the API is gone with no replacement.
3. Examples that drifted from idiomatic style

The skill teaches a Pythonic style: prefer plain strings to col(...) when a column reference is all you need; prefer raw Python values to lit(...) where auto-wrapping applies. Recent refactors (see the make-pythonic skill) keep moving more functions toward accepting native types.

For each code example in SKILL.md, check:

  • Does it use lit(value) where a raw value would work? Comparison RHS, arithmetic with a column, etc. all auto-wrap. (Reserve lit() for the cases listed in pitfall #2.)
  • Does it use col("name") where a plain string would work? select(...), aggregate([keys], ...), sort(...), sort_by(...) all accept plain name strings.
  • Do functions.py calls match the current pythonic signature for that function? If make-pythonic recently changed a signature (e.g., repeat(string, n: Expr | int)), the example should pass 3 rather than lit(3).
  • Does any example use a deprecated or removed parameter name?

For drift, propose the updated snippet. If the change is purely stylistic and the older form still works, mark the suggestion as non-blocking.

Show full SKILL.md (530 more words)Show less
4. Missing or stale version notes

When an API depends on a specific version, the skill should say so — otherwise an agent referencing the skill in an older project will write code that fails at import or at runtime.

datafusion-python shares its major version number with the upstream datafusion crate (e.g., datafusion-python 53.x tracks upstream datafusion 53). Always express version requirements in terms of datafusion-python only — there is no need to call out upstream and package versions separately.

Add a version note when:

  • A method or function shown in the skill was added in a specific release (e.g., a new DataFrame method that didn't exist before 53).
  • A breaking change altered behavior in a specific release (signature change, default-value change, new required argument).
  • A pitfall was fixed in a specific release. Either annotate the pitfall block with "fixed in datafusion-python NN, kept here for users on older versions" or remove it once the supported floor moves past that version.

Format for version notes (inline, italicized):

markdown
*Requires datafusion-python 53 or newer.*

For each missing/stale version note, propose the exact line and where it belongs.

How to discover changes since the last audit

If the user supplies a previous version or commit SHA where the audit was last run, diff against it:

bash
# Public-API-relevant changes since SHA <prev>
git log --oneline <prev>..HEAD -- python/datafusion/

# Whose signatures actually moved
git diff <prev>..HEAD -- python/datafusion/functions.py | grep '^[+-]def '

If no prior audit point is given, fall back to "since the last upstream sync" by inspecting commits that touch Cargo.toml's datafusion pin:

bash
git log --oneline -- Cargo.toml | grep -i datafusion | head -5

Output Format

Produce a report grouped by scope. Each finding is one bullet with a proposed action, so a maintainer can review the list quickly and apply edits in order.

## SKILL.md Audit (scope: <scope>)

Audited against:
- skills/datafusion_python/SKILL.md @ <git SHA / "working tree">
- datafusion-python <version>

### New APIs to cover
- `DataFrame.foo()` — added in datafusion-python 53. Insert in "DataFrame Operations Quick Reference" under <subsection>.
  Proposed snippet:
  ```python
  df.foo(...)
Stale mentions
  • "old_function_name" referenced in the categorized list (line N) — renamed to "new_function_name". Replace.
Drifted examples
  • "Filtering" section, df.filter(col("a") > lit(10)) — drop lit(10), auto-wrap applies. (non-blocking)
  • "Aggregation" section, df.aggregate([col("region")], ...) — pass "region" as a plain string per "Projection" guidance.
Version notes
  • DataFrame.foo() block needs Requires datafusion-python 53 or newer.
  • "Common Pitfalls" #N — fixed in datafusion-python 53; remove the pitfall and update the SQL-to-DataFrame row to no longer flag the workaround.
No-change confirmed
  • SessionContext data-loading section — all entries match current API.

If asked to apply the changes, edit `skills/datafusion_python/SKILL.md`
directly with `Edit` tool calls, one finding at a time, and re-run the
relevant doctest sanity check at the end:

```bash
pytest --doctest-modules python/datafusion -q

What NOT to flag

  • Internal helpers / underscored names. Private symbols are not part of the user-facing surface.
  • Functions intentionally omitted. Niche / advanced APIs (custom catalogs, raw FFI plumbing, low-level execution plan accessors) live in the API reference, not the skill. If an omission was deliberate and a comment / commit explains why, leave it out.
  • Style nits inside explanatory prose. The skill mixes example code and prose; only enforce the pythonic style on actual code blocks.
  • Function-by-function coverage of every functions.py symbol. The "Available Functions (Categorized)" list is curated by category, not exhaustive. Adding a single new aggregate to the aggregate list is enough — the user follows the pointer to the API reference for the rest.

Coordination with other skills

  • Run /check-upstream first to expose any missing upstream APIs into the Python layer. Without that, this skill cannot recommend documenting something that is not yet exposed.
  • Run /make-pythonic before this skill if a Pythonic-signature pass is planned for a release — that way this skill can update examples to the final signature in one shot rather than churning them twice.
  • The order during an upstream sync (PR 3 of dev/release/upstream-sync.md) is therefore: /check-upstream → /make-pythonic (optional) → /audit-skill-md.

© apache, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .ai/skills/audit-skill-md of apache/datafusion-python.

Open the folder on GitHubat commit 6c5d9ff

Compare with similar skills

Audit Skill Md next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit Skill Md compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit Skill Md this skillapache/datafusion-python606—~3.3kAutomated safety check: PassApache-2.0
Scikit LearnzLanqing/codex-claude-academic-skills4.6k17 repos~3.9kAutomated safety check: PassBSD-3-Clause
TimesFM Forecastinggoogle-research/timesfm34k—~4.7kAutomated safety check: PassApache-2.0
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
StatsmodelszLanqing/codex-claude-academic-skills4.6k16 repos~4.9kAutomated safety check: PassBSD-3-Clause
Scientific Figure MakingChenLiu-1996/figures4papers8.1k—~557Automated safety check: PassCustom licence

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • TimesFM Forecasting

    google-research/timesfm

    Forecasts any univariate time series zero-shot with Google's TimesFM model, returning point forecasts and calibrated prediction intervals without training.

    34k GitHub stars~4.7k tokensUpdated 8 days ago
    Data & AnalyticsAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Statsmodels

    zLanqing/codex-claude-academic-skills

    Statistical models library for Python. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 16 repos~4.9k tokens
    Data & AnalyticsAuto-check passed
  • Scientific Figure Making

    ChenLiu-1996/figures4papers

    Covers publication-ready matplotlib figures for academic papers, slides, and reports—bars, trends, scatter, heatmaps, and multi-panel layouts—with this…

    8.1k GitHub stars~557 tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Diagnose and fix ModuleNotFoundError in Nuitka standalone binaries caused by missing implicit imports.

    15k GitHub stars~519 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from apache/datafusion-python

  • Ffi Capsule Protocol

    apache/datafusion-python

    TRIGGER — read before adding, changing, or reviewing any datafusion capsule getter, any FFI export that asks for a TaskContextProvider or an extension codec, or any code that calls…

    606 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Check Upstream

    apache/datafusion-python

    Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project.

    606 GitHub stars~5.9k tokensUpdated yesterday
    Auto-check passed
  • Datafusion Python

    apache/datafusion-python

    A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.

    606 GitHub stars~7.8k tokensUpdated yesterday
    Auto-check passed
  • Make Pythonic

    apache/datafusion-python

    Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping.

    606 GitHub stars~5.8k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Audit Skill Md

What does Audit Skill Md do?

Audit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API. Audit Skill Md is an agent skill from apache/datafusion-python.md against the current public Python API.

When should I use Audit Skill Md?

Audit Skill Md fits situations like: data & Analytics work in your project.

How do I install Audit Skill Md in Claude Code?

Run `npx skills add apache/datafusion-python --skill audit-skill-md -a claude-code`. Or copy the skill folder (.ai/skills/audit-skill-md in apache/datafusion-python) into .claude/skills/audit-skill-md in your project. Claude Code loads it when a task matches its description.

How do I install Audit Skill Md in Codex?

Run `npx skills add apache/datafusion-python --skill audit-skill-md -a codex`. Or copy the skill folder (.ai/skills/audit-skill-md in apache/datafusion-python) into .agents/skills/audit-skill-md in your project. Codex loads it when a task matches its description.

Can I use Audit Skill Md in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add apache/datafusion-python --skill audit-skill-md -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-skill-md, .gemini/skills/audit-skill-md, .github/skills/audit-skill-md and .opencode/skills/audit-skill-md in your project.

What does Audit Skill Md need to run?

Going by SKILL.md and its folder, Audit Skill Md needs the command-line tools its instructions call (git, python and pytest). Our summary lists: Python 3.

Does Audit Skill Md access the network?

SKILL.md names 1 domain. As links in the text: apache.org. This is read from the text; nothing was executed.

Is Audit Skill Md safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audit Skill Md use?

Audit Skill Md is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit Skill Md use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audit Skill Md?

Skills that share tags, products or a category with Audit Skill Md: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), TimesFM Forecasting (google-research/timesfm, 34k stars), Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars) and Statsmodels (zLanqing/codex-claude-academic-skills, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit Skill Md?

apache (a GitHub organization) maintains it in apache/datafusion-python, which has 606 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 6, 2026.

Source: apache/datafusion-python on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.