Agent skill

Check Upstream

by apache in apache/datafusion-python

Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project.

Apache-2.0Auto-check passedData & Analytics

Install Check Upstream

skills CLI
$ npx skills add apache/datafusion-python --skill check-upstream -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install apache/datafusion-python check-upstream --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.ai/skills/check-upstream .claude/skills/check-upstream && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
check-upstream
GitHub stars
606
Token cost
~5.9k tokens
SKILL.md length
2,277 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project.

  • Works in 9 steps: Scalar Functions → Aggregate Functions → Window Functions → …
  • Adding missing functions
  • SKILL.md covers Compile-Signal Triggers, Areas to Check, Checking for Existing GitHub… and Output Format, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Check Upstream is an agent skill from apache/datafusion-python. Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project. Use when adding missing functions, auditing API coverage, or ensuring parity with upstream.

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering DataFrames. It works with Python and Rust. The repository describes itself as: Apache DataFusion Python Bindings. The licence is Apache-2.0.

When your agent uses it

  • Adding missing functions
  • Auditing API coverage
  • Ensuring parity with upstream

Example prompts

  • “/check-upstream”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Scalar Functions
  2. Aggregate Functions
  3. Window Functions
  4. Table Functions
  5. DataFrame Operations
  6. SessionContext Methods
  7. FFI Types (datafusion-ffi)
  8. Spark-Compatible Functions (datafusion-spark crate)
  9. all Hygiene (functions.py and functions/spark.py)

What it can do on your machine

Read from SKILL.md and the folder at commit 6c5d9ff. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are rust and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.rs
    • datafusion.apache.org
    • github.com
    • apache.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Check Upstream loads about 5.9k tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 2,277 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from apache/datafusion-python at commit 6c5d9ff, republished under its Apache-2.0 licence (© apache). 2,277 words, ~5,933 tokens.

Download SKILL.mdSave it as .claude/skills/check-upstream/SKILL.md (or your agent's skills folder).
name
check-upstream
description
Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project. Use when adding missing functions, auditing API coverage, or ensuring parity with upstream.
argument-hint
[area] (e.g., "scalar functions", "aggregate functions", "window functions", "dataframe", "session context", "ffi types", "all")
<!---
  Licensed to the Apache Software Foundation (ASF) under one
  or more contributor license agreements.  See the NOTICE file
  distributed with this work for additional information
  regarding copyright ownership.  The ASF licenses this file
  to you under the Apache License, Version 2.0 (the
  "License"); you may not use this file except in compliance
  with the License.  You may obtain a copy of the License at

    http://www.apache.org/licenses/LICENSE-2.0

  Unless required by applicable law or agreed to in writing,
  software distributed under the License is distributed on an
  "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
  KIND, either express or implied.  See the License for the
  specific language governing permissions and limitations
  under the License.
-->

Check Upstream DataFusion Feature Coverage

You are auditing the datafusion-python project to find features from the upstream Apache DataFusion Rust library that are not yet exposed in this Python binding project. Your goal is to identify gaps and, if asked, implement the missing bindings.

IMPORTANT: The Python API is the source of truth for coverage. A function or method is considered "exposed" if it exists in the Python API (e.g., python/datafusion/functions.py), even if there is no corresponding entry in the Rust bindings. Many upstream functions are aliases of other functions — the Python layer can expose these aliases by calling a different underlying Rust binding. Do NOT report a function as missing if it appears in the Python __all__ list and has a working implementation, regardless of whether a matching #[pyfunction] exists in Rust.

IMPORTANT: audit the total upstream surface, not the delta since the last pin. Gaps accumulate across syncs. A patch-release bump with a "bug fixes only" changelog does not mean there is nothing to find — pre-existing gaps from earlier majors still need to be surfaced. Always run the full comparison.

Compile-Signal Triggers

If a recent upstream bump required any of the following while fixing compile errors in crates/core/ or the FFI example, treat that as a hard signal that user-facing surface area grew and run this skill before considering the bump done. Each pattern corresponds to a class of gap that frequently shows up in the audit:

Signal during PR 1 compile fixLikely gap to check
New Expr::* variant added to a non-exhaustive match (HigherOrderFunction, Lambda, LambdaVariable, …)New lambda / higher-order scalar functions (any_match, array_transform, list_transform, …)
New ScalarValue::* variant (ListView, LargeListView, …)New scalar / array functions that consume or produce the type
New required trait method on ExecutionPlan / TableProvider / *UDFImpl (apply_expressions, …)Corresponding capability on the Python wrapper class
Renamed or restructured struct field (e.g. Cast.data_type → Cast.field: FieldRef)Any Python accessor / SKILL.md doc that read the old field
Newly deprecated trait method with a _with_args / _with_options replacementThe *_with_options variant frequently warrants a separate Python entry point

PR 1 of dev/release/upstream-sync.md asks you to log these signals as they appear. When you run this skill, use that log as a checklist: every entry must either show up in the audit output or be explicitly skipped with a reason.

Areas to Check

The user may specify an area via $ARGUMENTS. If no area is specified or "all" is given, check all areas.

1. Scalar Functions

Upstream source of truth:

Where they are exposed in this project:

  • Python API: python/datafusion/functions.py — each function wraps a call to datafusion._internal.functions
  • Rust bindings: crates/core/src/functions.rs — #[pyfunction] definitions registered via init_module()

Evaluated and not requiring separate Python exposure:

  • get_field_path — already covered by get_field(expr, *names), which takes a variadic field path and dispatches to the same underlying functions::core::get_field UDF as the upstream get_field_path helper.

How to check:

  1. Fetch the upstream scalar function documentation page
  2. Compare against functions listed in python/datafusion/functions.py (check the __all__ list and function definitions)
  3. A function is covered if it exists in the Python API — it does NOT need a dedicated Rust #[pyfunction]. Many functions are aliases that reuse another function's Rust binding.
  4. Check against the "evaluated and not requiring exposure" list before flagging as a gap
  5. Only report functions that are missing from the Python __all__ list / function definitions
2. Aggregate Functions

Upstream source of truth:

Where they are exposed in this project:

  • Python API: python/datafusion/functions.py (aggregate functions are mixed in with scalar functions)
  • Rust bindings: crates/core/src/functions.rs

Evaluated and not requiring separate Python exposure:

  • count_distinct — covered by count(expr, distinct=True). Both forms call count_udaf with distinct: bool = true and produce the same logical plan.
  • sum_distinct — covered by sum(expr, distinct=True).
  • avg_distinct — covered by avg(expr, distinct=True).

How to check:

  1. Fetch the upstream aggregate function documentation page
  2. Compare against aggregate functions in python/datafusion/functions.py (check __all__ list and function definitions)
  3. A function is covered if it exists in the Python API, even if it aliases another function's Rust binding
  4. Check against the "evaluated and not requiring exposure" list before flagging as a gap
  5. Report only functions missing from the Python API
3. Window Functions

Upstream source of truth:

Where they are exposed in this project:

  • Python API: python/datafusion/functions.py (window functions like rank, dense_rank, lag, lead, etc.)
  • Rust bindings: crates/core/src/functions.rs

How to check:

  1. Fetch the upstream window function documentation page
  2. Compare against window functions in python/datafusion/functions.py (check __all__ list and function definitions)
  3. A function is covered if it exists in the Python API, even if it aliases another function's Rust binding
  4. Report only functions missing from the Python API
4. Table Functions

Upstream source of truth:

Where they are exposed in this project:

  • Python API: python/datafusion/functions.py and python/datafusion/user_defined.py (TableFunction/udtf)
  • Rust bindings: crates/core/src/functions.rs and crates/core/src/udtf.rs

How to check:

  1. Fetch the upstream table function documentation
  2. Compare against what's available in the Python API
  3. A function is covered if it exists in the Python API, even if it aliases another function's Rust binding
  4. Report only functions missing from the Python API
5. DataFrame Operations

Upstream source of truth:

Where they are exposed in this project:

  • Python API: python/datafusion/dataframe.py — the DataFrame class
  • Rust bindings: crates/core/src/dataframe.rs — PyDataFrame with #[pymethods]

Evaluated and not requiring separate Python exposure:

  • show_limit — already covered by DataFrame.show(), which provides the same functionality with a simpler API
  • with_param_values — already covered by the param_values argument on SessionContext.sql(), which accomplishes the same thing more robustly
  • union_by_name_distinct — already covered by DataFrame.union_by_name(distinct=True), which provides a more Pythonic API
  • to_string — str(df) is the Pythonic way to get a string and already goes through __repr__ and the configurable formatter. A separate to_string() would either duplicate str(df) or render every row through a different path (session datafusion.format.* options, no formatter), giving a third text rendering alongside repr and show()

How to check:

  1. Fetch the upstream DataFrame documentation page listing all methods
  2. Compare against methods in python/datafusion/dataframe.py — this is the source of truth for coverage
  3. The Rust bindings (crates/core/src/dataframe.rs) may be consulted for context, but a method is covered if it exists in the Python API
  4. Check against the "evaluated and not requiring exposure" list before flagging as a gap
  5. Report only methods missing from the Python API
6. SessionContext Methods

Upstream source of truth:

Where they are exposed in this project:

  • Python API: python/datafusion/context.py — the SessionContext class
  • Rust bindings: crates/core/src/context.rs — PySessionContext with #[pymethods]

How to check:

  1. Fetch the upstream SessionContext documentation page listing all methods
  2. Compare against methods in python/datafusion/context.py — this is the source of truth for coverage
  3. The Rust bindings (crates/core/src/context.rs) may be consulted for context, but a method is covered if it exists in the Python API
  4. Report only methods missing from the Python API
7. FFI Types (datafusion-ffi)

Upstream source of truth:

Where they are exposed in this project:

  • Rust bindings: various files under crates/core/src/ and crates/util/src/
  • FFI example: examples/datafusion-ffi-example/src/
  • Dependency declared in root Cargo.toml and crates/core/Cargo.toml

Discovering currently supported FFI types: Grep for use datafusion_ffi:: in crates/core/src/ and crates/util/src/ to find all FFI types currently imported and used.

Evaluated and not requiring direct Python exposure: These upstream FFI types have been reviewed and do not need to be independently exposed to end users:

  • FFI_ExecutionPlan — already used indirectly through table providers; no need for direct exposure
  • FFI_PhysicalExpr / FFI_PhysicalSortExpr — internal physical planning types not expected to be needed by end users
  • FFI_RecordBatchStream — one level deeper than FFI_ExecutionPlan, used internally when execution plans stream results
  • FFI_SessionRef / ForeignSession — session sharing across FFI; Python manages sessions natively via SessionContext
  • FFI_SessionConfig — Python can configure sessions natively without FFI
  • FFI_ConfigOptions / FFI_TableOptions — internal configuration plumbing
  • FFI_PlanProperties / FFI_Boundedness / FFI_EmissionType — read from existing plans, not user-facing
  • FFI_Partitioning — supporting type for physical planning
  • Supporting/utility types (FFI_Option, FFI_Result, WrappedSchema, WrappedArray, FFI_ColumnarValue, FFI_Volatility, FFI_InsertOp, FFI_AccumulatorArgs, FFI_Accumulator, FFI_GroupsAccumulator, FFI_EmitTo, FFI_AggregateOrderSensitivity, FFI_PartitionEvaluator, FFI_PartitionEvaluatorArgs, FFI_Range, FFI_SortOptions, FFI_Distribution, FFI_ExprProperties, FFI_SortProperties, FFI_Interval, FFI_TableProviderFilterPushDown, FFI_TableType) — used as building blocks within the types above, not independently exposed

How to check:

  1. Discover currently supported types by grepping for use datafusion_ffi:: in crates/core/src/ and crates/util/src/, then compare against the upstream datafusion-ffi crate's lib.rs exports
  2. If new FFI types appear upstream, evaluate whether they represent a user-facing capability
  3. Check against the "evaluated and not requiring exposure" list before flagging as a gap
  4. Report any genuinely new types that enable user-facing functionality
  5. For each currently supported FFI type, verify the full pipeline is present using the checklist from "Adding a New FFI Type":
    • Rust PyO3 wrapper with from_pycapsule() method
    • Python Protocol type (e.g., ScalarUDFExportable) for FFI objects
    • Python wrapper class with full type hints on all public methods
    • ABC base class (if the type can be user-implemented)
    • Registered in Rust init_module() and Python __init__.py
    • FFI example in examples/datafusion-ffi-example/
    • Type appears in union type hints where accepted
Show full SKILL.md (791 more words)Show less
8. Spark-Compatible Functions (datafusion-spark crate)

Upstream source of truth:

Where they are exposed in this project:

  • Python API: python/datafusion/functions/spark.py — each function wraps a call to datafusion._internal.functions.spark; the public surface is the module's __all__ list.
  • Rust bindings: crates/core/src/spark_functions.rs — #[pyfunction] definitions registered via init_module() and re-exported under datafusion._internal.functions.spark.

Coverage policy: The spark namespace mirrors pyspark.sql.functions parameter names and shapes exactly so pyspark callers can paste code unchanged. Extras over pyspark are permitted as long as positional pyspark calls still work — for example, the spark avg / try_sum / collect_list / collect_set retain the distinct/filter/order_by/null_treatment kwargs from the main namespace while pyspark's single-positional form continues to work.

How to check:

  1. Fetch the upstream datafusion-spark function list from the crate source under datafusion/spark/src/function/ (each subdirectory is a category: string/, math/, datetime/, etc.). The crate's function.rs collects all ScalarUDF factories.
  2. Cross-reference against pyspark.sql.functions for the public-facing shape — pyspark is the contract this namespace is matching.
  3. Compare against the functions listed in python/datafusion/functions/spark.py's __all__. A function is covered if it exists in the Python spark namespace, even if it aliases another function's Rust binding.
  4. Report functions that are missing from the Python spark namespace.
9. __all__ Hygiene (functions.py and functions/spark.py)

Independent of upstream parity, also flag public def symbols in python/datafusion/functions.py and python/datafusion/functions/spark.py that are missing from that file's __all__. These are functions a user can call but that do not show up in from datafusion.functions import *, in tab-completion against the namespace, or in generated API docs — typically an oversight rather than an intentional omission.

How to check:

  1. Grep for ^def ([a-z_][a-z0-9_]*)\( in each file to enumerate every public function definition.
  2. Read the __all__ list at the top of the same file.
  3. Report any function in (1) that is not in (2). Skip private helpers (names starting with _).

A historical example: instr and position shipped as public defs but were absent from __all__ until the gap was caught here.

For each finding, propose adding the name to __all__ in alphabetical position with the existing entries.

Checking for Existing GitHub Issues

After identifying missing APIs, search the open issues at https://github.com/apache/datafusion-python/issues for each gap to see if an issue already exists requesting that API be exposed. Search using the function or method name as the query.

  • If an existing issue is found, include a link to it in the report. Do NOT create a new issue.
  • If no existing issue is found, note that no issue exists yet. If the user asks to create issues for missing APIs, each issue should specify that Python test coverage is required as part of the implementation.

Output Format

For each area checked, produce a report like:

## [Area Name] Coverage Report

### Currently Exposed (X functions/methods)
- list of what's already available

### Missing from Upstream (Y functions/methods)
- function_name — brief description of what it does (existing issue: #123)
- function_name — brief description of what it does (no existing issue)

### Notes
- Any relevant observations about partial implementations, naming differences, etc.

Implementation Pattern

If the user asks you to implement missing features, follow these patterns:

Adding a New Function (Scalar/Aggregate/Window)

Step 1: Rust binding in crates/core/src/functions.rs:

rust
#[pyfunction]
#[pyo3(signature = (arg1, arg2))]
fn new_function_name(arg1: PyExpr, arg2: PyExpr) -> PyResult<PyExpr> {
    Ok(datafusion::functions::module::expr_fn::new_function_name(arg1.expr, arg2.expr).into())
}

Then register in init_module():

rust
m.add_wrapped(wrap_pyfunction!(new_function_name))?;

Step 2: Python wrapper in python/datafusion/functions.py:

python
def new_function_name(arg1: Expr, arg2: Expr) -> Expr:
    """Description of what the function does.

    Args:
        arg1: Description of first argument.
        arg2: Description of second argument.

    Returns:
        Description of return value.
    """
    return Expr(f.new_function_name(arg1.expr, arg2.expr))

Add to __all__ list.

Adding a New DataFrame Method

Step 1: Rust binding in crates/core/src/dataframe.rs:

rust
#[pymethods]
impl PyDataFrame {
    fn new_method(&self, py: Python, param: PyExpr) -> PyDataFusionResult<Self> {
        let df = self.df.as_ref().clone().new_method(param.into())?;
        Ok(Self::new(df))
    }
}

Step 2: Python wrapper in python/datafusion/dataframe.py:

python
def new_method(self, param: Expr) -> DataFrame:
    """Description of the method."""
    return DataFrame(self.df.new_method(param.expr))
Adding a New SessionContext Method

Step 1: Rust binding in crates/core/src/context.rs:

rust
#[pymethods]
impl PySessionContext {
    pub fn new_method(&self, py: Python, param: String) -> PyDataFusionResult<PyDataFrame> {
        let df = wait_for_future(py, self.ctx.new_method(&param))?;
        Ok(PyDataFrame::new(df))
    }
}

Step 2: Python wrapper in python/datafusion/context.py:

python
def new_method(self, param: str) -> DataFrame:
    """Description of the method."""
    return DataFrame(self.ctx.new_method(param))
Adding a New FFI Type

FFI types require a full pipeline from C struct through to a typed Python wrapper. Each layer must be present.

Step 1: Rust PyO3 wrapper class in a new or existing file under crates/core/src/:

rust
use datafusion_ffi::new_type::FFI_NewType;

#[pyclass(from_py_object, frozen, name = "RawNewType", module = "datafusion.module_name", subclass)]
pub struct PyNewType {
    pub inner: Arc<dyn NewTypeTrait>,
}

#[pymethods]
impl PyNewType {
    #[staticmethod]
    fn from_pycapsule(obj: &Bound<'_, PyAny>) -> PyDataFusionResult<Self> {
        let capsule = obj
            .getattr("__datafusion_new_type__")?
            .call0()?
            .downcast::<PyCapsule>()?;
        let ffi_ptr = unsafe { capsule.reference::<FFI_NewType>() };
        let provider: Arc<dyn NewTypeTrait> = ffi_ptr.into();
        Ok(Self { inner: provider })
    }

    fn some_method(&self) -> PyResult<...> {
        // wrap inner trait method
    }
}

Register in the appropriate init_module():

rust
m.add_class::<PyNewType>()?;

Step 2: Python Protocol type in the appropriate Python module (e.g., python/datafusion/catalog.py):

python
class NewTypeExportable(Protocol):
    """Type hint for objects providing a __datafusion_new_type__ PyCapsule."""

    def __datafusion_new_type__(self) -> object: ...

Step 3: Python wrapper class in the same module:

python
class NewType:
    """Description of the type.

    This class wraps a DataFusion NewType, which can be created from a native
    Python implementation or imported from an FFI-compatible library.
    """

    def __init__(
        self,
        new_type: df_internal.module_name.RawNewType | NewTypeExportable,
    ) -> None:
        if isinstance(new_type, df_internal.module_name.RawNewType):
            self._raw = new_type
        else:
            self._raw = df_internal.module_name.RawNewType.from_pycapsule(new_type)

    def some_method(self) -> ReturnType:
        """Description of the method."""
        return self._raw.some_method()

Step 4: ABC base class (if users should be able to subclass and provide custom implementations in Python):

python
from abc import ABC, abstractmethod

class NewTypeProvider(ABC):
    """Abstract base class for implementing a custom NewType in Python."""

    @abstractmethod
    def some_method(self) -> ReturnType:
        """Description of the method."""
        ...

Step 5: Module exports — add to the appropriate __init__.py:

  • Add the wrapper class (NewType) to python/datafusion/__init__.py
  • Add the ABC (NewTypeProvider) if applicable
  • Add the Protocol type (NewTypeExportable) if it should be public

Step 6: FFI example — add an example implementation under examples/datafusion-ffi-example/src/:

rust
// examples/datafusion-ffi-example/src/new_type.rs
use datafusion_ffi::new_type::FFI_NewType;
// ... example showing how an external Rust library exposes this type via PyCapsule

Checklist for each FFI type:

  • Rust PyO3 wrapper with from_pycapsule() method
  • Python Protocol type (e.g., NewTypeExportable) for FFI objects
  • Python wrapper class with full type hints on all public methods
  • ABC base class (if the type can be user-implemented)
  • Registered in Rust init_module() and Python __init__.py
  • FFI example in examples/datafusion-ffi-example/
  • Type appears in union type hints where accepted (e.g., Table | TableProviderExportable)

Important Notes

  • The upstream DataFusion version used by this project is specified in crates/core/Cargo.toml — check the datafusion dependency version to ensure you're comparing against the right upstream version.
  • Some upstream features may intentionally not be exposed (e.g., internal-only APIs). Use judgment about what's user-facing.
  • When fetching upstream docs, prefer the published docs.rs documentation as it matches the crate version.
  • Function aliases (e.g., array_append / list_append) should both be exposed if upstream supports them.
  • Check the __all__ list in functions.py to see what's publicly exported vs just defined.

© apache, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .ai/skills/check-upstream of apache/datafusion-python.

Open the folder on GitHubat commit 6c5d9ff

Compare with similar skills

Check Upstream next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Check Upstream compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Check Upstream this skillapache/datafusion-python606—~5.9kAutomated safety check: PassApache-2.0
Chdb Datastorevemetric/vemetric3942 repos~1.4kAutomated safety check: PassApache-2.0
Polar Python SDKpolarsource/polar10k—~1.8kAutomated safety check: PassApache-2.0
CSV Data Summarizercoffeefuelbump/csv-data-summarizer-claude-skill4682 repos~1.4kAutomated safety check: PassNone
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Retentioneering Product Analyticsretentioneering/retentioneering-tools920—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Chdb Datastore

    vemetric/vemetric

    A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

    394 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Polar Python SDK

    polarsource/polar

    Integrate Polar billing in server-side Python applications using the versioned Polar and PolarAsync clients.

    10k GitHub stars~1.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • CSV Data Summarizer

    coffeefuelbump/csv-data-summarizer-claude-skill

    Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.

    468 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Retentioneering Product Analytics

    retentioneering/retentioneering-tools

    Analyze event logs, clickstreams, user paths, product funnels, retention, behavioral segments, transition graphs, step matrices, sequence patterns, and customer journeys using Retentioneering.

    920 GitHub stars~1.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed

More from apache/datafusion-python

  • Ffi Capsule Protocol

    apache/datafusion-python

    TRIGGER — read before adding, changing, or reviewing any datafusion capsule getter, any FFI export that asks for a TaskContextProvider or an extension codec, or any code that calls…

    606 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Datafusion Python

    apache/datafusion-python

    A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.

    606 GitHub stars~7.8k tokensUpdated yesterday
    Auto-check passed
  • Audit Skill Md

    apache/datafusion-python

    Audit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API.

    606 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Make Pythonic

    apache/datafusion-python

    Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping.

    606 GitHub stars~5.8k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Check Upstream

What does Check Upstream do?

Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project. Check Upstream is an agent skill from apache/datafusion-python. Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project.

When should I use Check Upstream?

Check Upstream fits situations like: adding missing functions; auditing API coverage; ensuring parity with upstream.

How do I install Check Upstream in Claude Code?

Run `npx skills add apache/datafusion-python --skill check-upstream -a claude-code`. Or copy the skill folder (.ai/skills/check-upstream in apache/datafusion-python) into .claude/skills/check-upstream in your project. Claude Code loads it when a task matches its description.

How do I install Check Upstream in Codex?

Run `npx skills add apache/datafusion-python --skill check-upstream -a codex`. Or copy the skill folder (.ai/skills/check-upstream in apache/datafusion-python) into .agents/skills/check-upstream in your project. Codex loads it when a task matches its description.

Can I use Check Upstream in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add apache/datafusion-python --skill check-upstream -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/check-upstream, .gemini/skills/check-upstream, .github/skills/check-upstream and .opencode/skills/check-upstream in your project.

What does Check Upstream need to run?

SKILL.md names no scripts, command-line tools or credentials: Check Upstream is instructions for the agent only. Our summary lists: Python 3.

Does Check Upstream access the network?

SKILL.md names 4 domains. As links in the text: docs.rs, datafusion.apache.org, github.com and apache.org. This is read from the text; nothing was executed.

Is Check Upstream safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Check Upstream use?

Check Upstream is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Check Upstream use?

About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Check Upstream?

Skills that share tags, products or a category with Check Upstream: Chdb Datastore (vemetric/vemetric, 394 stars), Polar Python SDK (polarsource/polar, 10k stars), CSV Data Summarizer (coffeefuelbump/csv-data-summarizer-claude-skill, 468 stars) and Python Executor (cortega26/chile-hub, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Check Upstream?

apache (a GitHub organization) maintains it in apache/datafusion-python, which has 606 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 6, 2026.

Source: apache/datafusion-python on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.