Chdb Datastore
vemetric/vemetric
A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.
Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project.
$ npx skills add apache/datafusion-python --skill check-upstream -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install apache/datafusion-python check-upstream --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.ai/skills/check-upstream .claude/skills/check-upstream && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "check-upstream" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/check-upstream into .claude/skills/check-upstream/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "check-upstream", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/apache/datafusion-python/tree/main/.ai/skills/check-upstreamType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add apache/datafusion-python --skill check-upstream -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install apache/datafusion-python check-upstream --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.ai/skills/check-upstream .agents/skills/check-upstream && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "check-upstream" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/check-upstream into .agents/skills/check-upstream/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "check-upstream", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add apache/datafusion-python --skill check-upstream -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install apache/datafusion-python check-upstream --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.ai/skills/check-upstream .cursor/skills/check-upstream && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "check-upstream" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/check-upstream into .cursor/skills/check-upstream/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "check-upstream", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/apache/datafusion-python.git --path .ai/skills/check-upstream--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add apache/datafusion-python --skill check-upstream -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install apache/datafusion-python check-upstream --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.ai/skills/check-upstream .gemini/skills/check-upstream && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "check-upstream" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/check-upstream into .gemini/skills/check-upstream/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "check-upstream", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install apache/datafusion-python check-upstreamInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add apache/datafusion-python --skill check-upstream -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .github/skills && cp -r skills-src/.ai/skills/check-upstream .github/skills/check-upstream && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "check-upstream" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/check-upstream into .github/skills/check-upstream/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "check-upstream", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add apache/datafusion-python --skill check-upstream -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install apache/datafusion-python check-upstream --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.ai/skills/check-upstream .opencode/skills/check-upstream && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "check-upstream" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/check-upstream into .opencode/skills/check-upstream/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "check-upstream", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
check-upstreamCheck if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project.
Check Upstream is an agent skill from apache/datafusion-python. Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project. Use when adding missing functions, auditing API coverage, or ensuring parity with upstream.
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering DataFrames. It works with Python and Rust. The repository describes itself as: Apache DataFusion Python Bindings. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6c5d9ff. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are rust and python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.rsdatafusion.apache.orggithub.comapache.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Check Upstream loads about 5.9k tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 2,277 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from apache/datafusion-python at commit 6c5d9ff, republished under its Apache-2.0 licence (© apache). 2,277 words, ~5,933 tokens.
.claude/skills/check-upstream/SKILL.md (or your agent's skills folder).<!---
Licensed to the Apache Software Foundation (ASF) under one
or more contributor license agreements. See the NOTICE file
distributed with this work for additional information
regarding copyright ownership. The ASF licenses this file
to you under the Apache License, Version 2.0 (the
"License"); you may not use this file except in compliance
with the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing,
software distributed under the License is distributed on an
"AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
KIND, either express or implied. See the License for the
specific language governing permissions and limitations
under the License.
-->
You are auditing the datafusion-python project to find features from the upstream Apache DataFusion Rust library that are not yet exposed in this Python binding project. Your goal is to identify gaps and, if asked, implement the missing bindings.
IMPORTANT: The Python API is the source of truth for coverage. A function or method is considered "exposed" if it exists in the Python API (e.g., python/datafusion/functions.py), even if there is no corresponding entry in the Rust bindings. Many upstream functions are aliases of other functions — the Python layer can expose these aliases by calling a different underlying Rust binding. Do NOT report a function as missing if it appears in the Python __all__ list and has a working implementation, regardless of whether a matching #[pyfunction] exists in Rust.
IMPORTANT: audit the total upstream surface, not the delta since the last pin. Gaps accumulate across syncs. A patch-release bump with a "bug fixes only" changelog does not mean there is nothing to find — pre-existing gaps from earlier majors still need to be surfaced. Always run the full comparison.
If a recent upstream bump required any of the following while fixing
compile errors in crates/core/ or the FFI example, treat that as a
hard signal that user-facing surface area grew and run this skill
before considering the bump done. Each pattern corresponds to a class of
gap that frequently shows up in the audit:
| Signal during PR 1 compile fix | Likely gap to check |
|---|---|
New Expr::* variant added to a non-exhaustive match (HigherOrderFunction, Lambda, LambdaVariable, …) | New lambda / higher-order scalar functions (any_match, array_transform, list_transform, …) |
New ScalarValue::* variant (ListView, LargeListView, …) | New scalar / array functions that consume or produce the type |
New required trait method on ExecutionPlan / TableProvider / *UDFImpl (apply_expressions, …) | Corresponding capability on the Python wrapper class |
Renamed or restructured struct field (e.g. Cast.data_type → Cast.field: FieldRef) | Any Python accessor / SKILL.md doc that read the old field |
Newly deprecated trait method with a _with_args / _with_options replacement | The *_with_options variant frequently warrants a separate Python entry point |
PR 1 of dev/release/upstream-sync.md asks you to log these signals as
they appear. When you run this skill, use that log as a checklist: every
entry must either show up in the audit output or be explicitly skipped
with a reason.
The user may specify an area via $ARGUMENTS. If no area is specified or "all" is given, check all areas.
Upstream source of truth:
Where they are exposed in this project:
python/datafusion/functions.py — each function wraps a call to datafusion._internal.functionscrates/core/src/functions.rs — #[pyfunction] definitions registered via init_module()Evaluated and not requiring separate Python exposure:
get_field_path — already covered by get_field(expr, *names), which takes a
variadic field path and dispatches to the same underlying
functions::core::get_field UDF as the upstream get_field_path helper.How to check:
python/datafusion/functions.py (check the __all__ list and function definitions)#[pyfunction]. Many functions are aliases that reuse another function's Rust binding.__all__ list / function definitionsUpstream source of truth:
Where they are exposed in this project:
python/datafusion/functions.py (aggregate functions are mixed in with scalar functions)crates/core/src/functions.rsEvaluated and not requiring separate Python exposure:
count_distinct — covered by count(expr, distinct=True). Both forms call
count_udaf with distinct: bool = true and produce the same logical plan.sum_distinct — covered by sum(expr, distinct=True).avg_distinct — covered by avg(expr, distinct=True).How to check:
python/datafusion/functions.py (check __all__ list and function definitions)Upstream source of truth:
Where they are exposed in this project:
python/datafusion/functions.py (window functions like rank, dense_rank, lag, lead, etc.)crates/core/src/functions.rsHow to check:
python/datafusion/functions.py (check __all__ list and function definitions)Upstream source of truth:
Where they are exposed in this project:
python/datafusion/functions.py and python/datafusion/user_defined.py (TableFunction/udtf)crates/core/src/functions.rs and crates/core/src/udtf.rsHow to check:
Upstream source of truth:
Where they are exposed in this project:
python/datafusion/dataframe.py — the DataFrame classcrates/core/src/dataframe.rs — PyDataFrame with #[pymethods]Evaluated and not requiring separate Python exposure:
show_limit — already covered by DataFrame.show(), which provides the same functionality with a simpler APIwith_param_values — already covered by the param_values argument on SessionContext.sql(), which accomplishes the same thing more robustlyunion_by_name_distinct — already covered by DataFrame.union_by_name(distinct=True), which provides a more Pythonic APIto_string — str(df) is the Pythonic way to get a string and already goes through __repr__ and the configurable formatter. A separate to_string() would either duplicate str(df) or render every row through a different path (session datafusion.format.* options, no formatter), giving a third text rendering alongside repr and show()How to check:
python/datafusion/dataframe.py — this is the source of truth for coveragecrates/core/src/dataframe.rs) may be consulted for context, but a method is covered if it exists in the Python APIUpstream source of truth:
Where they are exposed in this project:
python/datafusion/context.py — the SessionContext classcrates/core/src/context.rs — PySessionContext with #[pymethods]How to check:
python/datafusion/context.py — this is the source of truth for coveragecrates/core/src/context.rs) may be consulted for context, but a method is covered if it exists in the Python APIUpstream source of truth:
Where they are exposed in this project:
crates/core/src/ and crates/util/src/examples/datafusion-ffi-example/src/Cargo.toml and crates/core/Cargo.tomlDiscovering currently supported FFI types:
Grep for use datafusion_ffi:: in crates/core/src/ and crates/util/src/ to find all FFI types currently imported and used.
Evaluated and not requiring direct Python exposure: These upstream FFI types have been reviewed and do not need to be independently exposed to end users:
FFI_ExecutionPlan — already used indirectly through table providers; no need for direct exposureFFI_PhysicalExpr / FFI_PhysicalSortExpr — internal physical planning types not expected to be needed by end usersFFI_RecordBatchStream — one level deeper than FFI_ExecutionPlan, used internally when execution plans stream resultsFFI_SessionRef / ForeignSession — session sharing across FFI; Python manages sessions natively via SessionContextFFI_SessionConfig — Python can configure sessions natively without FFIFFI_ConfigOptions / FFI_TableOptions — internal configuration plumbingFFI_PlanProperties / FFI_Boundedness / FFI_EmissionType — read from existing plans, not user-facingFFI_Partitioning — supporting type for physical planningFFI_Option, FFI_Result, WrappedSchema, WrappedArray, FFI_ColumnarValue, FFI_Volatility, FFI_InsertOp, FFI_AccumulatorArgs, FFI_Accumulator, FFI_GroupsAccumulator, FFI_EmitTo, FFI_AggregateOrderSensitivity, FFI_PartitionEvaluator, FFI_PartitionEvaluatorArgs, FFI_Range, FFI_SortOptions, FFI_Distribution, FFI_ExprProperties, FFI_SortProperties, FFI_Interval, FFI_TableProviderFilterPushDown, FFI_TableType) — used as building blocks within the types above, not independently exposedHow to check:
use datafusion_ffi:: in crates/core/src/ and crates/util/src/, then compare against the upstream datafusion-ffi crate's lib.rs exportsfrom_pycapsule() methodScalarUDFExportable) for FFI objectsinit_module() and Python __init__.pyexamples/datafusion-ffi-example/datafusion-spark crate)Upstream source of truth:
Where they are exposed in this project:
python/datafusion/functions/spark.py — each function wraps
a call to datafusion._internal.functions.spark; the public surface is
the module's __all__ list.crates/core/src/spark_functions.rs — #[pyfunction]
definitions registered via init_module() and re-exported under
datafusion._internal.functions.spark.Coverage policy: The spark namespace mirrors
pyspark.sql.functions parameter names and shapes exactly so pyspark
callers can paste code unchanged. Extras over pyspark are permitted as
long as positional pyspark calls still work — for example, the spark
avg / try_sum / collect_list / collect_set retain the
distinct/filter/order_by/null_treatment kwargs from the main
namespace while pyspark's single-positional form continues to work.
How to check:
datafusion-spark function list from the crate
source under datafusion/spark/src/function/ (each subdirectory is a
category: string/, math/, datetime/, etc.). The crate's
function.rs collects all ScalarUDF factories.pyspark.sql.functions for the public-facing
shape — pyspark is the contract this namespace is matching.python/datafusion/functions/spark.py's __all__. A function is
covered if it exists in the Python spark namespace, even if it
aliases another function's Rust binding.__all__ Hygiene (functions.py and functions/spark.py)Independent of upstream parity, also flag public def symbols in
python/datafusion/functions.py and python/datafusion/functions/spark.py
that are missing from that file's __all__. These are functions a user
can call but that do not show up in
from datafusion.functions import *, in tab-completion against the
namespace, or in generated API docs — typically an oversight rather than
an intentional omission.
How to check:
^def ([a-z_][a-z0-9_]*)\( in each file to enumerate every
public function definition.__all__ list at the top of the same file._).A historical example: instr and position shipped as public defs but
were absent from __all__ until the gap was caught here.
For each finding, propose adding the name to __all__ in alphabetical
position with the existing entries.
After identifying missing APIs, search the open issues at https://github.com/apache/datafusion-python/issues for each gap to see if an issue already exists requesting that API be exposed. Search using the function or method name as the query.
For each area checked, produce a report like:
## [Area Name] Coverage Report
### Currently Exposed (X functions/methods)
- list of what's already available
### Missing from Upstream (Y functions/methods)
- function_name — brief description of what it does (existing issue: #123)
- function_name — brief description of what it does (no existing issue)
### Notes
- Any relevant observations about partial implementations, naming differences, etc.If the user asks you to implement missing features, follow these patterns:
Step 1: Rust binding in crates/core/src/functions.rs:
#[pyfunction]
#[pyo3(signature = (arg1, arg2))]
fn new_function_name(arg1: PyExpr, arg2: PyExpr) -> PyResult<PyExpr> {
Ok(datafusion::functions::module::expr_fn::new_function_name(arg1.expr, arg2.expr).into())
}Then register in init_module():
m.add_wrapped(wrap_pyfunction!(new_function_name))?;Step 2: Python wrapper in python/datafusion/functions.py:
def new_function_name(arg1: Expr, arg2: Expr) -> Expr:
"""Description of what the function does.
Args:
arg1: Description of first argument.
arg2: Description of second argument.
Returns:
Description of return value.
"""
return Expr(f.new_function_name(arg1.expr, arg2.expr))Add to __all__ list.
Step 1: Rust binding in crates/core/src/dataframe.rs:
#[pymethods]
impl PyDataFrame {
fn new_method(&self, py: Python, param: PyExpr) -> PyDataFusionResult<Self> {
let df = self.df.as_ref().clone().new_method(param.into())?;
Ok(Self::new(df))
}
}Step 2: Python wrapper in python/datafusion/dataframe.py:
def new_method(self, param: Expr) -> DataFrame:
"""Description of the method."""
return DataFrame(self.df.new_method(param.expr))Step 1: Rust binding in crates/core/src/context.rs:
#[pymethods]
impl PySessionContext {
pub fn new_method(&self, py: Python, param: String) -> PyDataFusionResult<PyDataFrame> {
let df = wait_for_future(py, self.ctx.new_method(¶m))?;
Ok(PyDataFrame::new(df))
}
}Step 2: Python wrapper in python/datafusion/context.py:
def new_method(self, param: str) -> DataFrame:
"""Description of the method."""
return DataFrame(self.ctx.new_method(param))FFI types require a full pipeline from C struct through to a typed Python wrapper. Each layer must be present.
Step 1: Rust PyO3 wrapper class in a new or existing file under crates/core/src/:
use datafusion_ffi::new_type::FFI_NewType;
#[pyclass(from_py_object, frozen, name = "RawNewType", module = "datafusion.module_name", subclass)]
pub struct PyNewType {
pub inner: Arc<dyn NewTypeTrait>,
}
#[pymethods]
impl PyNewType {
#[staticmethod]
fn from_pycapsule(obj: &Bound<'_, PyAny>) -> PyDataFusionResult<Self> {
let capsule = obj
.getattr("__datafusion_new_type__")?
.call0()?
.downcast::<PyCapsule>()?;
let ffi_ptr = unsafe { capsule.reference::<FFI_NewType>() };
let provider: Arc<dyn NewTypeTrait> = ffi_ptr.into();
Ok(Self { inner: provider })
}
fn some_method(&self) -> PyResult<...> {
// wrap inner trait method
}
}Register in the appropriate init_module():
m.add_class::<PyNewType>()?;Step 2: Python Protocol type in the appropriate Python module (e.g., python/datafusion/catalog.py):
class NewTypeExportable(Protocol):
"""Type hint for objects providing a __datafusion_new_type__ PyCapsule."""
def __datafusion_new_type__(self) -> object: ...Step 3: Python wrapper class in the same module:
class NewType:
"""Description of the type.
This class wraps a DataFusion NewType, which can be created from a native
Python implementation or imported from an FFI-compatible library.
"""
def __init__(
self,
new_type: df_internal.module_name.RawNewType | NewTypeExportable,
) -> None:
if isinstance(new_type, df_internal.module_name.RawNewType):
self._raw = new_type
else:
self._raw = df_internal.module_name.RawNewType.from_pycapsule(new_type)
def some_method(self) -> ReturnType:
"""Description of the method."""
return self._raw.some_method()Step 4: ABC base class (if users should be able to subclass and provide custom implementations in Python):
from abc import ABC, abstractmethod
class NewTypeProvider(ABC):
"""Abstract base class for implementing a custom NewType in Python."""
@abstractmethod
def some_method(self) -> ReturnType:
"""Description of the method."""
...Step 5: Module exports — add to the appropriate __init__.py:
NewType) to python/datafusion/__init__.pyNewTypeProvider) if applicableNewTypeExportable) if it should be publicStep 6: FFI example — add an example implementation under examples/datafusion-ffi-example/src/:
// examples/datafusion-ffi-example/src/new_type.rs
use datafusion_ffi::new_type::FFI_NewType;
// ... example showing how an external Rust library exposes this type via PyCapsuleChecklist for each FFI type:
from_pycapsule() methodNewTypeExportable) for FFI objectsinit_module() and Python __init__.pyexamples/datafusion-ffi-example/Table | TableProviderExportable)crates/core/Cargo.toml — check the datafusion dependency version to ensure you're comparing against the right upstream version.array_append / list_append) should both be exposed if upstream supports them.__all__ list in functions.py to see what's publicly exported vs just defined.© apache, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .ai/skills/check-upstream of apache/datafusion-python.
Open the folder on GitHubat commit 6c5d9ff
Check Upstream next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Check Upstream this skillapache/datafusion-python | 606 | — | ~5.9k | Automated safety check: Pass | Apache-2.0 | |
| Chdb Datastorevemetric/vemetric | 394 | 2 repos | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Polar Python SDKpolarsource/polar | 10k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| CSV Data Summarizercoffeefuelbump/csv-data-summarizer-claude-skill | 468 | 2 repos | ~1.4k | Automated safety check: Pass | None | |
| Python Executorcortega26/chile-hub | 113 | 2 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Retentioneering Product Analyticsretentioneering/retentioneering-tools | 920 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 |
vemetric/vemetric
A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.
polarsource/polar
Integrate Polar billing in server-side Python applications using the versioned Polar and PolarAsync clients.
coffeefuelbump/csv-data-summarizer-claude-skill
Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.
cortega26/chile-hub
Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).
retentioneering/retentioneering-tools
Analyze event logs, clickstreams, user paths, product funnels, retention, behavioral segments, transition graphs, step matrices, sequence patterns, and customer journeys using Retentioneering.
Jeffallan/claude-skills
Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.
apache/datafusion-python
TRIGGER — read before adding, changing, or reviewing any datafusion capsule getter, any FFI export that asks for a TaskContextProvider or an extension codec, or any code that calls…
apache/datafusion-python
A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.
apache/datafusion-python
Audit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API.
apache/datafusion-python
Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping.
Categories
Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project. Check Upstream is an agent skill from apache/datafusion-python. Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project.
Check Upstream fits situations like: adding missing functions; auditing API coverage; ensuring parity with upstream.
Run `npx skills add apache/datafusion-python --skill check-upstream -a claude-code`. Or copy the skill folder (.ai/skills/check-upstream in apache/datafusion-python) into .claude/skills/check-upstream in your project. Claude Code loads it when a task matches its description.
Run `npx skills add apache/datafusion-python --skill check-upstream -a codex`. Or copy the skill folder (.ai/skills/check-upstream in apache/datafusion-python) into .agents/skills/check-upstream in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add apache/datafusion-python --skill check-upstream -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/check-upstream, .gemini/skills/check-upstream, .github/skills/check-upstream and .opencode/skills/check-upstream in your project.
SKILL.md names no scripts, command-line tools or credentials: Check Upstream is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 4 domains. As links in the text: docs.rs, datafusion.apache.org, github.com and apache.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Check Upstream is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Check Upstream: Chdb Datastore (vemetric/vemetric, 394 stars), Polar Python SDK (polarsource/polar, 10k stars), CSV Data Summarizer (coffeefuelbump/csv-data-summarizer-claude-skill, 468 stars) and Python Executor (cortega26/chile-hub, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
apache (a GitHub organization) maintains it in apache/datafusion-python, which has 606 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 6, 2026.
Source: apache/datafusion-python on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.