Aic Collector Op Development
ai-dynamo/aiconfigurator
Design, add, review, or modify AIC Collector operations and their case population.
Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping.
$ npx skills add apache/datafusion-python --skill make-pythonic -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install apache/datafusion-python make-pythonic --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.ai/skills/make-pythonic .claude/skills/make-pythonic && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "make-pythonic" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/make-pythonic into .claude/skills/make-pythonic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "make-pythonic", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/apache/datafusion-python/tree/main/.ai/skills/make-pythonicType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add apache/datafusion-python --skill make-pythonic -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install apache/datafusion-python make-pythonic --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.ai/skills/make-pythonic .agents/skills/make-pythonic && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "make-pythonic" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/make-pythonic into .agents/skills/make-pythonic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "make-pythonic", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add apache/datafusion-python --skill make-pythonic -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install apache/datafusion-python make-pythonic --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.ai/skills/make-pythonic .cursor/skills/make-pythonic && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "make-pythonic" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/make-pythonic into .cursor/skills/make-pythonic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "make-pythonic", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/apache/datafusion-python.git --path .ai/skills/make-pythonic--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add apache/datafusion-python --skill make-pythonic -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install apache/datafusion-python make-pythonic --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.ai/skills/make-pythonic .gemini/skills/make-pythonic && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "make-pythonic" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/make-pythonic into .gemini/skills/make-pythonic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "make-pythonic", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install apache/datafusion-python make-pythonicInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add apache/datafusion-python --skill make-pythonic -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .github/skills && cp -r skills-src/.ai/skills/make-pythonic .github/skills/make-pythonic && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "make-pythonic" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/make-pythonic into .github/skills/make-pythonic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "make-pythonic", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add apache/datafusion-python --skill make-pythonic -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install apache/datafusion-python make-pythonic --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/apache/datafusion-python.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.ai/skills/make-pythonic .opencode/skills/make-pythonic && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "make-pythonic" agent skill from https://github.com/apache/datafusion-python/tree/main/.ai/skills/make-pythonic into .opencode/skills/make-pythonic/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "make-pythonic", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
make-pythonicAudit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping.
Make Pythonic is an agent skill from apache/datafusion-python. Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping. Analyzes function signatures, checks upstream Rust implementations for type constraints, and applies the appropriate coercion pattern.
Its SKILL.md is about 5.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics. It works with Python, Rust and Apache Spark. The repository describes itself as: Apache DataFusion Python Bindings. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6c5d9ff. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
apache.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Make Pythonic loads about 5.8k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 2,085 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from apache/datafusion-python at commit 6c5d9ff, republished under its Apache-2.0 licence (© apache). 2,085 words, ~5,750 tokens.
.claude/skills/make-pythonic/SKILL.md (or your agent's skills folder).<!---
Licensed to the Apache Software Foundation (ASF) under one
or more contributor license agreements. See the NOTICE file
distributed with this work for additional information
regarding copyright ownership. The ASF licenses this file
to you under the Apache License, Version 2.0 (the
"License"); you may not use this file except in compliance
with the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing,
software distributed under the License is distributed on an
"AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
KIND, either express or implied. See the License for the
specific language governing permissions and limitations
under the License.
-->
You are improving the datafusion-python API to feel more natural to Python users. The goal is to allow functions to accept native Python types (int, float, str, bool, etc.) for arguments that are contextually always or typically literal values, instead of requiring users to manually wrap them in lit().
Core principle: A Python user should be able to write split_part(col("a"), ",", 2) instead of split_part(col("a"), lit(","), lit(2)) when the arguments are contextually obvious literals.
functions vs functions.sparkBoth python/datafusion/functions/__init__.py and
python/datafusion/functions/spark.py are in scope. We want both to feel
pythonic — accept native Python types where the argument is contextually
a literal — but functions.spark carries an additional constraint:
every signature must remain compatible with pyspark.sql.functions.
Compatibility rules for the spark namespace:
spark.shiftleft(col=..., numBits=...)), so renames break
them. Do NOT rename a parameter just because it would be more pythonic
in the main namespace.Column or str (column name) for most args; we accept
Expr already, and widening to Expr | int / Expr | str for
literal-friendly arguments is on-brand because the int/str case is
exactly what a pyspark caller would also try. Just verify the widened
set is a superset of what pyspark accepts for that arg.None and pyspark's positional/keyword form still works (e.g. the
spark avg/try_sum/collect_list/collect_set retain DataFusion's
distinct/filter/order_by/null_treatment kwargs).Practical effect: in functions.spark, apply Categories A and (where
pyspark exposes the same arg as a non-Expr) B normally, but cross-check
each proposed signature against pyspark.sql.functions before landing
it. When pyspark's own type hint is Column | str for a "column name"
arg, prefer leaving the spark wrapper at Expr — Category C
("Expr | str meaning column name") is unusual in functions.py and
should remain so in functions.spark.
The user may specify a scope via $ARGUMENTS. If no scope is given or "all" is specified, audit all functions in python/datafusion/functions/__init__.py and python/datafusion/functions/spark.py. When updating a spark-namespace function, apply the compatibility rules from "Scope" above on top of the standard analysis.
For each function, determine if any parameter can accept native Python types by evaluating two complementary signals:
Some arguments are contextually always or almost always literal values based on what the function does:
| Context | Typical Arguments | Examples |
|---|---|---|
| String position/count | Character counts, indices, repetition counts | left(str, n), right(str, n), repeat(str, n), lpad(str, count, ...) |
| Delimiters/separators | Fixed separator characters | split_part(str, delim, idx), concat_ws(sep, ...) |
| Search/replace patterns | Literal search strings, replacements | replace(str, from, to), regexp_replace(str, pattern, replacement, flags) |
| Date/time parts | Part names from a fixed set | date_part(part, date), date_trunc(part, date) |
| Rounding precision | Decimal place counts | round(val, places), trunc(val, places) |
| Fill characters | Padding characters | lpad(str, count, fill), rpad(str, count, fill) |
Check the Rust binding in crates/core/src/functions.rs and the upstream DataFusion function implementation to determine type constraints. The upstream source is cached locally at:
~/.cargo/registry/src/index.crates.io-*/datafusion-functions-<VERSION>/src/Check the DataFusion version in crates/core/Cargo.toml to find the right directory. Key subdirectories: string/, datetime/, math/, regex/.
For aggregate functions, the upstream source is in a separate crate:
~/.cargo/registry/src/index.crates.io-*/datafusion-functions-aggregate-<VERSION>/src/There are five concrete techniques to check, in order of signal strength:
invoke_with_args() for literal-only enforcement (strongest signal)Some functions pattern-match on ColumnarValue::Scalar in their invoke_with_args() method and return an error if the argument is a column/array. This means the argument must be a literal — passing a column expression will fail at runtime.
Example from date_trunc.rs:
let granularity_str = if let ColumnarValue::Scalar(ScalarValue::Utf8(Some(v))) = granularity {
v.to_lowercase()
} else {
return exec_err!("Granularity of `date_trunc` must be non-null scalar Utf8");
};If you find this pattern: The argument is Category B — accept only the corresponding native Python type (e.g., str), not Expr. The function will error at runtime with a column expression anyway.
accumulator() for literal-only enforcement (aggregate functions)Technique 1 applies to scalar UDFs. Aggregate functions do not have invoke_with_args() — instead, they enforce literal-only arguments in their accumulator() (or create_accumulator()) method, which runs at planning time before any data is processed.
Look for these patterns inside accumulator():
get_scalar_value(expr) — evaluates the expression against an empty batch and errors if it's not a scalarvalidate_percentile_expr(expr) — specific helper used by percentile functionsdowncast_ref::<Literal>() — checks that the physical expression is a literal constantExample from approx_percentile_cont.rs:
fn accumulator(&self, args: AccumulatorArgs) -> Result<ApproxPercentileAccumulator> {
let percentile =
validate_percentile_expr(&args.exprs[1], "APPROX_PERCENTILE_CONT")?;
// ...
}Where validate_percentile_expr calls get_scalar_value and errors with "must be a literal".
Example from string_agg.rs:
fn accumulator(&self, acc_args: AccumulatorArgs) -> Result<Box<dyn Accumulator>> {
let Some(lit) = acc_args.exprs[1].as_any().downcast_ref::<Literal>() else {
return not_impl_err!(
"The second argument of the string_agg function must be a string literal"
);
};
// ...
}If you find this pattern: The argument is Category B — accept only the corresponding native Python type, not Expr. The function will error at planning time with a non-literal expression.
To discover which aggregate functions have literal-only arguments, search the upstream aggregate crate for get_scalar_value, validate_percentile_expr, and downcast_ref::<Literal>() inside accumulator() methods. For example, you should expect to find approx_percentile_cont (percentile) and string_agg (delimiter) among the results.
partition_evaluator() for literal-only enforcement (window functions)Window functions do not have invoke_with_args() or accumulator(). Instead, they enforce literal-only arguments in their partition_evaluator() method, which constructs the evaluator that processes each partition.
The upstream source is in a separate crate:
~/.cargo/registry/src/index.crates.io-*/datafusion-functions-window-<VERSION>/src/Look for get_scalar_value_from_args() calls inside partition_evaluator(). This helper (defined in the window crate's utils.rs) calls downcast_ref::<Literal>() and errors with "There is only support Literal types for field at idx: {index} in Window Function".
Example from ntile.rs:
fn partition_evaluator(
&self,
partition_evaluator_args: PartitionEvaluatorArgs,
) -> Result<Box<dyn PartitionEvaluator>> {
let scalar_n =
get_scalar_value_from_args(partition_evaluator_args.input_exprs(), 0)?
.ok_or_else(|| {
exec_datafusion_err!("NTILE requires a positive integer")
})?;
// ...
}If you find this pattern: The argument is Category B — accept only the corresponding native Python type, not Expr. The function will error at planning time with a non-literal expression.
To discover which window functions have literal-only arguments, search the upstream window crate for get_scalar_value_from_args inside partition_evaluator() methods. For example, you should expect to find ntile (n) and lead/lag (offset, default_value) among the results.
Signature for data type constraintsEach function defines a Signature::coercible(...) that specifies what data types each argument accepts, using Coercion entries. This tells you the expected data type even if it doesn't enforce literal-only.
Example from repeat.rs:
signature: Signature::coercible(
vec![
Coercion::new_exact(TypeSignatureClass::Native(logical_string())),
Coercion::new_implicit(
TypeSignatureClass::Native(logical_int64()),
vec![TypeSignatureClass::Integer],
NativeType::Int64,
),
],
Volatility::Immutable,
),This tells you arg 2 (n) must be an integer type coerced to Int64. Use this to choose the correct Python type (e.g., int not str or float).
Common mappings:
| Rust Type Constraint | Python Type |
|---|---|
logical_int64() / TypeSignatureClass::Integer | int |
logical_float64() / TypeSignatureClass::Numeric | int | float |
logical_string() / TypeSignatureClass::String | str |
LogicalType::Boolean | bool |
Important: In Python's type system (PEP 484), float already accepts int values, so int | float is redundant and will fail the ruff linter (rule PYI041). Use float alone when the Rust side accepts a float/numeric type — Python users can still pass integer literals like log(10, col("a")) or power(col("a"), 3) without issue. Only use int when the Rust side strictly requires an integer (e.g., logical_int64()).
return_field_from_args() for scalar_arguments usageFunctions that inspect literal values at query planning time use args.scalar_arguments.get(n) in their return_field_from_args() method. This indicates the argument is expected to be a literal for optimal behavior (e.g., to determine output type precision), but may still work as a column.
Example from round.rs:
let decimal_places: Option<i32> = match args.scalar_arguments.get(1) {
None => Some(0),
Some(None) => None, // argument is not a literal (column)
Some(Some(scalar)) if scalar.is_null() => Some(0),
Some(Some(scalar)) => Some(decimal_places_from_scalar(scalar)?),
};If you find this pattern: The argument is Category A — accept native types AND Expr. It works as a column but is primarily used as a literal.
What kind of function is this?
Scalar UDF:
Is argument rejected at runtime if not a literal?
(check invoke_with_args for ColumnarValue::Scalar-only match + exec_err!)
→ YES: Category B — accept only native type, no Expr
→ NO: continue below
Aggregate:
Is argument rejected at planning time if not a literal?
(check accumulator() for get_scalar_value / validate_percentile_expr /
downcast_ref::<Literal>() + error)
→ YES: Category B — accept only native type, no Expr
→ NO: continue below
Window:
Is argument rejected at planning time if not a literal?
(check partition_evaluator() for get_scalar_value_from_args /
downcast_ref::<Literal>() + error)
→ YES: Category B — accept only native type, no Expr
→ NO: continue below
Does the Signature constrain it to a specific data type?
→ YES: Category A — accept Expr | <native type matching the constraint>
→ NO: Leave as Expr onlyWhen making a function more pythonic, apply the correct coercion pattern based on what the argument represents:
These are arguments that are typically literals but could be column references in advanced use cases. For these, accept a union type and coerce native types to Expr.literal().
Type hint pattern: Expr | int, Expr | str, Expr | int | str, etc.
When to use: When the argument could plausibly come from a column in some use case (e.g., the repeat count might come from a column in a data-driven scenario).
def repeat(string: Expr, n: Expr | int) -> Expr:
"""Repeats the ``string`` to ``n`` times.
Examples:
>>> ctx = dfn.SessionContext()
>>> df = ctx.from_pydict({"a": ["ha"]})
>>> result = df.select(
... dfn.functions.repeat(dfn.col("a"), 3).alias("r"))
>>> result.collect_column("r")[0].as_py()
'hahaha'
"""
if not isinstance(n, Expr):
n = Expr.literal(n)
return Expr(f.repeat(string.expr, n.expr))These are arguments where an Expr never makes sense because the value must be a fixed literal known at query-planning time (not a per-row value). For these, accept only the native type(s) and wrap internally.
Type hint pattern: str, int, list[str], etc. (no Expr in the union)
When to use: When the argument is from a fixed enumeration or is always a compile-time constant, AND the parameter was not previously typed as Expr:
concat_ws (already typed as str in the Rust binding)array_position (already typed as int in the Rust binding)Backward compatibility rule: If a parameter was previously typed as Expr, you must keep Expr in the union even if the Rust side requires a literal. Removing Expr would break existing user code like date_part(lit("year"), col("a")). Use Category A instead — accept Expr | str — and let users who pass column expressions discover the runtime error from the Rust side. Never silently break backward compatibility.
def concat_ws(separator: str, *args: Expr) -> Expr:
"""Concatenates the list ``args`` with the separator.
``separator`` is already typed as ``str`` in the Rust binding, so
there is no backward-compatibility concern.
Examples:
>>> ctx = dfn.SessionContext()
>>> df = ctx.from_pydict({"a": ["hello"], "b": ["world"]})
>>> result = df.select(
... dfn.functions.concat_ws("-", dfn.col("a"), dfn.col("b")).alias("c"))
>>> result.collect_column("c")[0].as_py()
'hello-world'
"""
args = [arg.expr for arg in args]
return Expr(f.concat_ws(separator, args))In some contexts a string argument naturally refers to a column name rather than a literal. This is the pattern used by DataFrame methods.
Type hint pattern: Expr | str
When to use: Only when the string contextually means a column name (rare in functions.py, more common in DataFrame methods).
# Use _to_raw_expr() from expr.py for this pattern
from datafusion.expr import _to_raw_expr
def some_function(column: Expr | str) -> Expr:
raw = _to_raw_expr(column) # str -> col(str)
return Expr(f.some_function(raw))IMPORTANT: In functions.py, string arguments almost never mean column names. Functions operate on expressions, and column references should use col(). Category C applies mainly to DataFrame methods and context APIs, not to scalar/aggregate/window functions. Do NOT convert string arguments to column expressions in functions.py unless there is a very clear reason to do so.
For each function being updated:
python/datafusion/functions/__init__.pycrates/core/src/functions.rsExpr -> Expr | int)Expr must still workAfter updating a primary function, find all alias functions that delegate to it (e.g., instr and position delegate to strpos). Update each alias's parameter type hints to match the primary function's new signature. Do not add coercion logic to aliases — the primary function handles that.
Per the project's CLAUDE.md rules:
Update examples to demonstrate the pythonic calling convention:
# BEFORE (old style - still works but verbose)
dfn.functions.left(dfn.col("a"), dfn.lit(3))
# AFTER (new style - shown in examples)
dfn.functions.left(dfn.col("a"), 3)After making changes, run the doctests to verify:
python -m pytest --doctest-modules python/datafusion/functions/__init__.py -vUse the coercion helpers from datafusion.expr to convert native Python values to Expr. These are the complement of ensure_expr() — where ensure_expr rejects non-Expr values, the coercion helpers wrap them via Expr.literal().
For required parameters use coerce_to_expr:
from datafusion.expr import coerce_to_expr
def left(string: Expr, n: Expr | int) -> Expr:
n = coerce_to_expr(n)
return Expr(f.left(string.expr, n.expr))For optional nullable parameters use coerce_to_expr_or_none:
from datafusion.expr import coerce_to_expr, coerce_to_expr_or_none
def regexp_count(
string: Expr,
pattern: Expr | str,
start: Expr | int | None = None,
flags: Expr | str | None = None,
) -> Expr:
pattern = coerce_to_expr(pattern)
start = coerce_to_expr_or_none(start)
flags = coerce_to_expr_or_none(flags)
return Expr(
f.regexp_count(
string.expr,
pattern.expr,
start.expr if start is not None else None,
flags.expr if flags is not None else None,
)
)Both helpers are defined in python/datafusion/expr.py alongside ensure_expr. Import them in functions.py via:
from datafusion.expr import coerce_to_expr, coerce_to_expr_or_nonestring in left(string, n) or the array in array_sort(array)), it should remain Expr only. Users should use col() for column references.*args: Expr parameters. These represent multiple expressions and should stay as Expr.Expr and let the user be explicit.return other_function(...), the primary function handles coercion. However, you must update the alias's type hints to match the primary function's signature so that type checkers and documentation accurately reflect what the alias accepts.PyExpr.When auditing functions, process them in this order:
date_part, date_trunc, date_bin — these have the clearest literal argumentsleft, right, repeat, lpad, rpad, split_part, substring, replace, regexp_replace, regexp_match, regexp_count — common and verbose without coercionround, trunc, power — numeric literal argumentsarray_slice, array_position, array_remove_n, array_replace_n, array_resize, array_element — index and count argumentsFor each function analyzed, report:
## [Function Name]
**Current signature:** `function(arg1: Expr, arg2: Expr) -> Expr`
**Proposed signature:** `function(arg1: Expr, arg2: Expr | int) -> Expr`
**Category:** A (accepts native + Expr)
**Arguments changed:**
- `arg2`: Expr -> Expr | int (always a literal count)
**Rust binding:** Takes PyExpr, wraps to literal internally
**Status:** [Changed / Skipped / Needs Discussion]If asked to implement (not just audit), make the changes directly and show a summary of what was updated.
© apache, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .ai/skills/make-pythonic of apache/datafusion-python.
Open the folder on GitHubat commit 6c5d9ff
Make Pythonic next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Make Pythonic this skillapache/datafusion-python | 606 | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | |
| Aic Collector Op Developmentai-dynamo/aiconfigurator | 454 | — | ~3k | Automated safety check: Pass | Apache-2.0 | |
| Elodin DBelodin-sys/elodin | 547 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Apache Spark Optimizationwshobson/agents | 40k | 8 repos | ~789 | Automated safety check: Pass | MIT | |
| GeomasterLeonChaoX/qinyan-academic-skills | 937 | 1 repos | ~2.9k | Automated safety check: Pass | MIT | |
| GeomasterK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~2.8k | Automated safety check: Pass | MIT |
ai-dynamo/aiconfigurator
Design, add, review, or modify AIC Collector operations and their case population.
elodin-sys/elodin
Work with Elodin-DB, the time-series telemetry database. An agent skill from elodin-sys/elodin.
wshobson/agents
Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.
LeonChaoX/qinyan-academic-skills
Comprehensive geospatial science skill covering remote sensing, GIS, spatial analysis, machine learning for earth observation, and 30+ scientific domains.
K-Dense-AI/scientific-agent-skills
Supports geospatial research workflows for remote sensing, vector and raster GIS, spatial statistics, terrain and network analysis, and machine learning for Earth observation.
data-goblin/power-bi-agentic-development
Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access.
apache/datafusion-python
TRIGGER — read before adding, changing, or reviewing any datafusion capsule getter, any FFI export that asks for a TaskContextProvider or an extension codec, or any code that calls…
apache/datafusion-python
Check if upstream Apache DataFusion features (functions, DataFrame ops, SessionContext methods, FFI types) are exposed in this Python project.
apache/datafusion-python
A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.
apache/datafusion-python
Audit the user-facing skill at skills/datafusionpython/SKILL.md against the current public Python API.
Works with
Categories
Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping. Make Pythonic is an agent skill from apache/datafusion-python. Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping.
Make Pythonic fits situations like: data & Analytics work in your project.
Run `npx skills add apache/datafusion-python --skill make-pythonic -a claude-code`. Or copy the skill folder (.ai/skills/make-pythonic in apache/datafusion-python) into .claude/skills/make-pythonic in your project. Claude Code loads it when a task matches its description.
Run `npx skills add apache/datafusion-python --skill make-pythonic -a codex`. Or copy the skill folder (.ai/skills/make-pythonic in apache/datafusion-python) into .agents/skills/make-pythonic in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add apache/datafusion-python --skill make-pythonic -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/make-pythonic, .gemini/skills/make-pythonic, .github/skills/make-pythonic and .opencode/skills/make-pythonic in your project.
Going by SKILL.md and its folder, Make Pythonic needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: apache.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Make Pythonic is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.8k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Make Pythonic: Aic Collector Op Development (ai-dynamo/aiconfigurator, 454 stars), Elodin DB (elodin-sys/elodin, 547 stars), Apache Spark Optimization (wshobson/agents, 40k stars) and Geomaster (LeonChaoX/qinyan-academic-skills, 937 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
apache (a GitHub organization) maintains it in apache/datafusion-python, which has 606 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 6, 2026.
Source: apache/datafusion-python on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.