Datafusion Python
apache/datafusion-python
A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
$ npx skills add Jeffallan/claude-skills --skill spark-engineer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Jeffallan/claude-skills spark-engineer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/spark-engineer .claude/skills/spark-engineer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "spark-engineer" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/spark-engineer into .claude/skills/spark-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-engineer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Jeffallan/claude-skills/tree/main/skills/spark-engineerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Jeffallan/claude-skills --skill spark-engineer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Jeffallan/claude-skills spark-engineer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/spark-engineer .agents/skills/spark-engineer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "spark-engineer" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/spark-engineer into .agents/skills/spark-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-engineer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Jeffallan/claude-skills --skill spark-engineer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Jeffallan/claude-skills spark-engineer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/spark-engineer .cursor/skills/spark-engineer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "spark-engineer" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/spark-engineer into .cursor/skills/spark-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-engineer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Jeffallan/claude-skills.git --path skills/spark-engineer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Jeffallan/claude-skills --skill spark-engineer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Jeffallan/claude-skills spark-engineer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/spark-engineer .gemini/skills/spark-engineer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "spark-engineer" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/spark-engineer into .gemini/skills/spark-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-engineer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Jeffallan/claude-skills spark-engineerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Jeffallan/claude-skills --skill spark-engineer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/spark-engineer .github/skills/spark-engineer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "spark-engineer" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/spark-engineer into .github/skills/spark-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-engineer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Jeffallan/claude-skills --skill spark-engineer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Jeffallan/claude-skills spark-engineer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/spark-engineer .opencode/skills/spark-engineer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "spark-engineer" agent skill from https://github.com/Jeffallan/claude-skills/tree/main/skills/spark-engineer into .opencode/skills/spark-engineer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-engineer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
spark-engineerGuides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
The skill casts the agent as a senior Spark engineer and gives it a five-step workflow: analyze requirements such as data volume, latency and cluster resources, design the pipeline by choosing between DataFrames and RDDs and planning partitioning, implement with optimized transformations and caching, optimize using the Spark UI, and validate. Validation means checking the Spark UI for shuffle spill and verifying partition counts, then returning to optimization if spill or skew shows up.
Five reference files load by context: Spark SQL and DataFrames, RDD operations, partitioning and caching, performance tuning, and streaming patterns such as watermarks and stateful operations. The skill includes PySpark examples for a small pipeline, broadcast joins for small dimension tables, salting to handle data skew, and caching only DataFrames that are reused.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1be15d8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comsynergetic.solutionsjeffallan.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Apache Spark Engineer loads about 1.7k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 109 tokens; SKILL.md has 398 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Jeffallan/claude-skills at commit 1be15d8, republished under its MIT licence (© Jeffallan). 398 words, ~1,679 tokens.
.claude/skills/spark-engineer/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Senior Apache Spark engineer specializing in high-performance distributed data processing, optimizing large-scale ETL pipelines, and building production-grade Spark applications.
df.rdd.getNumPartitions(); if spill or skew detected, return to step 4; test with production-scale data, monitor resource usage, verify performance targetsLoad detailed guidance based on context:
| Topic | Reference | Load When |
|---|---|---|
| Spark SQL & DataFrames | references/spark-sql-dataframes.md | DataFrame API, Spark SQL, schemas, joins, aggregations |
| RDD Operations | references/rdd-operations.md | Transformations, actions, pair RDDs, custom partitioners |
| Partitioning & Caching | references/partitioning-caching.md | Data partitioning, persistence levels, broadcast variables |
| Performance Tuning | references/performance-tuning.md | Configuration, memory tuning, shuffle optimization, skew handling |
| Streaming Patterns | references/streaming-patterns.md | Structured Streaming, watermarks, stateful operations, sinks |
from pyspark.sql import SparkSession
from pyspark.sql import functions as F
from pyspark.sql.types import StructType, StructField, StringType, LongType, DoubleType
spark = SparkSession.builder \
.appName("example-pipeline") \
.config("spark.sql.shuffle.partitions", "400") \
.config("spark.sql.adaptive.enabled", "true") \
.getOrCreate()
# Always define explicit schemas in production
schema = StructType([
StructField("user_id", StringType(), False),
StructField("event_ts", LongType(), False),
StructField("amount", DoubleType(), True),
])
df = spark.read.schema(schema).parquet("s3://bucket/events/")
result = df \
.filter(F.col("amount").isNotNull()) \
.groupBy("user_id") \
.agg(F.sum("amount").alias("total_amount"), F.count("*").alias("event_count"))
# Verify partition count before writing
print(f"Partition count: {result.rdd.getNumPartitions()}")
result.write.mode("overwrite").parquet("s3://bucket/output/")from pyspark.sql.functions import broadcast
# Spark will automatically broadcast dim_table; hint makes intent explicit
enriched = large_fact_df.join(broadcast(dim_df), on="product_id", how="left")import pyspark.sql.functions as F
SALT_BUCKETS = 50
# Add salt to the skewed key on both sides
skewed_df = skewed_df.withColumn("salt", (F.rand() * SALT_BUCKETS).cast("int")) \
.withColumn("salted_key", F.concat(F.col("skewed_key"), F.lit("_"), F.col("salt")))
other_df = other_df.withColumn("salt", F.explode(F.array([F.lit(i) for i in range(SALT_BUCKETS)]))) \
.withColumn("salted_key", F.concat(F.col("skewed_key"), F.lit("_"), F.col("salt")))
result = skewed_df.join(other_df, on="salted_key", how="inner") \
.drop("salt", "salted_key")# Cache ONLY when the DataFrame is reused multiple times
df_cleaned = df.filter(...).withColumn(...).cache()
df_cleaned.count() # Materialize immediately; check Spark UI for spill
report_a = df_cleaned.groupBy("region").agg(...)
report_b = df_cleaned.groupBy("product").agg(...)
df_cleaned.unpersist() # Release when doneWhen implementing Spark solutions, provide:
Spark DataFrame API, Spark SQL, RDD transformations/actions, catalyst optimizer, tungsten execution engine, partitioning strategies, broadcast variables, accumulators, structured streaming, watermarks, checkpointing, Spark UI analysis, memory management, shuffle optimization
Maintained by @jeffallan, Principal Consultant at Synergetic Solutions
© Jeffallan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/spark-engineer of Jeffallan/claude-skills.
Open the folder on GitHubat commit 1be15d8
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Jeffallan/claude-skills, which our catalogue first saw on October 7, 2026.
Apache Spark Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Apache Spark Engineer this skillJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Datafusion Pythonapache/datafusion-python | 606 | — | ~7.8k | Automated safety check: Pass | Apache-2.0 | |
| Apache Spark Optimizationwshobson/agents | 40k | 8 repos | ~789 | Automated safety check: Pass | MIT | |
| Spark Version UpgradeOpenHands/extensions | 157 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Pyspark EtlMindrally/skills | 267 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Duckdb Experttheneoai/awesome-skills | 183 | — | ~4.2k | Automated safety check: Pass | MIT |
apache/datafusion-python
A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.
wshobson/agents
Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.
OpenHands/extensions
Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x).
Mindrally/skills
Best practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg.
theneoai/awesome-skills
DuckDB expert for embedded OLAP analytics, Parquet/CSV querying, and high-performance analytical SQL on local data.
theneoai/awesome-skills
Apache Spark expert: DataFrame API, Spark SQL, Spark Structured Streaming, performance tuning, AQE, and adaptive execution.
Jeffallan/claude-skills
Designs REST and GraphQL APIs from resource modeling to an OpenAPI 3.1 contract, with versioning, pagination and RFC 7807 error handling.
Jeffallan/claude-skills
Walks through designing, building and polishing a command-line tool: user workflow and command hierarchy, implementation in commander, click, typer or cobra, completions and cross-platform testing.
Jeffallan/claude-skills
Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.
Jeffallan/claude-skills
Designs GraphQL schemas and Apollo Federation graphs, with DataLoader resolvers, subscriptions, query complexity limits and caching.
Jeffallan/claude-skills
Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.
Jeffallan/claude-skills
Builds Laravel 10+ applications with Eloquent models, Sanctum authentication, Horizon queues, API resources and Livewire components, tested with Pest or PHPUnit.
Works with
Categories
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming. The skill casts the agent as a senior Spark engineer and gives it a five-step workflow: analyze requirements such as data volume, latency and cluster resources, design the pipeline by choosing between DataFrames and RDDs and planning partitioning, implement with optimized transformations and caching, optimize using the Spark UI, and validate. Validation means checking the Spark UI for shuffle spill and verifying partition counts, then returning to optimization if spill or skew shows up.
Apache Spark Engineer fits situations like: writing Spark jobs and DataFrame transformations; tuning shuffle partitions, executor memory or other cluster settings; fixing data skew or slow joins in a Spark pipeline; building a structured streaming job with watermarks and stateful logic.
Run `npx skills add Jeffallan/claude-skills --skill spark-engineer -a claude-code`. Or copy the skill folder (skills/spark-engineer in Jeffallan/claude-skills) into .claude/skills/spark-engineer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Jeffallan/claude-skills --skill spark-engineer -a codex`. Or copy the skill folder (skills/spark-engineer in Jeffallan/claude-skills) into .agents/skills/spark-engineer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jeffallan/claude-skills --skill spark-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-engineer, .gemini/skills/spark-engineer, .github/skills/spark-engineer and .opencode/skills/spark-engineer in your project.
SKILL.md names no scripts, command-line tools or credentials: Apache Spark Engineer is instructions for the agent only. Our summary lists: Apache Spark or PySpark.
SKILL.md names 3 domains. As links in the text: github.com, synergetic.solutions and jeffallan.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Apache Spark Engineer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Apache Spark Engineer: Datafusion Python (apache/datafusion-python, 606 stars), Apache Spark Optimization (wshobson/agents, 40k stars), Spark Version Upgrade (OpenHands/extensions, 157 stars) and Pyspark Etl (Mindrally/skills, 267 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Jeffallan (a GitHub user) maintains it in Jeffallan/claude-skills, which has 11,754 GitHub stars. The repository holds 58 skills in this directory. The repository was last updated on October 3, 2026.
Source: Jeffallan/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.