Apache Spark Engineer
Jeffallan/claude-skills
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
Best practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg.
$ npx skills add Mindrally/skills --skill pyspark-etl -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Mindrally/skills pyspark-etl --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Mindrally/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pyspark-etl .claude/skills/pyspark-etl && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pyspark-etl" agent skill from https://github.com/Mindrally/skills/tree/main/pyspark-etl into .claude/skills/pyspark-etl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyspark-etl", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Mindrally/skills/tree/main/pyspark-etlType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Mindrally/skills --skill pyspark-etl -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Mindrally/skills pyspark-etl --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mindrally/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/pyspark-etl .agents/skills/pyspark-etl && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pyspark-etl" agent skill from https://github.com/Mindrally/skills/tree/main/pyspark-etl into .agents/skills/pyspark-etl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyspark-etl", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mindrally/skills --skill pyspark-etl -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Mindrally/skills pyspark-etl --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mindrally/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/pyspark-etl .cursor/skills/pyspark-etl && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pyspark-etl" agent skill from https://github.com/Mindrally/skills/tree/main/pyspark-etl into .cursor/skills/pyspark-etl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyspark-etl", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Mindrally/skills.git --path pyspark-etl--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Mindrally/skills --skill pyspark-etl -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Mindrally/skills pyspark-etl --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mindrally/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/pyspark-etl .gemini/skills/pyspark-etl && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pyspark-etl" agent skill from https://github.com/Mindrally/skills/tree/main/pyspark-etl into .gemini/skills/pyspark-etl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyspark-etl", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Mindrally/skills pyspark-etlInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Mindrally/skills --skill pyspark-etl -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Mindrally/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/pyspark-etl .github/skills/pyspark-etl && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pyspark-etl" agent skill from https://github.com/Mindrally/skills/tree/main/pyspark-etl into .github/skills/pyspark-etl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyspark-etl", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mindrally/skills --skill pyspark-etl -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Mindrally/skills pyspark-etl --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mindrally/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/pyspark-etl .opencode/skills/pyspark-etl && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pyspark-etl" agent skill from https://github.com/Mindrally/skills/tree/main/pyspark-etl into .opencode/skills/pyspark-etl/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyspark-etl", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pyspark-etlBest practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg.
Pyspark Etl is an agent skill from Mindrally/skills. Best practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg. Use when writing or reviewing PySpark jobs, designing joins and window functions, working with map/array higher-order functions, or building idempotent cumulative/snapshot table merges.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data pipelines and ETL and SQL. It works with Apache Spark. The repository describes itself as: 265+ Claude Code skills for every major framework and language. Install with: npx skills add Mindrally/skills. The licence is Apache-2.0.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 7682ca7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Pyspark Etl loads about 2.5k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 853 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Mindrally/skills at commit 7682ca7, republished under its Apache-2.0 licence (© Mindrally). 853 words, ~2,458 tokens.
.claude/skills/pyspark-etl/SKILL.md (or your agent's skills folder).This skill covers patterns for building production-grade, testable ETL pipelines with PySpark, Spark SQL, and Apache Iceberg, including project structure, join and window-function idioms, and safe cumulative-table merge patterns.
SparkSession lifecycle, accepts an injectable session for testing, and exposes an abstract run_job method.sys.argv..transform() — Chain named methods (read_source().transform(self.enrich).transform(self.merge_with_existing)) so run_job stays pure orchestration.select over withColumn chains, explicit join types, explicit window frames, and native functions instead of UDFs..byName() when writing to Iceberg tables so column order doesn't matter.SparkSession with small, hand-built DataFrames.from abc import ABC, abstractmethod
import logging
from pyspark.sql import SparkSession
class BaseETL(ABC):
def __init__(self, config, app_name="ETL Job", spark_session=None):
self.spark = spark_session or SparkSession.builder.appName(app_name).getOrCreate()
self.config = config
self.logger = logging.getLogger(self.__class__.__name__)
@abstractmethod
def run_job(self): ...
def stop(self):
self.spark.stop()Keep the dataclass as pure data; put CLI parsing in a standalone factory so configs are easy to build in tests.
import argparse
from dataclasses import dataclass
@dataclass
class MyConfig:
read_date: int = 20260101
def create_config() -> MyConfig:
parser = argparse.ArgumentParser()
parser.add_argument("--read_date", type=int, default=20260101)
args = parser.parse_args()
return MyConfig(read_date=args.read_date)Build one generic reader for partition mechanics; keep domain-specific filters visible in the ETL, not buried in a one-off reader class.
import pyspark.sql.functions as F
class PartitionedReader:
@staticmethod
def read_latest(spark, table_name, partition_col):
row = spark.read.table(table_name).agg(F.max(partition_col)).first()
if row is None or row[0] is None:
return spark.createDataFrame([], spark.read.table(table_name).schema)
return spark.read.table(table_name).filter(F.col(partition_col) == row[0])
@staticmethod
def read_by_date(spark, table_name, partition_col, date_value):
return spark.read.table(table_name).filter(F.col(partition_col) == date_value)
events = PartitionedReader.read_by_date(spark, "catalog.my_table", "event_date", 20260319)
events = events.filter(F.col("event_type").isin("login", "purchase"))import pyspark.sql.functions as F and always use F.col() instead of df.colA attribute access — attribute access binds a column to a specific DataFrame variable and breaks after joins or reassignment..filter() or F.when() into named variables once it exceeds 3 expressions.select over chains of withColumn — select states the output schema in one pass, while each withColumn call adds a projection to the query plan..alias() instead of withColumnRenamed.select/filter chains from withColumn chains from join chains by operation type.# BAD — 3 intermediate DataFrames, one projection per call
df = df.withColumn("a", F.col("a").cast("double"))
df = df.withColumn("b", F.upper(F.col("b")))
df = df.withColumn("c", F.lit(1))
# GOOD — one DataFrame, explicit schema contract
df = df.select(
F.col("a").cast("double"),
F.upper(F.col("b")).alias("b"),
F.lit(1).alias("c"),
)how= explicitly — never rely on the default.how="right".withColumnRenamed.F.broadcast() when joining against a large fact table, especially after filters or transformations that prevent Spark from inferring the size automatically (spark.sql.autoBroadcastJoinThreshold only auto-broadcasts tables Spark can size, typically under 10MB). Confirm with df.explain() — look for BroadcastHashJoin vs SortMergeJoin..dropDuplicates() to mask unexpected duplicate rows — find the root cause; it also adds shuffle overhead.flights = flights.alias("f")
parking = parking.alias("p")
result = flights.join(F.broadcast(parking), "code", how="left").select(
F.col("f.start_time").alias("flight_start"),
F.col("p.total_time").alias("parking_total"),
)Use from pyspark.sql import Window as W.
F.sum().over(w) behaves differently depending on whether orderBy is present (running sum vs. total).row_number() + filter (drops rows, keeps the best one) and first() over a window (overwrites a column, keeps all rows).ignorenulls=True to F.first()/F.last() — otherwise a null in the first row of a partition nulls the entire partition's result.partitionBy(); it forces all data into a single partition. Use .agg() for global aggregations instead.w = W.partitionBy("key").orderBy("num").rowsBetween(W.unboundedPreceding, W.unboundedFollowing)
df = df.withColumn("version", F.first("version", ignorenulls=True).over(w))map_zip_with instead of map_concat when merging maps needs per-key conflict resolution (e.g., keep the entry with the later timestamp) rather than one side blindly winning.transform + array_max/array_min to extract values out of nested structs without a UDF.merged = F.map_zip_with(
new_map, existing_map,
lambda key, v1, v2: (
F.when(v1.isNull(), v2)
.when(v2.isNull(), v1)
.otherwise(F.when(v1.event_ts >= v2.event_ts, v1).otherwise(v2))
),
)coalesce argument order.F.lit(None) for empty columns — never empty strings or sentinel values like "NA"..otherwise() as a catch-all in F.when() chains for categorical mappings; an unmapped value should surface as null, not silently collapse into "Other"..show(), .collect(), or .printSchema() in production code — they force full materialization or add driver overhead. .count() is fine when used intentionally for row-count logging or to force materialization before a DAG fork..persist() only when a DataFrame is referenced by multiple subsequent actions. Choose the storage level deliberately: MEMORY_AND_DISK (safe default), MEMORY_ONLY (fastest, risks recompute on eviction), DISK_ONLY (for DataFrames too large for memory)..byName() when writing so Spark matches columns by name, not position — this keeps writes safe across schema evolution.df.write.byName().mode("overwrite").insertInto("catalog.my_table")__partitions Iceberg metadata table to find the latest snapshot instead of scanning the full table:partition_df = spark.read.table("catalog.my_table__partitions").select(
"partition.partition_date", "partition.partition_hour"
)
max_partition = partition_df.orderBy(
F.col("partition_date").desc(), F.col("partition_hour").desc()
).first()
if max_partition is None:
raise ValueError("No partitions found in catalog.my_table")write.distribution-mode deliberately: "none" (fastest, no re-shuffle, file sizes depend on upstream partitioning), "hash" (shuffles by partition key for evenly sized files), "range" (sorts before writing, best scan performance but most expensive).© Mindrally, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in pyspark-etl of Mindrally/skills.
Open the folder on GitHubat commit 7682ca7
Pyspark Etl next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Pyspark Etl this skillMindrally/skills | 271 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Apache Spark EngineerJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Optimizing Databricks SQLAltimateAI/data-engineering-skills | 128 | — | ~6.7k | Automated safety check: Pass | MIT | |
| Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | |
| Datafusion Pythonapache/datafusion-python | 607 | — | ~7.8k | Automated safety check: Pass | Apache-2.0 | |
| Dinobase Business Data Querieskappa90/dinobase | 263 | — | ~1.5k | Automated safety check: Pass | Custom licence |
Jeffallan/claude-skills
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
AltimateAI/data-engineering-skills
Analyze DBSQL queries, including SQL embedded in notebooks (spark.sql(...), %sql cells), for anti-patterns, lint issues, and performance problems, using Databricks-specific dialect and platform…
aws-samples/aws-glue-samples
Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.
apache/datafusion-python
A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.
kappa90/dinobase
Sets up Dinobase, a local DuckDB database that syncs data from 100+ business sources, then answers questions across them with SQL joins and previewed write-backs.
AltimateAI/data-engineering-skills
Creates or modifies dbt models in line with a project's own conventions, then runs dbt build and dbt show to check the output instead of stopping at compile.
Mindrally/skills
Best practices for analytics, data analysis, and visualization using Python, pandas, matplotlib, seaborn, and Jupyter notebooks.
Mindrally/skills
Best practices for AutoML and hyperparameter search with Optuna, Ray Tune, and PyCaret, covering search-space design, validation splits, and leakage prevention.
Mindrally/skills
Best practices for writing Blender Python add-ons using the bpy API, covering operators, panels, properties, registration, and API-safe scripting.
Mindrally/skills
Expert guidelines for Chrome extension development with Manifest V3, covering security, performance, and best practices.
Mindrally/skills
Clean, maintainable, human-readable code principles combined with anti-over-engineering discipline: naming, single responsibility, DRY, and scoping changes to exactly what was requested.
Mindrally/skills
Comprehensive design system guidelines for building consistent, accessible, and scalable component libraries.
Works with
Categories
Best practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg. Pyspark Etl is an agent skill from Mindrally/skills. Best practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg.
Pyspark Etl fits situations like: reviewing PySpark jobs; designing joins and window functions; working with map/array higher-order functions; building idempotent cumulative/snapshot table merges.
Run `npx skills add Mindrally/skills --skill pyspark-etl -a claude-code`. Or copy the skill folder (pyspark-etl in Mindrally/skills) into .claude/skills/pyspark-etl in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Mindrally/skills --skill pyspark-etl -a codex`. Or copy the skill folder (pyspark-etl in Mindrally/skills) into .agents/skills/pyspark-etl in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mindrally/skills --skill pyspark-etl -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pyspark-etl, .gemini/skills/pyspark-etl, .github/skills/pyspark-etl and .opencode/skills/pyspark-etl in your project.
SKILL.md names no scripts, command-line tools or credentials: Pyspark Etl is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Pyspark Etl is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Pyspark Etl: Apache Spark Engineer (Jeffallan/claude-skills, 12k stars), Optimizing Databricks SQL (AltimateAI/data-engineering-skills, 128 stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Datafusion Python (apache/datafusion-python, 607 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Mindrally (a GitHub organization) maintains it in Mindrally/skills, which has 271 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 8, 2026.
Source: Mindrally/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.