Apache Spark Engineer
Jeffallan/claude-skills
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.
$ npx skills add wshobson/agents --skill spark-optimization -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents spark-optimization --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/data-engineering/skills/spark-optimization .claude/skills/spark-optimization && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "spark-optimization" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/spark-optimization into .claude/skills/spark-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-optimization", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/spark-optimizationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill spark-optimization -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents spark-optimization --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/data-engineering/skills/spark-optimization .agents/skills/spark-optimization && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "spark-optimization" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/spark-optimization into .agents/skills/spark-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-optimization", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill spark-optimization -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents spark-optimization --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/data-engineering/skills/spark-optimization .cursor/skills/spark-optimization && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "spark-optimization" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/spark-optimization into .cursor/skills/spark-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-optimization", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/data-engineering/skills/spark-optimization--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill spark-optimization -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents spark-optimization --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/data-engineering/skills/spark-optimization .gemini/skills/spark-optimization && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "spark-optimization" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/spark-optimization into .gemini/skills/spark-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-optimization", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents spark-optimizationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill spark-optimization -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/data-engineering/skills/spark-optimization .github/skills/spark-optimization && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "spark-optimization" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/spark-optimization into .github/skills/spark-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-optimization", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill spark-optimization -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents spark-optimization --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/data-engineering/skills/spark-optimization .opencode/skills/spark-optimization && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "spark-optimization" agent skill from https://github.com/wshobson/agents/tree/main/plugins/data-engineering/skills/spark-optimization into .opencode/skills/spark-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-optimization", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
spark-optimizationSpeed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.
The skill is a tuning reference for Apache Spark jobs. It describes the execution model from driver program to jobs triggered by actions, and a table of performance factors with a remedy for each: shuffle (minimize wide transformations), data skew (salting and broadcast joins), serialization (Kryo and columnar formats), memory (tune executor memory) and partitions (size them sensibly).
The do list covers enabling adaptive query execution, using Parquet or Delta files, broadcasting small tables, watching the Spark UI for skew, spills and garbage collection, and sizing partitions at roughly 128MB to 256MB. The don't list warns against collecting large data to the driver, needless UDFs, over-caching, ignoring skew and calling count() just to test whether data exists. A PySpark quick start and a details reference are included.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Apache Spark Optimization loads about 789 tokens when it runs, and up to ~3.3k if it reads all its reference files. Until then it costs about 53 tokens; SKILL.md has 199 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 199 words, ~789 tokens.
.claude/skills/spark-optimization/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Production patterns for optimizing Apache Spark jobs including partitioning strategies, memory management, shuffle optimization, and performance tuning.
Driver Program
↓
Job (triggered by action)
↓
Stages (separated by shuffles)
↓
Tasks (one per partition)| Factor | Impact | Solution |
|---|---|---|
| Shuffle | Network I/O, disk I/O | Minimize wide transformations |
| Data Skew | Uneven task duration | Salting, broadcast joins |
| Serialization | CPU overhead | Use Kryo, columnar formats |
| Memory | GC pressure, spills | Tune executor memory |
| Partitions | Parallelism | Right-size partitions |
from pyspark.sql import SparkSession
from pyspark.sql import functions as F
# Create optimized Spark session
spark = (SparkSession.builder
.appName("OptimizedJob")
.config("spark.sql.adaptive.enabled", "true")
.config("spark.sql.adaptive.coalescePartitions.enabled", "true")
.config("spark.sql.adaptive.skewJoin.enabled", "true")
.config("spark.serializer", "org.apache.spark.serializer.KryoSerializer")
.config("spark.sql.shuffle.partitions", "200")
.getOrCreate())
# Read with optimized settings
df = (spark.read
.format("parquet")
.option("mergeSchema", "false")
.load("s3://bucket/data/"))
# Efficient transformations
result = (df
.filter(F.col("date") >= "2024-01-01")
.select("id", "amount", "category")
.groupBy("category")
.agg(F.sum("amount").alias("total")))
result.write.mode("overwrite").parquet("s3://bucket/output/")Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.
.count() for existence - Use .take(1) or .isEmpty()© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in plugins/data-engineering/skills/spark-optimization of wshobson/agents.
Open the folder on GitHubat commit 46891e7
We found 19 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 9 other GitHub owners. This page covers the copy in wshobson/agents, which our catalogue first saw on October 7, 2026.
Apache Spark Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Apache Spark Optimization this skillwshobson/agents | 40k | 9 repos | ~789 | Automated safety check: Pass | MIT | |
| Apache Spark EngineerJeffallan/claude-skills | 12k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Executing Sparkdata-goblin/power-bi-agentic-development | 1k | — | ~1.7k | Automated safety check: Pass | GPL-3.0 | |
| Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | |
| Monitor With HaolemeHaolemeApp/Haoleme | 157 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | |
| Tushare Plugin BuilderYourdaylight/stock_datasource | 188 | — | ~2.5k | Automated safety check: Pass | MIT |
Jeffallan/claude-skills
Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.
data-goblin/power-bi-agentic-development
Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access.
aws-samples/aws-glue-samples
Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.
HaolemeApp/Haoleme
Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.
Yourdaylight/stock_datasource
Turns a Tushare API doc URL into a full data plugin for the stock_datasource repo: extractor, ClickHouse schema, query service, config and curl examples.
astronomer/agents
Builds, packs, and deploys compiled Airflow Go SDK bundles so the ExecutableCoordinator can run them.
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
wshobson/agents
Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.
Works with
Categories
Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code. The skill is a tuning reference for Apache Spark jobs. It describes the execution model from driver program to jobs triggered by actions, and a table of performance factors with a remedy for each: shuffle (minimize wide transformations), data skew (salting and broadcast joins), serialization (Kryo and columnar formats), memory (tune executor memory) and partitions (size them sensibly).
Apache Spark Optimization fits situations like: speeding up a slow Spark job; tuning executor memory and partition sizes; diagnosing data skew or heavy shuffles in the Spark UI; scaling a Spark pipeline to larger datasets.
Run `npx skills add wshobson/agents --skill spark-optimization -a claude-code`. Or copy the skill folder (plugins/data-engineering/skills/spark-optimization in wshobson/agents) into .claude/skills/spark-optimization in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill spark-optimization -a codex`. Or copy the skill folder (plugins/data-engineering/skills/spark-optimization in wshobson/agents) into .agents/skills/spark-optimization in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill spark-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-optimization, .gemini/skills/spark-optimization, .github/skills/spark-optimization and .opencode/skills/spark-optimization in your project.
SKILL.md names no scripts, command-line tools or credentials: Apache Spark Optimization is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Apache Spark Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 789 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Apache Spark Optimization: Apache Spark Engineer (Jeffallan/claude-skills, 12k stars), Executing Spark (data-goblin/power-bi-agentic-development, 1k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,287 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.