Agent skill

Apache Spark Optimization

by wshobson in wshobson/agents

Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.

MITAuto-check passedData & Analytics

Install Apache Spark Optimization

skills CLI
$ npx skills add wshobson/agents --skill spark-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents spark-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/data-engineering/skills/spark-optimization .claude/skills/spark-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spark-optimization
GitHub stars
40k
Used in
9 other repos
Token cost
~789 tokens
SKILL.md length
199 words
Files
2 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.

  • Works in 2 steps: Spark Execution Model → Key Performance Factors
  • Speeding up a slow Spark job
  • SKILL.md covers When to Use This Skill, Core Concepts, Quick Start and Detailed patterns and worked…, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill is a tuning reference for Apache Spark jobs. It describes the execution model from driver program to jobs triggered by actions, and a table of performance factors with a remedy for each: shuffle (minimize wide transformations), data skew (salting and broadcast joins), serialization (Kryo and columnar formats), memory (tune executor memory) and partitions (size them sensibly).

The do list covers enabling adaptive query execution, using Parquet or Delta files, broadcasting small tables, watching the Spark UI for skew, spills and garbage collection, and sizing partitions at roughly 128MB to 256MB. The don't list warns against collecting large data to the driver, needless UDFs, over-caching, ignoring skew and calling count() just to test whether data exists. A PySpark quick start and a details reference are included.

When your agent uses it

  • Speeding up a slow Spark job
  • Tuning executor memory and partition sizes
  • Diagnosing data skew or heavy shuffles in the Spark UI
  • Scaling a Spark pipeline to larger datasets

Example prompts

  • “This PySpark job takes hours on the join between orders and customers. Find out why and speed it up.”
  • “Review our Spark config and suggest executor memory and partition settings.”
  • “Replace the UDFs in this transformation with built-in functions.”

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Spark Execution Model
  2. Key Performance Factors

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Apache Spark Optimization loads about 789 tokens when it runs, and up to ~3.3k if it reads all its reference files. Until then it costs about 53 tokens; SKILL.md has 199 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~789
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 199 words, ~789 tokens.

Download SKILL.mdSave it as .claude/skills/spark-optimization/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
spark-optimization
description
Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.

Apache Spark Optimization

Production patterns for optimizing Apache Spark jobs including partitioning strategies, memory management, shuffle optimization, and performance tuning.

When to Use This Skill

  • Optimizing slow Spark jobs
  • Tuning memory and executor configuration
  • Implementing efficient partitioning strategies
  • Debugging Spark performance issues
  • Scaling Spark pipelines for large datasets
  • Reducing shuffle and data skew

Core Concepts

1. Spark Execution Model
Driver Program
    ↓
Job (triggered by action)
    ↓
Stages (separated by shuffles)
    ↓
Tasks (one per partition)
2. Key Performance Factors
FactorImpactSolution
ShuffleNetwork I/O, disk I/OMinimize wide transformations
Data SkewUneven task durationSalting, broadcast joins
SerializationCPU overheadUse Kryo, columnar formats
MemoryGC pressure, spillsTune executor memory
PartitionsParallelismRight-size partitions

Quick Start

python
from pyspark.sql import SparkSession
from pyspark.sql import functions as F

# Create optimized Spark session
spark = (SparkSession.builder
    .appName("OptimizedJob")
    .config("spark.sql.adaptive.enabled", "true")
    .config("spark.sql.adaptive.coalescePartitions.enabled", "true")
    .config("spark.sql.adaptive.skewJoin.enabled", "true")
    .config("spark.serializer", "org.apache.spark.serializer.KryoSerializer")
    .config("spark.sql.shuffle.partitions", "200")
    .getOrCreate())

# Read with optimized settings
df = (spark.read
    .format("parquet")
    .option("mergeSchema", "false")
    .load("s3://bucket/data/"))

# Efficient transformations
result = (df
    .filter(F.col("date") >= "2024-01-01")
    .select("id", "amount", "category")
    .groupBy("category")
    .agg(F.sum("amount").alias("total")))

result.write.mode("overwrite").parquet("s3://bucket/output/")

Detailed patterns and worked examples

Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.

Best Practices

Do's
  • Enable AQE - Adaptive query execution handles many issues
  • Use Parquet/Delta - Columnar formats with compression
  • Broadcast small tables - Avoid shuffle for small joins
  • Monitor Spark UI - Check for skew, spills, GC
  • Right-size partitions - 128MB - 256MB per partition
Don'ts
  • Don't collect large data - Keep data distributed
  • Don't use UDFs unnecessarily - Use built-in functions
  • Don't over-cache - Memory is limited
  • Don't ignore data skew - It dominates job time
  • Don't use .count() for existence - Use .take(1) or .isEmpty()

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/data-engineering/skills/spark-optimization of wshobson/agents.

  • SKILL.md
  • references/details.md

Open the folder on GitHubat commit 46891e7

Used in 9 other repositories

We found 19 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 9 other GitHub owners. This page covers the copy in wshobson/agents, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Apache Spark Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Apache Spark Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Apache Spark Optimization this skillwshobson/agents40k9 repos~789Automated safety check: PassMIT
Apache Spark EngineerJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Executing Sparkdata-goblin/power-bi-agentic-development1k—~1.7kAutomated safety check: PassGPL-3.0
Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples1.5k—~3.6kAutomated safety check: PassMIT-0
Monitor With HaolemeHaolemeApp/Haoleme157—~1.3kAutomated safety check: PassAGPL-3.0
Tushare Plugin BuilderYourdaylight/stock_datasource188—~2.5kAutomated safety check: PassMIT

Similar skills

  • Apache Spark Engineer

    Jeffallan/claude-skills

    Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    Data & AnalyticsAuto-check passed
  • Executing Spark

    data-goblin/power-bi-agentic-development

    Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access.

    1k GitHub stars~1.7k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Official

    Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.

    1.5k GitHub stars~3.6k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Monitor With Haoleme

    HaolemeApp/Haoleme

    Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.

    157 GitHub stars~1.3k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Tushare Plugin Builder

    Yourdaylight/stock_datasource

    Turns a Tushare API doc URL into a full data plugin for the stock_datasource repo: extractor, ClickHouse schema, query service, config and curl examples.

    188 GitHub stars~2.5k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Deploying Go SDK Bundles

    astronomer/agents

    Builds, packs, and deploys compiled Airflow Go SDK bundles so the ExecutableCoordinator can run them.

    451 GitHub stars~1.8k tokensUpdated today
    Data & AnalyticsAuto-check: notes

More from wshobson/agents

All 142 skills in this repo
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 14 repos~473 tokens
    Auto-check passed
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Portfolio Risk Metrics

    wshobson/agents

    Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.

    40k GitHub starsUsed in 13 repos~502 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed

Questions about Apache Spark Optimization

What does Apache Spark Optimization do?

Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code. The skill is a tuning reference for Apache Spark jobs. It describes the execution model from driver program to jobs triggered by actions, and a table of performance factors with a remedy for each: shuffle (minimize wide transformations), data skew (salting and broadcast joins), serialization (Kryo and columnar formats), memory (tune executor memory) and partitions (size them sensibly).

When should I use Apache Spark Optimization?

Apache Spark Optimization fits situations like: speeding up a slow Spark job; tuning executor memory and partition sizes; diagnosing data skew or heavy shuffles in the Spark UI; scaling a Spark pipeline to larger datasets.

How do I install Apache Spark Optimization in Claude Code?

Run `npx skills add wshobson/agents --skill spark-optimization -a claude-code`. Or copy the skill folder (plugins/data-engineering/skills/spark-optimization in wshobson/agents) into .claude/skills/spark-optimization in your project. Claude Code loads it when a task matches its description.

How do I install Apache Spark Optimization in Codex?

Run `npx skills add wshobson/agents --skill spark-optimization -a codex`. Or copy the skill folder (plugins/data-engineering/skills/spark-optimization in wshobson/agents) into .agents/skills/spark-optimization in your project. Codex loads it when a task matches its description.

Can I use Apache Spark Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill spark-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-optimization, .gemini/skills/spark-optimization, .github/skills/spark-optimization and .opencode/skills/spark-optimization in your project.

What does Apache Spark Optimization need to run?

SKILL.md names no scripts, command-line tools or credentials: Apache Spark Optimization is instructions for the agent only.

Does Apache Spark Optimization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Apache Spark Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Apache Spark Optimization use?

Apache Spark Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Apache Spark Optimization use?

About 789 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Apache Spark Optimization?

Skills that share tags, products or a category with Apache Spark Optimization: Apache Spark Engineer (Jeffallan/claude-skills, 12k stars), Executing Spark (data-goblin/power-bi-agentic-development, 1k stars), Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars) and Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Apache Spark Optimization?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,287 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.