Agent skill

Apache Spark Engineer

by Jeffallan in Jeffallan/claude-skills

Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

MITAuto-check passedData & Analytics

Install Apache Spark Engineer

skills CLI
$ npx skills add Jeffallan/claude-skills --skill spark-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Jeffallan/claude-skills spark-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/spark-engineer .claude/skills/spark-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spark-engineer
GitHub stars
12k
Used in
1 other repo
Token cost
~1.7k tokens
SKILL.md length
398 words
Files
6 (incl. references)
Skills in repo
58
Repo updated
First seen
Licence
MIT

At a glance

Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

  • Works in 5 steps: Analyze requirements - Understand data… → Design pipeline - Choose DataFrame vs… → Implement - Write Spark code with… → …
  • Writing Spark jobs and DataFrame transformations
  • SKILL.md covers Core Workflow, Reference Guide, Code Examples and Constraints, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill casts the agent as a senior Spark engineer and gives it a five-step workflow: analyze requirements such as data volume, latency and cluster resources, design the pipeline by choosing between DataFrames and RDDs and planning partitioning, implement with optimized transformations and caching, optimize using the Spark UI, and validate. Validation means checking the Spark UI for shuffle spill and verifying partition counts, then returning to optimization if spill or skew shows up.

Five reference files load by context: Spark SQL and DataFrames, RDD operations, partitioning and caching, performance tuning, and streaming patterns such as watermarks and stateful operations. The skill includes PySpark examples for a small pipeline, broadcast joins for small dimension tables, salting to handle data skew, and caching only DataFrames that are reused.

When your agent uses it

  • Writing Spark jobs and DataFrame transformations
  • Tuning shuffle partitions, executor memory or other cluster settings
  • Fixing data skew or slow joins in a Spark pipeline
  • Building a structured streaming job with watermarks and stateful logic

Example prompts

  • “Rewrite this PySpark job to use a broadcast join for the country lookup table.”
  • “My Spark stage spills to disk during the shuffle. Help me tune partitions and memory.”
  • “Write a structured streaming job that reads new files from a landing folder and counts events per five-minute window with a watermark.”
  • “Change the nightly Parquet processing script to write partitioned output.”

Requirements

  • Apache Spark or PySpark

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Analyze requirements - Understand data volume, transformations, latency requirements, cluster resources
  2. Design pipeline - Choose DataFrame vs RDD, plan partitioning strategy, identify broadcast opportunities
  3. Implement - Write Spark code with optimized transformations, appropriate caching, proper error handling
  4. Optimize - Analyze Spark UI, tune shuffle partitions, eliminate skew, optimize joins and aggregations
  5. Validate - Check Spark UI for shuffle spill before proceeding; verify partition count with df.rdd.getNumPartitions(); if spill or skew…

What it can do on your machine

Read from SKILL.md and the folder at commit 1be15d8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • synergetic.solutions
    • jeffallan.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Apache Spark Engineer loads about 1.7k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 109 tokens; SKILL.md has 398 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~22k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Jeffallan/claude-skills at commit 1be15d8, republished under its MIT licence (© Jeffallan). 398 words, ~1,679 tokens.

Download SKILL.mdSave it as .claude/skills/spark-engineer/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
spark-engineer
description
Use when writing Spark jobs, debugging performance issues, or configuring cluster settings for Apache Spark applications, distributed data processing pipelines, or big data workloads. Invoke to write DataFrame transformations, optimize Spark SQL queries, implement RDD pipelines, tune shuffle operations, configure executor memory, process .parquet files, handle data partitioning, or build structured streaming analytics.
license
MIT
metadata.author
https://github.com/Jeffallan
metadata.company
https://synergetic.solutions
metadata.version
1.1.0
metadata.domain
data-ml
metadata.triggers
Apache Spark, PySpark, Spark SQL, distributed computing, big data, DataFrame API, RDD, Spark Streaming, structured streaming, data partitioning, Spark…
metadata.role
expert
metadata.scope
implementation
metadata.output-format
code
metadata.related-skills
python-pro, sql-pro, devops-engineer

Spark Engineer

Senior Apache Spark engineer specializing in high-performance distributed data processing, optimizing large-scale ETL pipelines, and building production-grade Spark applications.

Core Workflow

  1. Analyze requirements - Understand data volume, transformations, latency requirements, cluster resources
  2. Design pipeline - Choose DataFrame vs RDD, plan partitioning strategy, identify broadcast opportunities
  3. Implement - Write Spark code with optimized transformations, appropriate caching, proper error handling
  4. Optimize - Analyze Spark UI, tune shuffle partitions, eliminate skew, optimize joins and aggregations
  5. Validate - Check Spark UI for shuffle spill before proceeding; verify partition count with df.rdd.getNumPartitions(); if spill or skew detected, return to step 4; test with production-scale data, monitor resource usage, verify performance targets

Reference Guide

Load detailed guidance based on context:

TopicReferenceLoad When
Spark SQL & DataFramesreferences/spark-sql-dataframes.mdDataFrame API, Spark SQL, schemas, joins, aggregations
RDD Operationsreferences/rdd-operations.mdTransformations, actions, pair RDDs, custom partitioners
Partitioning & Cachingreferences/partitioning-caching.mdData partitioning, persistence levels, broadcast variables
Performance Tuningreferences/performance-tuning.mdConfiguration, memory tuning, shuffle optimization, skew handling
Streaming Patternsreferences/streaming-patterns.mdStructured Streaming, watermarks, stateful operations, sinks

Code Examples

Quick-Start Mini-Pipeline (PySpark)
python
from pyspark.sql import SparkSession
from pyspark.sql import functions as F
from pyspark.sql.types import StructType, StructField, StringType, LongType, DoubleType

spark = SparkSession.builder \
    .appName("example-pipeline") \
    .config("spark.sql.shuffle.partitions", "400") \
    .config("spark.sql.adaptive.enabled", "true") \
    .getOrCreate()

# Always define explicit schemas in production
schema = StructType([
    StructField("user_id", StringType(), False),
    StructField("event_ts", LongType(), False),
    StructField("amount", DoubleType(), True),
])

df = spark.read.schema(schema).parquet("s3://bucket/events/")

result = df \
    .filter(F.col("amount").isNotNull()) \
    .groupBy("user_id") \
    .agg(F.sum("amount").alias("total_amount"), F.count("*").alias("event_count"))

# Verify partition count before writing
print(f"Partition count: {result.rdd.getNumPartitions()}")

result.write.mode("overwrite").parquet("s3://bucket/output/")
Broadcast Join (small dimension table < 200 MB)
python
from pyspark.sql.functions import broadcast

# Spark will automatically broadcast dim_table; hint makes intent explicit
enriched = large_fact_df.join(broadcast(dim_df), on="product_id", how="left")
Handling Data Skew with Salting
python
import pyspark.sql.functions as F

SALT_BUCKETS = 50

# Add salt to the skewed key on both sides
skewed_df = skewed_df.withColumn("salt", (F.rand() * SALT_BUCKETS).cast("int")) \
    .withColumn("salted_key", F.concat(F.col("skewed_key"), F.lit("_"), F.col("salt")))

other_df = other_df.withColumn("salt", F.explode(F.array([F.lit(i) for i in range(SALT_BUCKETS)]))) \
    .withColumn("salted_key", F.concat(F.col("skewed_key"), F.lit("_"), F.col("salt")))

result = skewed_df.join(other_df, on="salted_key", how="inner") \
    .drop("salt", "salted_key")
Correct Caching Pattern
python
# Cache ONLY when the DataFrame is reused multiple times
df_cleaned = df.filter(...).withColumn(...).cache()
df_cleaned.count()  # Materialize immediately; check Spark UI for spill

report_a = df_cleaned.groupBy("region").agg(...)
report_b = df_cleaned.groupBy("product").agg(...)

df_cleaned.unpersist()  # Release when done

Constraints

MUST DO
  • Use DataFrame API over RDD for structured data processing
  • Define explicit schemas for production pipelines
  • Partition data appropriately (200-1000 partitions per executor core)
  • Cache intermediate results only when reused multiple times
  • Use broadcast joins for small dimension tables (<200MB)
  • Handle data skew with salting or custom partitioning
  • Monitor Spark UI for shuffle, spill, and GC metrics
  • Test with production-scale data volumes
Show full SKILL.md (146 more words)Show less
MUST NOT DO
  • Use collect() on large datasets (causes OOM)
  • Skip schema definition and rely on inference in production
  • Cache every DataFrame without measuring benefit
  • Ignore shuffle partition tuning (default 200 often wrong)
  • Use UDFs when built-in functions available (10-100x slower)
  • Process small files without coalescing (small file problem)
  • Run transformations without understanding lazy evaluation
  • Ignore data skew warnings in Spark UI

Output Templates

When implementing Spark solutions, provide:

  1. Complete Spark code (PySpark or Scala) with type hints/types
  2. Configuration recommendations (executors, memory, shuffle partitions)
  3. Partitioning strategy explanation
  4. Performance analysis (expected shuffle size, memory usage)
  5. Monitoring recommendations (key Spark UI metrics to watch)

Knowledge Reference

Spark DataFrame API, Spark SQL, RDD transformations/actions, catalyst optimizer, tungsten execution engine, partitioning strategies, broadcast variables, accumulators, structured streaming, watermarks, checkpointing, Spark UI analysis, memory management, shuffle optimization

Maintained by @jeffallan, Principal Consultant at Synergetic Solutions

Documentation

© Jeffallan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/spark-engineer of Jeffallan/claude-skills.

  • SKILL.md
  • references/partitioning-caching.md
  • references/performance-tuning.md
  • references/rdd-operations.md
  • references/spark-sql-dataframes.md
  • references/streaming-patterns.md

Open the folder on GitHubat commit 1be15d8

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Jeffallan/claude-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Apache Spark Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Apache Spark Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Apache Spark Engineer this skillJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Datafusion Pythonapache/datafusion-python606—~7.8kAutomated safety check: PassApache-2.0
Apache Spark Optimizationwshobson/agents40k8 repos~789Automated safety check: PassMIT
Spark Version UpgradeOpenHands/extensions157—~1.9kAutomated safety check: PassMIT
Pyspark EtlMindrally/skills267—~2.4kAutomated safety check: PassApache-2.0
Duckdb Experttheneoai/awesome-skills183—~4.2kAutomated safety check: PassMIT

Similar skills

  • Datafusion Python

    apache/datafusion-python

    A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.

    606 GitHub stars~7.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.

    40k GitHub starsUsed in 8 repos~789 tokens
    Data & AnalyticsAuto-check passed
  • Spark Version Upgrade

    OpenHands/extensions

    Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x).

    157 GitHub stars~1.9k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Pyspark Etl

    Mindrally/skills

    Best practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg.

    267 GitHub stars~2.4k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Duckdb Expert

    theneoai/awesome-skills

    DuckDB expert for embedded OLAP analytics, Parquet/CSV querying, and high-performance analytical SQL on local data.

    183 GitHub stars~4.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Spark Expert

    theneoai/awesome-skills

    Apache Spark expert: DataFrame API, Spark SQL, Spark Structured Streaming, performance tuning, AQE, and adaptive execution.

    183 GitHub stars~3.6k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed

More from Jeffallan/claude-skills

All 58 skills in this repo
  • API Designer

    Jeffallan/claude-skills

    Designs REST and GraphQL APIs from resource modeling to an OpenAPI 3.1 contract, with versioning, pagination and RFC 7807 error handling.

    12k GitHub starsUsed in 2 repos~2k tokens
    Auto-check passed
  • CLI Developer

    Jeffallan/claude-skills

    Walks through designing, building and polishing a command-line tool: user workflow and command hierarchy, implementation in commander, click, typer or cobra, completions and cross-platform testing.

    12k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • GraphQL Architect

    Jeffallan/claude-skills

    Designs GraphQL schemas and Apollo Federation graphs, with DataLoader resolvers, subscriptions, query complexity limits and caching.

    12k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Kubernetes Specialist

    Jeffallan/claude-skills

    Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Laravel Specialist

    Jeffallan/claude-skills

    Builds Laravel 10+ applications with Eloquent models, Sanctum authentication, Horizon queues, API resources and Livewire components, tested with Pest or PHPUnit.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed

Works with

Questions about Apache Spark Engineer

What does Apache Spark Engineer do?

Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming. The skill casts the agent as a senior Spark engineer and gives it a five-step workflow: analyze requirements such as data volume, latency and cluster resources, design the pipeline by choosing between DataFrames and RDDs and planning partitioning, implement with optimized transformations and caching, optimize using the Spark UI, and validate. Validation means checking the Spark UI for shuffle spill and verifying partition counts, then returning to optimization if spill or skew shows up.

When should I use Apache Spark Engineer?

Apache Spark Engineer fits situations like: writing Spark jobs and DataFrame transformations; tuning shuffle partitions, executor memory or other cluster settings; fixing data skew or slow joins in a Spark pipeline; building a structured streaming job with watermarks and stateful logic.

How do I install Apache Spark Engineer in Claude Code?

Run `npx skills add Jeffallan/claude-skills --skill spark-engineer -a claude-code`. Or copy the skill folder (skills/spark-engineer in Jeffallan/claude-skills) into .claude/skills/spark-engineer in your project. Claude Code loads it when a task matches its description.

How do I install Apache Spark Engineer in Codex?

Run `npx skills add Jeffallan/claude-skills --skill spark-engineer -a codex`. Or copy the skill folder (skills/spark-engineer in Jeffallan/claude-skills) into .agents/skills/spark-engineer in your project. Codex loads it when a task matches its description.

Can I use Apache Spark Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jeffallan/claude-skills --skill spark-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-engineer, .gemini/skills/spark-engineer, .github/skills/spark-engineer and .opencode/skills/spark-engineer in your project.

What does Apache Spark Engineer need to run?

SKILL.md names no scripts, command-line tools or credentials: Apache Spark Engineer is instructions for the agent only. Our summary lists: Apache Spark or PySpark.

Does Apache Spark Engineer access the network?

SKILL.md names 3 domains. As links in the text: github.com, synergetic.solutions and jeffallan.github.io. This is read from the text; nothing was executed.

Is Apache Spark Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Apache Spark Engineer use?

Apache Spark Engineer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Apache Spark Engineer use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.

What are the alternatives to Apache Spark Engineer?

Skills that share tags, products or a category with Apache Spark Engineer: Datafusion Python (apache/datafusion-python, 606 stars), Apache Spark Optimization (wshobson/agents, 40k stars), Spark Version Upgrade (OpenHands/extensions, 157 stars) and Pyspark Etl (Mindrally/skills, 267 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Apache Spark Engineer?

Jeffallan (a GitHub user) maintains it in Jeffallan/claude-skills, which has 11,754 GitHub stars. The repository holds 58 skills in this directory. The repository was last updated on October 3, 2026.

Source: Jeffallan/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.