Agent skill

Spark Version Upgrade

by OpenHands in OpenHands/extensions

Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x).

MITAuto-check passedData & Analytics

Install Spark Version Upgrade

skills CLI
$ npx skills add OpenHands/extensions --skill spark-version-upgrade -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenHands/extensions spark-version-upgrade --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenHands/extensions.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/spark-version-upgrade .claude/skills/spark-version-upgrade && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spark-version-upgrade
GitHub stars
157
Token cost
~1.9k tokens
SKILL.md length
655 words
Files
5
Skills in repo
78
Repo updated
First seen
Licence
MIT

At a glance

Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x).

  • Works in 6 steps: Inventory & Impact Analysis → Build File Updates → API Migration → …
  • Tasks that involve Code migrations
  • SKILL.md covers When to Use, Workflow Overview, Phase 1: Inventory & Impact… and Phase 2: Build File Updates, plus 5 more sections
  • Calls mvn

What it does

Spark Version Upgrade is an agent skill from OpenHands/extensions. Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x). Covers build files, deprecated APIs, configuration changes, SQL/DataFrame updates, and test validation.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `.plugin/plugin.json` and `README.md`). Compatibility notes: Requires Java 8+/11+/17+, Scala 2.12/2.13, Maven/Gradle/SBT, Apache Spark

It sits in Data & Analytics, covering Code migrations, DataFrames and SQL. It works with Apache Spark, SQL, Gradle and Java. The repository describes itself as: Public registry for OpenHands extensions. The licence is MIT.

When your agent uses it

  • Tasks that involve Code migrations
  • Tasks that involve DataFrames
  • Tasks that involve SQL

Example prompts

  • “/spark-version-upgrade”

Requirements

  • Compatibility (from SKILL.md): Requires Java 8+/11+/17+, Scala 2.12/2.13, Maven/Gradle/SBT, Apache Spark

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Inventory & Impact Analysis
  2. Build File Updates
  3. API Migration
  4. Configuration Migration
  5. SQL & DataFrame Migration
  6. Test Validation

What it can do on your machine

Read from SKILL.md and the folder at commit d008b81. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • mvn

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • spark.apache.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Java 8+/11+/17+, Scala 2.12/2.13, Maven/Gradle/SBT, Apache Spark

    From compatibility in the SKILL.md frontmatter.

Context cost

Spark Version Upgrade loads about 1.9k tokens when it runs. Until then it costs about 51 tokens; SKILL.md has 655 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OpenHands/extensions at commit d008b81, republished under its MIT licence (© OpenHands). 655 words, ~1,931 tokens.

Download SKILL.mdSave it as .claude/skills/spark-version-upgrade/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
spark-version-upgrade
description
Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x). Covers build files, deprecated APIs, configuration changes, SQL/DataFrame updates, and test validation.
compatibility
Requires Java 8+/11+/17+, Scala 2.12/2.13, Maven/Gradle/SBT, Apache Spark
license
MIT
triggers
spark upgrade, spark migration, spark version, upgrade spark, spark 3, spark 4, pyspark upgrade

Upgrade Apache Spark applications between major versions with a structured, phase-by-phase workflow.

When to Use

  • Migrating from Spark 2.x → 3.x or Spark 3.x → 4.x
  • Updating PySpark, Spark SQL, or Structured Streaming applications
  • Resolving deprecation warnings before a Spark version bump

Workflow Overview

  1. Inventory & Impact Analysis — Scan the codebase and assess scope
  2. Build File Updates — Bump Spark/Scala/Java dependencies
  3. API Migration — Replace deprecated and removed APIs
  4. Configuration Migration — Update Spark config properties
  5. SQL & DataFrame Migration — Fix query-level breaking changes
  6. Test Validation — Compile, run tests, verify results

Phase 1: Inventory & Impact Analysis

Before changing any code, assess what needs to change. Read the official Apache Spark migration guide for the target version — it documents every API removal, config rename, and behavioral change per release: https://spark.apache.org/docs/latest/migration-guide.html

Checklist
  • Read the migration guide section for the target Spark version
  • Identify current Spark version (check pom.xml, build.sbt, build.gradle, or requirements.txt)
  • Identify target Spark version
  • Search for deprecated APIs: grep -rn 'import org.apache.spark' --include='*.scala' --include='*.java' --include='*.py'
  • List all Spark config properties: grep -rn 'spark\.' --include='*.conf' --include='*.properties' --include='*.scala' --include='*.java' --include='*.py' | grep -v 'test'
  • On Windows PowerShell, use Get-ChildItem -Recurse -Include *.scala,*.java,*.py | Select-String 'import org.apache.spark' and adjust the extensions/pattern for config searches.
  • Check for custom SparkSession or SparkContext extensions
  • Identify connector dependencies (Hive, Kafka, Cassandra, Delta, Iceberg)
  • Document findings in spark_upgrade_impact.md
Output
spark_upgrade_impact.md   # Summary of affected files, APIs, and configs

Phase 2: Build File Updates

Update dependency versions and resolve compilation.

Maven (pom.xml)
xml
<!-- Update Spark version property -->
<spark.version>3.5.1</spark.version>    <!-- or 4.0.0 -->
<scala.version>2.13.12</scala.version>  <!-- Spark 3.x: 2.12/2.13; Spark 4.x: 2.13 -->

<!-- Update artifact IDs if Scala cross-version changed -->
<artifactId>spark-core_2.13</artifactId>
<artifactId>spark-sql_2.13</artifactId>
SBT (build.sbt)
scala
val sparkVersion = "3.5.1" // or "4.0.0"
scalaVersion := "2.13.12"

libraryDependencies += "org.apache.spark" %% "spark-core" % sparkVersion
libraryDependencies += "org.apache.spark" %% "spark-sql" % sparkVersion
Gradle (build.gradle)
groovy
ext {
    sparkVersion = '3.5.1' // or '4.0.0'
}
dependencies {
    implementation "org.apache.spark:spark-core_2.13:${sparkVersion}"
    implementation "org.apache.spark:spark-sql_2.13:${sparkVersion}"
}
PySpark (requirements.txt / pyproject.toml)
pyspark==3.5.1   # or 4.0.0
Checklist
  • Update Spark version in build file
  • Update Scala version if crossing 2.12→2.13 boundary
  • Update Java source/target level if required (Spark 4.x requires Java 17+)
  • Update connector library versions to match new Spark version
  • Resolve dependency conflicts (mvn dependency:tree / sbt dependencyTree)
  • Confirm project compiles (errors at this stage are expected — they guide Phase 3)

Phase 3: API Migration

Replace removed and deprecated APIs. Work through compiler errors systematically.

Common Patterns

Consult the official Apache Spark migration guide for the complete list of changes for each version: https://spark.apache.org/docs/latest/migration-guide.html

SparkSession Creation (2.x → 3.x)
scala
// BEFORE (Spark 1.x/2.x)
val sc = new SparkContext(conf)
val sqlContext = new SQLContext(sc)

// AFTER (Spark 2.x+/3.x)
val spark = SparkSession.builder()
  .config(conf)
  .enableHiveSupport() // if needed
  .getOrCreate()
val sc = spark.sparkContext
RDD to DataFrame (2.x → 3.x)
scala
// BEFORE
rdd.toDF()  // implicit from SQLContext

// AFTER
import spark.implicits._
rdd.toDF()  // implicit from SparkSession
Accumulator API (2.x → 3.x)
scala
// BEFORE
val acc = sc.accumulator(0)

// AFTER
val acc = sc.longAccumulator("name")
Checklist
  • Replace SQLContext / HiveContext with SparkSession
  • Replace deprecated Accumulator with AccumulatorV2
  • Update DataFrame → Dataset[Row] where needed
  • Replace removed RDD.mapPartitionsWithContext with mapPartitions
  • Fix SparkConf deprecated setters
  • Update custom UserDefinedFunction registration
  • Migrate Experimental / DeveloperApi usages that were removed
  • Verify all compilation errors from Phase 2 are resolved

Show full SKILL.md (272 more words)Show less

Phase 4: Configuration Migration

Spark renames and removes configuration properties between versions. The official migration guide documents every renamed and removed property per release: https://spark.apache.org/docs/latest/migration-guide.html

Checklist
  • Rename deprecated config keys (e.g., spark.shuffle.file.buffer.kb → spark.shuffle.file.buffer)
  • Update removed configs to their replacements
  • Review spark-defaults.conf, application code, and submit scripts
  • Check for hardcoded config values in test fixtures
  • Verify SparkSession.builder().config(...) calls use current property names

Phase 5: SQL & DataFrame Migration

Spark SQL behavior changes between versions can silently alter query results.

Key Breaking Changes (2.x → 3.x)
  • CAST to integer no longer truncates silently — set spark.sql.ansi.enabled if needed
  • FROM clause is required in SELECT (no more SELECT 1)
  • Column resolution order changed in subqueries
  • spark.sql.legacy.timeParserPolicy controls date/time parsing behavior
Key Breaking Changes (3.x → 4.x)
  • ANSI mode is default (spark.sql.ansi.enabled=true)
  • Stricter type coercion in comparisons
  • spark.sql.legacy.* flags removed
Checklist
  • Audit SQL strings and DataFrame expressions for changed behavior
  • Add explicit CAST where implicit coercion relied on legacy behavior
  • Update date/time format patterns to match new parser
  • Test SQL queries with representative data and compare output to pre-upgrade baseline
  • Set spark.sql.legacy.* flags temporarily if needed for phased migration

Phase 6: Test Validation

Checklist
  • All code compiles without errors
  • All existing unit tests pass
  • All existing integration tests pass
  • Run Spark jobs locally with sample data and compare output to pre-upgrade baseline
  • No deprecation warnings remain (or are documented with a migration timeline)
  • Update CI/CD pipeline to use new Spark version
  • Document any spark.sql.legacy.* flags that are set temporarily

Done When

✓ Project compiles against target Spark version ✓ All tests pass ✓ No removed APIs remain in code ✓ Configuration properties are current ✓ SQL queries produce correct results ✓ Upgrade impact documented in spark_upgrade_impact.md

© OpenHands, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/spark-version-upgrade of OpenHands/extensions.

  • SKILL.md
  • .claude-plugin
  • .codex-plugin
  • .plugin/plugin.json
  • README.md

Open the folder on GitHubat commit d008b81

Compare with similar skills

Spark Version Upgrade next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Spark Version Upgrade compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Spark Version Upgrade this skillOpenHands/extensions157—~1.9kAutomated safety check: PassMIT
Datafusion Pythonapache/datafusion-python606—~7.8kAutomated safety check: PassApache-2.0
Centia Snapshot Catalogmapcentia/geocloud2152—~2.3kAutomated safety check: PassAGPL-3.0
Querying Big Datasetsflyrank-bih/flyrank-ml-internship-starter140—~750Automated safety check: PassCustom licence
Apache Spark EngineerJeffallan/claude-skills12k1 repos~1.7kAutomated safety check: PassMIT
Neo4j Spark Skillneo4j-contrib/neo4j-skills114—~4.1kAutomated safety check: NotesMIT

Similar skills

  • Datafusion Python

    apache/datafusion-python

    A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.

    606 GitHub stars~7.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Centia Snapshot Catalog

    mapcentia/geocloud2

    Analyse GC2/Centia Parquet snapshots with DuckDB by walking the STAC catalog.json in the snapshot store — find datasets, decide whether a dataset has geometry (and in which CRS), read one snapshot…

    152 GitHub stars~2.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Querying Big Datasets

    flyrank-bih/flyrank-ml-internship-starter

    Works with datasets far too big to download or load in pandas — SQL over remote Parquet with DuckDB, aggregate-then-model, iterate on samples.

    140 GitHub stars~750 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Apache Spark Engineer

    Jeffallan/claude-skills

    Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    Data & AnalyticsAuto-check passed
  • Neo4j Spark Skill

    neo4j-contrib/neo4j-skills

    A skill your agent uses when reading from or writing to Neo4j with Apache Spark or Databricks using the Neo4j Connector for Apache Spark 6.0 (org.neo4j.connectors:spark) or 5.x…

    114 GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check: notes
  • Openfdd Cookbook Parity

    bbartling/open-fdd

    A skill your agent uses when editing rule cookbooks, parity matrix, or cookbook CI (cookbook-parity.yml, cookbookparitycheck.py).

    172 GitHub stars~382 tokensUpdated today
    Data & AnalyticsAuto-check passed

More from OpenHands/extensions

All 78 skills in this repo
  • Agent Readiness Report

    OpenHands/extensions

    Evaluate how well a codebase supports autonomous AI-assisted development.

    157 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Discord

    OpenHands/extensions

    Build and automate Discord integrations (bots, webhooks, slash commands, and REST API workflows).

    157 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • GitHub

    OpenHands/extensions

    Interact with GitHub repositories, pull requests, issues, and workflows using the GITHUBTOKEN environment variable and GitHub CLI.

    157 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • GitHub Issue To PR

    OpenHands/extensions

    Create an automation that implements GitHub issues when a configurable trigger label is applied.

    157 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • GitHub Repo Monitor

    OpenHands/extensions

    This skill should be used when the user asks to "monitor a GitHub repository", "watch GitHub for issues or PRs", "respond to @OpenHands mentions on GitHub", "set up an OpenHands GitHub integration"…

    157 GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • GitLab Issue To Mr

    OpenHands/extensions

    Create an automation that implements GitLab issues when a configurable trigger label is applied.

    157 GitHub stars~4.9k tokensUpdated today
    Auto-check passed

Questions about Spark Version Upgrade

What does Spark Version Upgrade do?

Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x). Spark Version Upgrade is an agent skill from OpenHands/extensions.x).

When should I use Spark Version Upgrade?

Spark Version Upgrade fits situations like: tasks that involve Code migrations; tasks that involve DataFrames; tasks that involve SQL.

How do I install Spark Version Upgrade in Claude Code?

Run `npx skills add OpenHands/extensions --skill spark-version-upgrade -a claude-code`. Or copy the skill folder (skills/spark-version-upgrade in OpenHands/extensions) into .claude/skills/spark-version-upgrade in your project. Claude Code loads it when a task matches its description.

How do I install Spark Version Upgrade in Codex?

Run `npx skills add OpenHands/extensions --skill spark-version-upgrade -a codex`. Or copy the skill folder (skills/spark-version-upgrade in OpenHands/extensions) into .agents/skills/spark-version-upgrade in your project. Codex loads it when a task matches its description.

Can I use Spark Version Upgrade in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenHands/extensions --skill spark-version-upgrade -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-version-upgrade, .gemini/skills/spark-version-upgrade, .github/skills/spark-version-upgrade and .opencode/skills/spark-version-upgrade in your project.

What does Spark Version Upgrade need to run?

Going by SKILL.md and its folder, Spark Version Upgrade needs the command-line tools its instructions call (mvn). Compatibility (from SKILL.md): Requires Java 8+/11+/17+, Scala 2.12/2.13, Maven/Gradle/SBT, Apache Spark.

Does Spark Version Upgrade access the network?

SKILL.md names 1 domain. As links in the text: spark.apache.org. This is read from the text; nothing was executed.

Is Spark Version Upgrade safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Spark Version Upgrade use?

Spark Version Upgrade is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Spark Version Upgrade use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Spark Version Upgrade?

Skills that share tags, products or a category with Spark Version Upgrade: Datafusion Python (apache/datafusion-python, 606 stars), Centia Snapshot Catalog (mapcentia/geocloud2, 152 stars), Querying Big Datasets (flyrank-bih/flyrank-ml-internship-starter, 140 stars) and Apache Spark Engineer (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Spark Version Upgrade?

OpenHands (a GitHub organization) maintains it in OpenHands/extensions, which has 157 GitHub stars. The repository holds 78 skills in this directory. The repository was last updated on October 6, 2026.

Source: OpenHands/extensions on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.