Platform

Apache Spark agent skills for Claude Code, Codex and other agents.

Unified analytics engine for large-scale data processing in batch and streaming.
skills
23
official
4
Type
Platform
Website
spark.apache.org
Official GitHub
apache
Reviews
See Apache Spark on Enlisted

Apache Spark skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Official (4 skills)

Official Apache Spark skills
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.

aws-samples/aws-glue-samples1.5k—~3.6kAutomated safety check: PassMIT-01 mo ago
2

Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine).

google/skills21k—~4.4kAutomated safety check: PassApache-2.0today
3

Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems.

databricks/databricks-agent-skills345—~2.1kAutomated safety check: PassUnknownyesterday
4

Use Databricks built-in AI Functions (aiclassify, aiextract, aisummarize, aimask, aitranslate, aifixgrammar, aigen, aianalyzesentiment, aisimilarity, aiparsedocument, aiprepsearch, aiquery…

databricks/databricks-agent-skills345—~3.9kAutomated safety check: PassUnknownyesterday

Community

Community Apache Spark skills
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
5

Create a new PySpark data source implementation. An agent skill from allisonwang-db/pyspark-data-sources.

allisonwang-db/pyspark-data-sources106—~635Automated safety check: PassApache-2.01 mo ago
6

A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.

apache/datafusion-python606—~7.8kAutomated safety check: PassApache-2.0today
7

Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

Jeffallan/claude-skills12k1 repo~1.7kAutomated safety check: PassMIT3 days ago
8

Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.

wshobson/agents40k8 repos~789Automated safety check: PassMIT2 days ago
9

Connect Databricks Apps to shared clusters or serverless compute using Databricks Connect.

databricks-solutions/databricks-apps-cookbook183—~790Automated safety check: PassUnknown2 days ago
10

Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping.

apache/datafusion-python606—~5.8kAutomated safety check: PassApache-2.0today
11

Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.

davila7/claude-code-templates32k7 repos~2.8kAutomated safety check: PassMITtoday
12

Analyze DBSQL queries, including SQL embedded in notebooks (spark.sql(...), %sql cells), for anti-patterns, lint issues, and performance problems, using Databricks-specific dialect and platform…

AltimateAI/data-engineering-skills127—~6.7kAutomated safety check: PassMIT5 days ago
13

Data engineering patterns for ETL pipelines, data warehousing, Apache Spark, and data quality validation

rohitg00/awesome-claude-code-toolkit2.7k—~1.7kAutomated safety check: PassApache-2.04 mo ago
14

Data pipeline expert for ETL, Apache Spark, Airflow, dbt, and data quality

RightNow-AI/openfang18k—~847Automated safety check: PassApache-2.03 mo ago
15

Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access.

data-goblin/power-bi-agentic-development1k—~1.7kAutomated safety check: PassGPL-3.02 days ago
16

A skill your agent uses when reading from or writing to Neo4j with Apache Spark or Databricks using the Neo4j Connector for Apache Spark 6.0 (org.neo4j.connectors:spark) or 5.x…

neo4j-contrib/neo4j-skills114—~4.1kAutomated safety check: NotesMITtoday
17

Expert guidance for working with the Apache Spark Catalyst query optimisation framework.

aehrc/pathling137—~5.4kAutomated safety check: PassApache-2.0today
18

byted-emr-skills提供管理火山引擎EMR(火山引擎 E-MapReduce(简称“EMR”)是开源Hadoop生态的企业级大数据分析系统,完全兼容开源)的技能,包括管理EMR on ECS集群、EMR on VKE集群、EMR serverless队列、计算组、作业模板/实例、日志、监控并提供 EMR Agent 智能诊断与知识问答能力。当用户提及“EMR on…

LeoYeAI/openclaw-master-skills2.2k—~2.7kAutomated safety check: PassMIT2 mo ago
19

Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x).

OpenHands/extensions157—~1.9kAutomated safety check: PassMITtoday
20

Assists with benchmarking and profiling the performance of an Apache Spark UDF on the GPU.

Kilo-Org/kilo-marketplace189—~802Automated safety check: PassProprietary8 days ago
21

Best practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg.

Mindrally/skills267—~2.4kAutomated safety check: PassApache-2.01 mo ago
22

Apache Spark expert: DataFrame API, Spark SQL, Spark Structured Streaming, performance tuning, AQE, and adaptive execution.

theneoai/awesome-skills183—~3.6kAutomated safety check: PassMIT4 mo ago
23

Apache Spark expertise covering RDD vs DataFrame vs Dataset APIs, partitioning strategies, shuffle optimization, broadcast joins, caching, Spark SQL, structured streaming, UDFs, cluster sizing…

FerroxLabs/wayland608—~4.2kAutomated safety check: PassApache-2.0yesterday

Questions, answered from the data.

What is the best Apache Spark skill?

Migrate Glue Devendpoint To Interactive Sessions (official) from aws-samples/aws-glue-samples ranks first of the 23 Apache Spark skills listed here, with the highest score: its repository has 1.5k GitHub stars, its SKILL.md loads about 3.6k tokens and it passes the automated safety check with no findings. Next come Google Cloud Solution Agentic Analytics Spark Knowledge Catalog and Spark Python Data Source.

Is there an official Apache Spark skill?

4 of the 23 Apache Spark skills are official, published by the vendor's own GitHub organization: Migrate Glue Devendpoint To Interactive Sessions, Google Cloud Solution Agentic Analytics Spark Knowledge Catalog, Spark Python Data Source and Databricks AI Functions.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.