Platform
Apache Spark agent skills for Claude Code, Codex and other agents.
- skills
- 23
- official
- 4
- Type
- Platform
- Website
- spark.apache.org
- Official GitHub
- apache
- Reviews
- See Apache Spark on Enlisted
Apache Spark skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
Official (4 skills)
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist. | aws-samples/ | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | 1 mo ago |
| 2 | Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine). | google/ | 21k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems. | databricks/ | 345 | — | ~2.1k | Automated safety check: Pass | Unknown | yesterday |
| 4 | Use Databricks built-in AI Functions (aiclassify, aiextract, aisummarize, aimask, aitranslate, aifixgrammar, aigen, aianalyzesentiment, aisimilarity, aiparsedocument, aiprepsearch, aiquery… | databricks/ | 345 | — | ~3.9k | Automated safety check: Pass | Unknown | yesterday |
Community
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 5 | Create a new PySpark data source implementation. An agent skill from allisonwang-db/pyspark-data-sources. | allisonwang-db/ | 106 | — | ~635 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 6 | A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code. | apache/ | 606 | — | ~7.8k | Automated safety check: Pass | Apache-2.0 | today |
| 7 | Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming. | Jeffallan/ | 12k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | 3 days ago |
| 8 | Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code. | wshobson/ | 40k | 8 repos | ~789 | Automated safety check: Pass | MIT | 2 days ago |
| 9 | Connect Databricks Apps to shared clusters or serverless compute using Databricks Connect. | databricks-solutions/ | 183 | — | ~790 | Automated safety check: Pass | Unknown | 2 days ago |
| 10 | Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping. | apache/ | 606 | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Build scalable data pipelines, modern data warehouses, and real-time streaming architectures. | davila7/ | 32k | 7 repos | ~2.8k | Automated safety check: Pass | MIT | today |
| 12 | Analyze DBSQL queries, including SQL embedded in notebooks (spark.sql(...), %sql cells), for anti-patterns, lint issues, and performance problems, using Databricks-specific dialect and platform… | AltimateAI/ | 127 | — | ~6.7k | Automated safety check: Pass | MIT | 5 days ago |
| 13 | Data engineering patterns for ETL pipelines, data warehousing, Apache Spark, and data quality validation | rohitg00/ | 2.7k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 4 mo ago |
| 14 | Data pipeline expert for ETL, Apache Spark, Airflow, dbt, and data quality | RightNow-AI/ | 18k | — | ~847 | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 15 | Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access. | data-goblin/ | 1k | — | ~1.7k | Automated safety check: Pass | GPL-3.0 | 2 days ago |
| 16 | A skill your agent uses when reading from or writing to Neo4j with Apache Spark or Databricks using the Neo4j Connector for Apache Spark 6.0 (org.neo4j.connectors:spark) or 5.x… | neo4j-contrib/ | 114 | — | ~4.1k | Automated safety check: Notes | MIT | today |
| 17 | Expert guidance for working with the Apache Spark Catalyst query optimisation framework. | aehrc/ | 137 | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | today |
| 18 | byted-emr-skills提供管理火山引擎EMR(火山引擎 E-MapReduce(简称“EMR”)是开源Hadoop生态的企业级大数据分析系统,完全兼容开源)的技能,包括管理EMR on ECS集群、EMR on VKE集群、EMR serverless队列、计算组、作业模板/实例、日志、监控并提供 EMR Agent 智能诊断与知识问答能力。当用户提及“EMR on… | LeoYeAI/ | 2.2k | — | ~2.7k | Automated safety check: Pass | MIT | 2 mo ago |
| 19 | Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x). | OpenHands/ | 157 | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 20 | Assists with benchmarking and profiling the performance of an Apache Spark UDF on the GPU. | Kilo-Org/ | 189 | — | ~802 | Automated safety check: Pass | Proprietary | 8 days ago |
| 21 | 21.Pyspark Etl Best practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg. | Mindrally/ | 267 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 22 | 22.Spark Expert Apache Spark expert: DataFrame API, Spark SQL, Spark Structured Streaming, performance tuning, AQE, and adaptive execution. | theneoai/ | 183 | — | ~3.6k | Automated safety check: Pass | MIT | 4 mo ago |
| 23 | Apache Spark expertise covering RDD vs DataFrame vs Dataset APIs, partitioning strategies, shuffle optimization, broadcast joins, caching, Spark SQL, structured streaming, UDFs, cluster sizing… | FerroxLabs/ | 608 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
Questions, answered from the data.
What is the best Apache Spark skill?
Migrate Glue Devendpoint To Interactive Sessions (official) from aws-samples/aws-glue-samples ranks first of the 23 Apache Spark skills listed here, with the highest score: its repository has 1.5k GitHub stars, its SKILL.md loads about 3.6k tokens and it passes the automated safety check with no findings. Next come Google Cloud Solution Agentic Analytics Spark Knowledge Catalog and Spark Python Data Source.
Is there an official Apache Spark skill?
4 of the 23 Apache Spark skills are official, published by the vendor's own GitHub organization: Migrate Glue Devendpoint To Interactive Sessions, Google Cloud Solution Agentic Analytics Spark Knowledge Catalog, Spark Python Data Source and Databricks AI Functions.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.