Search
Data & Analytics · Apache Spark
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist. | aws-samples/ | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | 1 mo ago |
| 2 | Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming. | Jeffallan/ | 12k | 1 repo | ~1.7k | Automated safety check: Pass | MIT | 6 days ago |
| 3 | Create a new PySpark data source implementation. An agent skill from allisonwang-db/pyspark-data-sources. | allisonwang-db/ | 106 | — | ~635 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 4 | A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code. | apache/ | 607 | — | ~7.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 5 | Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code. | wshobson/ | 40k | 9 repos | ~789 | Automated safety check: Pass | MIT | 5 days ago |
| 6 | Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping. | apache/ | 607 | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 7 | Build scalable data pipelines, modern data warehouses, and real-time streaming architectures. | davila7/ | 33k | 8 repos | ~2.8k | Automated safety check: Pass | MIT | today |
| 8 | Analyze DBSQL queries, including SQL embedded in notebooks (spark.sql(...), %sql cells), for anti-patterns, lint issues, and performance problems, using Databricks-specific dialect and platform… | AltimateAI/ | 128 | — | ~6.7k | Automated safety check: Pass | MIT | 2 days ago |
| 9 | Data pipeline expert for ETL, Apache Spark, Airflow, dbt, and data quality | RightNow-AI/ | 18k | — | ~847 | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 10 | Data engineering patterns for ETL pipelines, data warehousing, Apache Spark, and data quality validation | rohitg00/ | 2.7k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 5 mo ago |
| 11 | Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine). | google/ | 21k | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | today |
| 12 | Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access. | data-goblin/ | 1k | — | ~1.7k | Automated safety check: Pass | GPL-3.0 | 3 days ago |
| 13 | A skill your agent uses when reading from or writing to Neo4j with Apache Spark or Databricks using the Neo4j Connector for Apache Spark 6.0 (org.neo4j.connectors:spark) or 5.x… | neo4j-contrib/ | 114 | — | ~4.1k | Automated safety check: Notes | MIT | today |
| 14 | Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems. | databricks/ | 345 | — | ~2.1k | Automated safety check: Pass | Unknown | today |
| 15 | Expert guidance for working with the Apache Spark Catalyst query optimisation framework. | aehrc/ | 137 | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 16 | Transform pyspark transformer operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~566 | Automated safety check: Pass | MIT | today |
| 17 | Optimize spark sql optimizer operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~564 | Automated safety check: Pass | MIT | today |
| 18 | Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x). | OpenHands/ | 163 | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 19 | Assists with benchmarking and profiling the performance of an Apache Spark UDF on the GPU. | Kilo-Org/ | 190 | — | ~802 | Automated safety check: Pass | Proprietary | 11 days ago |
| 20 | 20.Pyspark Etl Best practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg. | Mindrally/ | 271 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | yesterday |