Search

Data & Analytics · Apache Spark

20 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.

aws-samples/aws-glue-samples1.5k—~3.6kAutomated safety check: PassMIT-01 mo ago
2

Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

Jeffallan/claude-skills12k1 repo~1.7kAutomated safety check: PassMIT6 days ago
3

Create a new PySpark data source implementation. An agent skill from allisonwang-db/pyspark-data-sources.

allisonwang-db/pyspark-data-sources106—~635Automated safety check: PassApache-2.01 mo ago
4

A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.

apache/datafusion-python607—~7.8kAutomated safety check: PassApache-2.0yesterday
5

Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.

wshobson/agents40k9 repos~789Automated safety check: PassMIT5 days ago
6

Audit and improve datafusion-python functions to accept native Python types (int, float, str, bool) instead of requiring explicit lit() or col() wrapping.

apache/datafusion-python607—~5.8kAutomated safety check: PassApache-2.0yesterday
7

Build scalable data pipelines, modern data warehouses, and real-time streaming architectures.

davila7/claude-code-templates33k8 repos~2.8kAutomated safety check: PassMITtoday
8

Analyze DBSQL queries, including SQL embedded in notebooks (spark.sql(...), %sql cells), for anti-patterns, lint issues, and performance problems, using Databricks-specific dialect and platform…

AltimateAI/data-engineering-skills128—~6.7kAutomated safety check: PassMIT2 days ago
9

Data pipeline expert for ETL, Apache Spark, Airflow, dbt, and data quality

RightNow-AI/openfang18k—~847Automated safety check: PassApache-2.03 mo ago
10

Data engineering patterns for ETL pipelines, data warehousing, Apache Spark, and data quality validation

rohitg00/awesome-claude-code-toolkit2.7k—~1.7kAutomated safety check: PassApache-2.05 mo ago
11

Discovers requirements and designs an end-to-end governed agentic analytics solution using Knowledge Catalog and Managed Service for Apache Spark (Lightning Engine).

google/skills21k—~4.4kAutomated safety check: PassApache-2.0today
12

Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access.

data-goblin/power-bi-agentic-development1k—~1.7kAutomated safety check: PassGPL-3.03 days ago
13

A skill your agent uses when reading from or writing to Neo4j with Apache Spark or Databricks using the Neo4j Connector for Apache Spark 6.0 (org.neo4j.connectors:spark) or 5.x…

neo4j-contrib/neo4j-skills114—~4.1kAutomated safety check: NotesMITtoday
14

Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems.

databricks/databricks-agent-skills345—~2.1kAutomated safety check: PassUnknowntoday
15

Expert guidance for working with the Apache Spark Catalyst query optimisation framework.

aehrc/pathling137—~5.4kAutomated safety check: PassApache-2.02 days ago
16

Transform pyspark transformer operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

jeremylongshore/tons-of-skills-marketplace2.8k—~566Automated safety check: PassMITtoday
17

Optimize spark sql optimizer operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

jeremylongshore/tons-of-skills-marketplace2.8k—~564Automated safety check: PassMITtoday
18

Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x).

OpenHands/extensions163—~1.9kAutomated safety check: PassMITtoday
19

Assists with benchmarking and profiling the performance of an Apache Spark UDF on the GPU.

Kilo-Org/kilo-marketplace190—~802Automated safety check: PassProprietary11 days ago
20

Best practices for building performant, testable PySpark ETL pipelines with Spark SQL and Apache Iceberg.

Mindrally/skills271—~2.5kAutomated safety check: PassApache-2.0yesterday