Apache Spark Optimization
wshobson/agents
Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.
Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access.
$ npx skills add data-goblin/power-bi-agentic-development --skill executing-spark -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install data-goblin/power-bi-agentic-development executing-spark --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/data-goblin/power-bi-agentic-development.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/etl/skills/executing-spark .claude/skills/executing-spark && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "executing-spark" agent skill from https://github.com/data-goblin/power-bi-agentic-development/tree/main/plugins/etl/skills/executing-spark into .claude/skills/executing-spark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "executing-spark", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/data-goblin/power-bi-agentic-development/tree/main/plugins/etl/skills/executing-sparkType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add data-goblin/power-bi-agentic-development --skill executing-spark -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install data-goblin/power-bi-agentic-development executing-spark --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/data-goblin/power-bi-agentic-development.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/etl/skills/executing-spark .agents/skills/executing-spark && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "executing-spark" agent skill from https://github.com/data-goblin/power-bi-agentic-development/tree/main/plugins/etl/skills/executing-spark into .agents/skills/executing-spark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "executing-spark", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add data-goblin/power-bi-agentic-development --skill executing-spark -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install data-goblin/power-bi-agentic-development executing-spark --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/data-goblin/power-bi-agentic-development.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/etl/skills/executing-spark .cursor/skills/executing-spark && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "executing-spark" agent skill from https://github.com/data-goblin/power-bi-agentic-development/tree/main/plugins/etl/skills/executing-spark into .cursor/skills/executing-spark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "executing-spark", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/data-goblin/power-bi-agentic-development.git --path plugins/etl/skills/executing-spark--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add data-goblin/power-bi-agentic-development --skill executing-spark -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install data-goblin/power-bi-agentic-development executing-spark --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/data-goblin/power-bi-agentic-development.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/etl/skills/executing-spark .gemini/skills/executing-spark && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "executing-spark" agent skill from https://github.com/data-goblin/power-bi-agentic-development/tree/main/plugins/etl/skills/executing-spark into .gemini/skills/executing-spark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "executing-spark", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install data-goblin/power-bi-agentic-development executing-sparkInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add data-goblin/power-bi-agentic-development --skill executing-spark -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/data-goblin/power-bi-agentic-development.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/etl/skills/executing-spark .github/skills/executing-spark && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "executing-spark" agent skill from https://github.com/data-goblin/power-bi-agentic-development/tree/main/plugins/etl/skills/executing-spark into .github/skills/executing-spark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "executing-spark", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add data-goblin/power-bi-agentic-development --skill executing-spark -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install data-goblin/power-bi-agentic-development executing-spark --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/data-goblin/power-bi-agentic-development.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/etl/skills/executing-spark .opencode/skills/executing-spark && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "executing-spark" agent skill from https://github.com/data-goblin/power-bi-agentic-development/tree/main/plugins/etl/skills/executing-spark into .opencode/skills/executing-spark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "executing-spark", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
executing-sparkExecute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access.
Executing Spark is an agent skill from data-goblin/power-bi-agentic-development. Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access. Automatically invoke when the user asks to "run PySpark in Fabric", "create a Livy session", "execute Python on Fabric compute", "run Spark without a notebook", "submit code to Fabric", "ephemeral Spark execution", "run ETL in Fabric".
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/example-script.md` and `references/livy-api.md`).
It sits in Data & Analytics, covering Data pipelines and ETL. It works with Apache Spark and Python. The repository describes itself as: Power BI AI skills and Power BI agents for Claude Code and GitHub Copilot: a plugin marketplace of Power BI skills, subagents, and hooks for semantic models, DAX, TMDL, reports… The licence is GPL-3.0.
Read from SKILL.md and the folder at commit 41886f2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
azFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.fabric.microsoft.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Executing Spark loads about 1.7k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 102 tokens; SKILL.md has 699 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from data-goblin/power-bi-agentic-development at commit 41886f2, republished under its GPL-3.0 licence (© data-goblin). 699 words, ~1,662 tokens.
.claude/skills/executing-spark/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Run arbitrary PySpark or Python code on Fabric Spark compute via the Livy API. No notebook artifact is created or persisted; sessions are ephemeral. Full read/write access to lakehouse Delta tables via Spark SQL.
az login)The Livy API requires a token from az account get-access-token --resource https://api.fabric.microsoft.com. Tokens from fab auth do not work for OneLake storage access inside the Spark session.
import subprocess, json
result = subprocess.run(
["az", "account", "get-access-token", "--resource", "https://api.fabric.microsoft.com"],
capture_output=True, text=True
)
token = json.loads(result.stdout)["accessToken"]Do not output or log the token. Pass it directly to the API call.
1. Create session POST .../sessions {"kind": "pyspark"}
2. Wait for idle GET .../sessions/{id} poll until state: "idle" (~30-90s)
3. Submit code POST .../sessions/{id}/statements {"code": "...", "kind": "pyspark"}
4. Get result GET .../sessions/{id}/statements/{n} poll until state: "available"
5. Delete session DELETE .../sessions/{id} ALWAYS do thisBase URL: https://api.fabric.microsoft.com/v1/workspaces/{wsId}/lakehouses/{lhId}/livyapi/versions/2023-12-01
CRITICAL: Always delete sessions when done. Idle sessions consume Fabric capacity units (CUs). A forgotten session burns compute until it times out (default: 20 minutes). In automation, wrap cleanup in a finally block.
WS_ID=$(fab get "Workspace.Workspace" -q "id" | tr -d '"')
LH_ID=$(fab get "Workspace.Workspace/Lakehouse.Lakehouse" -q "id" | tr -d '"')Submit PySpark or pure Python as statements. The spark object is available automatically.
# Statement payload
{"code": "df = spark.sql('SELECT * FROM products LIMIT 10')\ndf.show()", "kind": "pyspark"}Results are in output.data["text/plain"] when state: "available" and output.status: "ok".
spark.sql("SELECT ...") ; full Spark SQL against lakehouse tablesspark.sql("SHOW TABLES") ; metastore accessdf.write.mode("overwrite").saveAsTable(...) ; write Delta tablesdeltalake (delta-rs) is not pre-installed; use Spark SQL insteadnotebookutils has limited functionality (no FUSE mount at /lakehouse/default/)fab auth ; must use az CLI token| Scenario | Approach |
|---|---|
| Quick read-only exploration | DuckDB locally (fastest; see using-duckdb skill) |
| Write data back to lakehouse | Livy session or notebook |
| Ephemeral transform; no artifact | Livy session (this skill) |
| Complex multi-cell workflow | Notebook (nb exec or portal) |
| Scheduled ETL | Notebook via fab job run |
| Agent-driven compute (Dagster, orchestrators) | Livy session |
This skill is for ephemeral execution with no artifact. When you instead want to persist or change a notebook (deploy new code, iterate on an existing one), that is an item-definition change, and the poll interval is the single biggest performance lever. fab import, nb create, and nb cell edit take 25-60s because they poll the create/update long-running operation at the server's advertised Retry-After: 20; the work itself finishes in ~1s, and neither CLI lets you change that interval. Poll the LRO at ~0.3s and the same deploy takes ~1-2s. The fabric-cli skill ships scripts/deploy_notebook.py which does this (auto-detects create vs update, --poll-interval default 0.3s); strongly prefer it over fab import / nb for any notebook definition change.
A Livy session (this skill) is interactive: create it, submit statements, read output as it runs, delete it. It stays alive and you pay for idle time until you delete it or it times out (~20 min).
A Livy batch is one-shot: submit a single job (a file or inline job spec), poll it to a terminal state, done. No idle-CU footgun, nothing to remember to delete. For scheduled or fire-and-forget agent ETL, prefer a batch over a session; keep sessions for interactive, multi-statement work. Same base URL, /batches instead of /sessions -- see references/livy-api.md.
A Livy statement returns its result directly in the response (output.status = ok/error), so you always know whether it worked. A notebook run via fab job run does not -- its job status reports Completed even when the notebook caught an exception and exited a failure payload. If you run notebooks as batch jobs instead of Livy, you must read the notebook's exit value to get its real verdict. The fabric-cli skill (in the fabric-cli plugin) documents that endpoint and ships scripts/run_notebook_checked.py for it.
references/livy-api.md -- Full API reference with endpoints (sessions + batches), request/response formats, and error handlingreferences/example-script.md -- Complete working script that creates a session, queries data, writes results, and cleans upusing-duckdb skill (same etl plugin) -- read-only Delta querying, local or in-notebook, when you don't need Spark computefabric-cli skill (fabric-cli plugin) -- nb exec / fab job run for notebooks, reading a notebook's exit value, the SQL-endpoint metadata sync after a Spark write, and scripts/deploy_notebook.py for fast notebook definition changes (tight LRO polling)© data-goblin, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in plugins/etl/skills/executing-spark of data-goblin/power-bi-agentic-development.
Open the folder on GitHubat commit 41886f2
Executing Spark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Executing Spark this skilldata-goblin/power-bi-agentic-development | 1k | — | ~1.7k | Automated safety check: Pass | GPL-3.0 | |
| Apache Spark Optimizationwshobson/agents | 40k | 8 repos | ~789 | Automated safety check: Pass | MIT | |
| Spark EngineerFerroxLabs/wayland | 608 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Crawl4AI Web Scrapingsmallnest/goclaw | 598 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Migrate Glue Devendpoint To Interactive Sessionsaws-samples/aws-glue-samples | 1.5k | — | ~3.6k | Automated safety check: Pass | MIT-0 | |
| Monitor With HaolemeHaolemeApp/Haoleme | 157 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 |
wshobson/agents
Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.
FerroxLabs/wayland
Apache Spark expertise covering RDD vs DataFrame vs Dataset APIs, partitioning strategies, shuffle optimization, broadcast joins, caching, Spark SQL, structured streaming, UDFs, cluster sizing…
smallnest/goclaw
Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.
aws-samples/aws-glue-samples
Migrate a legacy AWS Glue development endpoint to a Glue interactive session, following the official AWS migration checklist.
HaolemeApp/Haoleme
Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.
Yourdaylight/stock_datasource
Turns a Tushare API doc URL into a full data plugin for the stock_datasource repo: extractor, ClickHouse schema, query service, config and curl examples.
data-goblin/power-bi-agentic-development
Author, validate, publish, and test Power BI paginated reports in the RDL format.
data-goblin/power-bi-agentic-development
Automatically invoke this skill whenever the user asks about Fabric tenant settings or Power BI tenant settings or auditing tenant settings.
data-goblin/power-bi-agentic-development
Interactive BPA rule generation for Power BI semantic models; guided discovery, model investigation, and expert rule authoring.
data-goblin/power-bi-agentic-development
Guidance for Power BI Project (PBIP) structure, thick and thin reports, project renames, forks, and validation.
data-goblin/power-bi-agentic-development
Actionable feedback on the quality, usage, and effectiveness of Power BI reports.
data-goblin/power-bi-agentic-development
This skill should be used whenever the user mentions a "semantic model", "data model", or "dataset", or asks to "build", "model", "design", "optimize", "review", or "audit" one, or to "add a…
Works with
Categories
Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access. Executing Spark is an agent skill from data-goblin/power-bi-agentic-development. Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access.
Executing Spark fits situations like: asks to run PySpark in Fabric; create a Livy session; execute Python on Fabric compute; run Spark without a notebook.
Run `npx skills add data-goblin/power-bi-agentic-development --skill executing-spark -a claude-code`. Or copy the skill folder (plugins/etl/skills/executing-spark in data-goblin/power-bi-agentic-development) into .claude/skills/executing-spark in your project. Claude Code loads it when a task matches its description.
Run `npx skills add data-goblin/power-bi-agentic-development --skill executing-spark -a codex`. Or copy the skill folder (plugins/etl/skills/executing-spark in data-goblin/power-bi-agentic-development) into .agents/skills/executing-spark in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add data-goblin/power-bi-agentic-development --skill executing-spark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/executing-spark, .gemini/skills/executing-spark, .github/skills/executing-spark and .opencode/skills/executing-spark in your project.
Going by SKILL.md and its folder, Executing Spark needs the command-line tools its instructions call (az). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: api.fabric.microsoft.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Executing Spark is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Executing Spark: Apache Spark Optimization (wshobson/agents, 40k stars), Spark Engineer (FerroxLabs/wayland, 608 stars), Crawl4AI Web Scraping (smallnest/goclaw, 598 stars) and Migrate Glue Devendpoint To Interactive Sessions (aws-samples/aws-glue-samples, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
data-goblin (a GitHub user) maintains it in data-goblin/power-bi-agentic-development, which has 1,026 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on October 5, 2026.
Source: data-goblin/power-bi-agentic-development on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.