Agent skill

Glue Diagnostics

by Kilo-Org in Kilo-Org/kilo-marketplace

A skill your agent uses to investigate and troubleshoot AWS Glue problems by analyzing ETL jobs, crawlers, connections, Data Catalog, DPU utilization, Spark execution, and job bookmarks following…

MITAuto-check passedData & Analytics

Install Glue Diagnostics

skills CLI
$ npx skills add Kilo-Org/kilo-marketplace --skill glue-diagnostics -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Kilo-Org/kilo-marketplace glue-diagnostics --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/glue-diagnostics .claude/skills/glue-diagnostics && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
glue-diagnostics
GitHub stars
189
Token cost
~2k tokens
SKILL.md length
753 words
Files
28 (incl. references)
Skills in repo
87
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses to investigate and troubleshoot AWS Glue problems by analyzing ETL jobs, crawlers, connections, Data Catalog, DPU utilization, Spark execution, and job bookmarks following…

  • Works in 2 steps: Collect and triage → Domain deep dive
  • Investigate and troubleshoot AWS Glue problems by analyzing ETL jobs
  • SKILL.md covers When to use, Investigation workflow, Tool quick reference and Gotchas: AWS Glue, plus 2 more sections
  • Calls aws

What it does

Glue Diagnostics is an agent skill from Kilo-Org/kilo-marketplace. Use this skill to investigate and troubleshoot AWS Glue problems by analyzing ETL jobs, crawlers, connections, Data Catalog, DPU utilization, Spark execution, and job bookmarks following structured runbooks. Activate when: job failures, job timeouts, OOM errors, Spark executor or driver crashes, crawler failures, schema detection issues, partition problems, JDBC connection failures, VPC/subnet connectivity, S3 endpoint access, Data Catalog sync issues, schema evolution conflicts, DPU sizing problems, shuffle…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 28 other files, including reference files (for example `README.md`, `references/A1-job-failures.md` and `references/A2-job-timeout.md`). Compatibility notes: Requires AWS CLI or SDK access with Glue, S3, CloudWatch Logs, IAM, EC2 (for VPC/connections), and optionally KMS permissions.

It sits in Data & Analytics, covering Web scraping, Data governance and Data pipelines and ETL. It works with Amazon Web Services. The repository describes itself as: Kilo Marketplace - A curated collection of Skills, MCP Servers, and Modes for enhancing AI agent capabilities across the Kilo ecosystem—including Kilo Code (VS Code extension)… The licence is MIT.

When your agent uses it

  • Investigate and troubleshoot AWS Glue problems by analyzing ETL jobs
  • DPU utilization
  • Spark execution
  • Job bookmarks following structured runbooks

Example prompts

  • “/glue-diagnostics”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires AWS CLI or SDK access with Glue, S3, CloudWatch Logs, IAM, EC2 (for VPC/connections), and optionally KMS permissions.

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Collect and triage
  2. Domain deep dive

What it can do on your machine

Read from SKILL.md and the folder at commit ff51758. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires AWS CLI or SDK access with Glue, S3, CloudWatch Logs, IAM, EC2 (for VPC/connections), and optionally KMS permissions.

    From compatibility in the SKILL.md frontmatter.

Context cost

Glue Diagnostics loads about 2k tokens when it runs, and up to ~38k if it reads all its reference files. Until then it costs about 200 tokens; SKILL.md has 753 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~200
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~38k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Kilo-Org/kilo-marketplace at commit ff51758, republished under its MIT licence (© Kilo-Org). 753 words, ~1,985 tokens.

Download SKILL.mdSave it as .claude/skills/glue-diagnostics/SKILL.md (or your agent's skills folder). This skill also uses 27 other files; get the full folder from GitHub.
name
glue-diagnostics
description
Use this skill to investigate and troubleshoot AWS Glue problems by analyzing ETL jobs, crawlers, connections, Data Catalog, DPU utilization, Spark execution, and job bookmarks following structured runbooks. Activate when: job failures, job timeouts, OOM errors, Spark executor or driver crashes, crawler failures, schema detection issues, partition problems, JDBC connection failures, VPC/subnet connectivity, S3 endpoint access, Data Catalog sync issues, schema evolution conflicts, DPU sizing problems, shuffle bottlenecks, data skew, transformation errors, bookmark issues, data quality failures, IAM permission errors, encryption problems, Glue Studio visual editor errors, job generation failures, or the user says something is wrong with Glue without naming specific symptoms.
compatibility
Requires AWS CLI or SDK access with Glue, S3, CloudWatch Logs, IAM, EC2 (for VPC/connections), and optionally KMS permissions.
metadata.category
data

Glue Diagnostics

When to use

Any AWS Glue investigation where the console alone is insufficient — job failures, OOM errors, Spark crashes, crawler schema misdetection, connection timeouts, Data Catalog drift, DPU under/over-provisioning, data skew, bookmark corruption, or Glue Studio generation errors.

Investigation workflow

Step 1 — Collect and triage
aws glue get-job --name <job-name>
aws glue get-job-run --job-name <job-name> --run-id <run-id>
aws glue batch-get-jobs --job-names <job1> <job2>
aws glue get-crawler --name <crawler-name>
aws glue get-connection --name <connection-name>
aws logs filter-log-events --log-group-name /aws-glue/jobs/logs-v2 --log-stream-name-prefix <run-id>
Step 2 — Domain deep dive
aws glue get-job-runs --job-name <job-name> --max-results 10
aws glue get-crawler-metrics --crawler-name-list <crawler-name>
aws glue get-databases
aws glue get-tables --database-name <db-name>
aws glue get-partitions --database-name <db-name> --table-name <table-name>
aws glue get-job-bookmark --job-name <job-name>
aws cloudwatch get-metric-statistics --namespace Glue --metric-name glue.driver.aggregate.bytesRead --dimensions Name=JobName,Value=<job-name> --start-time <iso> --end-time <iso> --period 300 --statistics Sum

Read references/glue-guardrails.md before concluding on any Glue issue.

Tool quick reference

Tool / APIWhen to use
glue get-jobJob configuration, Glue version, DPU, worker type
glue get-job-runSpecific run status, error message, execution time
glue batch-get-jobsRetrieve multiple job configs at once
glue get-job-runsJob run history, failure patterns
glue get-crawlerCrawler config, targets, schedule, schema change policy
glue get-crawler-metricsCrawler runtime stats, tables created/updated
glue get-connectionJDBC/network connection config, VPC, subnet
glue get-databases / get-tablesData Catalog metadata, schema definitions
glue get-partitionsPartition metadata, partition keys, storage location
glue get-job-bookmarkBookmark state for incremental processing
logs filter-log-eventsGlue job CloudWatch logs for Spark errors
cloudwatch get-metric-statisticsGlue job metrics (bytes read/written, DPU usage)

Gotchas: AWS Glue

  • DPU sizing matters: G.1X (1 DPU per worker, 16 GB memory), G.2X (2 DPU, 32 GB), G.4X (4 DPU, 64 GB), G.8X (8 DPU, 128 GB). Under-provisioning causes OOM; over-provisioning wastes cost.
  • Spark executor OOM vs driver OOM: executor OOM means data partitions are too large (repartition or increase worker type). Driver OOM means too much data collected to the driver (avoid collect(), reduce broadcast join size).
  • Job bookmarks track processed data for incremental loads. Bookmarks only work with S3 sources using job.init()/job.commit(). Resetting bookmarks reprocesses all data.
  • Crawler schema evolution: crawlers can add new columns but may not handle type changes gracefully. Schema change policy (UPDATE_IN_DATABASE vs LOG) controls behavior.
  • Glue connections for JDBC require VPC, subnet, and security group configuration. The subnet must have a NAT gateway or VPC endpoints for Glue service access.
  • Glue Data Catalog vs Hive metastore: Glue Data Catalog is the default metastore for Glue jobs. External Hive metastore requires explicit configuration and network connectivity.
  • Glue Studio visual editor has limitations: complex transformations may require custom code nodes. Not all PySpark/Scala operations are available as visual transforms.
  • Spark UI is available for Glue 2.0+ jobs via the Glue console. It provides DAG visualization, stage details, and executor metrics for debugging performance issues.
  • Job timeout defaults to 48 hours (2880 minutes). Long-running jobs may silently consume DPUs. Always set an explicit timeout.
  • Glue version compatibility: Glue 2.0 (Spark 2.4), Glue 3.0 (Spark 3.1), Glue 4.0 (Spark 3.3). Library availability and behavior differ across versions.
  • Partition management: too many small partitions cause excessive S3 LIST calls. Too few large partitions cause OOM. Aim for 128 MB–512 MB per partition.
  • S3 eventual consistency impact: S3 provides strong read-after-write consistency since December 2020, but Glue Data Catalog partition metadata updates may still lag behind S3 changes.
Show full SKILL.md (287 more words)Show less
Worker type comparison
Worker TypeDPUMemoryvCPUUse Case
G.1X116 GB4Standard ETL, small-medium datasets
G.2X232 GB8Memory-intensive transforms, large joins
G.4X464 GB16ML transforms, very large datasets
G.8X8128 GB32Massive datasets, complex aggregations
G.025X0.252 GB2Python shell jobs only
Z.2X232 GB8Ray jobs (Glue 4.0+)
Glue version comparison
VersionSparkPythonKey Features
Glue 2.02.43.7Spark UI, no startup overhead
Glue 3.03.13.7Optimized shuffle, auto-scaling
Glue 4.03.33.10Ray support, Python 3.10, improved performance

Anti-hallucination rules

  1. Always cite specific job run error messages, crawler metrics, or CloudWatch log entries as evidence.
  2. Never assume OOM is always executor-side. Check whether the error is on the driver or executor — the fix is different.
  3. Job bookmarks only work with supported sources (S3, JDBC) and require job.init()/job.commit() calls. Never claim bookmarks work automatically with all sources.
  4. Crawler schema changes depend on the SchemaChangePolicy. Never assume crawlers automatically update table schemas.
  5. Glue connections require VPC networking. Never suggest JDBC connections work without proper VPC, subnet, and security group configuration.
  6. Spend no more than 2 minutes on any single hypothesis. Pivot if inconclusive.

28 runbooks

CategoryIDsCovers
A — JobsA1–A4Job failures, timeout, OOM, Spark errors
B — CrawlersB1–B3Crawler failures, schema detection, partition issues
C — ConnectionsC1–C3JDBC connection failures, VPC/subnet, S3 endpoint
D — Data CatalogD1–D2Catalog sync issues, schema evolution
E — PerformanceE1–E3DPU sizing, shuffle issues, data skew
F — ETLF1–F3Transformation errors, bookmark issues, data quality
G — SecurityG1–G2IAM permissions, encryption
H — Glue StudioH1–H2Visual editor errors, job generation
Z — Catch-AllZ1General troubleshooting

© Kilo-Org, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 27 other files (references) in skills/glue-diagnostics of Kilo-Org/kilo-marketplace.

  • SKILL.md
  • LICENSE
  • README.md
  • references/A1-job-failures.md
  • references/A2-job-timeout.md
  • references/A3-oom-errors.md
  • references/A4-spark-errors.md
  • references/B1-crawler-failures.md
  • references/B2-schema-detection.md
  • references/B3-partition-issues.md
  • references/C1-jdbc-connection-failures.md
  • references/C2-vpc-subnet-issues.md
  • references/C3-s3-endpoint-access.md
  • references/D1-catalog-sync-issues.md
  • references/D2-schema-evolution.md
  • references/E1-dpu-sizing.md
  • references/E2-shuffle-issues.md
  • references/E3-data-skew.md
  • references/F1-transformation-errors.md
  • references/F2-bookmark-issues.md
  • … and 8 more

Open the folder on GitHubat commit ff51758

Compare with similar skills

Glue Diagnostics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Glue Diagnostics compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Glue Diagnostics this skillKilo-Org/kilo-marketplace189—~2kAutomated safety check: PassMIT
Querying AWS Sagemaker Catalogaws/agent-toolkit-for-aws2.8k—~2.6kAutomated safety check: PassApache-2.0
Authoritative Data Harvesteryushui2022/MathModel-Skill4521 repos~1.1kAutomated safety check: PassMIT
Data Quality Frameworkswshobson/agents40k10 repos~1.1kAutomated safety check: PassMIT
Ingesting Dataancoleman/ai-design-components526—~1.9kAutomated safety check: PassMIT
Exploring Data Catalogaws/agent-toolkit-for-aws2.8k1 repos~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Querying AWS Sagemaker Catalog

    aws/agent-toolkit-for-aws

    Official

    Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables.

    2.8k GitHub stars~2.6k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Authoritative Data Harvester

    yushui2022/MathModel-Skill

    Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.

    452 GitHub starsUsed in 1 repo~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.

    40k GitHub starsUsed in 10 repos~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Ingesting Data

    ancoleman/ai-design-components

    Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases.

    526 GitHub stars~1.9k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed
  • Exploring Data Catalog

    aws/agent-toolkit-for-aws

    Official

    Full inventory and audit of AWS Glue Data Catalog assets across S3 Tables, Redshift-federated, and remote Iceberg catalogs.

    2.8k GitHub starsUsed in 1 repo~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Authoring Mwaa Workflow

    aws/agent-toolkit-for-aws

    Official

    Authors and deploys MWAA workflow artifacts: Python Airflow DAGs for provisioned environments or YAML workflow files for Serverless.

    2.8k GitHub stars~2.8k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from Kilo-Org/kilo-marketplace

All 87 skills in this repo
  • AzureML Project Scaffolding

    Kilo-Org/kilo-marketplace

    Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible.

    189 GitHub stars~3.1k tokensUpdated 8 days ago
    Auto-check: notes
  • Jupyter Notebook Builder

    Kilo-Org/kilo-marketplace

    Creates, inspects, edits and runs Jupyter notebooks, scaffolding experiment or tutorial notebooks from templates and preferring a Jupyter MCP server over raw JSON edits.

    189 GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed
  • Tableau Dashboard Creator

    Kilo-Org/kilo-marketplace

    Takes a plain-language dashboard request through brand setup, data exploration, planning, an interactive HTML mock and a Tableau implementation spec.

    189 GitHub stars~3.8k tokensUpdated 8 days ago
    Auto-check: notes
  • Elasticsearch File Ingest

    Kilo-Org/kilo-marketplace

    Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

    189 GitHub stars~2.8k tokensUpdated 8 days ago
    Auto-check passed
  • Nifi Flow Layout

    Kilo-Org/kilo-marketplace

    A skill your agent uses when arranging Apache NiFi processors, process groups, ports, comments, numbering, crossing connections, dense fan-in/fan-out, or reusable readable canvas layouts.

    189 GitHub stars~1.5k tokensUpdated 8 days ago
    Auto-check passed
  • Splunk Ingest Processor Setup

    Kilo-Org/kilo-marketplace

    Render Cisco Data Fabric ingest-time routing workflows and Splunk Cloud Platform Ingest Processor setup plans with SPL2 pipelines, source types, destinations, lifecycle handoffs, queue and…

    189 GitHub stars~1.2k tokensUpdated 8 days ago
    Auto-check passed

Questions about Glue Diagnostics

What does Glue Diagnostics do?

A skill your agent uses to investigate and troubleshoot AWS Glue problems by analyzing ETL jobs, crawlers, connections, Data Catalog, DPU utilization, Spark execution, and job bookmarks following…. Glue Diagnostics is an agent skill from Kilo-Org/kilo-marketplace. Use this skill to investigate and troubleshoot AWS Glue problems by analyzing ETL jobs, crawlers, connections, Data Catalog, DPU utilization, Spark execution, and job bookmarks following structured runbooks.

When should I use Glue Diagnostics?

Glue Diagnostics fits situations like: investigate and troubleshoot AWS Glue problems by analyzing ETL jobs; DPU utilization; spark execution; job bookmarks following structured runbooks.

How do I install Glue Diagnostics in Claude Code?

Run `npx skills add Kilo-Org/kilo-marketplace --skill glue-diagnostics -a claude-code`. Or copy the skill folder (skills/glue-diagnostics in Kilo-Org/kilo-marketplace) into .claude/skills/glue-diagnostics in your project. Claude Code loads it when a task matches its description.

How do I install Glue Diagnostics in Codex?

Run `npx skills add Kilo-Org/kilo-marketplace --skill glue-diagnostics -a codex`. Or copy the skill folder (skills/glue-diagnostics in Kilo-Org/kilo-marketplace) into .agents/skills/glue-diagnostics in your project. Codex loads it when a task matches its description.

Can I use Glue Diagnostics in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kilo-Org/kilo-marketplace --skill glue-diagnostics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/glue-diagnostics, .gemini/skills/glue-diagnostics, .github/skills/glue-diagnostics and .opencode/skills/glue-diagnostics in your project.

What does Glue Diagnostics need to run?

Going by SKILL.md and its folder, Glue Diagnostics needs the command-line tools its instructions call (aws). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires AWS CLI or SDK access with Glue, S3, CloudWatch Logs, IAM, EC2 (for VPC/connections), and optionally KMS permissions. .

Does Glue Diagnostics access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Glue Diagnostics safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Glue Diagnostics use?

Glue Diagnostics is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Glue Diagnostics use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 36k tokens, read only when the agent opens those files.

What are the alternatives to Glue Diagnostics?

Skills that share tags, products or a category with Glue Diagnostics: Querying AWS Sagemaker Catalog (aws/agent-toolkit-for-aws, 2.8k stars), Authoritative Data Harvester (yushui2022/MathModel-Skill, 452 stars), Data Quality Frameworks (wshobson/agents, 40k stars) and Ingesting Data (ancoleman/ai-design-components, 526 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Glue Diagnostics?

Kilo-Org (a GitHub organization) maintains it in Kilo-Org/kilo-marketplace, which has 189 GitHub stars. The repository holds 87 skills in this directory. The repository was last updated on September 28, 2026.

Source: Kilo-Org/kilo-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.