Official agent skill

Querying Data Lake

by aws in aws/agent-toolkit-for-aws

Execute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift).

OfficialApache-2.0Auto-check passedDatabases

Install Querying Data Lake

skills CLI
$ npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws/agent-toolkit-for-aws querying-data-lake --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/aws-data-analytics/skills/querying-data-lake .claude/skills/querying-data-lake && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
querying-data-lake
GitHub stars
2.8k
Token cost
~1.9k tokens
SKILL.md length
888 words
Files
3 (incl. references)
Skills in repo
138
Repo updated
First seen
Licence
Apache-2.0

At a glance

Execute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift).

  • Works in 7 steps: Verify Dependencies → Resolve Workgroup → Resolve the Target Asset → …
  • Phrases like: query data
  • SKILL.md covers Overview, Common Tasks, Troubleshooting and Additional Resources
  • Calls aws

What it does

Querying Data Lake is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Execute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift). Triggers on phrases like: query data, run SQL, athena query, analyze table, SQL query, workgroup status, profile table, query Redshift catalog, query S3 Tables. Do NOT use for finding specific data assets (use finding-data-lake-assets), full catalog audits (use exploring-data-catalog), importing data (use ingesting-into-data-lake).

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/query-patterns.md` and `references/workgroup-selection.md`).

It sits in Databases, covering SQL, File uploads and storage and Data warehousing. It works with SQL, Amazon Web Services and Model Context Protocol. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.

When your agent uses it

  • Phrases like: query data
  • Workgroup status
  • Query Redshift catalog
  • Query S3 Tables

Example prompts

  • “/querying-data-lake”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Verify Dependencies
  2. Resolve Workgroup
  3. Resolve the Target Asset
  4. Discover Schema
  5. Build Query
  6. Classify and Execute
  7. Present and Recover

What it can do on your machine

Read from SKILL.md and the folder at commit 188af2f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.aws.amazon.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Querying Data Lake loads about 1.9k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 114 tokens; SKILL.md has 888 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~114
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aws/agent-toolkit-for-aws at commit 188af2f, republished under its Apache-2.0 licence (© aws). 888 words, ~1,934 tokens.

Download SKILL.mdSave it as .claude/skills/querying-data-lake/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
querying-data-lake
description
Execute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift). Triggers on phrases like: query data, run SQL, athena query, analyze table, SQL query, workgroup status, profile table, query Redshift catalog, query S3 Tables. Do NOT use for finding specific data assets (use finding-data-lake-assets), full catalog audits (use exploring-data-catalog), importing data (use ingesting-into-data-lake).
metadata.version
1
metadata.argument-hint
'[SQL-query|query-name|workgroup-name|catalog-name|''profile TABLE_NAME'']'

Query Data Lake

Execute SQL queries on Amazon Athena across default and federated catalogs (Glue, S3 Tables, Redshift) with workgroup selection, statement classification, and error recovery.

Overview

Executes and manages Athena SQL queries across default and federated catalogs. Selects a workgroup, resolves target assets (delegating fuzzy references to finding-data-lake-assets), classifies statements for safety, and reports cost and data scanned. Use the AWS MCP server for sandboxed execution and audit logging; the same AWS CLI commands work directly when the MCP server is not available.

Constraints for parameter acquisition:

  • You MUST accept a single optional argument: SQL text, a named-query name, a workgroup name, a catalog name, or profile TABLE_NAME
  • You MUST accept the argument as direct text or a pointer to a file containing SQL
  • You MUST ask the user for the target AWS region if not already set
  • You MUST confirm the output S3 location before executing any non-trivial query
  • You MUST respect the user's decision to abort at any step

Common Tasks

1. Verify Dependencies

Check for required tools and AWS access before running queries.

Constraints:

  • You MUST verify AWS MCP server tools are available (aws___call_aws) and run queries through them when present; fall back to AWS CLI only if the MCP server is unavailable
  • You MUST NOT fall back to shell or Bash for query execution — results must be captured via the MCP tool or aws athena CLI so output location and cost are tracked
  • You MUST confirm credentials with aws sts get-caller-identity and inform the user about any missing tools
2. Resolve Workgroup

Check caller identity, list workgroups, auto-select the best one (see workgroup-selection.md).

Constraints:

  • You MUST select a workgroup before submitting any query (prevents output-location errors)
  • You MUST present the selected workgroup and its output location to the user
  • You MUST NOT auto-escalate to a different workgroup on failure without user confirmation
3. Resolve the Target Asset

If the user refers to a table by name, by business concept ("our quarterly report", "the sales data"), by S3 path, or by catalog without specifying the table, delegate to finding-data-lake-assets to return the concrete database.table (and catalog if non-default).

Constraints:

  • You MUST NOT attempt to resolve fuzzy asset references with athena list-data-catalogs or by iterating get-tables — those miss federated catalogs and waste tokens
  • You SHOULD skip this step only when the user provides a fully-qualified reference (exact database.table) or raw SQL they want executed as-is
  • You MUST state the resolved asset explicitly before building the query: "Found [table] in [catalog]. Using this for the query."
  • You SHOULD default to the default Glue catalog unless the user mentions "federated", "Redshift", "S3 Tables", or finding-data-lake-assets returns a different catalog
4. Discover Schema

For analytical queries, You SHOULD profile the target table before building the final query. You MUST show sample rows (SELECT ... LIMIT 5) as part of profiling.

5. Build Query

Table addressing depends on catalog type:

  • Default Glue catalog: database.table (omit the catalog prefix for single-catalog queries). In cross-catalog queries, qualify default-catalog tables with "awsdatacatalog".database.table.
  • Registered data source: datasource.database.table
  • Unregistered Glue catalog: "catalog/subcatalog".database.table
Show full SKILL.md (383 more words)Show less
6. Classify and Execute

Classify the SQL statement before executing:

StatementBehavior
SELECT, SHOW, DESCRIBE, EXPLAINSafe — execute
INSERT, UPDATE, DELETE, DROP, ALTER, CREATE, TRUNCATE, MERGEDestructive — warn the user and require explicit confirmation
UnsureTreat as destructive; confirm

Example tool call (via AWS MCP server):

aws___call_aws(command="aws athena start-query-execution --work-group <WORKGROUP_NAME> --query-string '<sql>' --query-execution-context Database=<db>")

For federated or S3 Tables catalogs, also set Catalog=<CATALOG_PATH> in the execution context (e.g. Catalog=s3tablescatalog/<BUCKET_NAME>).

Constraints:

  • You MUST warn the user before executing when the target is Redshift-federated ("No partition pruning — every query scans the full table")
  • You MUST warn the user before executing a cross-catalog join ("Cross-catalog joins incur network overhead and may be slow")
  • You MUST confirm the output S3 location before executing
  • You MUST explain which tool is being called before executing
  • You MUST respect the user's decision to abort
7. Present and Recover

Present results with cost, data scanned, duration, and actionable insights. On failure, list available workgroups and let the user choose which to retry with.

Argument Routing

Resolve in this order; stop at the first match:

  1. Contains SQL keywords (SELECT, SHOW, DESCRIBE, INSERT, etc.) — SQL text, execute directly
  2. profile TABLE_NAME — run comprehensive table profiling (see query-patterns.md)
  3. Matches a known named query — look up and execute
  4. Matches a known workgroup — show workgroup status and recent queries
  5. Matches a known catalog — delegate to exploring-data-catalog to enumerate databases and tables
  6. No args — show recent query activity and available tables
Principles
  • Always select workgroup before executing (prevents output-location errors)
  • Profile unfamiliar tables before running analytical queries
  • Present cost alongside results so users build cost awareness
  • Suggest LIMIT for exploratory queries on large tables
  • Never ask domain questions with obvious answers, but always confirm security-relevant actions (workgroup switches, output location changes, non-SELECT statements)

Troubleshooting

ErrorCauseFix
Redshift identifier error with mixed caseRedshift-federated names are lowercase onlyLowercase the identifier
CatalogId validation failureARN passed instead of catalog namePass the catalog name, not the ARN
Cross-catalog information_schema returns nothingMissing catalog qualifierUse catalog-qualified path: "catalog".information_schema.tables
Query fails with output-location errorWorkgroup has no output location configuredSelect a different workgroup with an output location, or configure one
Destructive statement executed without confirmationStatement classification skippedAlways classify INSERT/UPDATE/DELETE/DROP/ALTER/CREATE/TRUNCATE/MERGE and confirm with the user

Additional Resources

© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in plugins/aws-data-analytics/skills/querying-data-lake of aws/agent-toolkit-for-aws.

  • SKILL.md
  • references/query-patterns.md
  • references/workgroup-selection.md

Open the folder on GitHubat commit 188af2f

Compare with similar skills

Querying Data Lake next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Querying Data Lake compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Querying Data Lake this skillaws/agent-toolkit-for-aws2.8k—~1.9kAutomated safety check: PassApache-2.0
Chdb SQLvemetric/vemetric3941 repos~1.2kAutomated safety check: PassApache-2.0
Modelersidequery/sidemantic129—~4.2kAutomated safety check: PassApache-2.0
Semantic Analystsidequery/sidemantic129—~982Automated safety check: PassAGPL-3.0
Pytorch Clickhousepytorch/test-infra113—~2.8kAutomated safety check: PassCustom licence
Imaging Data CommonsK-Dense-AI/scientific-agent-skills48k1 repos~7.8kAutomated safety check: PassMIT

Similar skills

  • Chdb SQL

    vemetric/vemetric

    A skill your agent uses when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse…

    394 GitHub starsUsed in 1 repo~1.2k tokens
    DatabasesAuto-check passed
  • Modeler

    sidequery/sidemantic

    Build, validate, and manage semantic models using Sidemantic.

    129 GitHub stars~4.2k tokensUpdated yesterday
    DatabasesAuto-check passed
  • Semantic Analyst

    sidequery/sidemantic

    Answer analytical, KPI, metric, trend, cohort, and business-performance questions through a Sidemantic semantic layer.

    129 GitHub stars~982 tokensUpdated yesterday
    DatabasesAuto-check passed
  • Pytorch Clickhouse

    pytorch/test-infra

    Load this FIRST whenever working with PyTorch CI data (any pytorch/ org repo), the torchci/HUD codebase, or the PyTorch HUD ClickHouse database.

    113 GitHub stars~2.8k tokensUpdated today
    DatabasesAuto-check passed
  • Imaging Data Commons

    K-Dense-AI/scientific-agent-skills

    Queries and downloads public cancer imaging data from NCI Imaging Data Commons.

    48k GitHub starsUsed in 1 repo~7.8k tokens
    DatabasesAuto-check passed
  • Ops Telemetry Query

    boundless-xyz/boundless

    Internal — for Boundless team members only. An agent skill from boundless-xyz/boundless.

    193 GitHub stars~3.8k tokensUpdated 1 mo ago
    DatabasesAuto-check passed

More from aws/agent-toolkit-for-aws

All 138 skills in this repo
  • Agent Advisor

    aws/agent-toolkit-for-aws

    Official

    Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.

    2.8k GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Agents Build

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.

    2.8k GitHub stars~2.3k tokensUpdated today
    Auto-check: notes
  • Launch With AWS

    aws/agent-toolkit-for-aws

    Official

    Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.

    2.8k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.

    2.8k GitHub stars~4k tokensUpdated today
    Auto-check passed
  • AWS Marketplace Metering

    aws/agent-toolkit-for-aws

    Official

    Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…

    2.8k GitHub stars~18k tokensUpdated today
    Auto-check passed
  • Agents Pay

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.

    2.8k GitHub stars~6.5k tokensUpdated today
    Auto-check: notes

Categories

Questions about Querying Data Lake

What does Querying Data Lake do?

Execute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift). Querying Data Lake is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Execute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift).

When should I use Querying Data Lake?

Querying Data Lake fits situations like: phrases like: query data; workgroup status; query Redshift catalog; query S3 Tables.

How do I install Querying Data Lake in Claude Code?

Run `npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a claude-code`. Or copy the skill folder (plugins/aws-data-analytics/skills/querying-data-lake in aws/agent-toolkit-for-aws) into .claude/skills/querying-data-lake in your project. Claude Code loads it when a task matches its description.

How do I install Querying Data Lake in Codex?

Run `npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a codex`. Or copy the skill folder (plugins/aws-data-analytics/skills/querying-data-lake in aws/agent-toolkit-for-aws) into .agents/skills/querying-data-lake in your project. Codex loads it when a task matches its description.

Can I use Querying Data Lake in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/querying-data-lake, .gemini/skills/querying-data-lake, .github/skills/querying-data-lake and .opencode/skills/querying-data-lake in your project.

What does Querying Data Lake need to run?

Going by SKILL.md and its folder, Querying Data Lake needs the command-line tools its instructions call (aws).

Does Querying Data Lake access the network?

SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.

Is Querying Data Lake safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Querying Data Lake use?

Querying Data Lake is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Querying Data Lake use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to Querying Data Lake?

Skills that share tags, products or a category with Querying Data Lake: Chdb SQL (vemetric/vemetric, 394 stars), Modeler (sidequery/sidemantic, 129 stars), Semantic Analyst (sidequery/sidemantic, 129 stars) and Pytorch Clickhouse (pytorch/test-infra, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Querying Data Lake?

aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,825 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 7, 2026.

Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.