Chdb SQL
vemetric/vemetric
A skill your agent uses when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse…
Execute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift).
$ npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aws/agent-toolkit-for-aws querying-data-lake --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/aws-data-analytics/skills/querying-data-lake .claude/skills/querying-data-lake && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "querying-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/querying-data-lake into .claude/skills/querying-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "querying-data-lake", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/querying-data-lakeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aws/agent-toolkit-for-aws querying-data-lake --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/aws-data-analytics/skills/querying-data-lake .agents/skills/querying-data-lake && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "querying-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/querying-data-lake into .agents/skills/querying-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "querying-data-lake", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aws/agent-toolkit-for-aws querying-data-lake --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/aws-data-analytics/skills/querying-data-lake .cursor/skills/querying-data-lake && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "querying-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/querying-data-lake into .cursor/skills/querying-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "querying-data-lake", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aws/agent-toolkit-for-aws.git --path plugins/aws-data-analytics/skills/querying-data-lake--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aws/agent-toolkit-for-aws querying-data-lake --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/aws-data-analytics/skills/querying-data-lake .gemini/skills/querying-data-lake && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "querying-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/querying-data-lake into .gemini/skills/querying-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "querying-data-lake", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aws/agent-toolkit-for-aws querying-data-lakeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/aws-data-analytics/skills/querying-data-lake .github/skills/querying-data-lake && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "querying-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/querying-data-lake into .github/skills/querying-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "querying-data-lake", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aws/agent-toolkit-for-aws querying-data-lake --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/aws-data-analytics/skills/querying-data-lake .opencode/skills/querying-data-lake && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "querying-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/querying-data-lake into .opencode/skills/querying-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "querying-data-lake", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
querying-data-lakeExecute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift).
Querying Data Lake is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Execute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift). Triggers on phrases like: query data, run SQL, athena query, analyze table, SQL query, workgroup status, profile table, query Redshift catalog, query S3 Tables. Do NOT use for finding specific data assets (use finding-data-lake-assets), full catalog audits (use exploring-data-catalog), importing data (use ingesting-into-data-lake).
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/query-patterns.md` and `references/workgroup-selection.md`).
It sits in Databases, covering SQL, File uploads and storage and Data warehousing. It works with SQL, Amazon Web Services and Model Context Protocol. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 188af2f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
awsFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.aws.amazon.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Querying Data Lake loads about 1.9k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 114 tokens; SKILL.md has 888 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from aws/agent-toolkit-for-aws at commit 188af2f, republished under its Apache-2.0 licence (© aws). 888 words, ~1,934 tokens.
.claude/skills/querying-data-lake/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Execute SQL queries on Amazon Athena across default and federated catalogs (Glue, S3 Tables, Redshift) with workgroup selection, statement classification, and error recovery.
Executes and manages Athena SQL queries across default and federated catalogs. Selects a workgroup, resolves target assets (delegating fuzzy references to finding-data-lake-assets), classifies statements for safety, and reports cost and data scanned. Use the AWS MCP server for sandboxed execution and audit logging; the same AWS CLI commands work directly when the MCP server is not available.
Constraints for parameter acquisition:
profile TABLE_NAMECheck for required tools and AWS access before running queries.
Constraints:
aws___call_aws) and run queries through them when present; fall back to AWS CLI only if the MCP server is unavailableaws athena CLI so output location and cost are trackedaws sts get-caller-identity and inform the user about any missing toolsCheck caller identity, list workgroups, auto-select the best one (see workgroup-selection.md).
Constraints:
If the user refers to a table by name, by business concept ("our quarterly report", "the sales data"), by S3 path, or by catalog without specifying the table, delegate to finding-data-lake-assets to return the concrete database.table (and catalog if non-default).
Constraints:
athena list-data-catalogs or by iterating get-tables — those miss federated catalogs and waste tokensdatabase.table) or raw SQL they want executed as-isfinding-data-lake-assets returns a different catalogFor analytical queries, You SHOULD profile the target table before building the final query. You MUST show sample rows (SELECT ... LIMIT 5) as part of profiling.
Table addressing depends on catalog type:
database.table (omit the catalog prefix for single-catalog queries). In cross-catalog queries, qualify default-catalog tables with "awsdatacatalog".database.table.datasource.database.table"catalog/subcatalog".database.tableClassify the SQL statement before executing:
| Statement | Behavior |
|---|---|
SELECT, SHOW, DESCRIBE, EXPLAIN | Safe — execute |
INSERT, UPDATE, DELETE, DROP, ALTER, CREATE, TRUNCATE, MERGE | Destructive — warn the user and require explicit confirmation |
| Unsure | Treat as destructive; confirm |
Example tool call (via AWS MCP server):
aws___call_aws(command="aws athena start-query-execution --work-group <WORKGROUP_NAME> --query-string '<sql>' --query-execution-context Database=<db>")For federated or S3 Tables catalogs, also set Catalog=<CATALOG_PATH> in the execution context (e.g. Catalog=s3tablescatalog/<BUCKET_NAME>).
Constraints:
Present results with cost, data scanned, duration, and actionable insights. On failure, list available workgroups and let the user choose which to retry with.
Resolve in this order; stop at the first match:
SELECT, SHOW, DESCRIBE, INSERT, etc.) — SQL text, execute directlyprofile TABLE_NAME — run comprehensive table profiling (see query-patterns.md)exploring-data-catalog to enumerate databases and tablesLIMIT for exploratory queries on large tables| Error | Cause | Fix |
|---|---|---|
| Redshift identifier error with mixed case | Redshift-federated names are lowercase only | Lowercase the identifier |
CatalogId validation failure | ARN passed instead of catalog name | Pass the catalog name, not the ARN |
Cross-catalog information_schema returns nothing | Missing catalog qualifier | Use catalog-qualified path: "catalog".information_schema.tables |
| Query fails with output-location error | Workgroup has no output location configured | Select a different workgroup with an output location, or configure one |
| Destructive statement executed without confirmation | Statement classification skipped | Always classify INSERT/UPDATE/DELETE/DROP/ALTER/CREATE/TRUNCATE/MERGE and confirm with the user |
© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in plugins/aws-data-analytics/skills/querying-data-lake of aws/agent-toolkit-for-aws.
Open the folder on GitHubat commit 188af2f
Querying Data Lake next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Querying Data Lake this skillaws/agent-toolkit-for-aws | 2.8k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Chdb SQLvemetric/vemetric | 394 | 1 repos | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Modelersidequery/sidemantic | 129 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Semantic Analystsidequery/sidemantic | 129 | — | ~982 | Automated safety check: Pass | AGPL-3.0 | |
| Pytorch Clickhousepytorch/test-infra | 113 | — | ~2.8k | Automated safety check: Pass | Custom licence | |
| Imaging Data CommonsK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~7.8k | Automated safety check: Pass | MIT |
vemetric/vemetric
A skill your agent uses when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse…
sidequery/sidemantic
Build, validate, and manage semantic models using Sidemantic.
sidequery/sidemantic
Answer analytical, KPI, metric, trend, cohort, and business-performance questions through a Sidemantic semantic layer.
pytorch/test-infra
Load this FIRST whenever working with PyTorch CI data (any pytorch/ org repo), the torchci/HUD codebase, or the PyTorch HUD ClickHouse database.
K-Dense-AI/scientific-agent-skills
Queries and downloads public cancer imaging data from NCI Imaging Data Commons.
boundless-xyz/boundless
Internal — for Boundless team members only. An agent skill from boundless-xyz/boundless.
aws/agent-toolkit-for-aws
Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.
aws/agent-toolkit-for-aws
A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.
aws/agent-toolkit-for-aws
Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.
aws/agent-toolkit-for-aws
Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.
aws/agent-toolkit-for-aws
Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…
aws/agent-toolkit-for-aws
A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.
Categories
Execute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift). Querying Data Lake is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Execute and manage Athena SQL queries across default and federated catalogs (Glue, S3 Tables, Redshift).
Querying Data Lake fits situations like: phrases like: query data; workgroup status; query Redshift catalog; query S3 Tables.
Run `npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a claude-code`. Or copy the skill folder (plugins/aws-data-analytics/skills/querying-data-lake in aws/agent-toolkit-for-aws) into .claude/skills/querying-data-lake in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a codex`. Or copy the skill folder (plugins/aws-data-analytics/skills/querying-data-lake in aws/agent-toolkit-for-aws) into .agents/skills/querying-data-lake in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill querying-data-lake -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/querying-data-lake, .gemini/skills/querying-data-lake, .github/skills/querying-data-lake and .opencode/skills/querying-data-lake in your project.
Going by SKILL.md and its folder, Querying Data Lake needs the command-line tools its instructions call (aws).
SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Querying Data Lake is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Querying Data Lake: Chdb SQL (vemetric/vemetric, 394 stars), Modeler (sidequery/sidemantic, 129 stars), Semantic Analyst (sidequery/sidemantic, 129 stars) and Pytorch Clickhouse (pytorch/test-infra, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,825 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 7, 2026.
Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.