Official agent skill

Querying AWS Sagemaker Catalog

by aws in aws/agent-toolkit-for-aws

Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables.

OfficialApache-2.0Auto-check passedData & Analytics

Install Querying AWS Sagemaker Catalog

skills CLI
$ npx skills add aws/agent-toolkit-for-aws --skill querying-aws-sagemaker-catalog -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws/agent-toolkit-for-aws querying-aws-sagemaker-catalog --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/specialized-skills/system-table-skills/querying-aws-sagemaker-catalog .claude/skills/querying-aws-sagemaker-catalog && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
querying-aws-sagemaker-catalog
GitHub stars
2.8k
Token cost
~2.6k tokens
SKILL.md length
774 words
Files
1
Skills in repo
138
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables.

  • Works in 4 steps: Check If Configured → Enable → Verify Permissions for Querying → …
  • Phrases: catalog inventory SQL
  • SKILL.md covers Overview, Decision Tree, Common Tasks and Key Behaviors, plus 3 more sections
  • Calls aws

What it does

Querying AWS Sagemaker Catalog is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables. Covers governance queries, asset growth tracking, ownership audits, time-travel over catalog state, and metadata quality analysis. Applies when querying catalog inventory, finding assets without descriptions, comparing catalog snapshots, or auditing data ownership. Trigger phrases: catalog inventory SQL, how many assets, assets without descriptions, asset growth over time, who owns this data, catalog governance…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data governance, File uploads and storage and SQL. It works with Amazon SageMaker, Amazon Web Services and SQL. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.

When your agent uses it

  • Phrases: catalog inventory SQL
  • How many assets
  • Assets without descriptions
  • Asset growth over time

Example prompts

  • “/querying-aws-sagemaker-catalog”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Check If Configured
  2. Enable
  3. Verify Permissions for Querying
  4. Query

What it can do on your machine

Read from SKILL.md and the folder at commit df2ab44. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.aws.amazon.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Querying AWS Sagemaker Catalog loads about 2.6k tokens when it runs. Until then it costs about 147 tokens; SKILL.md has 774 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aws/agent-toolkit-for-aws at commit df2ab44, republished under its Apache-2.0 licence (© aws). 774 words, ~2,630 tokens.

Download SKILL.mdSave it as .claude/skills/querying-aws-sagemaker-catalog/SKILL.md (or your agent's skills folder).
name
querying-aws-sagemaker-catalog
description
Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables. Covers governance queries, asset growth tracking, ownership audits, time-travel over catalog state, and metadata quality analysis. Applies when querying catalog inventory, finding assets without descriptions, comparing catalog snapshots, or auditing data ownership. Trigger phrases: catalog inventory SQL, how many assets, assets without descriptions, asset growth over time, who owns this data, catalog governance, data quality audit, catalog analytics.
version
1
argument-hint
[query|domain-id|'configure'|'status']

Query AWS SageMaker Catalog System Tables

Overview

Works best with the AWS MCP server for sandboxed execution and audit logging. All commands below use the AWS CLI and work in any environment with configured AWS credentials.

Amazon SageMaker Unified Studio (whose catalog feature is referred to below as SageMaker Catalog) exports asset metadata as a daily-snapshot Apache Iceberg table in the AWS-managed aws-sagemaker-catalog table bucket. This enables SQL queries over your entire data catalog inventory — asset counts, governance gaps, ownership audits, and historical comparisons — without building custom ETL.

Data is partitioned by snapshot_time and exported once daily (around midnight per region). The table is read-only.

Decision Tree

User intentUse this skill?Alternative
SQL analytics on catalog state (counts, governance, trends)Yes—
Historical comparison ("what changed in catalog last week")Yes — time travel via snapshot_time—
Find assets without owners or descriptionsYes—
Find a specific table by name or conceptNofinding-data-lake-assets or Glue Discovery search
Browse/enumerate catalog interactivelyNoexploring-data-catalog
Run a query on a table's dataNoquerying-data-lake
Manage catalog metadata (add descriptions, tags)NoGlue Discovery put-form-type / associate-glossary-terms

Common Tasks

1. Check If Configured
bash
aws datazone get-data-export-configuration \
  --domain-identifier <DOMAIN_ID> \
  --region <REGION>
  • If no domain exists: aws datazone list-domains --region <REGION>
  • If export not enabled: guide user to enable.
  • One domain per account per region.

Verify table bucket exists:

bash
aws s3tables list-table-buckets --region <REGION> \
  --query "tableBuckets[?name=='aws-sagemaker-catalog']"
2. Enable

With KMS encryption (recommended for production):

bash
aws datazone put-data-export-configuration \
  --domain-identifier <DOMAIN_ID> \
  --region <REGION> \
  --enable-export \
  --encryption-configuration kmsKeyArn=<KMS_KEY_ARN>,sseAlgorithm=aws:kms

Note: Encryption cannot be changed after creation. Always specify KMS for sensitive catalog data.

Without encryption (for quick testing only):

bash
aws datazone put-data-export-configuration \
  --domain-identifier <DOMAIN_ID> \
  --region <REGION> \
  --enable-export

First data available within 24 hours. See: Exporting asset metadata

3. Verify Permissions for Querying

Requires:

  • S3 Tables federated catalog registered in Glue (s3tablescatalog)
  • Lake Formation SELECT + DESCRIBE grants on the table

Grant access:

bash
aws lakeformation grant-permissions \
  --principal DataLakePrincipalIdentifier=<ROLE_ARN> \
  --resource '{"Table": {"CatalogId": "<ACCOUNT>:s3tablescatalog/aws-sagemaker-catalog", "DatabaseName": "asset_metadata", "Name": "asset"}}' \
  --permissions DESCRIBE SELECT \
  --region <REGION>
4. Query

Query syntax:

sql
"s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"

Constraints:

  • You MUST always filter by snapshot_time — without it, the query scans all historical snapshots and returns duplicates

  • You MUST confirm workgroup and output location before executing

  • Default to DATE(snapshot_time) = CURRENT_DATE for current state

  • You SHOULD use the key columns documented in this skill to build queries. If you need the full schema, run get-tables once:

    aws glue get-tables --catalog-id "<ACCOUNT>:s3tablescatalog/aws-sagemaker-catalog" --database-name "asset_metadata" --region <REGION>

Key columns:

ColumnWhat it holdsUsage
snapshot_timePartition key — daily snapshot timestampAlways filter on this
asset_idUnique catalog asset identifierPrimary key for lookups
resource_type_enumGlueTable, RedshiftTable, S3Collection, etc.Filter by asset type
resource_idARN or native identifierCross-reference with source systems
asset_nameBusiness-friendly nameDisplay, search
resource_nameTechnical name (table name, prefix)Filtering
business_descriptionBusiness context (NULL if not provided)Governance gaps
extended_metadatamap<string,string> — flexible key-value attributesUse bracket notation: extended_metadata['owningEntityId']
asset_created_timeWhen asset first appeared in catalogGrowth analysis
asset_updated_timeLast modification timeFreshness checks

Current catalog state:

sql
SELECT resource_type_enum, COUNT(*) as count
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
GROUP BY resource_type_enum
ORDER BY count DESC;

Assets without business descriptions:

sql
SELECT asset_name, resource_name, resource_type_enum, account_id
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
  AND business_description IS NULL;

Asset growth over last 30 days:

sql
SELECT DATE(snapshot_time) as date, COUNT(*) as total_assets
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) >= CURRENT_DATE - INTERVAL '30' DAY
GROUP BY DATE(snapshot_time)
ORDER BY date DESC;

Time travel — compare current vs 7 days ago (new descriptions added):

sql
SELECT t.asset_id, t.resource_name,
       p.business_description as before,
       t.business_description as now
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset" t
JOIN "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset" p
  ON t.asset_id = p.asset_id
WHERE DATE(t.snapshot_time) = CURRENT_DATE
  AND DATE(p.snapshot_time) = CURRENT_DATE - INTERVAL '7' DAY
  AND p.business_description IS NULL
  AND t.business_description IS NOT NULL;

Assets by owner:

sql
SELECT extended_metadata['owningEntityId'] as owner, COUNT(*) as count
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
  AND extended_metadata['owningEntityId'] IS NOT NULL
GROUP BY extended_metadata['owningEntityId']
ORDER BY count DESC;

Filter by metadata form field:

sql
SELECT *
FROM "s3tablescatalog/aws-sagemaker-catalog"."asset_metadata"."asset"
WHERE DATE(snapshot_time) = CURRENT_DATE
  AND extended_metadata['<metadata-form-name>.<field-name>'] = '<field-value>';
Show full SKILL.md (317 more words)Show less

Key Behaviors

  • Daily snapshots — exported around midnight per region
  • Always filter by snapshot_time — without it you get all history (duplicates, slow)
  • One domain per account per region — to switch domains, delete config first
  • No additional charge beyond S3 Tables storage + Athena queries
  • Read-only — to update asset metadata, use Glue Discovery APIs or SageMaker Unified Studio

Troubleshooting

ErrorCauseFix
aws-sagemaker-catalog bucket not foundExport not enabledRun put-data-export-configuration --enable-export
Empty results with CURRENT_DATEFirst export hasn't run yet (takes up to 24h)Wait; try yesterday's date
AccessDenied on queryMissing Lake Formation grantsGrant SELECT + DESCRIBE on the table
CATALOG_NOT_FOUNDS3 Tables not registered in GlueEnable integration: S3 console > Table buckets > Enable integration
Duplicate rows in resultsMissing snapshot_time filterAdd WHERE DATE(snapshot_time) = CURRENT_DATE
extended_metadata key returns NULLKey doesn't exist for that assetCheck available keys: SELECT DISTINCT key FROM ... CROSS JOIN UNNEST(map_keys(extended_metadata)) AS t(key) WHERE DATE(snapshot_time) = CURRENT_DATE
Cannot update export encryptionEncryption set at creation time onlyDelete and recreate export config

Security Considerations

Data sensitivity: Catalog metadata exposes organizational structure including asset names, ownership, account IDs, naming conventions, and internal resource identifiers. Treat query results as sensitive by default.

Encryption at rest: Always enable KMS encryption when creating the export configuration. Encryption cannot be changed after creation. Additionally, configure SSE-KMS on your Athena workgroup output bucket.

Least-privilege access: Grant Lake Formation SELECT + DESCRIBE only on the specific asset_metadata.asset table to roles that need catalog analytics. Avoid granting access to the entire aws-sagemaker-catalog bucket.

Audit trail: Enable CloudTrail logging for DataZone (PutDataExportConfiguration, GetDataExportConfiguration), Athena (StartQueryExecution, GetQueryResults), and S3 Tables API calls to track who queries catalog metadata.

Credential hygiene: Use IAM roles with temporary credentials for querying. Avoid long-lived access keys for users accessing catalog metadata. Scope down or rotate principals when access is no longer needed.

Additional Resources

© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/specialized-skills/system-table-skills/querying-aws-sagemaker-catalog of aws/agent-toolkit-for-aws.

Open the folder on GitHubat commit df2ab44

Compare with similar skills

Querying AWS Sagemaker Catalog next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Querying AWS Sagemaker Catalog compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Querying AWS Sagemaker Catalog this skillaws/agent-toolkit-for-aws2.8k—~2.6kAutomated safety check: PassApache-2.0
Glue DiagnosticsKilo-Org/kilo-marketplace190—~2kAutomated safety check: PassMIT
Querying Big Datasetsflyrank-bih/flyrank-ml-internship-starter140—~750Automated safety check: PassCustom licence
Performing Cloud Forensics With AWS Cloudtrailmukul975/Anthropic-Cybersecurity-Skills34k—~845Automated safety check: PassApache-2.0
Performing Cloud Log Forensics With Athenamukul975/Anthropic-Cybersecurity-Skills34k—~3.7kAutomated safety check: PassApache-2.0
Monte Carlo Preventsickn33/agentic-awesome-skills47k1 repos~3.3kAutomated safety check: PassMIT

Similar skills

  • Glue Diagnostics

    Kilo-Org/kilo-marketplace

    A skill your agent uses to investigate and troubleshoot AWS Glue problems by analyzing ETL jobs, crawlers, connections, Data Catalog, DPU utilization, Spark execution, and job bookmarks following…

    190 GitHub stars~2k tokensUpdated 10 days ago
    Data & AnalyticsAuto-check passed
  • Querying Big Datasets

    flyrank-bih/flyrank-ml-internship-starter

    Works with datasets far too big to download or load in pandas — SQL over remote Parquet with DuckDB, aggregate-then-model, iterate on samples.

    140 GitHub stars~750 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Performing Cloud Forensics With AWS Cloudtrail

    mukul975/Anthropic-Cybersecurity-Skills

    Investigate AWS account compromise by querying CloudTrail with boto3's LookupEvents or AWS Athena SQL over S3-delivered logs, filtering on suspicious user agents, source IPs, and event names to…

    34k GitHub stars~845 tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Performing Cloud Log Forensics With Athena

    mukul975/Anthropic-Cybersecurity-Skills

    Uses AWS Athena to query CloudTrail, VPC Flow Logs, S3 access logs, and ALB logs for forensic investigation.

    34k GitHub stars~3.7k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Monte Carlo Prevent

    sickn33/agentic-awesome-skills

    Surfaces Monte Carlo data observability context (table health, alerts, lineage, blast radius) before SQL/dbt edits.

    47k GitHub starsUsed in 1 repo~3.3k tokens
    Data & AnalyticsAuto-check passed
  • Data Quality Checks

    mohitagw15856/pm-claude-skills

    Design the data quality checks for a table or pipeline across the standard dimensions.

    1.4k GitHub stars~919 tokensUpdated today
    Data & AnalyticsAuto-check passed

More from aws/agent-toolkit-for-aws

All 138 skills in this repo
  • Agent Advisor

    aws/agent-toolkit-for-aws

    Official

    Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.

    2.8k GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Agents Build

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.

    2.8k GitHub stars~2.3k tokensUpdated today
    Auto-check: notes
  • Launch With AWS

    aws/agent-toolkit-for-aws

    Official

    Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.

    2.8k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.

    2.8k GitHub stars~4k tokensUpdated today
    Auto-check passed
  • AWS Marketplace Metering

    aws/agent-toolkit-for-aws

    Official

    Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…

    2.8k GitHub stars~18k tokensUpdated today
    Auto-check passed
  • Agents Pay

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.

    2.8k GitHub stars~6.5k tokensUpdated today
    Auto-check: notes

Questions about Querying AWS Sagemaker Catalog

What does Querying AWS Sagemaker Catalog do?

Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables. Querying AWS Sagemaker Catalog is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Runs SQL analytics on SageMaker Catalog asset metadata tables exported as Apache Iceberg in S3 Tables.

When should I use Querying AWS Sagemaker Catalog?

Querying AWS Sagemaker Catalog fits situations like: phrases: catalog inventory SQL; how many assets; assets without descriptions; asset growth over time.

How do I install Querying AWS Sagemaker Catalog in Claude Code?

Run `npx skills add aws/agent-toolkit-for-aws --skill querying-aws-sagemaker-catalog -a claude-code`. Or copy the skill folder (skills/specialized-skills/system-table-skills/querying-aws-sagemaker-catalog in aws/agent-toolkit-for-aws) into .claude/skills/querying-aws-sagemaker-catalog in your project. Claude Code loads it when a task matches its description.

How do I install Querying AWS Sagemaker Catalog in Codex?

Run `npx skills add aws/agent-toolkit-for-aws --skill querying-aws-sagemaker-catalog -a codex`. Or copy the skill folder (skills/specialized-skills/system-table-skills/querying-aws-sagemaker-catalog in aws/agent-toolkit-for-aws) into .agents/skills/querying-aws-sagemaker-catalog in your project. Codex loads it when a task matches its description.

Can I use Querying AWS Sagemaker Catalog in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill querying-aws-sagemaker-catalog -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/querying-aws-sagemaker-catalog, .gemini/skills/querying-aws-sagemaker-catalog, .github/skills/querying-aws-sagemaker-catalog and .opencode/skills/querying-aws-sagemaker-catalog in your project.

What does Querying AWS Sagemaker Catalog need to run?

Going by SKILL.md and its folder, Querying AWS Sagemaker Catalog needs the command-line tools its instructions call (aws).

Does Querying AWS Sagemaker Catalog access the network?

SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.

Is Querying AWS Sagemaker Catalog safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Querying AWS Sagemaker Catalog use?

Querying AWS Sagemaker Catalog is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Querying AWS Sagemaker Catalog use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Querying AWS Sagemaker Catalog?

Skills that share tags, products or a category with Querying AWS Sagemaker Catalog: Glue Diagnostics (Kilo-Org/kilo-marketplace, 190 stars), Querying Big Datasets (flyrank-bih/flyrank-ml-internship-starter, 140 stars), Performing Cloud Forensics With AWS Cloudtrail (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Performing Cloud Log Forensics With Athena (mukul975/Anthropic-Cybersecurity-Skills, 34k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Querying AWS Sagemaker Catalog?

aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,830 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 9, 2026.

Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.