Official agent skill

Creating Data Lake Table

by aws in aws/agent-toolkit-for-aws

Create managed Iceberg tables using Amazon S3 Tables (s3tables API namespace) with automatic compaction and snapshot management.

OfficialApache-2.0Auto-check passedBackend & APIs

Install Creating Data Lake Table

skills CLI
$ npx skills add aws/agent-toolkit-for-aws --skill creating-data-lake-table -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws/agent-toolkit-for-aws creating-data-lake-table --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/aws-data-analytics/skills/creating-data-lake-table .claude/skills/creating-data-lake-table && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
creating-data-lake-table
GitHub stars
2.8k
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
735 words
Files
5 (incl. references)
Skills in repo
138
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create managed Iceberg tables using Amazon S3 Tables (s3tables API namespace) with automatic compaction and snapshot management.

  • Works in 8 steps: Verify Dependencies → Understand the Schema → Create Table Bucket → …
  • Data lake table
  • SKILL.md covers Overview, Common Tasks, Decision Guide and Troubleshooting, plus 1 more section
  • Calls aws

What it does

Creating Data Lake Table is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Create managed Iceberg tables using Amazon S3 Tables (s3tables API namespace) with automatic compaction and snapshot management. Sets up table bucket, namespace, table, schema, Glue catalog registration, partitioning, IAM access control. Triggers on: create table, data lake table, analytics table, structured data storage, S3 Tables, Iceberg, Athena table, partitioning strategy, access permissions. Do NOT use for: importing files (use ingesting-into-data-lake), vector storage (use storing-and-querying-vectors)…

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/access-control.md`, `references/athena-ddl-path.md` and `references/best-practices.md`).

It sits in Backend & APIs, covering Authorization and RBAC, File uploads and storage and Schema markup. It works with Amazon S3, Amazon Web Services and Model Context Protocol. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.

When your agent uses it

  • Data lake table
  • Analytics table
  • Structured data storage
  • Partitioning strategy

Example prompts

  • “/creating-data-lake-table”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Verify Dependencies
  2. Understand the Schema
  3. Create Table Bucket
  4. Create Namespace
  5. Create Glue Data Catalog Integration
  6. Configure Access Control
  7. Create the Table
  8. Verify and Confirm

What it can do on your machine

Read from SKILL.md and the folder at commit bd49cc8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Creating Data Lake Table loads about 2.1k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 163 tokens; SKILL.md has 735 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~163
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aws/agent-toolkit-for-aws at commit bd49cc8, republished under its Apache-2.0 licence (© aws). 735 words, ~2,073 tokens.

Download SKILL.mdSave it as .claude/skills/creating-data-lake-table/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
creating-data-lake-table
description
Create managed Iceberg tables using Amazon S3 Tables (s3tables API namespace) with automatic compaction and snapshot management. Sets up table bucket, namespace, table, schema, Glue catalog registration, partitioning, IAM access control. Triggers on: create table, data lake table, analytics table, structured data storage, S3 Tables, Iceberg, Athena table, partitioning strategy, access permissions. Do NOT use for: importing files (use ingesting-into-data-lake), vector storage (use storing-and-querying-vectors), querying existing tables (use querying-data-lake), or locating existing table (use finding-data-lake-assets).
metadata.version
1
metadata.argument-hint
'[table-description|schema-spec]'

Create Data Lake Tables with Amazon S3 Tables

Overview

Amazon S3 Tables provides managed Iceberg tables with automatic compaction and snapshot management. Queryable via Athena and Iceberg-compatible engines.

Common Tasks

You MUST use AWS MCP server tools when connected, they provide command validation, sandboxed execution, and audit logging. Fall back to AWS CLI if MCP unavailable.

Decision Guide

Before creating, You MUST check what exists:

You MUST run aws glue get-tables --database-name <NAME> when user mentions a database.

What you findAction
Fuzzy database name ("our analytics db")You MUST STOP. Delegate to finding-data-lake-assets to resolve.
Non-S3-Tables table with matching nameYou MUST STOP. Delegate to finding-data-lake-assets. You MUST NOT create until user confirms.
Existing S3 Tables table with matching nameYou MUST check schema match. Reuse if compatible, recreate only if user confirms.
No matching tablesProceed with creation (Steps 1-8).
User explicitly requests new S3 Tables tableSkip checks, proceed with creation.

Creation paths:

  • Existing data in S3: Create empty table (Steps 1-8), then use ingesting-into-data-lake skill.
  • Glue ETL pipeline: Read references/table-creation-glue-etl.md first, then Steps 1-6.
  • Lake Formation access control: Search AWS docs for "S3 Tables integration with Lake Formation".
1. Verify Dependencies

Constraints:

  • You MUST check whether AWS MCP server tools or AWS CLI are available and inform user if missing
  • You MUST confirm target AWS region and verify credentials with aws sts get-caller-identity
2. Understand the Schema
  • Explicit schema: Validate Iceberg types.
  • Loose description: Ask columns, types, grain. Propose and confirm.
  • Existing S3 data: Infer schema from file headers only. Create empty table first, then use ingesting-into-data-lake skill.

Constraints:

  • You MUST read references/best-practices.md for Iceberg type mapping, partitions, and naming.
  • You MUST ask for all required parameters upfront: table name, columns, types, partition strategy. For schema evolution, see references/athena-ddl-path.md.
  • You MUST use all lowercase names -- Glue rejects mixed case with GENERIC_INTERNAL_ERROR. Namespace and table names MUST NOT contain hyphens.
  • You SHOULD suggest partition columns based on access patterns.
3. Create Table Bucket

Names: 3-63 chars, lowercase, numbers, hyphens.

bash
aws s3tables create-table-bucket --name <BUCKET_NAME> --region <REGION>

Capture table-bucket-arn. Encryption (SSE-S3 default, SSE-KMS) and storage class (STANDARD, INTELLIGENT_TIERING) set at creation. See references/best-practices.md.

Constraints:

  • You MUST check existing buckets with aws s3tables list-table-buckets and ask user to select or create new.
  • If using SSE-KMS, KMS key policy MUST allow S3 Tables maintenance service principal to read data. Search AWS docs for "S3 Tables KMS key policy" for required policy.
  • If bucket creation fails, see references/best-practices.md for common errors.
4. Create Namespace
bash
aws s3tables create-namespace --table-bucket-arn <ARN> --namespace <NAMESPACE>

Constraints:

  • You MUST list existing namespaces first and suggest reusing if relevant
  • You MUST use lowercase names with no hyphens
Show full SKILL.md (309 more words)Show less
5. Create Glue Data Catalog Integration

Check if s3tablescatalog exists (create once per region per account):

bash
aws glue get-catalog --catalog-id s3tablescatalog

If not found, create (requires glue:CreateCatalog, glue:passConnection):

bash
aws glue create-catalog --name "s3tablescatalog" --catalog-input '{
  "FederatedCatalog": {
    "Identifier": "arn:aws:s3tables:<REGION>:<ACCOUNT_ID>:bucket/*",
    "ConnectionName": "aws:s3tables"
  },
  "CreateDatabaseDefaultPermissions": [{"Principal": {"DataLakePrincipalIdentifier": "IAM_ALLOWED_PRINCIPALS"}, "Permissions": ["ALL"]}],
  "CreateTableDefaultPermissions": [{"Principal": {"DataLakePrincipalIdentifier": "IAM_ALLOWED_PRINCIPALS"}, "Permissions": ["ALL"]}],
  "AllowFullTableExternalDataAccess": "True"
}'

Verify with aws glue get-catalogs --parent-catalog-id s3tablescatalog.

6. Configure Access Control

S3 Tables uses s3tables:* IAM namespace (not s3:*).

Querying principal permissions (bucket policy):

  • s3tables:GetTableBucket, s3tables:GetNamespace, s3tables:GetTable, s3tables:GetTableMetadataLocation, s3tables:GetTableData

Querying principal permissions (IAM policy):

  • glue:GetCatalog, glue:GetDatabase, glue:GetTable

You MUST scope to correct ARN patterns. You MUST read references/access-control.md for exact resource ARNs.

Constraints:

  • You MUST ask user for querying principal ARN
  • You MUST NOT grant broader permissions than necessary
  • You MUST NOT create IAM roles automatically, verify existing and guide user
7. Create the Table
ContextPath
Default (any user)S3 Tables API (below)
User specifically wants SQL DDLAthena DDL (see references/athena-ddl-path.md)
Glue ETL pipelineSpark DDL via --conf job args (not spark.conf.set()). You MUST read references/table-creation-glue-etl.md for the --conf string.

Default: S3 Tables API:

bash
aws s3tables create-table \
  --table-bucket-arn <ARN> \
  --namespace <NAMESPACE> \
  --name <TABLE_NAME> \
  --format ICEBERG \
  --metadata '<METADATA_JSON>'

Metadata JSON MUST nest under "iceberg" key:

json
{"iceberg":{"schema":{"fields":[
  {"name":"order_date","type":"date","required":true},
  {"name":"customer_id","type":"string","required":true},
  {"name":"amount","type":"double","required":false}
]},
"partitionSpec":{"fields":[
  {"sourceId":1,"fieldId":1000,"transform":"month","name":"order_date_month"}
]}}}

Constraints:

  • partitionSpec.sourceId MUST reference a valid schema field ID
  • For schema evolution after creation, use Athena DDL. See references/athena-ddl-path.md
  • You MUST use schemaV2 for complex types (list, map, struct) with explicit field IDs. See references/best-practices.md.
  • You SHOULD search AWS docs for "IcebergPartitionField S3 Tables" for supported partition transforms
8. Verify and Confirm

You MUST verify with aws s3tables get-table and confirm queryability with DESCRIBE <table_name> via Athena using --query-execution-context '{"Catalog":"s3tablescatalog/<BUCKET_NAME>","Database":"<NAMESPACE>"}'. Do NOT put catalog in SQL. Present summary: bucket ARN, namespace, table, schema, partitions.

Troubleshooting

ErrorCauseFix
"Table location can not be specified"LOCATION in CREATE TABLERemove LOCATION clause. S3 Tables manages storage automatically.
AccessDeniedException with s3:* policyUsing s3:* not s3tables:*S3 Tables uses s3tables:* namespace. Update IAM policy.

Additional Resources

© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in plugins/aws-data-analytics/skills/creating-data-lake-table of aws/agent-toolkit-for-aws.

  • SKILL.md
  • references/access-control.md
  • references/athena-ddl-path.md
  • references/best-practices.md
  • references/table-creation-glue-etl.md

Open the folder on GitHubat commit bd49cc8

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in aws/agent-toolkit-for-aws, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Creating Data Lake Table next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Creating Data Lake Table compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Creating Data Lake Table this skillaws/agent-toolkit-for-aws2.8k1 repos~2.1kAutomated safety check: PassApache-2.0
S3itsmostafa/aws-agent-skills1.2k—~2.3kAutomated safety check: PassMIT
AWS Essentialsericrisco/rsc-harness156—~2.9kAutomated safety check: NotesMIT
FoundatioFoundatioFx/Foundatio2.1k—~3.9kAutomated safety check: PassApache-2.0
Django Storages for S3Jeffallan/claude-skills12k—~1.9kAutomated safety check: PassMIT
AWS S3sickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: PassMIT

Similar skills

  • S3

    itsmostafa/aws-agent-skills

    AWS S3 object storage for bucket management, object operations, and access control.

    1.2k GitHub stars~2.3k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • AWS Essentials

    ericrisco/rsc-harness

    A skill your agent uses when standing up the core AWS surface a small product needs: hardening a fresh account, a private S3 bucket, encrypted RDS Postgres, ECS Fargate vs EC2, CloudFront + OAC, or…

    156 GitHub stars~2.9k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Foundatio

    FoundatioFx/Foundatio

    A skill your agent uses when working with Foundatio infrastructure abstractions for .NET -- caching, queuing, messaging, file storage, distributed locking, or background jobs.

    2.1k GitHub stars~3.9k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Django Storages for S3

    Jeffallan/claude-skills

    Sets up Django 4.2+ to keep static and media files on AWS S3 through django-storages, with public and private backends, presigned URLs and CloudFront.

    12k GitHub stars~1.9k tokensUpdated 4 days ago
    Backend & APIsAuto-check passed
  • AWS S3

    sickn33/agentic-awesome-skills

    Configure S3 buckets, policies, and lifecycle rules. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~3.1k tokens
    Backend & APIsAuto-check passed
  • Remediating S3 Bucket Misconfiguration

    mukul975/Anthropic-Cybersecurity-Skills

    Provides step-by-step procedures for remediating Amazon S3 bucket misconfigurations that expose sensitive data: enabling S3 Block Public Access, auditing bucket policies and ACLs, enforcing…

    34k GitHub stars~3k tokensUpdated 1 mo ago
    Backend & APIsAuto-check passed

More from aws/agent-toolkit-for-aws

All 138 skills in this repo
  • Agent Advisor

    aws/agent-toolkit-for-aws

    Official

    Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.

    2.8k GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Agents Build

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.

    2.8k GitHub stars~2.3k tokensUpdated today
    Auto-check: notes
  • Launch With AWS

    aws/agent-toolkit-for-aws

    Official

    Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.

    2.8k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.

    2.8k GitHub stars~4k tokensUpdated today
    Auto-check passed
  • AWS Marketplace Metering

    aws/agent-toolkit-for-aws

    Official

    Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…

    2.8k GitHub stars~18k tokensUpdated today
    Auto-check passed
  • Agents Pay

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.

    2.8k GitHub stars~6.5k tokensUpdated today
    Auto-check: notes

Categories

Questions about Creating Data Lake Table

What does Creating Data Lake Table do?

Create managed Iceberg tables using Amazon S3 Tables (s3tables API namespace) with automatic compaction and snapshot management. Creating Data Lake Table is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Create managed Iceberg tables using Amazon S3 Tables (s3tables API namespace) with automatic compaction and snapshot management.

When should I use Creating Data Lake Table?

Creating Data Lake Table fits situations like: data lake table; analytics table; structured data storage; partitioning strategy.

How do I install Creating Data Lake Table in Claude Code?

Run `npx skills add aws/agent-toolkit-for-aws --skill creating-data-lake-table -a claude-code`. Or copy the skill folder (plugins/aws-data-analytics/skills/creating-data-lake-table in aws/agent-toolkit-for-aws) into .claude/skills/creating-data-lake-table in your project. Claude Code loads it when a task matches its description.

How do I install Creating Data Lake Table in Codex?

Run `npx skills add aws/agent-toolkit-for-aws --skill creating-data-lake-table -a codex`. Or copy the skill folder (plugins/aws-data-analytics/skills/creating-data-lake-table in aws/agent-toolkit-for-aws) into .agents/skills/creating-data-lake-table in your project. Codex loads it when a task matches its description.

Can I use Creating Data Lake Table in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill creating-data-lake-table -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/creating-data-lake-table, .gemini/skills/creating-data-lake-table, .github/skills/creating-data-lake-table and .opencode/skills/creating-data-lake-table in your project.

What does Creating Data Lake Table need to run?

Going by SKILL.md and its folder, Creating Data Lake Table needs the command-line tools its instructions call (aws).

Does Creating Data Lake Table access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Creating Data Lake Table safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Creating Data Lake Table use?

Creating Data Lake Table is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Creating Data Lake Table use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Creating Data Lake Table?

Skills that share tags, products or a category with Creating Data Lake Table: S3 (itsmostafa/aws-agent-skills, 1.2k stars), AWS Essentials (ericrisco/rsc-harness, 156 stars), Foundatio (FoundatioFx/Foundatio, 2.1k stars) and Django Storages for S3 (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Creating Data Lake Table?

aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,816 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 7, 2026.

Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.