Official agent skill

Ingesting Into Data Lake

by aws in aws/agent-toolkit-for-aws

Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue…

OfficialApache-2.0Auto-check passedDatabases

Install Ingesting Into Data Lake

skills CLI
$ npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws/agent-toolkit-for-aws ingesting-into-data-lake --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/aws-data-analytics/skills/ingesting-into-data-lake .claude/skills/ingesting-into-data-lake && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ingesting-into-data-lake
GitHub stars
2.8k
Used in
1 other repo
Token cost
~2.8k tokens
SKILL.md length
1,064 words
Files
26 (incl. references)
Skills in repo
138
Repo updated
First seen
Licence
Apache-2.0

At a glance

Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue…

  • Works in 7 steps: Verify Dependencies and Context → Classify the Source → Confirm Connection Exists (if applicable) → …
  • Move data to AWS
  • SKILL.md covers Philosophy, Common Tasks, Workflow and Argument Routing, plus 3 more sections
  • Calls aws

What it does

Ingesting Into Data Lake is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration). Default target is S3 Tables; standard Iceberg on a general purpose bucket is supported where S3 Tables is not adopted. Handles one-time loads, recurring pipelines, migrations. Triggers on: import data, load data, ingest, sync database, migrate table, move data to AWS, set up pipeline, ETL…

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including reference files (for example `references/athena-loading.md`, `references/bigquery-ingest.md` and `references/catalog-migration.md`).

It sits in Databases, covering Data warehousing, NoSQL databases and File uploads and storage. It works with Amazon Web Services, Amazon DynamoDB, Snowflake and Google BigQuery. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.

When your agent uses it

  • Move data to AWS
  • Set up pipeline
  • Pull from Snowflake
  • Query BigQuery into S3

Example prompts

  • “/ingesting-into-data-lake”

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Verify Dependencies and Context
  2. Classify the Source
  3. Confirm Connection Exists (if applicable)
  4. Clarify the Target
  5. Execute Source Workflow
  6. Validate
  7. Schedule (if recurring)

What it can do on your machine

Read from SKILL.md and the folder at commit bd49cc8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ingesting Into Data Lake loads about 2.8k tokens when it runs, and up to ~50k if it reads all its reference files. Until then it costs about 249 tokens; SKILL.md has 1,064 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~249
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~50k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aws/agent-toolkit-for-aws at commit bd49cc8, republished under its Apache-2.0 licence (© aws). 1,064 words, ~2,769 tokens.

Download SKILL.mdSave it as .claude/skills/ingesting-into-data-lake/SKILL.md (or your agent's skills folder). This skill also uses 25 other files; get the full folder from GitHub.
name
ingesting-into-data-lake
description
Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration). Default target is S3 Tables; standard Iceberg on a general purpose bucket is supported where S3 Tables is not adopted. Handles one-time loads, recurring pipelines, migrations. Triggers on: import data, load data, ingest, sync database, migrate table, move data to AWS, set up pipeline, ETL, pull from Snowflake, query BigQuery into S3, export DynamoDB, CTAS, convert to Iceberg. Do NOT use for setting up or troubleshooting Glue connections (use connecting-to-data-source), creating empty tables (use creating-data-lake-table), running queries (use querying-data-lake), finding tables by fuzzy name (use finding-data-lake-assets), catalog audit (use exploring-data-catalog), or SaaS platforms like Salesforce, ServiceNow, SAP, MongoDB, Kafka.
metadata.version
1
metadata.argument-hint
'[source-path|connection-name|table-name] [--target s3-tables|iceberg|parquet]'

Ingest into Data Lake

Move data from a source into a queryable table in the data lake. This skill assumes the source connection (if one is needed) already exists. For Glue connection setup or troubleshooting, delegate to connecting-to-data-source.

Philosophy

Default to S3 Tables unless the environment says otherwise. S3 Tables is the recommended target for new data lake work. If the user's catalog inventory shows they haven't adopted S3 Tables, recommend standard Iceberg on their existing general-purpose bucket instead of forcing them to change posture.

Common Tasks

You MUST execute commands using AWS MCP server tools when connected -- they provide validation, sandboxed execution, and audit logging. Fall back to AWS CLI only if MCP is unavailable. You MUST explain each step before executing.

Workflow

1. Verify Dependencies and Context
  • You MUST check whether AWS MCP tools or AWS CLI are available and inform the user if missing
  • You MUST confirm target AWS region and verify credentials with aws sts get-caller-identity
  • For SageMaker Unified Studio project roles, note that target tables and connections may be scoped to the project. See the caller ARN detection pattern in querying-data-lake.
2. Classify the Source
User says...Source typeReference
"upload my file", "local CSV", "move to S3"Local filelocal-upload.md
"load from S3", "import CSV/JSON/Parquet from s3://"S3 filess3-files.md
"import from Oracle/Postgres/MySQL/SQL Server/Redshift/RDS/Aurora"JDBCjdbc-ingest.md
"pull from Snowflake", "Snowflake table to S3"Snowflakesnowflake-ingest.md
"import from BigQuery", "GCP analytics to S3"BigQuerybigquery-ingest.md
"export DynamoDB", "DynamoDB to data lake"DynamoDBdynamodb-ingest.md
"migrate Glue table", "convert Hive to Iceberg"Catalog migrationcatalog-migration.md

If the user names Salesforce, ServiceNow, SAP, MongoDB, Kafka, or another SaaS/streaming source, decline -- these are not supported in this release.

If the source table is referenced by a fuzzy or business name ("migrate our orders table", "pull from the sales warehouse"), delegate to finding-data-lake-assets to resolve before proceeding.

3. Confirm Connection Exists (if applicable)

For JDBC, Snowflake, and BigQuery sources, a Glue connection is required. Check:

bash
aws glue get-connection --name <CONNECTION_NAME> --region <REGION>

If the connection does not exist, stop and delegate to connecting-to-data-source to create and test it. Do not proceed with ingest until the connection is verified.

Local files, S3 files, DynamoDB, and catalog migration do not need a Glue connection.

4. Clarify the Target

You MUST ask the user (or suggest based on catalog inventory) before creating or writing to any table:

  • Database/namespace: Does a specific target database exist? Or should one be created?
  • Table: Existing table (append/merge) or new table (delegate to creating-data-lake-table)?
  • Format: S3 Tables (default), standard Iceberg, or raw Parquet?

Inventory-aware defaults:

If you have already run exploring-data-catalog or can quickly check, use what exists:

  • Account has an s3tablescatalog federated catalog and active table buckets: recommend S3 Tables
  • Account has general-purpose buckets with Iceberg tables and no S3 Tables usage: recommend standard Iceberg on their existing bucket
  • Account uses Parquet/ORC on S3 without Iceberg metadata: ask whether to adopt Iceberg now (recommend yes) or continue with raw files

Do not force S3 Tables on customers who haven't adopted it. See iceberg-catalog-config-and-usage.md.

Delegations from this step:

  • Target table doesn't exist -> creating-data-lake-table
  • Target database named by fuzzy term -> finding-data-lake-assets
  • User doesn't know what exists -> exploring-data-catalog
5. Execute Source Workflow

Read the source-specific reference and follow its phases. Each is self-contained with job templates, gotchas, and troubleshooting:

  • Local / S3 / JDBC / Snowflake / BigQuery / DynamoDB / catalog migration -- one reference per source

Common Glue 5.1 or higher job configuration and PySpark templates are shared in glue-job-config.md and glue-job-scripts.md.

6. Validate

Run all three, do not skip:

  1. Row count matches expected (source vs target)
  2. Null check on critical columns
  3. Spot-check 3-5 sample rows

See data-quality-validation.md.

7. Schedule (if recurring)

For recurring pipelines, create a Glue Trigger with a cron schedule. See testing-and-scheduling.md. Simple single-step pipelines use Glue Triggers; multi-step with branching uses MWAA.

Show full SKILL.md (437 more words)Show less

Argument Routing

  • S3 path only: Infer one-time load, start Step 2 with S3 files
  • Connection name: Start Step 3 with the named connection
  • Table name: Start Step 4, ask whether this is source or target
  • --target flag: Pre-fill the target format in Step 4
  • No args: Walk through interactively

Gotchas

  • S3 Tables requires Glue 5.1 or higher and --datalake-formats iceberg job argument
  • All spark.sql.catalog.* config MUST go in --conf job arguments, never in spark.conf.set(). Glue 5.x throws AnalysisException: Cannot modify the value of a static config otherwise. See iceberg-catalog-config-and-usage.md for correct catalog configs.
  • The warehouse parameter is required in S3 Tables catalog config. Without it Spark fails with "Cannot derive default warehouse location".
  • Table and column names in S3 Tables MUST be all lowercase
  • overwritePartitions() only replaces partitions present in the DataFrame -- for full refresh with deletes, use createOrReplace()
  • Standard Iceberg targets MUST include a LOCATION clause; S3 Tables MUST NOT
  • DynamoDB does not need a Glue connection -- do not attempt to create one
  • Connection failures during ingest delegate back to connecting-to-data-source; do not debug network/credentials in this skill
  • For target tables in SageMaker Unified Studio projects, ensure the project role has write access to the target namespace before the Glue job runs

Troubleshooting

ErrorLikely causeAction
Access Denied on S3Missing IAM permissionsCheck Glue role has s3:GetObject, s3:PutObject
Access Denied on S3 TablesMissing s3tables:* permissionsAdd S3 Tables inline policy to Glue role
CTAS timeoutDataset too large for AthenaSwitch to Glue ETL or batch with WHERE filters
JDBC connection timeout/auth failureConnection-level issueDelegate to connecting-to-data-source
Throughput exceeded (DynamoDB)Read percent too highLower read.percent or use native export

See error-handling.md for the full catalog.

References

Source-specific
Cross-cutting
Migration-specific
JDBC-specific

© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 25 other files (references) in plugins/aws-data-analytics/skills/ingesting-into-data-lake of aws/agent-toolkit-for-aws.

  • SKILL.md
  • references/athena-loading.md
  • references/bigquery-ingest.md
  • references/catalog-migration.md
  • references/ctas-patterns.md
  • references/data-quality-validation.md
  • references/dynamodb-ingest.md
  • references/error-handling.md
  • references/format-specific-loading.md
  • references/glue-etl-migration.md
  • references/glue-job-config.md
  • references/glue-job-scripts.md
  • references/iceberg-catalog-config-and-usage.md
  • references/incremental-loading.md
  • references/jdbc-ingest.md
  • references/jdbc-performance.md
  • references/jdbc-schema-discovery.md
  • references/local-upload.md
  • references/migration-troubleshooting.md
  • references/migration-validation.md
  • … and 6 more

Open the folder on GitHubat commit bd49cc8

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in aws/agent-toolkit-for-aws, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Ingesting Into Data Lake next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ingesting Into Data Lake compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ingesting Into Data Lake this skillaws/agent-toolkit-for-aws2.8k1 repos~2.8kAutomated safety check: PassApache-2.0
Mfs Findzilliztech/mfs150—~4kAutomated safety check: PassApache-2.0
Mfs Ingestzilliztech/mfs150—~4.7kAutomated safety check: PassApache-2.0
AWS CLI Beastgiuseppe-trisciuoglio/developer-kit355—~1.7kAutomated safety check: NotesMIT
AWSkid-sid/claude-spellbook189—~4.9kAutomated safety check: WarnMIT
SQL Query Explainermohitagw15856/pm-claude-skills1.4k—~1.6kAutomated safety check: PassMIT

Similar skills

  • Mfs Find

    zilliztech/mfs

    Search, grep, browse, and read across registered MFS data sources via the mfs CLI — codebases, docs, PDFs, web crawls, databases (postgres/mysql/mongo/snowflake/bigquery), issue trackers…

    150 GitHub stars~4k tokensUpdated 2 mo ago
    DatabasesAuto-check passed
  • Mfs Ingest

    zilliztech/mfs

    Register, update, or re-sync data sources for MFS so they become searchable — postgres / mysql / mongo / snowflake / bigquery, github / jira / linear / notion / hubspot / zendesk, slack / discord /…

    150 GitHub stars~4.7k tokensUpdated 2 mo ago
    DatabasesAuto-check passed
  • AWS CLI Beast

    giuseppe-trisciuoglio/developer-kit

    Provides advanced AWS CLI patterns for managing EC2, Lambda, S3, DynamoDB, RDS, VPC, IAM, and CloudWatch.

    355 GitHub stars~1.7k tokensUpdated 27 days ago
    DatabasesAuto-check: notes
  • AWS

    kid-sid/claude-spellbook

    A skill your agent uses when writing boto3 or AWS SDK v3 code — configuring IAM auth, reading/writing S3, designing DynamoDB access patterns, writing Lambda handlers, processing SQS batches, or…

    189 GitHub stars~4.9k tokensUpdated 2 mo ago
    DatabasesAuto-check: warnings
  • SQL Query Explainer

    mohitagw15856/pm-claude-skills

    Explains, optimises, writes, and documents SQL queries. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~1.6k tokensUpdated yesterday
    DatabasesAuto-check passed
  • AWS Cloud Patterns

    rohitg00/awesome-claude-code-toolkit

    AWS cloud patterns for Lambda, ECS, S3, DynamoDB, and Infrastructure as Code with CDK/Terraform

    2.7k GitHub stars~1.1k tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed

More from aws/agent-toolkit-for-aws

All 138 skills in this repo
  • Agent Advisor

    aws/agent-toolkit-for-aws

    Official

    Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.

    2.8k GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Agents Build

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.

    2.8k GitHub stars~2.3k tokensUpdated today
    Auto-check: notes
  • Launch With AWS

    aws/agent-toolkit-for-aws

    Official

    Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.

    2.8k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.

    2.8k GitHub stars~4k tokensUpdated today
    Auto-check passed
  • AWS Marketplace Metering

    aws/agent-toolkit-for-aws

    Official

    Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…

    2.8k GitHub stars~18k tokensUpdated today
    Auto-check passed
  • Agents Pay

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.

    2.8k GitHub stars~6.5k tokensUpdated today
    Auto-check: notes

Categories

Questions about Ingesting Into Data Lake

What does Ingesting Into Data Lake do?

Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue…. Ingesting Into Data Lake is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration).

When should I use Ingesting Into Data Lake?

Ingesting Into Data Lake fits situations like: move data to AWS; set up pipeline; pull from Snowflake; query BigQuery into S3.

How do I install Ingesting Into Data Lake in Claude Code?

Run `npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a claude-code`. Or copy the skill folder (plugins/aws-data-analytics/skills/ingesting-into-data-lake in aws/agent-toolkit-for-aws) into .claude/skills/ingesting-into-data-lake in your project. Claude Code loads it when a task matches its description.

How do I install Ingesting Into Data Lake in Codex?

Run `npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a codex`. Or copy the skill folder (plugins/aws-data-analytics/skills/ingesting-into-data-lake in aws/agent-toolkit-for-aws) into .agents/skills/ingesting-into-data-lake in your project. Codex loads it when a task matches its description.

Can I use Ingesting Into Data Lake in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ingesting-into-data-lake, .gemini/skills/ingesting-into-data-lake, .github/skills/ingesting-into-data-lake and .opencode/skills/ingesting-into-data-lake in your project.

What does Ingesting Into Data Lake need to run?

Going by SKILL.md and its folder, Ingesting Into Data Lake needs the command-line tools its instructions call (aws).

Does Ingesting Into Data Lake access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ingesting Into Data Lake safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ingesting Into Data Lake use?

Ingesting Into Data Lake is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ingesting Into Data Lake use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 47k tokens, read only when the agent opens those files.

What are the alternatives to Ingesting Into Data Lake?

Skills that share tags, products or a category with Ingesting Into Data Lake: Mfs Find (zilliztech/mfs, 150 stars), Mfs Ingest (zilliztech/mfs, 150 stars), AWS CLI Beast (giuseppe-trisciuoglio/developer-kit, 355 stars) and AWS (kid-sid/claude-spellbook, 189 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ingesting Into Data Lake?

aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,816 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 7, 2026.

Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.