Mfs Find
zilliztech/mfs
Search, grep, browse, and read across registered MFS data sources via the mfs CLI — codebases, docs, PDFs, web crawls, databases (postgres/mysql/mongo/snowflake/bigquery), issue trackers…
Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue…
$ npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aws/agent-toolkit-for-aws ingesting-into-data-lake --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/aws-data-analytics/skills/ingesting-into-data-lake .claude/skills/ingesting-into-data-lake && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ingesting-into-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/ingesting-into-data-lake into .claude/skills/ingesting-into-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ingesting-into-data-lake", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/ingesting-into-data-lakeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aws/agent-toolkit-for-aws ingesting-into-data-lake --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/aws-data-analytics/skills/ingesting-into-data-lake .agents/skills/ingesting-into-data-lake && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ingesting-into-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/ingesting-into-data-lake into .agents/skills/ingesting-into-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ingesting-into-data-lake", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aws/agent-toolkit-for-aws ingesting-into-data-lake --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/aws-data-analytics/skills/ingesting-into-data-lake .cursor/skills/ingesting-into-data-lake && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ingesting-into-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/ingesting-into-data-lake into .cursor/skills/ingesting-into-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ingesting-into-data-lake", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aws/agent-toolkit-for-aws.git --path plugins/aws-data-analytics/skills/ingesting-into-data-lake--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aws/agent-toolkit-for-aws ingesting-into-data-lake --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/aws-data-analytics/skills/ingesting-into-data-lake .gemini/skills/ingesting-into-data-lake && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ingesting-into-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/ingesting-into-data-lake into .gemini/skills/ingesting-into-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ingesting-into-data-lake", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aws/agent-toolkit-for-aws ingesting-into-data-lakeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/aws-data-analytics/skills/ingesting-into-data-lake .github/skills/ingesting-into-data-lake && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ingesting-into-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/ingesting-into-data-lake into .github/skills/ingesting-into-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ingesting-into-data-lake", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aws/agent-toolkit-for-aws ingesting-into-data-lake --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/aws-data-analytics/skills/ingesting-into-data-lake .opencode/skills/ingesting-into-data-lake && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ingesting-into-data-lake" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/plugins/aws-data-analytics/skills/ingesting-into-data-lake into .opencode/skills/ingesting-into-data-lake/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ingesting-into-data-lake", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ingesting-into-data-lakeImport data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue…
Ingesting Into Data Lake is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration). Default target is S3 Tables; standard Iceberg on a general purpose bucket is supported where S3 Tables is not adopted. Handles one-time loads, recurring pipelines, migrations. Triggers on: import data, load data, ingest, sync database, migrate table, move data to AWS, set up pipeline, ETL…
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including reference files (for example `references/athena-loading.md`, `references/bigquery-ingest.md` and `references/catalog-migration.md`).
It sits in Databases, covering Data warehousing, NoSQL databases and File uploads and storage. It works with Amazon Web Services, Amazon DynamoDB, Snowflake and Google BigQuery. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit bd49cc8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
awsFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ingesting Into Data Lake loads about 2.8k tokens when it runs, and up to ~50k if it reads all its reference files. Until then it costs about 249 tokens; SKILL.md has 1,064 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from aws/agent-toolkit-for-aws at commit bd49cc8, republished under its Apache-2.0 licence (© aws). 1,064 words, ~2,769 tokens.
.claude/skills/ingesting-into-data-lake/SKILL.md (or your agent's skills folder). This skill also uses 25 other files; get the full folder from GitHub.Move data from a source into a queryable table in the data lake. This skill assumes the source connection (if one is needed) already exists. For Glue connection setup or troubleshooting, delegate to connecting-to-data-source.
Default to S3 Tables unless the environment says otherwise. S3 Tables is the recommended target for new data lake work. If the user's catalog inventory shows they haven't adopted S3 Tables, recommend standard Iceberg on their existing general-purpose bucket instead of forcing them to change posture.
You MUST execute commands using AWS MCP server tools when connected -- they provide validation, sandboxed execution, and audit logging. Fall back to AWS CLI only if MCP is unavailable. You MUST explain each step before executing.
aws sts get-caller-identityquerying-data-lake.| User says... | Source type | Reference |
|---|---|---|
| "upload my file", "local CSV", "move to S3" | Local file | local-upload.md |
| "load from S3", "import CSV/JSON/Parquet from s3://" | S3 files | s3-files.md |
| "import from Oracle/Postgres/MySQL/SQL Server/Redshift/RDS/Aurora" | JDBC | jdbc-ingest.md |
| "pull from Snowflake", "Snowflake table to S3" | Snowflake | snowflake-ingest.md |
| "import from BigQuery", "GCP analytics to S3" | BigQuery | bigquery-ingest.md |
| "export DynamoDB", "DynamoDB to data lake" | DynamoDB | dynamodb-ingest.md |
| "migrate Glue table", "convert Hive to Iceberg" | Catalog migration | catalog-migration.md |
If the user names Salesforce, ServiceNow, SAP, MongoDB, Kafka, or another SaaS/streaming source, decline -- these are not supported in this release.
If the source table is referenced by a fuzzy or business name ("migrate our orders table", "pull from the sales warehouse"), delegate to finding-data-lake-assets to resolve before proceeding.
For JDBC, Snowflake, and BigQuery sources, a Glue connection is required. Check:
aws glue get-connection --name <CONNECTION_NAME> --region <REGION>If the connection does not exist, stop and delegate to connecting-to-data-source to create and test it. Do not proceed with ingest until the connection is verified.
Local files, S3 files, DynamoDB, and catalog migration do not need a Glue connection.
You MUST ask the user (or suggest based on catalog inventory) before creating or writing to any table:
creating-data-lake-table)?Inventory-aware defaults:
If you have already run exploring-data-catalog or can quickly check, use what exists:
s3tablescatalog federated catalog and active table buckets: recommend S3 TablesDo not force S3 Tables on customers who haven't adopted it. See iceberg-catalog-config-and-usage.md.
Delegations from this step:
creating-data-lake-tablefinding-data-lake-assetsexploring-data-catalogRead the source-specific reference and follow its phases. Each is self-contained with job templates, gotchas, and troubleshooting:
Common Glue 5.1 or higher job configuration and PySpark templates are shared in glue-job-config.md and glue-job-scripts.md.
Run all three, do not skip:
See data-quality-validation.md.
For recurring pipelines, create a Glue Trigger with a cron schedule. See testing-and-scheduling.md. Simple single-step pipelines use Glue Triggers; multi-step with branching uses MWAA.
--target flag: Pre-fill the target format in Step 4--datalake-formats iceberg job argumentspark.sql.catalog.* config MUST go in --conf job arguments, never in spark.conf.set(). Glue 5.x throws AnalysisException: Cannot modify the value of a static config otherwise. See iceberg-catalog-config-and-usage.md for correct catalog configs.warehouse parameter is required in S3 Tables catalog config. Without it Spark fails with "Cannot derive default warehouse location".overwritePartitions() only replaces partitions present in the DataFrame -- for full refresh with deletes, use createOrReplace()connecting-to-data-source; do not debug network/credentials in this skill| Error | Likely cause | Action |
|---|---|---|
| Access Denied on S3 | Missing IAM permissions | Check Glue role has s3:GetObject, s3:PutObject |
| Access Denied on S3 Tables | Missing s3tables:* permissions | Add S3 Tables inline policy to Glue role |
| CTAS timeout | Dataset too large for Athena | Switch to Glue ETL or batch with WHERE filters |
| JDBC connection timeout/auth failure | Connection-level issue | Delegate to connecting-to-data-source |
| Throughput exceeded (DynamoDB) | Read percent too high | Lower read.percent or use native export |
See error-handling.md for the full catalog.
© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 25 other files (references) in plugins/aws-data-analytics/skills/ingesting-into-data-lake of aws/agent-toolkit-for-aws.
Open the folder on GitHubat commit bd49cc8
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in aws/agent-toolkit-for-aws, which our catalogue first saw on October 7, 2026.
Ingesting Into Data Lake next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ingesting Into Data Lake this skillaws/agent-toolkit-for-aws | 2.8k | 1 repos | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Mfs Findzilliztech/mfs | 150 | — | ~4k | Automated safety check: Pass | Apache-2.0 | |
| Mfs Ingestzilliztech/mfs | 150 | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | |
| AWS CLI Beastgiuseppe-trisciuoglio/developer-kit | 355 | — | ~1.7k | Automated safety check: Notes | MIT | |
| AWSkid-sid/claude-spellbook | 189 | — | ~4.9k | Automated safety check: Warn | MIT | |
| SQL Query Explainermohitagw15856/pm-claude-skills | 1.4k | — | ~1.6k | Automated safety check: Pass | MIT |
zilliztech/mfs
Search, grep, browse, and read across registered MFS data sources via the mfs CLI — codebases, docs, PDFs, web crawls, databases (postgres/mysql/mongo/snowflake/bigquery), issue trackers…
zilliztech/mfs
Register, update, or re-sync data sources for MFS so they become searchable — postgres / mysql / mongo / snowflake / bigquery, github / jira / linear / notion / hubspot / zendesk, slack / discord /…
giuseppe-trisciuoglio/developer-kit
Provides advanced AWS CLI patterns for managing EC2, Lambda, S3, DynamoDB, RDS, VPC, IAM, and CloudWatch.
kid-sid/claude-spellbook
A skill your agent uses when writing boto3 or AWS SDK v3 code — configuring IAM auth, reading/writing S3, designing DynamoDB access patterns, writing Lambda handlers, processing SQS batches, or…
mohitagw15856/pm-claude-skills
Explains, optimises, writes, and documents SQL queries. An agent skill from mohitagw15856/pm-claude-skills.
rohitg00/awesome-claude-code-toolkit
AWS cloud patterns for Lambda, ECS, S3, DynamoDB, and Infrastructure as Code with CDK/Terraform
aws/agent-toolkit-for-aws
Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.
aws/agent-toolkit-for-aws
A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.
aws/agent-toolkit-for-aws
Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.
aws/agent-toolkit-for-aws
Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.
aws/agent-toolkit-for-aws
Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…
aws/agent-toolkit-for-aws
A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.
Categories
Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue…. Ingesting Into Data Lake is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration).
Ingesting Into Data Lake fits situations like: move data to AWS; set up pipeline; pull from Snowflake; query BigQuery into S3.
Run `npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a claude-code`. Or copy the skill folder (plugins/aws-data-analytics/skills/ingesting-into-data-lake in aws/agent-toolkit-for-aws) into .claude/skills/ingesting-into-data-lake in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a codex`. Or copy the skill folder (plugins/aws-data-analytics/skills/ingesting-into-data-lake in aws/agent-toolkit-for-aws) into .agents/skills/ingesting-into-data-lake in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill ingesting-into-data-lake -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ingesting-into-data-lake, .gemini/skills/ingesting-into-data-lake, .github/skills/ingesting-into-data-lake and .opencode/skills/ingesting-into-data-lake in your project.
Going by SKILL.md and its folder, Ingesting Into Data Lake needs the command-line tools its instructions call (aws).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ingesting Into Data Lake is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 47k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ingesting Into Data Lake: Mfs Find (zilliztech/mfs, 150 stars), Mfs Ingest (zilliztech/mfs, 150 stars), AWS CLI Beast (giuseppe-trisciuoglio/developer-kit, 355 stars) and AWS (kid-sid/claude-spellbook, 189 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,816 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 7, 2026.
Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.