Agent skill

Ingesting Data

by ancoleman in ancoleman/ai-design-components

Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases.

MITAuto-check passedData & Analytics

Install Ingesting Data

skills CLI
$ npx skills add ancoleman/ai-design-components --skill ingesting-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ancoleman/ai-design-components ingesting-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ingesting-data .claude/skills/ingesting-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ingesting-data
GitHub stars
526
Token cost
~1.9k tokens
SKILL.md length
329 words
Files
11 (incl. scripts, references)
Skills in repo
75
Repo updated
First seen
Licence
MIT

At a glance

Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases.

  • Works in 4 steps: Batch Ingestion (Files/Storage) → Streaming Ingestion (Real-time) → API Polling (Feeds) → …
  • Importing CSV/JSON/Parquet files
  • SKILL.md covers When to Use This Skill, Ingestion Pattern Decision Tree, Quick Start by Language and Ingestion Patterns, plus 4 more sections
  • Runs Python scripts from its folder; reaches api.github.com

What it does

Ingesting Data is an agent skill from ancoleman/ai-design-components. Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases. Use when importing CSV/JSON/Parquet files, pulling from S3/GCS buckets, consuming API feeds, or building ETL pipelines.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts and reference files (for example `outputs.yaml`, `references/api-feeds.md` and `references/cloud-storage.md`).

It sits in Data & Analytics, covering Data pipelines and ETL, DataFrames and File uploads and storage. It works with Python, Microsoft Excel, Amazon Web Services and Polars. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.

When your agent uses it

  • Importing CSV/JSON/Parquet files
  • Pulling from S3/GCS buckets
  • Consuming API feeds
  • Building ETL pipelines

Example prompts

  • “/ingesting-data”

Requirements

  • Python 3
  • Node.js

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Batch Ingestion (Files/Storage)
  2. Streaming Ingestion (Real-time)
  3. API Polling (Feeds)
  4. Change Data Capture (CDC)

What it can do on your machine

Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ingesting Data loads about 1.9k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 329 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 329 words, ~1,904 tokens.

Download SKILL.mdSave it as .claude/skills/ingesting-data/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
ingesting-data
description
Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases. Use when importing CSV/JSON/Parquet files, pulling from S3/GCS buckets, consuming API feeds, or building ETL pipelines.

Data Ingestion Patterns

This skill provides patterns for getting data INTO systems from external sources.

When to Use This Skill

  • Importing CSV, JSON, Parquet, or Excel files
  • Loading data from S3, GCS, or Azure Blob storage
  • Consuming REST/GraphQL API feeds
  • Building ETL/ELT pipelines
  • Database migration and CDC (Change Data Capture)
  • Streaming data ingestion from Kafka/Kinesis

Ingestion Pattern Decision Tree

What is your data source?
├── Cloud Storage (S3, GCS, Azure) → See cloud-storage.md
├── Files (CSV, JSON, Parquet) → See file-formats.md
├── REST/GraphQL APIs → See api-feeds.md
├── Streaming (Kafka, Kinesis) → See streaming-sources.md
├── Legacy Database → See database-migration.md
└── Need full ETL framework → See etl-tools.md

Quick Start by Language

dlt (data load tool) - Modern Python ETL:

python
import dlt

# Define a source
@dlt.source
def github_source(repo: str):
    @dlt.resource(write_disposition="merge", primary_key="id")
    def issues():
        response = requests.get(f"https://api.github.com/repos/{repo}/issues")
        yield response.json()
    return issues

# Load to destination
pipeline = dlt.pipeline(
    pipeline_name="github_issues",
    destination="postgres",  # or duckdb, bigquery, snowflake
    dataset_name="github_data"
)

load_info = pipeline.run(github_source("owner/repo"))
print(load_info)

Polars for file processing (faster than pandas):

python
import polars as pl

# Read CSV with schema inference
df = pl.read_csv("data.csv")

# Read Parquet (columnar, efficient)
df = pl.read_parquet("s3://bucket/data.parquet")

# Read JSON lines
df = pl.read_ndjson("events.jsonl")

# Write to database
df.write_database(
    table_name="events",
    connection="postgresql://user:pass@localhost/db",
    if_table_exists="append"
)
TypeScript/Node.js

S3 ingestion:

typescript
import { S3Client, GetObjectCommand } from "@aws-sdk/client-s3";
import { parse } from "csv-parse/sync";

const s3 = new S3Client({ region: "us-east-1" });

async function ingestFromS3(bucket: string, key: string) {
  const response = await s3.send(new GetObjectCommand({ Bucket: bucket, Key: key }));
  const body = await response.Body?.transformToString();

  // Parse CSV
  const records = parse(body, { columns: true, skip_empty_lines: true });

  // Insert to database
  await db.insert(eventsTable).values(records);
}

API feed polling:

typescript
import { Hono } from "hono";

// Webhook receiver for real-time ingestion
const app = new Hono();

app.post("/webhooks/stripe", async (c) => {
  const event = await c.req.json();

  // Validate webhook signature
  const signature = c.req.header("stripe-signature");
  // ... validation logic

  // Ingest event
  await db.insert(stripeEventsTable).values({
    eventId: event.id,
    type: event.type,
    data: event.data,
    receivedAt: new Date()
  });

  return c.json({ received: true });
});
Rust

High-performance file ingestion:

rust
use polars::prelude::*;
use aws_sdk_s3::Client;

async fn ingest_parquet(client: &Client, bucket: &str, key: &str) -> Result<DataFrame> {
    // Download from S3
    let resp = client.get_object()
        .bucket(bucket)
        .key(key)
        .send()
        .await?;

    let bytes = resp.body.collect().await?.into_bytes();

    // Parse with Polars
    let df = ParquetReader::new(Cursor::new(bytes))
        .finish()?;

    Ok(df)
}
Go

Concurrent file processing:

go
package main

import (
    "context"
    "encoding/csv"
    "github.com/aws/aws-sdk-go-v2/service/s3"
)

func ingestCSV(ctx context.Context, client *s3.Client, bucket, key string) error {
    resp, err := client.GetObject(ctx, &s3.GetObjectInput{
        Bucket: &bucket,
        Key:    &key,
    })
    if err != nil {
        return err
    }
    defer resp.Body.Close()

    reader := csv.NewReader(resp.Body)
    records, err := reader.ReadAll()
    if err != nil {
        return err
    }

    // Batch insert to database
    return batchInsert(ctx, records)
}

Ingestion Patterns

1. Batch Ingestion (Files/Storage)

For periodic bulk loads:

Source → Extract → Transform → Load → Validate
  ↓         ↓          ↓         ↓        ↓
 S3      Download   Clean/Map  Insert   Count check

Key considerations:

  • Use chunked reading for large files (>100MB)
  • Implement idempotency with checksums
  • Track file processing state
  • Handle partial failures
2. Streaming Ingestion (Real-time)

For continuous data flow:

Source → Buffer → Process → Load → Ack
  ↓        ↓         ↓        ↓      ↓
Kafka   In-memory  Transform  DB   Commit offset

Key considerations:

  • At-least-once vs exactly-once semantics
  • Backpressure handling
  • Dead letter queues for failures
  • Checkpoint management
3. API Polling (Feeds)

For external API data:

Schedule → Fetch → Dedupe → Load → Update cursor
   ↓         ↓        ↓       ↓         ↓
 Cron     API call  By ID   Insert   Last timestamp

Key considerations:

  • Rate limiting and backoff
  • Incremental loading (cursors, timestamps)
  • API pagination handling
  • Retry with exponential backoff
4. Change Data Capture (CDC)

For database replication:

Source DB → Capture changes → Transform → Target DB
    ↓             ↓               ↓            ↓
 Postgres    Debezium/WAL      Map schema   Insert/Update

Key considerations:

  • Initial snapshot + streaming changes
  • Schema evolution handling
  • Ordering guarantees
  • Conflict resolution

Library Recommendations

Use CasePythonTypeScriptRustGo
ETL Frameworkdlt, Meltano, Dagster---
Cloud Storageboto3, gcsfs, adlfs@aws-sdk/, @google-cloud/aws-sdk-s3, object_storeaws-sdk-go-v2
File Processingpolars, pandas, pyarrowpapaparse, xlsx, parquetjspolars-rs, arrow-rsencoding/csv, parquet-go
Streamingconfluent-kafka, aiokafkakafkajsrdkafka-rsfranz-go, sarama
CDCDebezium, pg_logical---

Reference Documentation

  • references/cloud-storage.md - S3, GCS, Azure Blob patterns
  • references/file-formats.md - CSV, JSON, Parquet, Excel handling
  • references/api-feeds.md - REST polling, webhooks, GraphQL subscriptions
  • references/streaming-sources.md - Kafka, Kinesis, Pub/Sub
  • references/database-migration.md - Schema migration, CDC patterns
  • references/etl-tools.md - dlt, Meltano, Airbyte, Fivetran

Scripts

  • scripts/validate_csv_schema.py - Validate CSV against expected schema
  • scripts/test_s3_connection.py - Test S3 bucket connectivity
  • scripts/generate_dlt_pipeline.py - Generate dlt pipeline scaffold

Chaining with Database Skills

After ingestion, chain to appropriate database skill:

DestinationChain to Skill
PostgreSQL, MySQLdatabases-relational
MongoDB, DynamoDBdatabases-document
Qdrant, Pineconedatabases-vector (after embedding)
ClickHouse, TimescaleDBdatabases-timeseries
Neo4jdatabases-graph

For vector databases, chain through ai-data-engineering for embedding:

ingesting-data → ai-data-engineering → databases-vector

© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts, references) in skills/ingesting-data of ancoleman/ai-design-components.

  • SKILL.md
  • outputs.yaml
  • references/api-feeds.md
  • references/cloud-storage.md
  • references/database-migration.md
  • references/etl-tools.md
  • references/file-formats.md
  • references/streaming-sources.md
  • scripts/generate_dlt_pipeline.py
  • scripts/test_s3_connection.py
  • scripts/validate_csv_schema.py

Open the folder on GitHubat commit 76551b7

Compare with similar skills

Ingesting Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ingesting Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ingesting Data this skillancoleman/ai-design-components526—~1.9kAutomated safety check: PassMIT
Python Pipelinejamditis/claude-skills-journalism416—~4.8kAutomated safety check: PassMIT
Authoring Mwaa Workflowaws/agent-toolkit-for-aws2.8k—~2.8kAutomated safety check: PassApache-2.0
Raccoon DataanalysisSenseTime-Copilot/raccoon-dataanalysis-skill137—~1.9kAutomated safety check: PassNone
Credit Risk Data Cleaninggithub/awesome-copilot40k1 repos~1.5kAutomated safety check: PassMIT
Ray Data for ML PipelinesOrchestra-Research/AI-Research-SKILLs13k3 repos~1.8kAutomated safety check: PassMIT

Similar skills

  • Python Pipeline

    jamditis/claude-skills-journalism

    Python data pipelines with modular architecture. An agent skill from jamditis/claude-skills-journalism.

    416 GitHub stars~4.8k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Authoring Mwaa Workflow

    aws/agent-toolkit-for-aws

    Official

    Authors and deploys MWAA workflow artifacts: Python Airflow DAGs for provisioned environments or YAML workflow files for Serverless.

    2.8k GitHub stars~2.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Raccoon Dataanalysis

    SenseTime-Copilot/raccoon-dataanalysis-skill

    Raccoon (小浣熊) Data Analysis - Remote code interpreter and data visualization service powered by SenseTime.

    137 GitHub stars~1.9k tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Ray Data for ML Pipelines

    Orchestra-Research/AI-Research-SKILLs

    Uses Ray Data to read, transform and write large datasets across a cluster for ML training and batch inference, with streaming execution and optional GPU steps.

    13k GitHub starsUsed in 3 repos~1.8k tokens
    Data & AnalyticsAuto-check passed
  • Hybrid-Engine Data Analysis

    code-yeongyu/oh-my-openagent

    Analyzes CSV, Parquet and JSON data with DuckDB, Polars, numpy and matplotlib, preferring a persistent kernel over repeated one-shot processes.

    70k GitHub stars~1.4k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from ancoleman/ai-design-components

All 75 skills in this repo
  • Building AI Chat

    ancoleman/ai-design-components

    Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.

    526 GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check passed
  • Building Forms

    ancoleman/ai-design-components

    Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.

    526 GitHub stars~3.7k tokensUpdated 10 mo ago
    Auto-check passed
  • Building Tables

    ancoleman/ai-design-components

    Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.

    526 GitHub stars~1.8k tokensUpdated 10 mo ago
    Auto-check passed
  • Creating Dashboards

    ancoleman/ai-design-components

    Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.

    526 GitHub stars~3.5k tokensUpdated 10 mo ago
    Auto-check passed
  • Designing Layouts

    ancoleman/ai-design-components

    Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.

    526 GitHub stars~1.7k tokensUpdated 10 mo ago
    Auto-check passed
  • Displaying Timelines

    ancoleman/ai-design-components

    Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.

    526 GitHub stars~2.7k tokensUpdated 10 mo ago
    Auto-check passed

Questions about Ingesting Data

What does Ingesting Data do?

Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases. Ingesting Data is an agent skill from ancoleman/ai-design-components. Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases.

When should I use Ingesting Data?

Ingesting Data fits situations like: importing CSV/JSON/Parquet files; pulling from S3/GCS buckets; consuming API feeds; building ETL pipelines.

How do I install Ingesting Data in Claude Code?

Run `npx skills add ancoleman/ai-design-components --skill ingesting-data -a claude-code`. Or copy the skill folder (skills/ingesting-data in ancoleman/ai-design-components) into .claude/skills/ingesting-data in your project. Claude Code loads it when a task matches its description.

How do I install Ingesting Data in Codex?

Run `npx skills add ancoleman/ai-design-components --skill ingesting-data -a codex`. Or copy the skill folder (skills/ingesting-data in ancoleman/ai-design-components) into .agents/skills/ingesting-data in your project. Codex loads it when a task matches its description.

Can I use Ingesting Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill ingesting-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ingesting-data, .gemini/skills/ingesting-data, .github/skills/ingesting-data and .opencode/skills/ingesting-data in your project.

What does Ingesting Data need to run?

Going by SKILL.md and its folder, Ingesting Data needs Python for the scripts in its folder. Our summary lists: Python 3; Node.js.

Does Ingesting Data access the network?

SKILL.md names 1 domain. In commands or code: api.github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Ingesting Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ingesting Data use?

Ingesting Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ingesting Data use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.3k tokens, read only when the agent opens those files.

What are the alternatives to Ingesting Data?

Skills that share tags, products or a category with Ingesting Data: Python Pipeline (jamditis/claude-skills-journalism, 416 stars), Authoring Mwaa Workflow (aws/agent-toolkit-for-aws, 2.8k stars), Raccoon Dataanalysis (SenseTime-Copilot/raccoon-dataanalysis-skill, 137 stars) and Credit Risk Data Cleaning (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ingesting Data?

ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 526 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.

Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.