Agent skill

Elasticsearch File Ingest

by Kilo-Org in Kilo-Org/kilo-marketplace

Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

Apache-2.0Auto-check passedBackend & APIs

Install Elasticsearch File Ingest

skills CLI
$ npx skills add Kilo-Org/kilo-marketplace --skill elasticsearch-file-ingest -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Kilo-Org/kilo-marketplace elasticsearch-file-ingest --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Kilo-Org/kilo-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/elasticsearch-file-ingest .claude/skills/elasticsearch-file-ingest && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
elasticsearch-file-ingest
GitHub stars
190
Token cost
~2.8k tokens
SKILL.md length
702 words
Files
11 (incl. scripts, references)
Skills in repo
86
Repo updated
First seen
Licence
Apache-2.0

At a glance

Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

  • Batch importing data — not for reindexing
  • SKILL.md covers Features & Use Cases, Prerequisites, Setup and Test Connection, plus 5 more sections
  • Runs JavaScript scripts from its folder; calls node and npm; needs ELASTICSEARCH_API_KEY and ELASTICSEARCH_PASSWORD
  • General ingest pipeline design

What it does

Elasticsearch File Ingest is an agent skill from Kilo-Org/kilo-marketplace. Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms. Use when loading files or batch importing data — not for reindexing, general ingest pipeline design, or bulk API patterns.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `examples/mappings.json`, `examples/skip-transform.js` and `examples/split-transform.js`).

It sits in Backend & APIs, covering Search implementation, DataFrames and CSV and tabular files. It works with Elasticsearch. The repository describes itself as: Kilo Marketplace - A curated collection of Skills, MCP Servers, and Modes for enhancing AI agent capabilities across the Kilo ecosystem—including Kilo Code (VS Code extension)… The licence is Apache-2.0.

When your agent uses it

  • Batch importing data — not for reindexing
  • General ingest pipeline design
  • Bulk API patterns

Example prompts

  • “/elasticsearch-file-ingest”

Requirements

  • Node.js
  • A credential in ELASTICSEARCH_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit ff51758. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • elastic.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ELASTICSEARCH_API_KEY
    • ELASTICSEARCH_PASSWORD

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Elasticsearch File Ingest loads about 2.8k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 702 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Kilo-Org/kilo-marketplace at commit ff51758, republished under its Apache-2.0 licence (© Kilo-Org). 702 words, ~2,786 tokens.

Download SKILL.mdSave it as .claude/skills/elasticsearch-file-ingest/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
elasticsearch-file-ingest
description
Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms. Use when loading files or batch importing data — not for reindexing, general ingest pipeline design, or bulk API patterns.
metadata.category
data

Elasticsearch File Ingest

Stream-based ingestion and transformation of large data files (NDJSON, CSV, Parquet, Arrow IPC) into Elasticsearch.

Features & Use Cases

  • Stream-based: Handle large files without running out of memory
  • High throughput: 50k+ documents/second on commodity hardware
  • Formats: NDJSON, CSV, Parquet, Arrow IPC
  • Transformations: Apply custom JavaScript transforms during ingestion (enrich, split, filter)
  • Batch processing: Ingest multiple files matching a pattern (e.g., logs/*.json)
  • Document splitting: Transform one source document into multiple targets

Prerequisites

  • Elasticsearch 8.x or 9.x accessible (local or remote)
  • Node.js 22+ installed

Setup

This skill is self-contained. The scripts/ folder and package.json live in this skill's directory. Run all commands from this directory. Use absolute paths when referencing data files located elsewhere.

Before first use, install dependencies:

bash
npm install
Environment Configuration

Elasticsearch connection is configured by users exclusively via environment variables. Never pass credentials as command-line arguments. If the test fails, output the setup options below to the user, then stop. Do not proceed with ingestion until a successful connection test.

bash
export ELASTICSEARCH_CLOUD_ID="<your-cloud-id>"
export ELASTICSEARCH_API_KEY="<your-api-key>"
Option 2: Direct URL with API Key
bash
export ELASTICSEARCH_URL="https://elasticsearch:9200"
export ELASTICSEARCH_API_KEY="<your-api-key>"
Option 3: Basic Authentication
bash
export ELASTICSEARCH_URL="https://elasticsearch:9200"
export ELASTICSEARCH_USERNAME="<your-username>"
export ELASTICSEARCH_PASSWORD="<your-password>"
Option 4: Local Development

For local development and testing, see Run Elasticsearch locally to spin up Elasticsearch and Kibana. After setup, export the connection variables (URL and API key or credentials) as shown in Option 2 or Option 3 above.

Private CA certificates

Keep TLS verification enabled. For a development cluster signed by a private CA, point Node.js at the reviewed CA bundle:

bash
export NODE_EXTRA_CA_CERTS="/path/to/private-ca-bundle.pem"

Test Connection

Verify the Elasticsearch connection before ingesting data:

bash
node scripts/ingest.js test

Always run this first. If the test fails, resolve the connection issue before proceeding.

Examples

Ingest a JSON file
bash
node scripts/ingest.js ingest --file /absolute/path/to/data.json --target my-index
Stream NDJSON/CSV via stdin
bash
# NDJSON
cat /absolute/path/to/data.ndjson | node scripts/ingest.js ingest --stdin --target my-index

# CSV
cat /absolute/path/to/data.csv | node scripts/ingest.js ingest --stdin --source-format csv --target my-index
Ingest CSV directly
bash
node scripts/ingest.js ingest --file /absolute/path/to/users.csv --source-format csv --target users
Ingest Parquet directly
bash
node scripts/ingest.js ingest --file /absolute/path/to/users.parquet --source-format parquet --target users
Ingest Arrow IPC directly
bash
node scripts/ingest.js ingest --file /absolute/path/to/users.arrow --source-format arrow --target users
Ingest CSV with parser options
bash
# csv-options.json
# {
#   "columns": true,
#   "delimiter": ";",
#   "trim": true
# }

node scripts/ingest.js ingest --file /absolute/path/to/users.csv --source-format csv --csv-options csv-options.json --target users
Infer mappings/pipeline from CSV

When using --infer-mappings, do not combine it with --source-format csv. Inference sends a raw sample to Elasticsearch's _text_structure/find_structure endpoint, which returns both mappings and an ingest pipeline with a CSV processor. If --source-format csv is also set, CSV is parsed client-side and server-side, resulting in an empty index. Let --infer-mappings handle everything:

bash
node scripts/ingest.js ingest --file /absolute/path/to/users.csv --infer-mappings --target users
Infer mappings with options
bash
# infer-options.json
# {
#   "sampleBytes": 200000,
#   "lines_to_sample": 2000
# }

node scripts/ingest.js ingest --file /absolute/path/to/users.csv --infer-mappings --infer-mappings-options infer-options.json --target users
Ingest with custom mappings
bash
node scripts/ingest.js ingest --file /absolute/path/to/data.json --target my-index --mappings mappings.json
Ingest with transformation
bash
node scripts/ingest.js ingest --file /absolute/path/to/data.json --target my-index --transform transform.js

Command Reference

Required Options
bash
--target <index>         # Target index name
Source Options (choose one)
bash
--file <path>            # Source file (supports wildcards, e.g., logs/*.json)
--stdin                  # Read NDJSON/CSV from stdin
Index Configuration
bash
--mappings <file.json>          # Mappings file
--infer-mappings                # Infer mappings/pipeline from file/stream (do NOT combine with --source-format)
--infer-mappings-options <file> # Options for inference (JSON file)
--delete-index                  # Delete target index if exists
--pipeline <name>               # Ingest pipeline name
Processing
bash
--transform <file.js>    # Transform function (export as default or module.exports)
--source-format <fmt>    # Source format: ndjson|csv|parquet|arrow (default: ndjson)
--csv-options <file>     # CSV parser options (JSON file)
--skip-header            # Skip first line (e.g., CSV header)
Performance
bash
--buffer-size <kb>       # Buffer size in KB (default: 5120)
--total-docs <n>         # Total docs for progress bar (file/stream)
--stall-warn-seconds <n> # Stall warning threshold (default: 30)
--progress-mode <mode>   # Progress output: auto|line|newline (default: auto)
--debug-events           # Log pause/resume/stall events
--quiet                  # Disable progress bars

Transform Functions

Transform functions let you modify documents during ingestion. Create a JavaScript file that exports a transform function:

Basic Transform (transform.js)
javascript
// ES modules (default)
export default function transform(doc) {
  return {
    ...doc,
    full_name: `${doc.first_name} ${doc.last_name}`,
    timestamp: new Date().toISOString(),
  };
}

// Or CommonJS
module.exports = function transform(doc) {
  return {
    ...doc,
    full_name: `${doc.first_name} ${doc.last_name}`,
  };
};
Skip Documents

Return null or undefined to skip a document:

javascript
export default function transform(doc) {
  // Skip invalid documents
  if (!doc.email || !doc.email.includes("@")) {
    return null;
  }
  return doc;
}
Split Documents

Return an array to create multiple target documents from one source:

javascript
export default function transform(doc) {
  // Split a tweet into multiple hashtag documents
  const hashtags = doc.text.match(/#\w+/g) || [];
  return hashtags.map((tag) => ({
    hashtag: tag,
    tweet_id: doc.id,
    created_at: doc.created_at,
  }));
}

Mappings

Custom Mappings (mappings.json)
json
{
  "properties": {
    "@timestamp": { "type": "date" },
    "message": { "type": "text" },
    "user": {
      "properties": {
        "name": { "type": "keyword" },
        "email": { "type": "keyword" }
      }
    }
  }
}
bash
node scripts/ingest.js ingest --file /absolute/path/to/data.json --target my-index --mappings mappings.json
Show full SKILL.md (285 more words)Show less

Boundaries

  • Never echo, print, log, or otherwise reveal the values of credential environment variables ($ELASTICSEARCH_API_KEY, $ELASTICSEARCH_PASSWORD, $ELASTICSEARCH_CLOUD_ID, etc.). Do not run shell commands whose output would expose secret values (e.g., echo $ELASTICSEARCH_API_KEY, env | grep KEY, printenv). Exporting these variables and running scripts that read them internally is expected and safe — the restriction is on surfacing secret values in command output. The only way to verify connectivity is node scripts/ingest.js test. If the test fails, ask the user to check their environment configuration — do not attempt to diagnose credentials yourself.
  • Never run destructive commands (such as using the --delete-index flag or deleting existing indices and data) without explicit user confirmation.

Guidelines

  • Test first: Always run node scripts/ingest.js test before ingesting data. If the connection fails, ask the user to verify their environment configuration and re-test. Do not attempt ingestion until the test passes.
  • Never combine --infer-mappings with --source-format. Inference creates a server-side ingest pipeline that handles parsing (e.g., CSV processor). Using --source-format csv parses client-side as well, causing double-parsing and an empty index. Use --infer-mappings alone for automatic detection, or --source-format with explicit --mappings for manual control.
  • Use --source-format csv with --mappings when you want client-side CSV parsing with known field types.
  • Use --infer-mappings alone when you want Elasticsearch to detect the format, infer field types, and create an ingest pipeline automatically.

When NOT to Use

Consider alternatives for:

  • Reindexing or index migration: Use the elasticsearch-reindex skill for copying, migrating, or transforming existing Elasticsearch indices
  • Real-time ingestion: Use Filebeat or Elastic Agent
  • Enterprise pipelines: Use Logstash
  • Built-in transforms: Use Elasticsearch Transforms

Additional Resources

References

© Kilo-Org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts, references) in skills/elasticsearch-file-ingest of Kilo-Org/kilo-marketplace.

  • SKILL.md
  • LICENSE
  • examples/mappings.json
  • examples/skip-transform.js
  • examples/split-transform.js
  • examples/transform.js
  • local.patch
  • package.json
  • references/patterns.md
  • references/troubleshooting.md
  • scripts/ingest.js

Open the folder on GitHubat commit ff51758

Compare with similar skills

Elasticsearch File Ingest next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Elasticsearch File Ingest compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Elasticsearch File Ingest this skillKilo-Org/kilo-marketplace190—~2.8kAutomated safety check: PassApache-2.0
Elasticsearch File Ingestaspectrr/deer405—~684Automated safety check: PassMIT
Elasticsearch Ingestelastic/agent-skills592—~2.7kAutomated safety check: PassApache-2.0
Product Full-Text Searchlobehub/lobehub83k—~4.1kAutomated safety check: PassCustom licence
Foundatio Repositoriesexceptionless/Exceptionless2.5k—~1.9kAutomated safety check: PassApache-2.0
Elasticsearch Authnaspectrr/deer405—~1.2kAutomated safety check: NotesMIT

Similar skills

  • Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

    405 GitHub stars~684 tokensUpdated 5 mo ago
    Backend & APIsAuto-check passed
  • Elasticsearch Ingest

    elastic/agent-skills

    Official

    Load CSV and JSON files into Elasticsearch indices using the bulk API and explicit mappings when field types matter.

    592 GitHub stars~2.7k tokensUpdated 3 days ago
    Backend & APIsAuto-check passed
  • Guides work on LobeHub's own product search: the shared search repository, provider choice, Elasticsearch mappings, change syncing and reindexing.

    83k GitHub stars~4.1k tokensUpdated today
    Backend & APIsAuto-check passed
  • Foundatio Repositories

    exceptionless/Exceptionless

    Query, aggregate, patch, or paginate Exceptionless data through its Elasticsearch repository abstractions.

    2.5k GitHub stars~1.9k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Elasticsearch Authn

    aspectrr/deer

    Authenticate to Elasticsearch using native, file-based, LDAP/AD, SAML, OIDC, Kerberos, JWT, or certificate realms.

    405 GitHub stars~1.2k tokensUpdated 5 mo ago
    Backend & APIsAuto-check: notes
  • Elasticsearch Authz

    aspectrr/deer

    Manage Elasticsearch RBAC: native users, roles, role mappings, document- and field-level security.

    405 GitHub stars~1.8k tokensUpdated 5 mo ago
    Backend & APIsAuto-check passed

More from Kilo-Org/kilo-marketplace

All 86 skills in this repo
  • AzureML Project Scaffolding

    Kilo-Org/kilo-marketplace

    Sets up and maintains AzureML-ready Python projects as uv workspaces with devcontainers, a Makefile and job YAML, so local runs match cloud jobs and experiments stay reproducible.

    190 GitHub stars~3.1k tokensUpdated 12 days ago
    Auto-check: notes
  • Jupyter Notebook Builder

    Kilo-Org/kilo-marketplace

    Creates, inspects, edits and runs Jupyter notebooks, scaffolding experiment or tutorial notebooks from templates and preferring a Jupyter MCP server over raw JSON edits.

    190 GitHub stars~1.3k tokensUpdated 12 days ago
    Auto-check passed
  • Tableau Dashboard Creator

    Kilo-Org/kilo-marketplace

    Takes a plain-language dashboard request through brand setup, data exploration, planning, an interactive HTML mock and a Tableau implementation spec.

    190 GitHub stars~3.8k tokensUpdated 12 days ago
    Auto-check: notes
  • Nifi Flow Layout

    Kilo-Org/kilo-marketplace

    A skill your agent uses when arranging Apache NiFi processors, process groups, ports, comments, numbering, crossing connections, dense fan-in/fan-out, or reusable readable canvas layouts.

    190 GitHub stars~1.5k tokensUpdated 12 days ago
    Auto-check passed
  • Splunk Ingest Processor Setup

    Kilo-Org/kilo-marketplace

    Render Cisco Data Fabric ingest-time routing workflows and Splunk Cloud Platform Ingest Processor setup plans with SPL2 pipelines, source types, destinations, lifecycle handoffs, queue and…

    190 GitHub stars~1.2k tokensUpdated 12 days ago
    Auto-check passed
  • Databricks Jobs

    Kilo-Org/kilo-marketplace

    Develop and deploy Lakeflow Jobs on Databricks via DABs, Python SDK, or the CLI.

    190 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed

Works with

Questions about Elasticsearch File Ingest

What does Elasticsearch File Ingest do?

Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms. Elasticsearch File Ingest is an agent skill from Kilo-Org/kilo-marketplace. Ingest and transform data files (CSV/JSON/Parquet/Arrow IPC) into Elasticsearch with stream processing and custom transforms.

When should I use Elasticsearch File Ingest?

Elasticsearch File Ingest fits situations like: batch importing data — not for reindexing; general ingest pipeline design; bulk API patterns.

How do I install Elasticsearch File Ingest in Claude Code?

Run `npx skills add Kilo-Org/kilo-marketplace --skill elasticsearch-file-ingest -a claude-code`. Or copy the skill folder (skills/elasticsearch-file-ingest in Kilo-Org/kilo-marketplace) into .claude/skills/elasticsearch-file-ingest in your project. Claude Code loads it when a task matches its description.

How do I install Elasticsearch File Ingest in Codex?

Run `npx skills add Kilo-Org/kilo-marketplace --skill elasticsearch-file-ingest -a codex`. Or copy the skill folder (skills/elasticsearch-file-ingest in Kilo-Org/kilo-marketplace) into .agents/skills/elasticsearch-file-ingest in your project. Codex loads it when a task matches its description.

Can I use Elasticsearch File Ingest in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Kilo-Org/kilo-marketplace --skill elasticsearch-file-ingest -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/elasticsearch-file-ingest, .gemini/skills/elasticsearch-file-ingest, .github/skills/elasticsearch-file-ingest and .opencode/skills/elasticsearch-file-ingest in your project.

What does Elasticsearch File Ingest need to run?

Going by SKILL.md and its folder, Elasticsearch File Ingest needs JavaScript for the scripts in its folder, the command-line tools its instructions call (node and npm) and credentials named ELASTICSEARCH_API_KEY and ELASTICSEARCH_PASSWORD. Our summary lists: Node.js; A credential in ELASTICSEARCH_API_KEY.

Does Elasticsearch File Ingest access the network?

SKILL.md names 1 domain. As links in the text: elastic.co. This is read from the text; nothing was executed.

Is Elasticsearch File Ingest safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Elasticsearch File Ingest use?

Elasticsearch File Ingest is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Elasticsearch File Ingest use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 939 tokens, read only when the agent opens those files.

What are the alternatives to Elasticsearch File Ingest?

Skills that share tags, products or a category with Elasticsearch File Ingest: Elasticsearch File Ingest (aspectrr/deer, 405 stars), Elasticsearch Ingest (elastic/agent-skills, 592 stars), Product Full-Text Search (lobehub/lobehub, 83k stars) and Foundatio Repositories (exceptionless/Exceptionless, 2.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Elasticsearch File Ingest?

Kilo-Org (a GitHub organization) maintains it in Kilo-Org/kilo-marketplace, which has 190 GitHub stars. The repository holds 86 skills in this directory. The repository was last updated on September 28, 2026.

Source: Kilo-Org/kilo-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.