Agent skill

Create Ingestor Plugin

by NomaDamas in NomaDamas/AutoRAG-Research

Guide developers through creating a custom data ingestor plugin for AutoRAG-Research.

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Create Ingestor Plugin

skills CLI
$ npx skills add NomaDamas/AutoRAG-Research --skill create-ingestor-plugin -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NomaDamas/AutoRAG-Research create-ingestor-plugin --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NomaDamas/AutoRAG-Research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/create-ingestor-plugin .claude/skills/create-ingestor-plugin && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
create-ingestor-plugin
GitHub stars
149
Token cost
~1.2k tokens
SKILL.md length
348 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guide developers through creating a custom data ingestor plugin for AutoRAG-Research.

  • Works in 4 steps: Scaffold → Implement the ingestor → Database Schema (critical) → …
  • Ingesting a new dataset format into AutoRAG-Research
  • SKILL.md covers Workflow, Testing, Key Files and Examples
  • Calls pip

What it does

Create Ingestor Plugin is an agent skill from NomaDamas/AutoRAG-Research. Guide developers through creating a custom data ingestor plugin for AutoRAG-Research. Ingestors load external datasets (HuggingFace, local files, APIs) into the database. Uses @registeringestor decorator for automatic CLI parameter extraction. Use when ingesting a new dataset format into AutoRAG-Research.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Model hubs and datasets. It works with Hugging Face. The repository describes itself as: Automate your RAG research. The licence is Apache-2.0.

When your agent uses it

  • Ingesting a new dataset format into AutoRAG-Research
  • Tasks that involve Model hubs and datasets

Example prompts

  • “/create-ingestor-plugin”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Scaffold
  2. Implement the ingestor
  3. Database Schema (critical)
  4. Install and verify

What it can do on your machine

Read from SKILL.md and the folder at commit a473cf0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Create Ingestor Plugin loads about 1.2k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 348 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NomaDamas/AutoRAG-Research at commit a473cf0, republished under its Apache-2.0 licence (© NomaDamas). 348 words, ~1,188 tokens.

Download SKILL.mdSave it as .claude/skills/create-ingestor-plugin/SKILL.md (or your agent's skills folder).
name
create-ingestor-plugin
description
Guide developers through creating a custom data ingestor plugin for AutoRAG-Research. Ingestors load external datasets (HuggingFace, local files, APIs) into the database. Uses @register_ingestor decorator for automatic CLI parameter extraction. Use when ingesting a new dataset format into AutoRAG-Research.
allowed-tools
Bash, Read, Write, Edit

Create Ingestor Plugin

Workflow

1. Scaffold
bash
autorag-research plugin create my_dataset --type=ingestor

Read the generated ingestor.py, pyproject.toml, and test file to understand the structure.

The generated pyproject.toml registers the autorag_research.ingestors entry point. The @register_ingestor decorator handles automatic CLI parameter extraction from __init__ type hints.

2. Implement the ingestor

For the code-level implementation rules that are shared with the agent workflows, read:

  • ai_instructions/implementation_specialist.md
  • ai_instructions/schema_architect.md
  • ai_instructions/test_writer.md

Required methods:

  • __init__(embedding_model, ...) — accept embedding model + dataset-specific params
  • detect_primary_key_type() → "bigint" or "string"
  • ingest(subset, query_limit, min_corpus_cnt) — load data and save via self.service

__init__ type hints drive CLI generation automatically:

Type HintCLI Behavior
Literal["a", "b"]--param with choices, required
str--param, required
int = 100--param, optional with default
bool = False--param/--no-param flag

Parameters named embedding_model or late_interaction_embedding_model are auto-skipped (injected by CLI).

self.service is injected after construction via set_service(). Read existing ingestors for exact service method signatures.

3. Database Schema (critical)

Ingestors must populate the correct entity hierarchy:

Document → Page → Chunk (text)
                → ImageChunk (images)
  • Document — top-level container (e.g., a Wikipedia article, a PDF)
  • Page — subdivision within a document (linked via document_id)
  • Chunk — text passage with embedding vector (linked to Page via PageChunkRelation)
  • ImageChunk — image binary with embedding vector (linked to Page via PageChunkRelation)
  • Query — search query with generation_gt: list[str] | None (ground truth answers)

RetrievalRelation — links queries to relevant chunks using AND/OR group structure:

RetrievalRelation(query_id, chunk_id, group_index, group_order, score)

group_index = AND group number
group_order = OR position within the group

Example: query needs (chunk_A OR chunk_B) AND chunk_C
  → (query, chunk_A, group_index=0, group_order=0)
  → (query, chunk_B, group_index=0, group_order=1)
  → (query, chunk_C, group_index=1, group_order=0)

This AND/OR structure is critical for multi-hop queries. See ai_instructions/db_schema.md for the full DBML schema.

4. Install and verify
bash
cd my_dataset_plugin
pip install -e .   # or: uv pip install -e .

No plugin sync needed — ingestors are discovered automatically via entry points.

bash
autorag-research ingest my_dataset --dataset-name subset_a

Testing

Use ingestor_test_utils for integration tests against a real PostgreSQL database:

  • IngestorTestConfig — declare expected counts (queries, chunks, image_chunks), relation checks, primary key type
  • create_test_database(config) — context manager that creates/drops an isolated test DB
  • IngestorTestVerifier — runs all configured checks: count verification, format validation, retrieval relation checks, generation_gt checks, content hash verification

See tests/autorag_research/data/ingestor_test_utils.py for full API and usage examples in the module docstring.

Key Files

PurposePath
Base classesautorag_research/data/base.py → TextEmbeddingDataIngestor, MultiModalEmbeddingDataIngestor
Registration decoratorautorag_research/data/registry.py → @register_ingestor
Text ingestion serviceautorag_research/orm/service/text_ingestion.py
Multi-modal ingestion serviceautorag_research/orm/service/multi_modal_ingestion.py
DB schema referenceai_instructions/db_schema.md
Test utilitiestests/autorag_research/data/ingestor_test_utils.py

Examples

Study these existing implementations for patterns:

  • autorag_research/data/beir.py — BEIR benchmark (simple, good starting point)
  • autorag_research/data/bright.py — BRIGHT dataset
  • autorag_research/data/mrtydi.py — Mr. TyDi multilingual dataset
  • autorag_research/data/ragbench.py — RAGBench dataset

© NomaDamas, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/create-ingestor-plugin of NomaDamas/AutoRAG-Research.

Open the folder on GitHubat commit a473cf0

Compare with similar skills

Create Ingestor Plugin next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Create Ingestor Plugin compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Create Ingestor Plugin this skillNomaDamas/AutoRAG-Research149—~1.2kAutomated safety check: NotesApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Dataset Transformationawslabs/agent-plugins9122 repos~3.5kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Dataset Transformation

    awslabs/agent-plugins

    Official

    Generates code that transforms datasets between ML schemas for model training or evaluation.

    912 GitHub starsUsed in 2 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from NomaDamas/AutoRAG-Research

  • Autorag Query

    NomaDamas/AutoRAG-Research

    Query AutoRAG-Research pipeline results using natural language.

    149 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • Create Generation Plugin

    NomaDamas/AutoRAG-Research

    Guide developers through creating a custom generation pipeline plugin for AutoRAG-Research.

    149 GitHub stars~871 tokensUpdated 2 mo ago
    Auto-check: notes
  • Create Metric Plugin

    NomaDamas/AutoRAG-Research

    Guide developers through creating a custom evaluation metric plugin for AutoRAG-Research.

    149 GitHub stars~859 tokensUpdated 2 mo ago
    Auto-check: notes
  • Create Retrieval Plugin

    NomaDamas/AutoRAG-Research

    Guide developers through creating a custom retrieval pipeline plugin for AutoRAG-Research.

    149 GitHub stars~728 tokensUpdated 2 mo ago
    Auto-check: notes
  • Refactor

    NomaDamas/AutoRAG-Research

    Orchestrate a 3-agent PR code review debate using Claude Code Teams.

    149 GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check: notes
  • Resolve Conversation

    NomaDamas/AutoRAG-Research

    Process [APPROVE] and [IGNORE] replies on /refactor review threads.

    149 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check: notes

Works with

Questions about Create Ingestor Plugin

What does Create Ingestor Plugin do?

Guide developers through creating a custom data ingestor plugin for AutoRAG-Research. Create Ingestor Plugin is an agent skill from NomaDamas/AutoRAG-Research. Guide developers through creating a custom data ingestor plugin for AutoRAG-Research.

When should I use Create Ingestor Plugin?

Create Ingestor Plugin fits situations like: ingesting a new dataset format into AutoRAG-Research; tasks that involve Model hubs and datasets.

How do I install Create Ingestor Plugin in Claude Code?

Run `npx skills add NomaDamas/AutoRAG-Research --skill create-ingestor-plugin -a claude-code`. Or copy the skill folder (.agents/skills/create-ingestor-plugin in NomaDamas/AutoRAG-Research) into .claude/skills/create-ingestor-plugin in your project. Claude Code loads it when a task matches its description.

How do I install Create Ingestor Plugin in Codex?

Run `npx skills add NomaDamas/AutoRAG-Research --skill create-ingestor-plugin -a codex`. Or copy the skill folder (.agents/skills/create-ingestor-plugin in NomaDamas/AutoRAG-Research) into .agents/skills/create-ingestor-plugin in your project. Codex loads it when a task matches its description.

Can I use Create Ingestor Plugin in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NomaDamas/AutoRAG-Research --skill create-ingestor-plugin -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/create-ingestor-plugin, .gemini/skills/create-ingestor-plugin, .github/skills/create-ingestor-plugin and .opencode/skills/create-ingestor-plugin in your project.

What does Create Ingestor Plugin need to run?

Going by SKILL.md and its folder, Create Ingestor Plugin needs the command-line tools its instructions call (pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.

Does Create Ingestor Plugin access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Create Ingestor Plugin safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Create Ingestor Plugin use?

Create Ingestor Plugin is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Create Ingestor Plugin use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Create Ingestor Plugin?

Skills that share tags, products or a category with Create Ingestor Plugin: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars), Dataset Transformation (awslabs/agent-plugins, 912 stars) and Hugging Face Local Model Evals (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Create Ingestor Plugin?

NomaDamas (a GitHub organization) maintains it in NomaDamas/AutoRAG-Research, which has 149 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on August 9, 2026.

Source: NomaDamas/AutoRAG-Research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.