Adds a new input format loader (e.g., Parquet, Excel) to the IO readers module.

MITAuto-check passedDocuments & Office

Install Add Dataset Format

skills CLI
$ npx skills add FabioYanezRomero/Knowledge-Graph-Builder --skill add-dataset-format -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FabioYanezRomero/Knowledge-Graph-Builder add-dataset-format --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FabioYanezRomero/Knowledge-Graph-Builder.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agent/skills/add-dataset-format .claude/skills/add-dataset-format && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-dataset-format
GitHub stars
103
Token cost
~1.9k tokens
SKILL.md length
286 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Adds a new input format loader (e.g., Parquet, Excel) to the IO readers module.

  • Works in 4 steps: Understand the Interface → Implement Your Loader → Register in Dispatcher → …
  • Tasks that involve Excel spreadsheets
  • SKILL.md covers Overview, Architecture, Dependencies and DataLoadError Class, plus 10 more sections
  • Calls pip and python

What it does

Add Dataset Format is an agent skill from FabioYanezRomero/Knowledge-Graph-Builder. Adds a new input format loader (e.g., Parquet, Excel) to the IO readers module.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Excel spreadsheets and DataFrames. It works with Microsoft Excel. The repository describes itself as: Repository for building knowledge graphs from specific datasets using generative language model through ollama. The licence is MIT.

When your agent uses it

  • Tasks that involve Excel spreadsheets
  • Tasks that involve DataFrames

Example prompts

  • “/add-dataset-format”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Understand the Interface
  2. Implement Your Loader
  3. Register in Dispatcher
  4. Verify

What it can do on your machine

Read from SKILL.md and the folder at commit 588f0d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Dataset Format loads about 1.9k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 286 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FabioYanezRomero/Knowledge-Graph-Builder at commit 588f0d9, republished under its MIT licence (© FabioYanezRomero). 286 words, ~1,923 tokens.

Download SKILL.mdSave it as .claude/skills/add-dataset-format/SKILL.md (or your agent's skills folder).
name
add-dataset-format
description
Adds a new input format loader (e.g., Parquet, Excel) to the IO readers module.

Adding a Dataset Format

This skill documents how to add a new input format loader to kgb/io/readers/.

Overview

Dataset loaders parse input files and return normalized records for extraction. The system provides:

  • JSONL for large datasets (streaming-friendly)
  • JSON for small datasets
  • CSV for spreadsheet exports
  • Extensible architecture for custom formats

Architecture

                        IO Module
    ┌───────────────────────────────────────────────────────────┐
    │                                                           │
    │  io/readers/__init__.py                                   │
    │  ├─ DataLoadError         # Custom exception              │
    │  ├─ detect_format()       # Auto-detect from extension    │
    │  ├─ load_records()        # Main dispatcher + normalizer  │
    │  │                                                        │
    │  ├─ _load_jsonl()         # JSONL loader (one obj/line)   │
    │  ├─ _load_json()          # JSON loader (array of objs)   │
    │  ├─ _load_csv()           # CSV loader (tabular)          │
    │  └─ _load_<format>()      # Your new loader               │
    │                                                           │
    │  io/writers/                                              │
    │  └─ graphml.py            # Output conversion (separate)  │
    │                                                           │
    └───────────────────────────────────────────────────────────┘

Data Flow:
  File → detect_format() → _load_<format>() → load_records() → Normalized Records

Dependencies

FormatRequired LibraryInstallation
JSONLjson (stdlib)Built-in
JSONjson (stdlib)Built-in
CSVcsv (stdlib)Built-in
Parquetpyarrow>=10.0pip install pyarrow
Excelopenpyxl>=3.0pip install openpyxl

DataLoadError Class

The custom exception used by all loaders (defined in kgb/io/readers/__init__.py):

python
class DataLoadError(Exception):
    """Raised when data loading fails."""
    def __init__(self, message: str, path: Path, line_number: int | None = None):
        super().__init__(message)
        self.path = path
        self.line_number = line_number

Step 1: Understand the Interface

All loaders follow this pattern:

python
def _load_<format>(path: Path) -> list[dict[str, Any]]:
    """Load records from <Format> file.

    Args:
        path: Path to input file

    Returns:
        List of record dictionaries (unnormalized)

    Raises:
        DataLoadError: If file cannot be parsed
    """

Loaders return raw records. Field normalization (mapping custom field names to text/id) happens in load_records().

Step 2: Implement Your Loader

Add to kgb/io/readers/__init__.py:

python
def _load_parquet(path: Path) -> list[dict[str, Any]]:
    """Load records from Parquet file.

    Requires: pip install pandas pyarrow
    """
    try:
        import pandas as pd
    except ImportError:
        raise DataLoadError(
            "Parquet support requires pandas: pip install pandas pyarrow",
            path
        )

    try:
        df = pd.read_parquet(path)
    except Exception as e:
        raise DataLoadError(f"Failed to read Parquet file: {e}", path) from e

    if df.empty:
        raise DataLoadError("Parquet file is empty", path)

    return df.to_dict('records')

Step 3: Register in Dispatcher

Update detect_format()
python
def detect_format(path: Path) -> str:
    suffix = path.suffix.lower()
    if suffix == '.jsonl':
        return 'jsonl'
    elif suffix == '.json':
        return 'json'
    elif suffix == '.csv':
        return 'csv'
    elif suffix in ('.parquet', '.pq'):  # Add new extensions
        return 'parquet'
    else:
        raise DataLoadError(
            f"Unknown file format: {suffix}. Supported: .jsonl, .json, .csv, .parquet",
            path
        )
Update load_records() Dispatcher

In the existing load_records() function, add the new format branch:

python
if format_type == 'jsonl':
    records = _load_jsonl(path)
elif format_type == 'json':
    records = _load_json(path)
elif format_type == 'csv':
    records = _load_csv(path)
elif format_type == 'parquet':  # Add new format
    records = _load_parquet(path)

Step 4: Verify

Check Import
bash
python -c "from kgb.io import load_records; print('OK')"
Unit Tests
python
def test_load_parquet_basic(tmp_path):
    import pandas as pd
    from kgb.io import load_records

    df = pd.DataFrame({
        "doc_id": ["1", "2"],
        "content": ["Text A", "Text B"],
        "meta": ["X", "Y"]
    })

    parquet_file = tmp_path / "test.parquet"
    df.to_parquet(parquet_file)

    records = load_records(
        parquet_file,
        text_field="content",
        id_field="doc_id"
    )

    assert len(records) == 2
    assert records[0]["text"] == "Text A"  # Normalized
    assert records[0]["id"] == "1"


def test_load_parquet_missing_field(tmp_path):
    import pandas as pd
    from kgb.io import load_records, DataLoadError
    import pytest

    df = pd.DataFrame({"wrong_field": ["A", "B"]})
    parquet_file = tmp_path / "bad.parquet"
    df.to_parquet(parquet_file)

    with pytest.raises(DataLoadError, match="Missing text field"):
        load_records(parquet_file)


def test_detect_format_parquet():
    from kgb.io.readers import detect_format
    from pathlib import Path

    assert detect_format(Path("data.parquet")) == "parquet"
    assert detect_format(Path("data.pq")) == "parquet"

Field Normalization

All loaders return raw records. load_records() normalizes field names:

Input (data.csv):

csv
doc_id,body,author
1,"Sample text","Alice"

Load Call:

python
records = load_records(Path("data.csv"), text_field="body", id_field="doc_id")

Output:

python
[{"text": "Sample text", "id": "1", "body": "Sample text", "doc_id": "1", "author": "Alice"}]

Original fields are preserved alongside normalized text and id keys.

CLI Usage

bash
# Auto-detect format from extension
kgb extract --input data.parquet --domain legal

# Custom field names
kgb extract --input data.csv --text-field content --id-field doc_id

# Filter specific records
kgb extract --input large.jsonl --domain legal --record-ids REC-001,REC-002

Key Principles

PrincipleImplementation
Raw RecordsLoaders return unnormalized dicts
Lazy ImportImport heavy libs inside function
Clear ErrorsUse DataLoadError with path and context
EncodingUse utf-8 or utf-8-sig (BOM) for CSV

Error Handling

ExceptionWhenAction
DataLoadErrorParse failure, empty file, missing fieldFail with path
FileNotFoundErrorFile doesn't existRaised by load_records via DataLoadError
ImportErrorMissing optional libraryCatch, raise DataLoadError

Files to Modify

FileAction
kgb/io/readers/__init__.pyModify — add _load_<format>(), update detect_format() and load_records()

Verification Checklist

  • Loader returns list[dict[str, Any]]
  • Uses DataLoadError for all failures
  • Lazy imports for optional dependencies
  • Extension(s) added to detect_format()
  • Dispatcher updated in load_records()
  • Tests for happy path, errors, empty file

© FabioYanezRomero, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agent/skills/add-dataset-format of FabioYanezRomero/Knowledge-Graph-Builder.

Open the folder on GitHubat commit 588f0d9

Compare with similar skills

Add Dataset Format next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Dataset Format compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Dataset Format this skillFabioYanezRomero/Knowledge-Graph-Builder103—~1.9kAutomated safety check: PassMIT
Convert Fileduckdb/duckdb-skills6031 repos~720Automated safety check: NotesMIT
Excel ParserHarryoung/efka104—~2.3kAutomated safety check: PassApache-2.0
Sn Da Image CaptionMichaelYang-lyx/AIDABench1111 repos~2kAutomated safety check: PassNone
Sn Da Excel WorkflowMichaelYang-lyx/AIDABench1111 repos~2.5kAutomated safety check: PassNone
Generate CodebookAperivue/medsci-skills333—~1.1kAutomated safety check: PassMIT

Similar skills

  • Convert File

    duckdb/duckdb-skills

    Official

    Convert any data file to another format: CSV, Parquet, JSON, Excel, GeoJSON, and more.

    603 GitHub starsUsed in 1 repo~720 tokens
    Documents & OfficeAuto-check: notes
  • Excel Parser

    Harryoung/efka

    Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.

    104 GitHub stars~2.3k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • Sn Da Image Caption

    MichaelYang-lyx/AIDABench

    图片理解与数据提取 skill。当图片文件(.png/.jpg/.jpeg/.gif/.webp/.bmp)是主要输入且用户需要理解、提取数据或分析图片内容时使用。提供预配置的 caption 脚本(scripts/caption.py),通过 vision 模型将图片转为文本描述,无需额外配置 API Key。覆盖:(1) 通过 scripts/caption.py…

    111 GitHub starsUsed in 1 repo~2k tokens
    Documents & OfficeAuto-check passed
  • Sn Da Excel Workflow

    MichaelYang-lyx/AIDABench

    Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子…

    111 GitHub starsUsed in 1 repo~2.5k tokens
    Documents & OfficeAuto-check passed
  • Generate Codebook

    Aperivue/medsci-skills

    A skill your agent uses when a tabular dataset (CSV, Excel, Parquet, Stata, SAS) needs a data dictionary.

    333 GitHub stars~1.1k tokensUpdated 6 days ago
    Documents & OfficeAuto-check passed
  • Sn Da Large File Analysis

    MichaelYang-lyx/AIDABench

    万行以上 Excel 数据集的高性能分析引擎。提供 openpyxl readonly 流式读取(iterrows 支持 10 万行以上)、Parquet 转换加速、内存优化、分块处理和大文件写入模式。遇到以下任一情况就主动使用本 skill:①数据行数 ≥ 10k(由 sn-da-excel-workflow 的行数评估步骤触发);②用户出现触发词:大文件 / 大数据量 / 性能优化 /…

    111 GitHub starsUsed in 1 repo~3.1k tokens
    Documents & OfficeAuto-check passed

More from FabioYanezRomero/Knowledge-Graph-Builder

  • Add LLM Client

    FabioYanezRomero/Knowledge-Graph-Builder

    Adds a new LLM client provider to the clients module. An agent skill from FabioYanezRomero/Knowledge-Graph-Builder.

    103 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Add Augmentation Strategy

    FabioYanezRomero/Knowledge-Graph-Builder

    Adds a new iterative augmentation strategy (e.g., enrichment, summarization) to the builder module.

    103 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Add Converter

    FabioYanezRomero/Knowledge-Graph-Builder

    Adds a new output format converter (e.g., CSV, RDF) to the IO writers module.

    103 GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Add Domain

    FabioYanezRomero/Knowledge-Graph-Builder

    Manage knowledge domains (e.g., Medical, Finance). An agent skill from FabioYanezRomero/Knowledge-Graph-Builder.

    103 GitHub stars~3.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Add Visualization

    FabioYanezRomero/Knowledge-Graph-Builder

    Adds a new visualization engine or style to the visualization module.

    103 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Add Dataset Format

What does Add Dataset Format do?

Adds a new input format loader (e.g., Parquet, Excel) to the IO readers module. Add Dataset Format is an agent skill from FabioYanezRomero/Knowledge-Graph-Builder., Parquet, Excel) to the IO readers module.

When should I use Add Dataset Format?

Add Dataset Format fits situations like: tasks that involve Excel spreadsheets; tasks that involve DataFrames.

How do I install Add Dataset Format in Claude Code?

Run `npx skills add FabioYanezRomero/Knowledge-Graph-Builder --skill add-dataset-format -a claude-code`. Or copy the skill folder (.agent/skills/add-dataset-format in FabioYanezRomero/Knowledge-Graph-Builder) into .claude/skills/add-dataset-format in your project. Claude Code loads it when a task matches its description.

How do I install Add Dataset Format in Codex?

Run `npx skills add FabioYanezRomero/Knowledge-Graph-Builder --skill add-dataset-format -a codex`. Or copy the skill folder (.agent/skills/add-dataset-format in FabioYanezRomero/Knowledge-Graph-Builder) into .agents/skills/add-dataset-format in your project. Codex loads it when a task matches its description.

Can I use Add Dataset Format in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FabioYanezRomero/Knowledge-Graph-Builder --skill add-dataset-format -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-dataset-format, .gemini/skills/add-dataset-format, .github/skills/add-dataset-format and .opencode/skills/add-dataset-format in your project.

What does Add Dataset Format need to run?

Going by SKILL.md and its folder, Add Dataset Format needs the command-line tools its instructions call (pip and python). Our summary lists: Python 3.

Does Add Dataset Format access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Add Dataset Format safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Dataset Format use?

Add Dataset Format is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Dataset Format use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Dataset Format?

Skills that share tags, products or a category with Add Dataset Format: Convert File (duckdb/duckdb-skills, 603 stars), Excel Parser (Harryoung/efka, 104 stars), Sn Da Image Caption (MichaelYang-lyx/AIDABench, 111 stars) and Sn Da Excel Workflow (MichaelYang-lyx/AIDABench, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Dataset Format?

FabioYanezRomero (a GitHub user) maintains it in FabioYanezRomero/Knowledge-Graph-Builder, which has 103 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on August 28, 2026.

Source: FabioYanezRomero/Knowledge-Graph-Builder on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.