Agent skill

Multi Source Data Integration Extraction

by Drchronx in Drchronx/ai-agent-research-starter-kit

Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.

Custom licenceAuto-check passedDocuments & Office

Install Multi Source Data Integration Extraction

skills CLI
$ npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Drchronx/ai-agent-research-starter-kit multi-source-data-integration-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Drchronx/ai-agent-research-starter-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction' .claude/skills/multi-source-data-integration-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multi-source-data-integration-extraction
GitHub stars
134
Token cost
~671 tokens
SKILL.md length
156 words
Files
7 (incl. scripts, assets)
Skills in repo
53
Repo updated
First seen
Licence
Custom licence

At a glance

Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.

  • Works in 3 steps: 安装依赖 → 基本用法 → 输出文件
  • Tasks that involve Excel spreadsheets
  • SKILL.md covers 功能特性, 快速开始, 支持的格式 and 使用示例, plus 2 more sections
  • Runs Python scripts from its folder; calls python and pip

What it does

Multi Source Data Integration Extraction is an agent skill from Drchronx/ai-agent-research-starter-kit. Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents. Supports OCR for scanned PDFs. Use when the user wants 多源数据整合、Excel/CSV 自动合并、PDF 表格提取、年报表格提取、政策文件表格抽取、OCR 文字识别,or needs a directory of mixed tabular sources transformed into consolidated CSV outputs.

Its SKILL.md is about 670 tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and assets (for example `agents/openai.yaml` and `scripts/merge_extract_sources.py`).

It sits in Documents & Office, covering Excel spreadsheets, CSV and tabular files and Data pipelines and ETL. It works with Microsoft Excel and pandas. The repository describes itself as: AI Agent 科研全流程教学包,能教学生从零部署 Codex、Claude Code、OpenClaw、Hermes 等 Agent,学会使用 Skills、飞书、AMiner、AI4Scholar、Zotero、Obsidian…

When your agent uses it

  • Tasks that involve Excel spreadsheets
  • Tasks that involve CSV and tabular files
  • Tasks that involve Data pipelines and ETL

Example prompts

  • “/multi-source-data-integration-extraction”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. 安装依赖
  2. 基本用法
  3. 输出文件

What it can do on your machine

Read from SKILL.md and the folder at commit aab1133. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multi Source Data Integration Extraction loads about 671 tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 156 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~671

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 156 words (~671 tokens).

name
multi-source-data-integration-extraction

Read the full SKILL.md on GitHub

Files

SKILL.md and 6 other files (scripts, assets) in 综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction of Drchronx/ai-agent-research-starter-kit.

  • SKILL.md
  • agents/openai.yaml
  • assets/company_panel_a.csv
  • assets/company_panel_b.csv
  • assets/policy_tables.txt
  • scripts/merge_extract_sources.py
  • scripts/requirements.txt

Open the folder on GitHubat commit aab1133

Compare with similar skills

Multi Source Data Integration Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multi Source Data Integration Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multi Source Data Integration Extraction this skillDrchronx/ai-agent-research-starter-kit134—~671Automated safety check: PassCustom licence
File Format Conversionpipeshub-ai/pipeshub-ai3.8k—~865Automated safety check: PassApache-2.0
Streaming Export Safetydoccker/cc-use-exp1.1k—~1.9kAutomated safety check: PassCustom licence
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Research Integrity Auditxuzhougeng/wisp-science1k—~2.6kAutomated safety check: PassAGPL-3.0
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT

Similar skills

  • File Format Conversion

    pipeshub-ai/pipeshub-ai

    Picks the right library for converting between CSV, XLSX, JSON, images, DOCX and PDF text, and lists the conversions that are not supported so the agent does not attempt them.

    3.8k GitHub stars~865 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Streaming Export Safety

    doccker/cc-use-exp

    当代码涉及 Excel/CSV/JSON/PDF 大文件导出、批量序列化、内存里构建大对象时触发。防止 OOM、临时文件残留、同步导出阻塞 HTTP 线程等内存安全陷阱。

    1.1k GitHub stars~1.9k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Research Integrity Audit

    xuzhougeng/wisp-science

    学术审查 / research-integrity screening of a manuscript's figures and reported numbers.

    1k GitHub stars~2.6k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • Excel Parser

    Harryoung/efka

    Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.

    104 GitHub stars~2.3k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed

More from Drchronx/ai-agent-research-starter-kit

All 53 skills in this repo
  • Literature Reviewer Skill

    Drchronx/ai-agent-research-starter-kit

    Build high-quality literature reviews from a research topic using a 10-phase workflow.

    134 GitHub stars~2.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Scenario Experiment Analysis

    Drchronx/ai-agent-research-starter-kit

    Analyze scenario/vignette experiment datasets for behavioral research.

    134 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Scenario Experiment Benchmark Mining

    Drchronx/ai-agent-research-starter-kit

    Mine and synthesize real top-journal scenario/vignette experiment patterns for behavioral research.

    134 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Web Browsing

    Drchronx/ai-agent-research-starter-kit

    Browse and summarize websites, extract content from URLs, search the web for information.

    134 GitHub stars~593 tokensUpdated 4 mo ago
    Auto-check passed
  • Scientific Slides

    Drchronx/ai-agent-research-starter-kit

    Build slide decks and presentations for research talks using Nano Banana Pro AI.

    134 GitHub stars~9.8k tokensUpdated 4 mo ago
    Auto-check: notes
  • Academic Research

    Drchronx/ai-agent-research-starter-kit

    Search academic papers and conduct literature reviews using OpenAlex API (free, no key needed).

    134 GitHub stars~783 tokensUpdated 4 mo ago
    Auto-check passed

Questions about Multi Source Data Integration Extraction

What does Multi Source Data Integration Extraction do?

Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents. Multi Source Data Integration Extraction is an agent skill from Drchronx/ai-agent-research-starter-kit. Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.

When should I use Multi Source Data Integration Extraction?

Multi Source Data Integration Extraction fits situations like: tasks that involve Excel spreadsheets; tasks that involve CSV and tabular files; tasks that involve Data pipelines and ETL.

How do I install Multi Source Data Integration Extraction in Claude Code?

Run `npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a claude-code`. Or copy the skill folder (综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction in Drchronx/ai-agent-research-starter-kit) into .claude/skills/multi-source-data-integration-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Multi Source Data Integration Extraction in Codex?

Run `npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a codex`. Or copy the skill folder (综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction in Drchronx/ai-agent-research-starter-kit) into .agents/skills/multi-source-data-integration-extraction in your project. Codex loads it when a task matches its description.

Can I use Multi Source Data Integration Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multi-source-data-integration-extraction, .gemini/skills/multi-source-data-integration-extraction, .github/skills/multi-source-data-integration-extraction and .opencode/skills/multi-source-data-integration-extraction in your project.

What does Multi Source Data Integration Extraction need to run?

Going by SKILL.md and its folder, Multi Source Data Integration Extraction needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3.

Does Multi Source Data Integration Extraction access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Multi Source Data Integration Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Multi Source Data Integration Extraction use?

Multi Source Data Integration Extraction has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Multi Source Data Integration Extraction use?

About 671 tokens (SKILL.md is roughly 2.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Multi Source Data Integration Extraction?

Skills that share tags, products or a category with Multi Source Data Integration Extraction: File Format Conversion (pipeshub-ai/pipeshub-ai, 3.8k stars), Streaming Export Safety (doccker/cc-use-exp, 1.1k stars), Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars) and Research Integrity Audit (xuzhougeng/wisp-science, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multi Source Data Integration Extraction?

Drchronx (a GitHub user) maintains it in Drchronx/ai-agent-research-starter-kit, which has 134 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on May 19, 2026.

Source: Drchronx/ai-agent-research-starter-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.