File Format Conversion
pipeshub-ai/pipeshub-ai
Picks the right library for converting between CSV, XLSX, JSON, images, DOCX and PDF text, and lists the conversions that are not supported so the agent does not attempt them.
Agent skill
Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.
$ npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Drchronx/ai-agent-research-starter-kit multi-source-data-integration-extraction --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Drchronx/ai-agent-research-starter-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction' .claude/skills/multi-source-data-integration-extraction && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "multi-source-data-integration-extraction" agent skill from https://github.com/Drchronx/ai-agent-research-starter-kit/tree/main/%E7%BB%BC%E5%90%88%E5%AD%A6%E6%9C%AF%E9%83%A8%E7%BD%B2%E5%8C%85/04_%E6%95%B0%E6%8D%AE%E9%87%87%E9%9B%86%E6%95%B4%E5%90%88%E4%B8%8E%E5%AE%9E%E8%AF%81Skills/multi-source-data-integration-extraction into .claude/skills/multi-source-data-integration-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multi-source-data-integration-extraction", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Drchronx/ai-agent-research-starter-kit/tree/main/%E7%BB%BC%E5%90%88%E5%AD%A6%E6%9C%AF%E9%83%A8%E7%BD%B2%E5%8C%85/04_%E6%95%B0%E6%8D%AE%E9%87%87%E9%9B%86%E6%95%B4%E5%90%88%E4%B8%8E%E5%AE%9E%E8%AF%81Skills/multi-source-data-integration-extractionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Drchronx/ai-agent-research-starter-kit multi-source-data-integration-extraction --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Drchronx/ai-agent-research-starter-kit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/'综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction' .agents/skills/multi-source-data-integration-extraction && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "multi-source-data-integration-extraction" agent skill from https://github.com/Drchronx/ai-agent-research-starter-kit/tree/main/%E7%BB%BC%E5%90%88%E5%AD%A6%E6%9C%AF%E9%83%A8%E7%BD%B2%E5%8C%85/04_%E6%95%B0%E6%8D%AE%E9%87%87%E9%9B%86%E6%95%B4%E5%90%88%E4%B8%8E%E5%AE%9E%E8%AF%81Skills/multi-source-data-integration-extraction into .agents/skills/multi-source-data-integration-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multi-source-data-integration-extraction", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Drchronx/ai-agent-research-starter-kit multi-source-data-integration-extraction --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Drchronx/ai-agent-research-starter-kit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/'综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction' .cursor/skills/multi-source-data-integration-extraction && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "multi-source-data-integration-extraction" agent skill from https://github.com/Drchronx/ai-agent-research-starter-kit/tree/main/%E7%BB%BC%E5%90%88%E5%AD%A6%E6%9C%AF%E9%83%A8%E7%BD%B2%E5%8C%85/04_%E6%95%B0%E6%8D%AE%E9%87%87%E9%9B%86%E6%95%B4%E5%90%88%E4%B8%8E%E5%AE%9E%E8%AF%81Skills/multi-source-data-integration-extraction into .cursor/skills/multi-source-data-integration-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multi-source-data-integration-extraction", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Drchronx/ai-agent-research-starter-kit.git --path '综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction'--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Drchronx/ai-agent-research-starter-kit multi-source-data-integration-extraction --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Drchronx/ai-agent-research-starter-kit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/'综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction' .gemini/skills/multi-source-data-integration-extraction && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "multi-source-data-integration-extraction" agent skill from https://github.com/Drchronx/ai-agent-research-starter-kit/tree/main/%E7%BB%BC%E5%90%88%E5%AD%A6%E6%9C%AF%E9%83%A8%E7%BD%B2%E5%8C%85/04_%E6%95%B0%E6%8D%AE%E9%87%87%E9%9B%86%E6%95%B4%E5%90%88%E4%B8%8E%E5%AE%9E%E8%AF%81Skills/multi-source-data-integration-extraction into .gemini/skills/multi-source-data-integration-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multi-source-data-integration-extraction", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Drchronx/ai-agent-research-starter-kit multi-source-data-integration-extractionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Drchronx/ai-agent-research-starter-kit.git skills-src && mkdir -p .github/skills && cp -r skills-src/'综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction' .github/skills/multi-source-data-integration-extraction && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "multi-source-data-integration-extraction" agent skill from https://github.com/Drchronx/ai-agent-research-starter-kit/tree/main/%E7%BB%BC%E5%90%88%E5%AD%A6%E6%9C%AF%E9%83%A8%E7%BD%B2%E5%8C%85/04_%E6%95%B0%E6%8D%AE%E9%87%87%E9%9B%86%E6%95%B4%E5%90%88%E4%B8%8E%E5%AE%9E%E8%AF%81Skills/multi-source-data-integration-extraction into .github/skills/multi-source-data-integration-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multi-source-data-integration-extraction", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Drchronx/ai-agent-research-starter-kit multi-source-data-integration-extraction --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Drchronx/ai-agent-research-starter-kit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/'综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction' .opencode/skills/multi-source-data-integration-extraction && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "multi-source-data-integration-extraction" agent skill from https://github.com/Drchronx/ai-agent-research-starter-kit/tree/main/%E7%BB%BC%E5%90%88%E5%AD%A6%E6%9C%AF%E9%83%A8%E7%BD%B2%E5%8C%85/04_%E6%95%B0%E6%8D%AE%E9%87%87%E9%9B%86%E6%95%B4%E5%90%88%E4%B8%8E%E5%AE%9E%E8%AF%81Skills/multi-source-data-integration-extraction into .opencode/skills/multi-source-data-integration-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multi-source-data-integration-extraction", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
multi-source-data-integration-extractionAutomatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.
Multi Source Data Integration Extraction is an agent skill from Drchronx/ai-agent-research-starter-kit. Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents. Supports OCR for scanned PDFs. Use when the user wants 多源数据整合、Excel/CSV 自动合并、PDF 表格提取、年报表格提取、政策文件表格抽取、OCR 文字识别,or needs a directory of mixed tabular sources transformed into consolidated CSV outputs.
Its SKILL.md is about 670 tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and assets (for example `agents/openai.yaml` and `scripts/merge_extract_sources.py`).
It sits in Documents & Office, covering Excel spreadsheets, CSV and tabular files and Data pipelines and ETL. It works with Microsoft Excel and pandas. The repository describes itself as: AI Agent 科研全流程教学包,能教学生从零部署 Codex、Claude Code、OpenClaw、Hermes 等 Agent,学会使用 Skills、飞书、AMiner、AI4Scholar、Zotero、Obsidian…
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit aab1133. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonpipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Multi Source Data Integration Extraction loads about 671 tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 156 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 156 words (~671 tokens).
SKILL.md and 6 other files (scripts, assets) in 综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction of Drchronx/ai-agent-research-starter-kit.
Open the folder on GitHubat commit aab1133
Multi Source Data Integration Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Multi Source Data Integration Extraction this skillDrchronx/ai-agent-research-starter-kit | 134 | — | ~671 | Automated safety check: Pass | Custom licence | |
| File Format Conversionpipeshub-ai/pipeshub-ai | 3.8k | — | ~865 | Automated safety check: Pass | Apache-2.0 | |
| Streaming Export Safetydoccker/cc-use-exp | 1.1k | — | ~1.9k | Automated safety check: Pass | Custom licence | |
| Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences | 274 | 2 repos | ~2.7k | Automated safety check: Pass | Apache-2.0 | |
| Research Integrity Auditxuzhougeng/wisp-science | 1k | — | ~2.6k | Automated safety check: Pass | AGPL-3.0 | |
| MineruNebutra/MinerU-Skill | 122 | — | ~504 | Automated safety check: Pass | MIT |
pipeshub-ai/pipeshub-ai
Picks the right library for converting between CSV, XLSX, JSON, images, DOCX and PDF text, and lists the conversions that are not supported so the agent does not attempt them.
doccker/cc-use-exp
当代码涉及 Excel/CSV/JSON/PDF 大文件导出、批量序列化、内存里构建大对象时触发。防止 OOM、临时文件残留、同步导出阻塞 HTTP 线程等内存安全陷阱。
aws-samples/amazon-bedrock-agents-healthcare-lifesciences
Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.
xuzhougeng/wisp-science
学术审查 / research-integrity screening of a manuscript's figures and reported numbers.
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.
Harryoung/efka
Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.
Drchronx/ai-agent-research-starter-kit
Build high-quality literature reviews from a research topic using a 10-phase workflow.
Drchronx/ai-agent-research-starter-kit
Analyze scenario/vignette experiment datasets for behavioral research.
Drchronx/ai-agent-research-starter-kit
Mine and synthesize real top-journal scenario/vignette experiment patterns for behavioral research.
Drchronx/ai-agent-research-starter-kit
Browse and summarize websites, extract content from URLs, search the web for information.
Drchronx/ai-agent-research-starter-kit
Build slide decks and presentations for research talks using Nano Banana Pro AI.
Drchronx/ai-agent-research-starter-kit
Search academic papers and conduct literature reviews using OpenAlex API (free, no key needed).
Works with
Categories
Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents. Multi Source Data Integration Extraction is an agent skill from Drchronx/ai-agent-research-starter-kit. Automatically merge scattered Excel and CSV files, normalize column names, and extract structured tables from PDF, HTML, TXT, or Markdown documents.
Multi Source Data Integration Extraction fits situations like: tasks that involve Excel spreadsheets; tasks that involve CSV and tabular files; tasks that involve Data pipelines and ETL.
Run `npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a claude-code`. Or copy the skill folder (综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction in Drchronx/ai-agent-research-starter-kit) into .claude/skills/multi-source-data-integration-extraction in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a codex`. Or copy the skill folder (综合学术部署包/04_数据采集整合与实证Skills/multi-source-data-integration-extraction in Drchronx/ai-agent-research-starter-kit) into .agents/skills/multi-source-data-integration-extraction in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Drchronx/ai-agent-research-starter-kit --skill multi-source-data-integration-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multi-source-data-integration-extraction, .gemini/skills/multi-source-data-integration-extraction, .github/skills/multi-source-data-integration-extraction and .opencode/skills/multi-source-data-integration-extraction in your project.
Going by SKILL.md and its folder, Multi Source Data Integration Extraction needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Multi Source Data Integration Extraction has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.
About 671 tokens (SKILL.md is roughly 2.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Multi Source Data Integration Extraction: File Format Conversion (pipeshub-ai/pipeshub-ai, 3.8k stars), Streaming Export Safety (doccker/cc-use-exp, 1.1k stars), Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars) and Research Integrity Audit (xuzhougeng/wisp-science, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Drchronx (a GitHub user) maintains it in Drchronx/ai-agent-research-starter-kit, which has 134 GitHub stars. The repository holds 53 skills in this directory. The repository was last updated on May 19, 2026.
Source: Drchronx/ai-agent-research-starter-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.