Agent skill

Ma Data Extraction

by htlin222 in htlin222/meta-pipe

Define extraction schema, extract study data from full texts, and store it in a structured database for meta-analysis.

Custom licenceAuto-check passedResearch & Science

Install Ma Data Extraction

skills CLI
$ npx skills add htlin222/meta-pipe --skill ma-data-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install htlin222/meta-pipe ma-data-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/htlin222/meta-pipe.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ma-data-extraction .claude/skills/ma-data-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ma-data-extraction
GitHub stars
134
Token cost
~1.6k tokens
SKILL.md length
576 words
Files
11 (incl. scripts, references)
Skills in repo
15
Repo updated
First seen
Licence
Custom licence

At a glance

Define extraction schema, extract study data from full texts, and store it in a structured database for meta-analysis.

  • Works in 2 steps: Web-Based Extraction (Default — Run First) → PDF-Based Extraction (Only for Gaps)
  • Moving from full-text collection to statistical analysis
  • SKILL.md covers Overview, Inputs, Outputs and Workflow (Web-First Hybrid —…, plus 5 more sections
  • Runs Python scripts from its folder; calls uv; reaches pubmed.ncbi.nlm.nih.gov and clinicaltrials.gov

What it does

Ma Data Extraction is an agent skill from htlin222/meta-pipe. Define extraction schema, extract study data from full texts, and store it in a structured database for meta-analysis. Use when moving from full-text collection to statistical analysis.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `references/data-dictionary-template.md`, `references/open-data-sources.md` and `references/web-extraction.md`).

It sits in Research & Science, covering Statistics. It works with SQLite. The repository describes itself as: Claude Code-powered end-to-end meta-analysis automation: AI-assisted literature review, screening, extraction, analysis, and manuscript generation for systematic reviews and….

When your agent uses it

  • Moving from full-text collection to statistical analysis
  • Tasks that involve Statistics

Example prompts

  • “/ma-data-extraction”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Web-Based Extraction (Default — Run First)
  2. PDF-Based Extraction (Only for Gaps)

What it can do on your machine

Read from SKILL.md and the folder at commit 5c5c3f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • pubmed.ncbi.nlm.nih.gov
    • clinicaltrials.gov
    • europepmc.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ma Data Extraction loads about 1.6k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 51 tokens; SKILL.md has 576 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~51
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 576 words (~1,629 tokens).

“Extract consistent data, capture provenance, and build a clean analysis dataset.”

— opening of SKILL.md by htlin222, Custom licence
name
ma-data-extraction

Read the full SKILL.md on GitHub

Files

SKILL.md and 10 other files (scripts, references) in ma-data-extraction of htlin222/meta-pipe.

  • SKILL.md
  • references/data-dictionary-template.md
  • references/open-data-sources.md
  • references/source-template.csv
  • references/study-map-template.csv
  • references/web-extraction.md
  • scripts/init_extraction_db.py
  • scripts/llm_extract.py
  • scripts/validate_sources.py
  • templates/prognostic_factor.md
  • templates/prognostic_factor.yaml

Open the folder on GitHubat commit 5c5c3f0

Compare with similar skills

Ma Data Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ma Data Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ma Data Extraction this skillhtlin222/meta-pipe134—~1.6kAutomated safety check: PassCustom licence
Oleafly Review ManuscriptOleafly/Oleafly206—~2.3kAutomated safety check: PassMIT
Peer ReviewK-Dense-AI/claude-scientific-writer2.4k2 repos~3.1kAutomated safety check: NotesMIT
Manuscript Statistics AuditYuan1z0825/nature-skills46k2 repos~2.1kAutomated safety check: PassApache-2.0
Experimental DesignOleafly/Oleafly2064 repos~3.5kAutomated safety check: NotesMIT
Manuscript Writing Reviewlabarba/sciwrite852—~2.4kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Run a referee-grade review of the manuscript in the open project without changing a single manuscript file.

    206 GitHub stars~2.3k tokensUpdated today
    Research & ScienceAuto-check passed
  • Peer Review

    K-Dense-AI/claude-scientific-writer

    Prepare evidence-bounded, constructive peer-review drafts and structured manuscript assessments.

    2.4k GitHub starsUsed in 2 repos~3.1k tokens
    Research & ScienceAuto-check: notes
  • Manuscript Statistics Audit

    Yuan1z0825/nature-skills

    Audits or rewrites the statistical reporting in a manuscript: experimental units, replication, tests, uncertainty and figure legends, without inventing missing details.

    46k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Experimental Design

    Oleafly/Oleafly

    Design experiments and studies BEFORE data is collected — choosing a design, randomizing, blocking, and laying out treatment combinations so results are interpretable.

    206 GitHub starsUsed in 4 repos~3.5k tokens
    Research & ScienceAuto-check: notes
  • A skill your agent uses when asked to review, edit, or improve the writing quality of a scientific or engineering manuscript.

    852 GitHub stars~2.4k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Math Tools

    ananddtyagi/cc-marketplace

    Deterministic mathematical computation using SymPy. An agent skill from ananddtyagi/cc-marketplace.

    687 GitHub starsUsed in 2 repos~1.3k tokens
    Research & ScienceAuto-check passed

More from htlin222/meta-pipe

All 15 skills in this repo
  • Ma End To End

    htlin222/meta-pipe

    End-to-end AI-assisted meta-analysis pipeline orchestration from TOPIC.txt to final manuscript and reviewer responses.

    134 GitHub stars~2.3k tokensUpdated 15 days ago
    Auto-check passed
  • Ma Fulltext Management

    htlin222/meta-pipe

    Collect and manage full-text PDFs for included studies, track provenance, and prepare documents for extraction.

    134 GitHub stars~1.7k tokensUpdated 15 days ago
    Auto-check: notes
  • Ma Screening Quality

    htlin222/meta-pipe

    Perform title and abstract screening, apply inclusion and exclusion criteria, and assess study quality or risk of bias.

    134 GitHub stars~2.2k tokensUpdated 15 days ago
    Auto-check passed
  • Ma Search Bibliography

    htlin222/meta-pipe

    Conduct literature searches for meta-analysis using Python with uv, query PubMed and other databases, deduplicate results, and store round-based bibliographies with notes.

    134 GitHub stars~2.1k tokensUpdated 15 days ago
    Auto-check: notes
  • Ma Manuscript Quarto

    htlin222/meta-pipe

    Draft and render a meta-analysis manuscript with Quarto using an IMRaD structure and embedded figures/tables.

    134 GitHub stars~7.9k tokensUpdated 15 days ago
    Auto-check passed
  • Ma Meta Analysis

    htlin222/meta-pipe

    Run statistical meta-analysis in R with renv, generate effect estimates, heterogeneity, and publication bias diagnostics, and export figures and tables.

    134 GitHub stars~985 tokensUpdated 15 days ago
    Auto-check passed

Works with

Questions about Ma Data Extraction

What does Ma Data Extraction do?

Define extraction schema, extract study data from full texts, and store it in a structured database for meta-analysis. Ma Data Extraction is an agent skill from htlin222/meta-pipe. Define extraction schema, extract study data from full texts, and store it in a structured database for meta-analysis.

When should I use Ma Data Extraction?

Ma Data Extraction fits situations like: moving from full-text collection to statistical analysis; tasks that involve Statistics.

How do I install Ma Data Extraction in Claude Code?

Run `npx skills add htlin222/meta-pipe --skill ma-data-extraction -a claude-code`. Or copy the skill folder (ma-data-extraction in htlin222/meta-pipe) into .claude/skills/ma-data-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Ma Data Extraction in Codex?

Run `npx skills add htlin222/meta-pipe --skill ma-data-extraction -a codex`. Or copy the skill folder (ma-data-extraction in htlin222/meta-pipe) into .agents/skills/ma-data-extraction in your project. Codex loads it when a task matches its description.

Can I use Ma Data Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add htlin222/meta-pipe --skill ma-data-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ma-data-extraction, .gemini/skills/ma-data-extraction, .github/skills/ma-data-extraction and .opencode/skills/ma-data-extraction in your project.

What does Ma Data Extraction need to run?

Going by SKILL.md and its folder, Ma Data Extraction needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Ma Data Extraction access the network?

SKILL.md names 3 domains. In commands or code: pubmed.ncbi.nlm.nih.gov, clinicaltrials.gov and europepmc.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Ma Data Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ma Data Extraction use?

Ma Data Extraction has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Ma Data Extraction use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.

What are the alternatives to Ma Data Extraction?

Skills that share tags, products or a category with Ma Data Extraction: Oleafly Review Manuscript (Oleafly/Oleafly, 206 stars), Peer Review (K-Dense-AI/claude-scientific-writer, 2.4k stars), Manuscript Statistics Audit (Yuan1z0825/nature-skills, 46k stars) and Experimental Design (Oleafly/Oleafly, 206 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ma Data Extraction?

htlin222 (a GitHub user) maintains it in htlin222/meta-pipe, which has 134 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on September 23, 2026.

Source: htlin222/meta-pipe on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.