Agent skill

Unstructured Medical Text Miner

by aipoch in aipoch/medical-research-skills

Mine unstructured clinical text from MIMIC-IV to extract diagnostic logic.

MITAuto-check passedResearch & Science

Install Unstructured Medical Text Miner

skills CLI
$ npx skills add aipoch/medical-research-skills --skill unstructured-medical-text-miner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills unstructured-medical-text-miner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Evidence Insight/unstructured-medical-text-miner' .claude/skills/unstructured-medical-text-miner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
unstructured-medical-text-miner
GitHub stars
1.9k
Token cost
~2.9k tokens
SKILL.md length
1,040 words
Files
8 (incl. scripts, references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

Mine unstructured clinical text from MIMIC-IV to extract diagnostic logic.

  • Works in 4 steps: Text Extraction → Information Extraction → Clinical Logic Parsing → …
  • Research & Science work in your project
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 16 more sections
  • Runs Python scripts from its folder; calls python

What it does

Unstructured Medical Text Miner is an agent skill from aipoch/medical-research-skills. Mine unstructured clinical text from MIMIC-IV to extract diagnostic logic.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `config.yaml`, `example.py` and `references/audit-reference.md`).

It sits in Research & Science. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Research & Science work in your project

Example prompts

  • “/unstructured-medical-text-miner”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Text Extraction
  2. Information Extraction
  3. Clinical Logic Parsing
  4. Structured Output

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • physionet.org
    • allenai.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Unstructured Medical Text Miner loads about 2.9k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 27 tokens; SKILL.md has 1,040 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 1,040 words, ~2,868 tokens.

Download SKILL.mdSave it as .claude/skills/unstructured-medical-text-miner/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
unstructured-medical-text-miner
description
Mine unstructured clinical text from MIMIC-IV to extract diagnostic logic.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

Unstructured Medical Text Miner (ID: 213)

When to Use

  • Use this skill when the task needs Mine unstructured clinical text from MIMIC-IV to extract diagnostic logic.
  • Use this skill for evidence insight tasks that require explicit assumptions, bounded scope, and a reproducible output format.
  • Use this skill when you need a documented fallback path for missing inputs, execution errors, or partial evidence.

Key Features

See ## Features above for related details.

  • Scope-focused workflow aligned to: Mine unstructured clinical text from MIMIC-IV to extract diagnostic logic.
  • Packaged executable path(s): scripts/__init__.py plus 1 additional script(s).
  • Reference material available in references/ for task-specific guidance.
  • Structured execution path designed to keep outputs consistent and reviewable.

Dependencies

pandas>=1.3.0
spacy>=3.4.0
scispacy>=0.5.1
radlex (for radiology terminology)
negspacy (for negation detection)

Example Usage

See ## Usage above for related details.

bash
cd "20260318/scientific-skills/Evidence Insight/unstructured-medical-text-miner"
python -m py_compile scripts/main.py
python scripts/main.py --help

Example run plan:

  1. Confirm the user input, output path, and any required config values.
  2. Edit the in-file CONFIG block or documented parameters if the script uses fixed settings.
  3. Run python scripts/main.py with the validated inputs.
  4. Review the generated output and return the final artifact with any assumptions called out.

Implementation Details

See ## Workflow above for related details.

  • Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
  • Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
  • Primary implementation surface: scripts/__init__.py with additional helper scripts under scripts/.
  • Reference guidance: references/ contains supporting rules, prompts, or checklists.
  • Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
  • Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.

Quick Check

Use this command to verify that the packaged script entry point can be parsed before deeper execution.

bash
python -m py_compile scripts/main.py

Audit-Ready Commands

Use these concrete commands for validation. They are intentionally self-contained and avoid placeholder paths.

bash
python -m py_compile scripts/main.py
python scripts/main.py --help
python scripts/main.py -h

Workflow

  1. Confirm the user objective, required inputs, and non-negotiable constraints before doing detailed work.
  2. Validate that the request matches the documented scope and stop early if the task would require unsupported assumptions.
  3. Use the packaged script path or the documented reasoning path with only the inputs that are actually available.
  4. Return a structured result that separates assumptions, deliverables, risks, and unresolved items.
  5. If execution fails or inputs are incomplete, switch to the fallback path and state exactly what blocked full completion.

Overview

Mine "text data" that has been long overlooked in MIMIC-IV, extracting unstructured diagnostic logic, order details, and progress notes.

Purpose

The MIMIC-IV database contains large amounts of structured data (vital signs, laboratory results, etc.), but its true clinical value is often hidden in unstructured text:

  • Diagnostic reasoning chains in discharge summaries
  • Subtle finding descriptions in imaging reports
  • Treatment decision logic in progress notes
  • Personalized medication considerations in orders

This Skill provides a complete text mining toolchain to transform raw medical text into analyzable structured insights.

Features

1. Text Extraction
  • NOTEEVENTS: Extract clinical notes from MIMIC-IV NOTE module
  • Radiology Reports: Extract imaging diagnostic text
  • ECG Reports: Parse ECG interpretation text
  • Discharge Summaries: Extract complete diagnostic and treatment course
2. Information Extraction
  • Entity Recognition: Diseases, symptoms, medications, procedures, anatomical sites
  • Relation Extraction: Medication-disease treatment relationships, symptom-disease diagnostic relationships
  • Timeline Extraction: Event occurrence times, disease progression sequence
  • Negation Detection: Identify negated clinical findings (e.g., "no fever")
3. Clinical Logic Parsing
  • Diagnostic Reasoning Chain: Reasoning path from symptoms → examination → diagnosis
  • Treatment Decision Tree: Clinical basis for medication selection and dosage adjustment
  • Disease Progression: Disease progression and outcome descriptions
4. Structured Output
  • FHIR-compatible clinical document format
  • Knowledge graph-friendly triple format
  • Temporal event sequences

Usage

python
from skills.unstructured_medical_text_miner.scripts.main import MedicalTextMiner

# Initialize miner
miner = MedicalTextMiner()

# Load MIMIC-IV note data
miner.load_notes(notes_path="path/to/noteevents.csv")

# Extract all text records for a specific patient
patient_texts = miner.get_patient_texts(subject_id=10000032)

# Execute complete information extraction
insights = miner.extract_insights(
    text=patient_texts,
    extract_entities=True,
    extract_relations=True,
    extract_timeline=True
)

Input

Data Sources
  • MIMIC-IV NOTEEVENTS table (csv/parquet format)
  • Discharge summary files
  • Imaging report files
  • Custom medical text
Field Requirements
Field NameDescriptionRequired
subject_idPatient unique identifierYes
hadm_idHospital admission record identifierNo
note_typeNote type (DS/RR/ECG, etc.)Yes
note_textNote text contentYes
charttimeRecord timeNo
Show full SKILL.md (410 more words)Show less

Output

Entity Extraction Results
json
{
  "entities": [
    {
      "text": "acute myocardial infarction",
      "type": "DISEASE",
      "start": 156,
      "end": 183,
      "confidence": 0.94
    },
    {
      "text": "aspirin 81mg",
      "type": "MEDICATION",
      "start": 245,
      "end": 257,
      "attributes": {
        "dose": "81mg",
        "frequency": "daily"
      }
    }
  ]
}
Clinical Logic Graph
json
{
  "clinical_logic": {
    "presenting_complaint": "chest pain",
    "differential_diagnoses": ["ACS", "PE", "aortic dissection"],
    "workup": ["ECG", "troponin", "CTA chest"],
    "final_diagnosis": "STEMI",
    "treatment_plan": ["PCI", "dual antiplatelet"]
  }
}
Temporal Events
json
{
  "timeline": [
    {
      "time": "2020-03-15 08:30",
      "event": "admission",
      "description": "presented with chest pain"
    },
    {
      "time": "2020-03-15 09:15",
      "event": "ECG",
      "description": "ST elevation in V1-V4"
    }
  ]
}

Configuration

yaml

# config.yaml
extraction:
  entity_types: ["DISEASE", "SYMPTOM", "MEDICATION", "PROCEDURE", "ANATOMY"]
  relation_types: ["TREATS", "CAUSES", "CONTRAINDICATED_WITH"]
  enable_negation_detection: true
  
models:
  ner_model: "en_core_sci_lg"  # or "en_core_sci_scibert"
  relation_model: "custom_relation_extractor"
  
output:
  format: "json"  # json/fhir/kg
  include_raw_text: false

CLI Usage

text

# Process single file
python -m skills.unstructured_medical_text_miner.scripts.main \
  --input notes.csv \
  --output extracted.json \
  --extract all

# Process specific patient
python -m skills.unstructured_medical_text_miner.scripts.main \
  --subject-id 10000032 \
  --db-path mimic_iv.db \
  --output patient_insights.json

References

  1. MIMIC-IV Clinical Database: https://physionet.org/content/mimiciv/
  2. scispacy: https://allenai.github.io/scispacy/
  3. NegEx/negspacy for negation detection
  4. FHIR Clinical Document specifications

Author

Skill ID: 213 Category: Medical Data Mining Complexity: Advanced

Risk Assessment

Risk IndicatorAssessmentLevel
Code ExecutionPython/R scripts executed locallyMedium
Network AccessNo external API callsLow
File System AccessRead input files, write output filesMedium
Instruction TamperingStandard prompt guidelinesLow
Data ExposureOutput files saved to workspaceLow

Security Checklist

  • No hardcoded credentials or API keys
  • No unauthorized file system access (../)
  • Output does not expose sensitive information
  • Prompt injection protections in place
  • Input file paths validated (no ../ traversal)
  • Output directory restricted to workspace
  • Script execution in sandboxed environment
  • Error messages sanitized (no stack traces exposed)
  • Dependencies audited

Prerequisites

text

# Python dependencies
pip install -r requirements.txt

Evaluation Criteria

Success Metrics
  • Successfully executes main functionality
  • Output meets quality standards
  • Handles edge cases gracefully
  • Performance is acceptable
Test Cases
  1. Basic Functionality: Standard input → Expected output
  2. Edge Case: Invalid input → Graceful error handling
  3. Performance: Large dataset → Acceptable processing time

Output Requirements

Every final response should make these items explicit when they are relevant:

  • Objective or requested deliverable
  • Inputs used and assumptions introduced
  • Workflow or decision path
  • Core result, recommendation, or artifact
  • Constraints, risks, caveats, or validation needs
  • Unresolved items and next-step checks

Error Handling

  • If required inputs are missing, state exactly which fields are missing and request only the minimum additional information.
  • If the task goes outside the documented scope, stop instead of guessing or silently widening the assignment.
  • If scripts/main.py fails, report the failure point, summarize what still can be completed safely, and provide a manual fallback.
  • Do not fabricate files, citations, data, search results, or execution outcomes.

Input Validation

This skill accepts requests that match the documented purpose of unstructured-medical-text-miner and include enough context to complete the workflow safely.

Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:

unstructured-medical-text-miner only handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.

References

Response Template

Use the following fixed structure for non-trivial requests:

  1. Objective
  2. Inputs Received
  3. Assumptions
  4. Workflow
  5. Deliverable
  6. Risks and Limits
  7. Next Checks

If the request is simple, you may compress the structure, but still keep assumptions and limits explicit when they affect correctness.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in scientific-skills/Evidence Insight/unstructured-medical-text-miner of aipoch/medical-research-skills.

  • SKILL.md
  • config.yaml
  • example.py
  • references/audit-reference.md
  • requirements.txt
  • scripts/__init__.py
  • scripts/main.py
  • unstructured-medical-text-miner_audit_result_v2.json

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Unstructured Medical Text Miner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Unstructured Medical Text Miner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Unstructured Medical Text Miner this skillaipoch/medical-research-skills1.9k—~2.9kAutomated safety check: PassMIT
Hypothesis Generationspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: NotesMIT
GitHub Deep Researchbytedance/deer-flow84k4 repos~1.3kAutomated safety check: PassMIT
Nature Paper CardYuan1z0825/nature-skills47k2 repos~2.1kAutomated safety check: PassApache-2.0
Content Research Writerweapp-tailwindcss/weapp-tailwindcss1.9k25 repos~3.5kAutomated safety check: PassMIT
Last30daysmvanhorn/last30days-skill64k—~7.9kAutomated safety check: NotesMIT

Similar skills

  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    84k GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    47k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Content Research Writer

    weapp-tailwindcss/weapp-tailwindcss

    Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.

    1.9k GitHub starsUsed in 25 repos~3.5k tokens
    Research & ScienceAuto-check passed
  • Last30days

    mvanhorn/last30days-skill

    Research what people actually say about any topic in the last 30 days.

    64k GitHub stars~7.9k tokensUpdated yesterday
    Research & ScienceAuto-check: notes
  • Peer Review

    spacering-net/codeg

    Structured manuscript/grant review with checklist-based evaluation.

    3.9k GitHub starsUsed in 17 repos~5.9k tokens
    Research & ScienceAuto-check: notes

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    1.9k GitHub stars~2.2k tokensUpdated 24 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    1.9k GitHub stars~1.4k tokensUpdated 24 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    1.9k GitHub stars~3.7k tokensUpdated 24 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    1.9k GitHub stars~1.8k tokensUpdated 24 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    1.9k GitHub stars~1.7k tokensUpdated 24 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    1.9k GitHub stars~1.3k tokensUpdated 24 days ago
    Auto-check passed

Questions about Unstructured Medical Text Miner

What does Unstructured Medical Text Miner do?

Mine unstructured clinical text from MIMIC-IV to extract diagnostic logic. Unstructured Medical Text Miner is an agent skill from aipoch/medical-research-skills. Mine unstructured clinical text from MIMIC-IV to extract diagnostic logic.

When should I use Unstructured Medical Text Miner?

Unstructured Medical Text Miner fits situations like: research & Science work in your project.

How do I install Unstructured Medical Text Miner in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill unstructured-medical-text-miner -a claude-code`. Or copy the skill folder (scientific-skills/Evidence Insight/unstructured-medical-text-miner in aipoch/medical-research-skills) into .claude/skills/unstructured-medical-text-miner in your project. Claude Code loads it when a task matches its description.

How do I install Unstructured Medical Text Miner in Codex?

Run `npx skills add aipoch/medical-research-skills --skill unstructured-medical-text-miner -a codex`. Or copy the skill folder (scientific-skills/Evidence Insight/unstructured-medical-text-miner in aipoch/medical-research-skills) into .agents/skills/unstructured-medical-text-miner in your project. Codex loads it when a task matches its description.

Can I use Unstructured Medical Text Miner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill unstructured-medical-text-miner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/unstructured-medical-text-miner, .gemini/skills/unstructured-medical-text-miner, .github/skills/unstructured-medical-text-miner and .opencode/skills/unstructured-medical-text-miner in your project.

What does Unstructured Medical Text Miner need to run?

Going by SKILL.md and its folder, Unstructured Medical Text Miner needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Unstructured Medical Text Miner access the network?

SKILL.md names 2 domains. As links in the text: physionet.org and allenai.github.io. This is read from the text; nothing was executed.

Is Unstructured Medical Text Miner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Unstructured Medical Text Miner use?

Unstructured Medical Text Miner is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Unstructured Medical Text Miner use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 173 tokens, read only when the agent opens those files.

What are the alternatives to Unstructured Medical Text Miner?

Skills that share tags, products or a category with Unstructured Medical Text Miner: Hypothesis Generation (spacering-net/codeg, 3.9k stars), GitHub Deep Research (bytedance/deer-flow, 84k stars), Nature Paper Card (Yuan1z0825/nature-skills, 47k stars) and Content Research Writer (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Unstructured Medical Text Miner?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,937 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.