Agent skill

Extracting Dicom Metadata

by maziyarpanahi in maziyarpanahi/openmed

Reads DICOM file headers and DICOM-SR (Structured Report) content to pull study/series metadata and embedded report text, and flags PHI carried in header tags.

Apache-2.0Auto-check passedResearch & Science

Install Extracting Dicom Metadata

skills CLI
$ npx skills add maziyarpanahi/openmed --skill extracting-dicom-metadata -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed extracting-dicom-metadata --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/extracting-dicom-metadata .claude/skills/extracting-dicom-metadata && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
extracting-dicom-metadata
GitHub stars
5.5k
Token cost
~1.8k tokens
SKILL.md length
622 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Reads DICOM file headers and DICOM-SR (Structured Report) content to pull study/series metadata and embedded report text, and flags PHI carried in header tags.

  • Works in 5 steps: Read the dataset with pydicom.dcmread… → Walk the SR content tree.… → Inventory PHI tags. Flag the standard… → …
  • Keywords: DICOM
  • SKILL.md covers When to use, DICOM headers in one minute, Quick start and Workflow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Extracting Dicom Metadata is an agent skill from maziyarpanahi/openmed. Reads DICOM file headers and DICOM-SR (Structured Report) content to pull study/series metadata and embedded report text, and flags PHI carried in header tags. Use before OpenMed processing when ingesting imaging data (CT/MR/CR/US, radiology SR) and you need the report narrative de-identified and analyzed, plus a list of header tags that must be scrubbed. Hand SR/report text to openmed.deidentify and openmed.analyzetext; use pydicom to read tags. Trigger keywords: DICOM, pydicom, DICOM-SR, structured report…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Clinical and healthcare research. The repository describes itself as: Local-first healthcare AI: clinical NER and HIPAA PII de-identification on hardware you control. 2,200+ medical models, 35 model-backed PII languages, and Python, MLX, Android… The licence is Apache-2.0.

When your agent uses it

  • Keywords: DICOM
  • Structured report
  • Radiology report

Example prompts

  • “Use the extracting-dicom-metadata skill to read DICOM file headers and DICOM-SR (Structured Report) content to pull study/series metadata and…”
  • “/extracting-dicom-metadata”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Read the dataset with pydicom.dcmread (use stop_before_pixels=True
  2. Walk the SR content tree. ContentSequence nests CONTAINER, TEXT,
  3. Inventory PHI tags. Flag the standard identifier tags and free-text
  4. De-identify → analyze the report narrative with OpenMed.
  5. Scrub the header before any image export using a DICOM de-identification

What it can do on your machine

Read from SKILL.md and the folder at commit 34d7b8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • dicom.nema.org
    • dicomstandard.org
    • pydicom.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Extracting Dicom Metadata loads about 1.8k tokens when it runs. Until then it costs about 150 tokens; SKILL.md has 622 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~150
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 34d7b8c, republished under its Apache-2.0 licence (© maziyarpanahi). 622 words, ~1,797 tokens.

Download SKILL.mdSave it as .claude/skills/extracting-dicom-metadata/SKILL.md (or your agent's skills folder).
name
extracting-dicom-metadata
description
Reads DICOM file headers and DICOM-SR (Structured Report) content to pull study/series metadata and embedded report text, and flags PHI carried in header tags. Use before OpenMed processing when ingesting imaging data (CT/MR/CR/US, radiology SR) and you need the report narrative de-identified and analyzed, plus a list of header tags that must be scrubbed. Hand SR/report text to openmed.deidentify and openmed.analyze_text; use pydicom to read tags. Trigger keywords: DICOM, pydicom, DICOM-SR, structured report, PatientName, study metadata, PACS, radiology report, PS3.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
data-ingestion
metadata.pairs
before
metadata.version
1.0

Extracting DICOM Metadata & Report Text for OpenMed

DICOM (Digital Imaging and Communications in Medicine) files carry far more than pixels: a header of tagged attributes (patient, study, series, equipment) and, for DICOM-SR (Structured Reports), a content tree holding the actual radiology/cardiology report text. Two jobs sit here: pull the report narrative for NLP, and flag the PHI in the header so it gets scrubbed. This skill does both, then hands narrative to OpenMed. Header tags are read with pydicom (external, MIT-licensed); de-identification of the extracted text is OpenMed's.

When to use

  • You ingest DICOM from PACS/VNA or a research archive and want the SR report text mined with clinical NLP.
  • You must enumerate PHI-bearing header tags before sharing/exporting images.
  • You have DICOM-SR objects (e.g. radiology measurements + impression) whose content tree contains the dictated report.

DICOM headers in one minute

Every attribute has a tag (gggg,eeee) (group, element), a VR (value representation, e.g. PN person name, DA date, UI UID), and a value. PHI clusters in well-known tags:

TagNameVRNotes
(0010,0010)PatientNamePNdirect identifier
(0010,0020)PatientIDLOMRN
(0010,0030)PatientBirthDateDADOB
(0010,1040)PatientAddressLOaddress
(0008,0090)ReferringPhysicianNamePNprovider
(0008,0020/0030)StudyDate / StudyTimeDA/TMdates
(0008,0050)AccessionNumberSHorder id
(0008,103E)SeriesDescriptionLOfree text — may leak PHI
(0020,4000)ImageCommentsLTfree text — may leak PHI
(0040,A730)ContentSequenceSQDICOM-SR report tree

Quick start

Read the header, pull SR report text, flag PHI tags, hand off to OpenMed:

python
import pydicom
import openmed

ds = pydicom.dcmread("study.dcm")

# 1) Enumerate PHI-bearing header tags (report, do not log values).
PHI_TAGS = [
    (0x0010, 0x0010), (0x0010, 0x0020), (0x0010, 0x0030), (0x0010, 0x1040),
    (0x0008, 0x0090), (0x0008, 0x0050), (0x0008, 0x0020), (0x0008, 0x0030),
]
present_phi = [hex_pair for hex_pair in PHI_TAGS if hex_pair in ds]

# 2) Extract report text from a DICOM-SR content tree (recursively).
def sr_text(dataset):
    chunks = []
    for item in dataset.get("ContentSequence", []):
        vt = item.get("ValueType")
        if vt == "TEXT" and "TextValue" in item:
            chunks.append(item.TextValue)
        if "ContentSequence" in item:          # nested CONTAINER
            chunks.append(sr_text(item))
    return "\n".join(c for c in chunks if c)

report = sr_text(ds)
# Some modalities stash narrative in free-text header tags too:
for tag in ("ImageComments", "SeriesDescription", "StudyDescription"):
    if tag in ds and isinstance(ds.get(tag), str):
        report += "\n" + ds.get(tag)

# 3) De-identify the narrative, then run NER.
if report.strip():
    deid = openmed.deidentify(report, method="replace", policy="hipaa_safe_harbor")
    result = openmed.analyze_text(deid.text, output_format="dict")

pydicom reads tags by keyword (ds.PatientName) or by (group, element). DICOM-SR text lives in the recursive ContentSequence content tree.

Workflow

  1. Read the dataset with pydicom.dcmread (use stop_before_pixels=True for header-only/metadata work — faster, avoids loading pixels).
  2. Walk the SR content tree. ContentSequence nests CONTAINER, TEXT, CODE, NUM, PNAME nodes; concatenate TEXT.TextValue (and relevant CODE/NUM measurements) in document order to reconstruct the report.
  3. Inventory PHI tags. Flag the standard identifier tags and free-text tags (ImageComments, *Description) that frequently leak PHI. Report tag presence — never echo the values into logs.
  4. De-identify → analyze the report narrative with OpenMed.
  5. Scrub the header before any image export using a DICOM de-identification profile (PS3.15 Annex E / Basic Application Level Confidentiality). OpenMed de-identifies the narrative; header scrubbing is a separate DICOM step.
Show full SKILL.md (254 more words)Show less

Hand-off to / from OpenMed

  • To OpenMed: SR report text (and free-text header tags) → openmed.deidentify → openmed.analyze_text.
  • Header de-id is out of scope for OpenMed — OpenMed handles the text narrative; use a DICOM-native de-identifier (pydicom + PS3.15 profile, or a PACS de-id node) to scrub (0010,xxxx) and burned-in-pixel PHI. This skill's job is to flag those tags so they aren't missed.
  • Re-link by UID, not PHI. Carry StudyInstanceUID/SeriesInstanceUID as rejoin keys; these are not identifiers but should be re-mapped consistently if the profile requires UID remapping.

Edge cases & gotchas

  • Pixel-burned PHI. Ultrasound and secondary-capture images often burn name/ MRN/date into the pixels — header scrubbing alone is insufficient; flag modalities (US, SC, XC) for pixel review/OCR. OpenMed's multimodal/OCR intake can read burned-in text for redaction screening.
  • Private tags. Vendor (gggg,eeee) odd-group private tags can hide PHI; PS3.15 requires removing or whitelisting them — don't trust unknown tags.
  • Date shifting must be consistent. If you date-shift StudyDate, shift all related dates by the same offset to preserve temporal relationships.
  • SR value types. Not all SR content is narrative — NUM (measurements), CODE (coded findings), PNAME (person names, PHI!) need different handling; don't dump PNAME into NLP text.
  • Character sets. Honor SpecificCharacterSet (0008,0005); non-Latin patient names need correct decoding before de-id.
  • Read-only intake. Treat source DICOM as immutable; write de-identified copies, never overwrite originals.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/extracting-dicom-metadata of maziyarpanahi/openmed.

Open the folder on GitHubat commit 34d7b8c

Compare with similar skills

Extracting Dicom Metadata next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Extracting Dicom Metadata compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Extracting Dicom Metadata this skillmaziyarpanahi/openmed5.5k—~1.8kAutomated safety check: PassApache-2.0
Clinical Trials Databasegoogle-deepmind/science-skills3.2k2 repos~3.2kAutomated safety check: PassApache-2.0
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Biomedical Analysis Dispatchxjtulyc/MedgeClaw6171 repos~2kAutomated safety check: PassNone
Research Paperluwill/research-skills862—~1.9kAutomated safety check: PassNone
Research Proposalluwill/research-skills862—~4.5kAutomated safety check: NotesNone

Similar skills

  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 2 repos~3.2k tokens
    Research & ScienceAuto-check passed
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.

    617 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Research Paper

    luwill/research-skills

    A skill your agent uses when the user asks to write or draft an ORIGINAL RESEARCH ARTICLE — IMRaD paper, conference paper, short/workshop paper, 研究论文/期刊论文/会议论文 — reporting their own completed…

    862 GitHub stars~1.9k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Research Proposal

    luwill/research-skills

    A skill your agent uses when the user asks to write or draft a PhD / doctoral research proposal, research plan, 研究计划书, or 开题报告 — a forward-looking plan of background, gap, research questions…

    862 GitHub stars~4.5k tokensUpdated yesterday
    Research & ScienceAuto-check: notes
  • Medical Imaging Review

    LeonChaoX/qinyan-academic-skills

    Write comprehensive literature reviews for medical imaging AI research.

    944 GitHub starsUsed in 3 repos~1.1k tokens
    Research & ScienceAuto-check: notes

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Extracting Dicom Metadata

What does Extracting Dicom Metadata do?

Reads DICOM file headers and DICOM-SR (Structured Report) content to pull study/series metadata and embedded report text, and flags PHI carried in header tags. Extracting Dicom Metadata is an agent skill from maziyarpanahi/openmed. Reads DICOM file headers and DICOM-SR (Structured Report) content to pull study/series metadata and embedded report text, and flags PHI carried in header tags.

When should I use Extracting Dicom Metadata?

Extracting Dicom Metadata fits situations like: keywords: DICOM; structured report; radiology report.

How do I install Extracting Dicom Metadata in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill extracting-dicom-metadata -a claude-code`. Or copy the skill folder (skills/extracting-dicom-metadata in maziyarpanahi/openmed) into .claude/skills/extracting-dicom-metadata in your project. Claude Code loads it when a task matches its description.

How do I install Extracting Dicom Metadata in Codex?

Run `npx skills add maziyarpanahi/openmed --skill extracting-dicom-metadata -a codex`. Or copy the skill folder (skills/extracting-dicom-metadata in maziyarpanahi/openmed) into .agents/skills/extracting-dicom-metadata in your project. Codex loads it when a task matches its description.

Can I use Extracting Dicom Metadata in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill extracting-dicom-metadata -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extracting-dicom-metadata, .gemini/skills/extracting-dicom-metadata, .github/skills/extracting-dicom-metadata and .opencode/skills/extracting-dicom-metadata in your project.

What does Extracting Dicom Metadata need to run?

SKILL.md names no scripts, command-line tools or credentials: Extracting Dicom Metadata is instructions for the agent only. Our summary lists: Python 3.

Does Extracting Dicom Metadata access the network?

SKILL.md names 3 domains. As links in the text: dicom.nema.org, dicomstandard.org and pydicom.github.io. This is read from the text; nothing was executed.

Is Extracting Dicom Metadata safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Extracting Dicom Metadata use?

Extracting Dicom Metadata is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Extracting Dicom Metadata use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Extracting Dicom Metadata?

Skills that share tags, products or a category with Extracting Dicom Metadata: Clinical Trials Database (google-deepmind/science-skills, 3.2k stars), CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars), Biomedical Analysis Dispatch (xjtulyc/MedgeClaw, 617 stars) and Research Paper (luwill/research-skills, 862 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Extracting Dicom Metadata?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,506 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 11, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.