Semantic Analyst
sidequery/sidemantic
Answer analytical, KPI, metric, trend, cohort, and business-performance questions through a Sidemantic semantic layer.
Queries and downloads public cancer imaging data from NCI Imaging Data Commons.
$ npx skills add K-Dense-AI/scientific-agent-skills --skill imaging-data-commons -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install K-Dense-AI/scientific-agent-skills imaging-data-commons --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/imaging-data-commons .claude/skills/imaging-data-commons && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "imaging-data-commons" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/imaging-data-commons into .claude/skills/imaging-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "imaging-data-commons", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/imaging-data-commonsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add K-Dense-AI/scientific-agent-skills --skill imaging-data-commons -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install K-Dense-AI/scientific-agent-skills imaging-data-commons --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/imaging-data-commons .agents/skills/imaging-data-commons && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "imaging-data-commons" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/imaging-data-commons into .agents/skills/imaging-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "imaging-data-commons", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add K-Dense-AI/scientific-agent-skills --skill imaging-data-commons -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install K-Dense-AI/scientific-agent-skills imaging-data-commons --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/imaging-data-commons .cursor/skills/imaging-data-commons && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "imaging-data-commons" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/imaging-data-commons into .cursor/skills/imaging-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "imaging-data-commons", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/K-Dense-AI/scientific-agent-skills.git --path skills/imaging-data-commons--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add K-Dense-AI/scientific-agent-skills --skill imaging-data-commons -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install K-Dense-AI/scientific-agent-skills imaging-data-commons --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/imaging-data-commons .gemini/skills/imaging-data-commons && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "imaging-data-commons" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/imaging-data-commons into .gemini/skills/imaging-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "imaging-data-commons", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install K-Dense-AI/scientific-agent-skills imaging-data-commonsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add K-Dense-AI/scientific-agent-skills --skill imaging-data-commons -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/imaging-data-commons .github/skills/imaging-data-commons && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "imaging-data-commons" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/imaging-data-commons into .github/skills/imaging-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "imaging-data-commons", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add K-Dense-AI/scientific-agent-skills --skill imaging-data-commons -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install K-Dense-AI/scientific-agent-skills imaging-data-commons --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/imaging-data-commons .opencode/skills/imaging-data-commons && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "imaging-data-commons" agent skill from https://github.com/K-Dense-AI/scientific-agent-skills/tree/main/skills/imaging-data-commons into .opencode/skills/imaging-data-commons/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "imaging-data-commons", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
imaging-data-commonsQueries and downloads public cancer imaging data from NCI Imaging Data Commons.
Imaging Data Commons is an agent skill from K-Dense-AI/scientific-agent-skills. Queries and downloads public cancer imaging data from NCI Imaging Data Commons. Supports IDC collection discovery, DICOM access, radiology (CT, MR, PET) and pathology AI datasets, metadata SQL, visualization, licensing, and citations. Uses public metadata and download routes without authentication; optional BigQuery and Google Healthcare routes require Google credentials.
Its SKILL.md is about 7.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including scripts and reference files (for example `references/bigquery_guide.md`, `references/cli_guide.md` and `references/clinical_data_guide.md`). Compatibility notes: Requires network access for hosted APIs, index fetching, citations, and downloads. Local Python workflows target idc-index 0.12.5; BigQuery and Google…
It sits in Databases, covering Clinical and healthcare research, Data warehousing and Citation management. It works with Google BigQuery, SQL, Model Context Protocol and Google Cloud. The repository describes itself as: Turn any AI agent into an AI Scientist. The 1 Agent Skills library for science, used by 250,000+ scientists worldwide. 177 ready-to-use validated skills plus 100+ scientific… The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 92ace75. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
curlpythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.imaging.datacommons.cancer.govAlso links to:
github.comportal.imaging.datacommons.cancer.govpydicom.github.iolearn.canceridc.devdiscourse.canceridc.devidc-index.readthedocs.iodoi.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires network access for hosted APIs, index fetching, citations, and downloads. Local Python workflows target idc-index 0.12.5; BigQuery and Google Healthcare require Google credentials.
From compatibility in the SKILL.md frontmatter.
Imaging Data Commons loads about 7.8k tokens when it runs, and up to ~60k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 3,322 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from K-Dense-AI/scientific-agent-skills at commit 92ace75, republished under its MIT licence (© K-Dense-AI). 3,322 words, ~7,833 tokens.
.claude/skills/imaging-data-commons/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.Query and download public cancer imaging data from the National Cancer Institute Imaging Data Commons (IDC). No authentication required for data access.
Expected network access: IDC metadata is reachable three ways — a bundled local DuckDB index (offline after installation; additional indices are fetched from GitHub and clinical tables from S3), or the hosted IDC service over MCP or REST (api.imaging.datacommons.cancer.gov, no authentication). File downloads use public GCS (storage.googleapis.com) and AWS S3 (s3.amazonaws.com) — no authentication required. DICOMweb access uses either the public IDC proxy (proxy.imaging.datacommons.cancer.gov, no auth) or the Google Cloud Healthcare API (healthcare.googleapis.com, requires GCP authentication). Optional BigQuery queries (bigquery.googleapis.com) also require GCP authentication. Citation resolution contacts DOI services. Public IDC routes require no credentials; optional Google clients use Application Default Credentials.
Reviewed 2026-09-30: idc-index 0.12.5, idc-index-data 24.2.2, IDC v24; hosted API 3.0.0b3. Recheck at use time.
Choose the access path first. There is no single default: the cheapest correct path depends on the session and the task.
idc-index installed and current? Run python scripts/check_version.py. If it passes,
use idc-index for everything.curl; do not install anything. Installing costs ~77 MB of packaged index data plus
pandas, pyarrow, and duckdb, which a metadata question does not need. See Data Access
Options.idc-index: check_version.py exits non-zero and prints
the exact install command for the running interpreter. Prefer a virtual environment, then
restart Python.idc-index (GitHub) is still the most
capable Python path, with query and download helpers in one client. check_version.py never installs anything itself — it also flags a newer
idc-index or skill release when one exists.
Setup for the idc-index path: use the intended interpreter for scripts/check_version.py
and confirm its meets pinned minimum message before continuing. A launch failure is not a pass.
from idc_index import IDCClient
client = IDCClient()
# Verify IDC data version (should be "v24")
print(f"IDC data version: {client.get_idc_version()}")Download and image-processing examples are illustrative unless noted; see the reference review notes for verification scope.
Core workflow: query metadata with client.sql_query() → download with
client.download_from_selection() → visualize with client.get_viewer_URL(). Python examples
below assume this client; Data Access Options has the REST equivalents. For current data
scale, run the summary query in references/sql_patterns.md or GET /v3/stats.
IDC operates a hosted MCP server at https://api.imaging.datacommons.cancer.gov/mcp
(streamable HTTP, no authentication). Where it is available it complements — it does not
replace — the idc-index workflow below.
Identify it by the MCP resource idc://guide, or by three or more of the tool names
build_cohort, get_cohort_urls, list_analysis_results, and get_idc_version. Generic
names such as run_sql are not evidence on their own. If identification is ambiguous, use
idc-index.
If this session has the server, treat it as authoritative for discovery and metadata —
IDC version, counts, attribute values, cohort building, metadata SQL — and follow the
server's own instructions rather than re-deriving them from this file. Its data version is
whatever the server reports: call get_idc_version instead of relying on the version pinned
in this file.
Return here for what the server does not do: downloading files, local pandas/notebook
analysis, DICOMweb, BigQuery, digital pathology tiling, and reproducible scripts. Hand off by
passing SeriesInstanceUIDs from the server to client.download_from_selection(...), and run
scripts/check_version.py at that point.
If it is not available, the identical service is reachable with no configuration as a REST
API at https://api.imaging.datacommons.cancer.gov/v3 — use it for read-only metadata rather
than installing idc-index, per the routing gate in Overview. Suggest connecting the MCP
server at most once, only for repeated interactive discovery, and never change the user's
configuration yourself.
See references/mcp_guide.md for the tool inventory, handoff patterns, and per-host notes.
Inline below: the MCP/REST routing rules, the IDC data model, the index tables and how they join, the core API patterns (query, download, visualize, license, cite), best practices, and troubleshooting.
Reference Guides (load on demand):
| Guide | When to Load |
|---|---|
index_tables_guide.md | Complex JOINs, schema discovery, DataFrame access |
use_cases.md | End-to-end workflows: training datasets, batch downloads, DICOM reading with pydicom/SimpleITK, pipeline integration |
sql_patterns.md | Quick SQL patterns for filter discovery, annotations, size estimation |
clinical_data_guide.md | Clinical/tabular data, imaging+clinical joins, value mapping |
licensing_and_citation.md | Commercial-use questions, mixed-license cohorts, citation formats |
cloud_storage_guide.md | Direct S3/GCS access, versioning, UUID mapping |
dicomweb_guide.md | DICOMweb endpoints, PACS integration |
digital_pathology_guide.md | Slide microscopy (SM), annotations (ANN), pathology workflows |
bigquery_guide.md | Full DICOM metadata, private elements (requires GCP) |
cli_guide.md | Command-line tools (idc download, manifest files) |
parquet_access_guide.md | Direct Parquet queries via GCS (no idc-index install needed) |
mcp_guide.md | Hosted IDC MCP server: tool inventory, identification, handoff to idc-index |
rest_api_guide.md | Hosted IDC REST API: endpoints, filter syntax, SQL over HTTP, manifests |
IDC adds two grouping levels above the standard DICOM hierarchy (Patient → Study → Series → Instance):
tcga_luad, nlst). Treat (collection_id, PatientID) as the patient key; do not assume PatientID is globally unique.collection_id finds original imaging data (which may itself include deposited annotations).Key identifiers for queries:
| Identifier | Scope | Use for |
|---|---|---|
collection_id | Dataset grouping | Filtering by project/study |
PatientID | Patient | Grouping images by patient |
StudyInstanceUID | DICOM study | Grouping of related series, visualization |
SeriesInstanceUID | DICOM series | Grouping of related series, visualization |
The idc-index package provides multiple metadata index tables, accessible via SQL or as pandas DataFrames. The REST API exposes the same tables through GET /tables and POST /sql.
Important: client.indices_overview is the authoritative source for current table descriptions, available columns, and their types — query it when writing SQL or exploring data structure. It also answers "which table contains column X"; see references/index_tables_guide.md for that search pattern and full schema discovery.
Always call client.fetch_index("table_name") before querying any index table — it is safe and idempotent for all tables, including those loaded automatically at startup.
| Family | Tables | Granularity |
|---|---|---|
| Core | index (primary metadata for all current data), collections_index, analysis_results_index | series / collection / analysis result |
| Modality acquisition parameters | ct_index, mr_index, pt_index, contrast_index | 1 row = 1 series of that modality |
| Derived objects | seg_index, rtstruct_index, ann_index, ann_group_index | 1 row = 1 series (or annotation group) |
| Microscopy | sm_index, sm_instance_index | 1 row = 1 SM series / instance |
| Geometry, clinical, history | volume_geometry_index, clinical_index, version_metadata_index, prior_versions_index | see guide |
references/index_tables_guide.md has the full inventory with each table's columns and
contents — load it when you need to know what a specialized table actually holds.
prior_versions_index contains historical series versions, including revised versions
whose DICOM SeriesInstanceUID still occurs in index. Pin crdc_series_uuid for historical
content. For "what's new" in the current release use
series_init_idc_version / series_revised_idc_version in the main index table, which are
not equivalent to this table's min_idc_version / max_idc_version.
SeriesInstanceUID is the universal join key for all series-level specialized tables: sm_index, sm_instance_index, seg_index, ann_index, ann_group_index, contrast_index, volume_geometry_index, rtstruct_index, ct_index, mr_index, pt_index. Always join these to index on SeriesInstanceUID. The exceptions below use different column names.
| Join Column | Tables | Use Case |
|---|---|---|
collection_id | index, prior_versions_index, collections_index, clinical_index | Link series to collection metadata or clinical data |
analysis_result_id | index, analysis_results_index | Link series to analysis result metadata (annotations, segmentations) |
source_DOI | index, analysis_results_index | Link by publication DOI |
segmented_SeriesInstanceUID | seg_index → index | Link segmentation to its source image series (seg_index.segmented_SeriesInstanceUID = index.SeriesInstanceUID) |
referenced_SeriesInstanceUID | ann_index → index, rtstruct_index → index | Link annotation or RTSTRUCT to its source image series |
Note: subjects, updated, and description appear in multiple tables but have different meanings (counts vs identifiers, different update contexts). A UID-only join to prior_versions_index can match several historical revisions; use CRDC UUIDs to distinguish them.
For detailed join examples, schema discovery patterns, key columns reference, and DataFrame access, see references/index_tables_guide.md.
Clinical (non-imaging) attributes — staging, demographics, therapy — live in per-collection
tables. client.fetch_index("clinical_index") loads the dictionary mapping columns to
collections; client.get_clinical_table(name) returns one table as a DataFrame.
See references/clinical_data_guide.md for the discovery workflow, coded-value mapping, and
joining clinical data with imaging.
| Method | Auth | Best For | Reference |
|---|---|---|---|
idc-index | No | Downloads, pandas analysis, unbounded queries — the most capable path | This document |
| IDC MCP server | No | Discovery, cohort building, metadata when the session already has it | mcp_guide.md |
| IDC REST API | No | Metadata with no install, from any language or shell — the default when idc-index is absent | rest_api_guide.md |
| Direct Parquet (GCS) | No | Version-pinned queries, or results past the REST row cap | parquet_access_guide.md |
| Cloud storage (S3/GCS) | No | Direct file access, bulk transfer, custom pipelines | cloud_storage_guide.md |
| DICOMweb via IDC proxy | No | Tool and PACS integration; daily quota, so testing and moderate use | dicomweb_guide.md |
| DICOMweb via Google Healthcare | Yes (GCP) | The same DICOMweb API at production volume, without the proxy quota | dicomweb_guide.md |
| SlicerIDCBrowser | No | 3D visualization and analysis in 3D Slicer | https://github.com/ImagingDataCommons/SlicerIDCBrowser |
| BigQuery | Yes (GCP) | Full DICOM metadata, private elements, SR measurements — last resort | bigquery_guide.md |
The IDC Portal (https://portal.imaging.datacommons.cancer.gov/) is interactive only — browser-based exploration, manual cohort selection, and download. Unlike every option above it has no programmatic interface, so point a user there to browse or click through data themselves; never use it as a step in a script or workflow.
REST API — the no-install metadata path
https://api.imaging.datacommons.cancer.gov/v3, no authentication: discovery, cohort counts and
manifests, read-only SQL, clinical tables, viewer URLs, licenses, citations. It is the same
service as the MCP server over plain HTTP, so it needs no configuration. It never moves image
bytes — switch to idc-index to download, to get a DataFrame, or for results past 10 000 rows.
B=https://api.imaging.datacommons.cancer.gov/v3
curl -s $B/version # idc_version, idc_index_data_version, api_version
curl -s $B/stats # collections, patients, studies, series, instances, size_TB
curl -s "$B/attributes/Modality/values?limit=5" # real filter values, with counts
curl -s $B/sql -H 'content-type: application/json' \
-d '{"sql":"SELECT collection_id, COUNT(*) n FROM index GROUP BY 1 ORDER BY n DESC LIMIT 3"}'
curl -s $B/cohort/counts -H 'content-type: application/json' \
-d '{"filters":{"terms":{"collection_id":["rider_pilot"]}}}'The filter object always goes under filters — on cohort/counts, cohort/manifest,
cohort/manifest.txt, licenses, and citations alike. A bare filter or an unrecognized key is
a 422 naming the fix; an unfiltered series-enumerating request is a 400, not the whole archive.
Filtered JSON responses echo filters_applied and warnings (inside counts for manifests) — read them, because they name
any predicate the server dropped. A zero count with empty warnings therefore means the filter
matched nothing, not that a value was miscased; miscasing produces a warning that says so.
POST /sql takes one read-only SELECT/WITH over the tables idc-index exposes plus
clinical.<table>; max_rows defaults to 5 000, caps at 10 000, and truncated flags clipping.
GET /attributes lists the 19 filterable attributes — clinical values, segmented anatomy, and
acquisition parameters are not among them and need SQL. Request limits still apply. Use
v3 only: V1 and V2 are superseded and scheduled for shutdown, so port any /v1/- or
Modality_btw-style example a user brings rather than extending it.
Both sides build on idc-index-data, so compare the API's idc_index_data_version against local
idc_index_data.__version__ before mixing them: the major is the IDC data release (24.x.y
serves v24); minor/patch index builds may correct metadata. If the API is a whole
release ahead, idc-index cannot download the extra series — mixed selections can omit
unrecognized UIDs, while wholly unmatched selections raise — so either upgrade it (run scripts/check_version.py for the right command)
or transfer directly from the bucket with s5cmd --no-sign-request.
See references/rest_api_guide.md for the endpoint reference, filter grounding, limits, and the
manifest-based download flow.
Cloud storage organization
All DICOM files live in public buckets mirrored between AWS S3 and GCS, organized by CRDC UUIDs
(not DICOM UIDs) to support versioning, as <crdc_series_uuid>/<crdc_instance_uuid>.dcm. Access
is free (no egress fees) via AWS CLI, gsutil, or s5cmd with anonymous access; use the
series_aws_url column for S3 URLs. Bucket names do not establish a license; query license_short_name for each selected series. See references/cloud_storage_guide.md for the full
bucket list and UUID mapping.
DICOMweb access
IDC data is available via DICOMweb (Google Cloud Healthcare API) for PACS integration and
DICOMweb-compatible tools: a public proxy (no auth, daily quota) for testing and moderate
queries, or Google Healthcare (GCP auth) for production volumes. See
references/dicomweb_guide.md.
Direct Parquet access
The idc-index metadata tables are also published as Parquet on a public GCS bucket
(idc-index-data-artifacts), queryable with DuckDB or pandas. This needs DuckDB installed
For ad-hoc metadata prefer REST /sql; choose Parquet for pinned versions or large results.
A separate public S3 export also includes clinical tables and full BigQuery metadata. See
references/parquet_access_guide.md.
The patterns below are the ones that go wrong when recalled from memory rather than checked. Worked examples for each area live in the reference guides named inline.
Filtering on a guessed Modality or BodyPartExamined string is the most common cause of an
empty result set. Enumerate first:
modalities = client.sql_query("""
SELECT DISTINCT Modality, COUNT(*) as series_count
FROM index
GROUP BY Modality
ORDER BY series_count DESC
""")
print(modalities)The same pattern works for any filter column, optionally narrowed by another —
BodyPartExamined within a Modality, Manufacturer, collection_id. On the REST path this
grounding is a single call — GET /attributes/{attr}/values returns values with counts — and the
cohort endpoints report a miscased value in warnings rather than as an empty result.
Two indices carry curated collection-level metadata the primary index does not, both
requiring client.fetch_index(...) first: collections_index (cancer types, tumor locations,
species, subject counts) and analysis_results_index (derived datasets — AI segmentations,
expert annotations, radiomics — with their source collections and modalities).
Cancer type lives in collections_index.cancer_types, not in index — filtering by
cancer type requires a join:
client.fetch_index("collections_index")
results = client.sql_query("""
SELECT i.collection_id, i.PatientID, i.SeriesInstanceUID, i.Modality
FROM index i
JOIN collections_index c ON i.collection_id = c.collection_id
WHERE c.cancer_types LIKE '%Breast%'
AND i.Modality = 'MR'
LIMIT 20
""")client.sql_query() returns a pandas DataFrame. Confirm column names with
client.get_index_schema('index') or client.indices_overview before writing a query rather
than assuming them.
See references/sql_patterns.md for filter-value discovery, annotation and segmentation
queries, size estimation, clinical linking, and version tracking ("what's new in vX" — use
series_init_idc_version / series_revised_idc_version in index; historical objects use
prior_versions_index).
The two download methods take their first two arguments in opposite order. This is the most common source of broken IDC code — check it rather than recalling it:
| Method | First arg | Second arg | Use when |
|---|---|---|---|
download_from_selection | downloadDir (required) | filter kwargs (optional) | Filtering by collection, patient, study, or series |
download_dicom_series | seriesInstanceUID (required) | downloadDir (required) | Downloading specific series by UID only |
download_from_selection takes filter keyword arguments, NOT a DataFrame. The name
"from_selection" refers to filtering the IDC index by criteria — not to accepting a pandas
DataFrame. To download query results, extract the UIDs into a list first:
# Step 1: Query for series UIDs
series_df = client.sql_query("""
SELECT SeriesInstanceUID
FROM index
WHERE Modality = 'CT'
AND BodyPartExamined = 'CHEST'
AND collection_id = 'nlst'
LIMIT 5
""")
# Step 2: Extract UIDs as a list from the DataFrame
uids = list(series_df['SeriesInstanceUID'].values)
# Step 3: Pass the list to download_from_selection (NOT the DataFrame itself)
client.download_from_selection(
downloadDir="./data/lung_ct",
seriesInstanceUID=uids # list of strings, not a DataFrame
)
# Alternative: download_dicom_series has seriesInstanceUID as FIRST arg (different order!)
client.download_dicom_series(
seriesInstanceUID=uids, # FIRST arg here
downloadDir="./data/lung_ct"
)
# Whole collection: downloadDir is still the FIRST positional argument
client.download_from_selection(downloadDir="./data/rider", collection_id="rider_pilot")Both default to AWS; use source_bucket_location="gcs" for Google. In 0.12.5 the most-specific
selector wins, so use SQL first for intersecting criteria and verify every requested UID exists.
Downloaded files are named <crdc_instance_uuid>.dcm, not by SOPInstanceUID. The DICOM
UIDs are preserved inside the file metadata, not in the filename. Read DICOM headers for the
series UID; crdc_instance_uuid is not a column of the series-level index.
idc download <collection|series-uid|manifest> --download-dir ./data does the same from a
shell. See references/cli_guide.md for the dirTemplate hierarchy options (Python default:
%collection_id/%PatientID/%StudyInstanceUID/%Modality_%SeriesInstanceUID; dirTemplate=""
flattens), manifest downloads with resume, and dry-run size estimation.
viewer_url = client.get_viewer_URL(seriesInstanceUID=uid) # one series
viewer_url = client.get_viewer_URL(studyInstanceUID=study_uid) # all series in a studyReturns a browser URL — nothing is downloaded. The method selects OHIF v3 for radiology or SLIM for slide microscopy automatically. Viewing by study is useful when a single DICOM Study holds several Series (T1, T2, and DWI from one MRI session).
IDC data carries license terms and attribution requirements that follow it into any downstream publication or product, and neither is inferable from the pixel data. Check the license before use, and generate citations for whatever you download.
# License breakdown for a selection
licenses = client.sql_query("""
SELECT DISTINCT collection_id, license_short_name,
COUNT(DISTINCT SeriesInstanceUID) as series_count
FROM index GROUP BY collection_id, license_short_name
""")
# Citations for the same selection you downloaded (APA by default)
for citation in client.citations_from_selection(collection_id="rider_pilot"):
print(citation)In the reviewed v24 snapshot, about 97% of IDC data by size is CC BY (commercial use allowed with attribution) and about 3% is CC BY-NC (non-commercial only). Licenses attach to series, not collections — 39 of 176 collections carry more than one — so check the selection you actually intend to use, and note that each component retains its license obligations.
Both tasks are available from all three access paths, so stay on whichever one the session is
already using: idc-index as above, POST /v3/licenses and POST /v3/citations over REST,
or the get_licenses and get_citations MCP tools. See
references/licensing_and_citation.md for the full license inventory, all three routes, the
citation formats (APA, BibTeX, CSL JSON, RDF Turtle), and what to include when publishing.
Pick the access path with the routing gate in Overview; Data Access Options above is the full routing table.
Before reaching for BigQuery (which needs a Google Cloud project and access), check whether a
specialized index table already has the column you want: search client.indices_overview,
then client.fetch_index(...) and query locally for free. Full instance metadata, per-segment
rows, and SR measurement tables require BigQuery or its public Parquet exports; these are
outside the compact idc-index tables.
client.get_index_schema('index') (reads cached metadata, no SQL executed) or client.indices_overview to see all available columns and their descriptions. The version-tracking columns series_init_idc_version and series_revised_idc_version in the main index table directly answer "what's new / when was this added" questions without touching prior_versions_index.client.sql_query() locally or POST /v3/sql over HTTP. Web sources (release notes, blog posts, documentation pages) are frequently out of date and will produce incorrect answers. The index is the authoritative source; use it even when web search is available.client.get_idc_version(), GET /v3/version, or the MCP get_idc_version tool, depending on the path in use (currently v24). For a stale local index, run scripts/check_version.py and use the upgrade command it printslicense_short_name and respect CC BY vs CC BY-NC terms; use citations_from_selection() to produce citations from source_DOI for publicationsLIMIT (or a low max_rows) while exploring, and check collection size before downloading — some collections are terabytes. See references/cli_guide.mddirTemplate (e.g. %collection_id/%PatientID/%Modality) and save the Series UIDs or manifest behind any dataset you buildIssue: ModuleNotFoundError: No module named 'idc_index'
scripts/check_version.py and use the install command it prints, which targets the running interpreter and pins the vetted version. For data analysis also add pandas, numpy, and pydicom (tested with pandas>=1.5, numpy>=1.23, pydicom>=2.3)Issue: Download fails with connection timeout
references/cli_guide.md for
--use-s5cmd-sync resume and retry guidanceIssue: BigQuery quota exceeded or billing errors
references/bigquery_guide.md for cost optimization tipsIssue: Series UID not found or no data returned
LIMIT 5 first, check field names against client.indices_overview,
and confirm the series is in the current version (some old data is deprecated)Issue: Column not found in index table (e.g., SliceThickness, PixelSpacing, KVP, EchoTime, InjectedDose)
index table contains series-level metadata only; modality-specific acquisition and reconstruction parameters live in dedicated tables (ct_index, mr_index, pt_index)client.indices_overview for the column to find its table — the loop is under Finding which table contains a column in references/index_tables_guide.md — then fetch and join on SeriesInstanceUID:client.fetch_index("ct_index")
result = client.sql_query("""
SELECT i.SeriesInstanceUID, i.Modality, c.SliceThickness, c.KVP, c.PixelSpacing_row_mm
FROM index i
JOIN ct_index c USING (SeriesInstanceUID)
WHERE i.collection_id = 'your_collection'
""")Issue: Downloaded DICOM files won't open
Modality, SOPClassUID, and transfer syntax with normal pydicom.dcmread;
forced parsing is not validation.
Check download integrity and decoder/viewer support before re-downloading; reserve force=True for diagnosed non-Part-10 inputs.Reference guides and their decision triggers are listed in Quick Navigation above.
© K-Dense-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 14 other files (scripts, references) in skills/imaging-data-commons of K-Dense-AI/scientific-agent-skills.
Open the folder on GitHubat commit 92ace75
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in K-Dense-AI/scientific-agent-skills, which our catalogue first saw on October 7, 2026.
Imaging Data Commons next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Imaging Data Commons this skillK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~7.8k | Automated safety check: Pass | MIT | |
| Semantic Analystsidequery/sidemantic | 129 | — | ~982 | Automated safety check: Pass | AGPL-3.0 | |
| Bigquery Observabilitygoogle/skills | 21k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Bigquery Optimizationgoogle/skills | 21k | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Cloud Monitoring Metric Selectiongoogle/skills | 21k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Ga4 Bigquery Exportjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~2.8k | Automated safety check: Pass | MIT |
sidequery/sidemantic
Answer analytical, KPI, metric, trend, cohort, and business-performance questions through a Sidemantic semantic layer.
google/skills
Provides data-retrieval best practices, tool selection guidance, and performant SQL query syntax for BigQuery telemetry across INFORMATIONSCHEMA, Cloud Monitoring, and the REST API.
google/skills
Provides workflows to optimize BigQuery environments (capacity planning, editions), storage assets (partitioning, clustering, storage lifecycles, billing models), and SQL queries.
google/skills
Retrieve, query, and identify relevant Cloud Monitoring metric descriptors on Google Cloud for a service or resource (such as Compute Engine, Spanner, BigQuery, Cloud Run, Cloud SQL, Pub/Sub, Cloud…
jeremylongshore/tons-of-skills-marketplace
Wire GA4 → BigQuery for unsampled, queryable event-level data.
google/skills
Translates Snowflake dbt SQL models into standardized BigQuery SQL, keeping Jinja constructs and tracking progress in a migration tasks file.
K-Dense-AI/scientific-agent-skills
Runs Cantera constant-volume or constant-pressure ignition simulations and reports temperature-based ignition delay with mechanism provenance and checks.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
K-Dense-AI/scientific-agent-skills
Plans, runs, and documents analytical method validation, verification, or transfer studies under ICH Q2(R2)/Q14, USP, ICH M10, CLSI EP, or ISO/IEC 17025.
K-Dense-AI/scientific-agent-skills
Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.
K-Dense-AI/scientific-agent-skills
Plans and audits runs of the HypoGeniC and HypoRefine packages, which propose hypotheses from labeled text datasets, with local checks before any model call.
K-Dense-AI/scientific-agent-skills
Organizes scope, controlled documents, risk files and traceability into draft evidence for human review against ISO 13485, 14971, 17025 and 15189.
Categories
Queries and downloads public cancer imaging data from NCI Imaging Data Commons. Imaging Data Commons is an agent skill from K-Dense-AI/scientific-agent-skills. Queries and downloads public cancer imaging data from NCI Imaging Data Commons.
Imaging Data Commons fits situations like: tasks that involve Clinical and healthcare research; tasks that involve Data warehousing; tasks that involve Citation management.
Run `npx skills add K-Dense-AI/scientific-agent-skills --skill imaging-data-commons -a claude-code`. Or copy the skill folder (skills/imaging-data-commons in K-Dense-AI/scientific-agent-skills) into .claude/skills/imaging-data-commons in your project. Claude Code loads it when a task matches its description.
Run `npx skills add K-Dense-AI/scientific-agent-skills --skill imaging-data-commons -a codex`. Or copy the skill folder (skills/imaging-data-commons in K-Dense-AI/scientific-agent-skills) into .agents/skills/imaging-data-commons in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add K-Dense-AI/scientific-agent-skills --skill imaging-data-commons -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/imaging-data-commons, .gemini/skills/imaging-data-commons, .github/skills/imaging-data-commons and .opencode/skills/imaging-data-commons in your project.
Going by SKILL.md and its folder, Imaging Data Commons needs Python for the scripts in its folder and the command-line tools its instructions call (curl and python). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires network access for hosted APIs, index fetching, citations, and downloads. Local Python workflows target idc-index 0.12.5; BigQuery and Google Healthcare require Google credentials..
SKILL.md names 8 domains. In commands or code: api.imaging.datacommons.cancer.gov; the agent is likely to contact it when it follows the instructions. As links in the text: github.com, portal.imaging.datacommons.cancer.gov, pydicom.github.io, learn.canceridc.dev, discourse.canceridc.dev, idc-index.readthedocs.io and doi.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Imaging Data Commons is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.8k tokens (SKILL.md is roughly 31k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 52k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Imaging Data Commons: Semantic Analyst (sidequery/sidemantic, 129 stars), Bigquery Observability (google/skills, 21k stars), Bigquery Optimization (google/skills, 21k stars) and Cloud Monitoring Metric Selection (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
K-Dense-AI (a GitHub organization) maintains it in K-Dense-AI/scientific-agent-skills, which has 47,942 GitHub stars. The repository holds 152 skills in this directory. The repository was last updated on October 5, 2026.
Source: K-Dense-AI/scientific-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.