Agent skill

Public Data Access

by xuzhougeng in xuzhougeng/wisp-science

Plan, validate, and document public-bioinformatics data acquisition for GEO/GSE/GSM/GPL/GDS, SRA/ENA, TCGA/GDC, GTEx, and DepMap.

AGPL-3.0Auto-check passedResearch & Science

Install Public Data Access

skills CLI
$ npx skills add xuzhougeng/wisp-science --skill public-data-access -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install xuzhougeng/wisp-science public-data-access --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/xuzhougeng/wisp-science.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/public-data-access .claude/skills/public-data-access && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
public-data-access
GitHub stars
1k
Token cost
~1.7k tokens
SKILL.md length
754 words
Files
5 (incl. scripts, references)
Skills in repo
25
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Plan, validate, and document public-bioinformatics data acquisition for GEO/GSE/GSM/GPL/GDS, SRA/ENA, TCGA/GDC, GTEx, and DepMap.

  • Works in 7 steps: Clarify the dataset contract. Identify… → Inspect before transfer. Use available… → Write a provider-neutral plan. Run… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Workflow, Provider routing, Create and validate a plan and Generate a manifest, plus 2 more sections
  • Runs Python scripts from its folder; calls python

What it does

Public Data Access is an agent skill from xuzhougeng/wisp-science. Plan, validate, and document public-bioinformatics data acquisition for GEO/GSE/GSM/GPL/GDS, SRA/ENA, TCGA/GDC, GTEx, and DepMap. Covers expression matrices, raw reads, download manifests, caches, and optional geokit SOFT/Series Matrix acquisition for R workflows.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/download-plan-schema.md`, `references/geokit.md` and `references/provider-routing.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: Open-source, local-first desktop AI research workbench for scientific computing with Python/R, MCP bioinformatics tools, SSH/WSL/GPU runtimes, and OpenAI/Anthropic models. The licence is AGPL-3.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/public-data-access”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Clarify the dataset contract. Identify the provider, accession/project,
  2. Inspect before transfer. Use available MCP/connectors or official
  3. Write a provider-neutral plan. Run scripts/public_data_plan.py init.
  4. Validate and review. Run validate, show the user the resolved provider,
  5. Select the adapter at runtime. Prefer an already available Wisp MCP tool
  6. Acquire safely. Reuse existing valid files, resume partial transfers when
  7. Verify and hand off. Check expected files, byte sizes, checksums when

What it can do on your machine

Read from SKILL.md and the folder at commit 2ba143b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Public Data Access loads about 1.7k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 754 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from xuzhougeng/wisp-science at commit 2ba143b, republished under its AGPL-3.0 licence (© xuzhougeng). 754 words, ~1,728 tokens.

Download SKILL.mdSave it as .claude/skills/public-data-access/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
public-data-access
description
Plan, validate, and document public-bioinformatics data acquisition for GEO/GSE/GSM/GPL/GDS, SRA/ENA, TCGA/GDC, GTEx, and DepMap. Covers expression matrices, raw reads, download manifests, caches, and optional geokit SOFT/Series Matrix acquisition for R workflows.

Public Bioinformatics Data Access

Build a reproducible acquisition plan before downloading. Treat GEO, SRA/ENA, GDC, GTEx, and DepMap as independent providers behind one provider-neutral workflow. Do not require a provider-specific toolkit or machine-specific checkout.

Workflow

  1. Clarify the dataset contract. Identify the provider, accession/project, release, modality, smallest useful data product, filters, target directory, expected scale, and downstream analysis.
  2. Inspect before transfer. Use available MCP/connectors or official metadata endpoints to list releases, files, samples, sizes, and checksums. Do not start a bulk transfer during discovery.
  3. Write a provider-neutral plan. Run scripts/public_data_plan.py init. Store the plan next to the future dataset as download-plan.json.
  4. Validate and review. Run validate, show the user the resolved provider, transport, filters, limits, output location, and known size. For a large, paid, authenticated, or overwrite-capable job, confirm that the user's authorization covers this concrete transfer; ask only when it does not.
  5. Select the adapter at runtime. Prefer an already available Wisp MCP tool for metadata and small queries. Prefer official HTTPS/FTP or provider clients for bulk files. Use an external project only when it is installed and record its version in the plan/manifest.
  6. Acquire safely. Reuse existing valid files, resume partial transfers when supported, keep raw files immutable, and never place credentials in the plan.
  7. Verify and hand off. Check expected files, byte sizes, checksums when available, and sample/file counts. Generate manifest.json with the script.

Provider routing

ProviderDiscovery and small queriesBulk acquisitionTypical products
GEOGEO metadata connector, NCBI E-utilitiesNCBI GEO HTTPS/FTP; optional geokit in Rseries matrix, SOFT, supplementary files
SRA/ENARunInfo or ENA Portal APIENA HTTPS/FTP or SRA ToolkitFASTQ, run metadata
GDCGDC files/cases APImanifest + gdc-client, or HTTPS for bounded filesexpression, mutation, CNV, clinical, methylation
GTExGTEx expression connector/APIofficial release files for matricesgene/tissue queries, median or sample expression
DepMapDepMap model/release metadataofficial release file endpointmodel metadata, expression, mutation, dependency
customUser-provided catalog/APIexplicit HTTPS/FTP URLsprovider-specific files

Read references/provider-routing.md before implementing or changing a provider adapter. DepMap-specific flags or release semantics must stay inside the DepMap adapter; they must not shape the common plan schema.

For GEO SOFT/Series Matrix parsing, sample metadata preparation, or ExpressionSet acquisition in an R workflow, read references/geokit.md. geokit is optional; ordinary GEO discovery does not require R or package installation.

Show full SKILL.md (368 more words)Show less

Create and validate a plan

Resolve scripts/public_data_plan.py against this skill's directory (the use_skill result lists its path), and invoke that resolved script with a Python 3.10+ interpreter. Keep the working directory at the project root so relative plan/output paths belong to the project. The examples below abbreviate the script path; quote the resolved path when it contains spaces. In an SSH/WSL context, stage the helper there or use an existing copy in that context; a desktop skill path is not automatically available remotely.

bash
python scripts/public_data_plan.py init \
  --provider geo \
  --identifier GSE12345 \
  --data-type series-matrix \
  --output-dir data/public/geo/GSE12345 \
  --plan data/public/geo/GSE12345/download-plan.json

python scripts/public_data_plan.py validate \
  data/public/geo/GSE12345/download-plan.json

Filters are provider-specific but encoded uniformly as repeated key=value pairs:

bash
python scripts/public_data_plan.py init \
  --provider gdc \
  --identifier TCGA-BRCA \
  --data-type expression \
  --filter workflow_type="STAR - Counts" \
  --filter sample_type="Primary Tumor" \
  --max-files 20 \
  --transport gdc-client \
  --plan data/public/gdc/TCGA-BRCA/download-plan.json

The planner does not download data. It produces a reviewable contract. See references/download-plan-schema.md for the complete schema. Validation checks the plan structure; it does not probe URLs, enforce transfer limits, verify installed packages, or approve a pending transfer. The selected adapter must honor the plan's limits and resume behavior.

Generate a manifest

After acquisition:

bash
python scripts/public_data_plan.py manifest \
  data/public/geo/GSE12345/download-plan.json \
  --scan-dir data/public/geo/GSE12345 \
  --output data/public/geo/GSE12345/manifest.json

Use SHA-256 for modest datasets and provider checksums for large archives. For very large datasets, --checksum none is acceptable only when official checksums or immutable object identifiers are recorded elsewhere.

Safety and reproducibility rules

  • Default to overwrite=false, resume=true, and the minimum useful subset.
  • Never translate an exploratory request into “download everything.”
  • Keep provider metadata, query/filter payloads, release/version, transport, tool version, URLs/object identifiers, and validation results.
  • Separate immutable source files from normalized/derived outputs.
  • Do not treat a successful HTTP response as a valid dataset; verify content.
  • Do not embed API keys, cookies, signed URLs, SSH keys, or bearer tokens.
  • Use structured runs or a remote execution context for long transfers rather than extending an interactive shell timeout.
  • If an adapter or connector cannot perform the requested transfer, stop after producing the validated plan and report the missing capability explicitly.

Wisp Science integration

  • Discover the live connector/tool catalog instead of assuming exact MCP tool names; installations can expose different provider adapters.
  • Use connectors for discovery and bounded queries, then official transfer mechanisms for large files.
  • Keep outputs under the active project, normally data/public/<provider>/....
  • Invoke the planner as a standalone CLI; no Python REPL helper loading is required. R-based acquisition can use geokit independently of the planner.
  • Treat this skill as an acquisition/orchestration layer. Downstream QC, statistics, annotation, and visualization belong to other skills.

© xuzhougeng, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/public-data-access of xuzhougeng/wisp-science.

  • SKILL.md
  • references/download-plan-schema.md
  • references/geokit.md
  • references/provider-routing.md
  • scripts/public_data_plan.py

Open the folder on GitHubat commit 2ba143b

Compare with similar skills

Public Data Access next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Public Data Access compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Public Data Access this skillxuzhougeng/wisp-science1k—~1.7kAutomated safety check: PassAGPL-3.0
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from xuzhougeng/wisp-science

All 25 skills in this repo
  • Singlecell Qc

    xuzhougeng/wisp-science

    A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.

    1k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Research Integrity Audit

    xuzhougeng/wisp-science

    学术审查 / research-integrity screening of a manuscript's figures and reported numbers.

    1k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Distill Concept Books

    xuzhougeng/wisp-science

    将概念、理论或分析方法类图书蒸馏为证据可追溯、经人工门禁审核且不暴露书名、作者、出版社等来源身份的任务型 Skill 候选。用于新建或恢复图书蒸馏、以本地 Tesseract 扫描 DOCX 全部内嵌图像或 Poppler 渲染的扫描 PDF 全页、建立 source map 与 evidence/claim/relation/capability…

    1k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Skill Creator

    xuzhougeng/wisp-science

    Create, update, validate, and evaluate Wisp skills. An agent skill from xuzhougeng/wisp-science.

    1k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Word Zotero Citations

    xuzhougeng/wisp-science

    Build, audit, authorize, recover, or finalize dynamic Zotero citations and bibliographies in Microsoft Word DOCX files with a protected-source, digest-bound workflow.

    1k GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Compute Env Setup

    xuzhougeng/wisp-science

    Set up and validate a reproducible Python or R environment on a Wisp execution context.

    1k GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Public Data Access

What does Public Data Access do?

Plan, validate, and document public-bioinformatics data acquisition for GEO/GSE/GSM/GPL/GDS, SRA/ENA, TCGA/GDC, GTEx, and DepMap. Public Data Access is an agent skill from xuzhougeng/wisp-science. Plan, validate, and document public-bioinformatics data acquisition for GEO/GSE/GSM/GPL/GDS, SRA/ENA, TCGA/GDC, GTEx, and DepMap.

When should I use Public Data Access?

Public Data Access fits situations like: tasks that involve Bioinformatics.

How do I install Public Data Access in Claude Code?

Run `npx skills add xuzhougeng/wisp-science --skill public-data-access -a claude-code`. Or copy the skill folder (skills/public-data-access in xuzhougeng/wisp-science) into .claude/skills/public-data-access in your project. Claude Code loads it when a task matches its description.

How do I install Public Data Access in Codex?

Run `npx skills add xuzhougeng/wisp-science --skill public-data-access -a codex`. Or copy the skill folder (skills/public-data-access in xuzhougeng/wisp-science) into .agents/skills/public-data-access in your project. Codex loads it when a task matches its description.

Can I use Public Data Access in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add xuzhougeng/wisp-science --skill public-data-access -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/public-data-access, .gemini/skills/public-data-access, .github/skills/public-data-access and .opencode/skills/public-data-access in your project.

What does Public Data Access need to run?

Going by SKILL.md and its folder, Public Data Access needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Public Data Access access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Public Data Access safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Public Data Access use?

Public Data Access is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Public Data Access use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.

What are the alternatives to Public Data Access?

Skills that share tags, products or a category with Public Data Access: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Public Data Access?

xuzhougeng (a GitHub user) maintains it in xuzhougeng/wisp-science, which has 1,026 GitHub stars. The repository holds 25 skills in this directory. The repository was last updated on October 10, 2026.

Source: xuzhougeng/wisp-science on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.