ETE Toolkit for Phylogenetic Trees
davila7/claude-code-templates
Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.
Query and download from NCBI Gene Expression Omnibus (GEO) and EMBL-EBI's BioStudies/ArrayExpress mirror.
$ npx skills add GPTomics/bioSkills --skill bio-geo-data -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-geo-data --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/database-access/geo-data .claude/skills/bio-geo-data && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-geo-data" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/geo-data into .claude/skills/bio-geo-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-geo-data", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/database-access/geo-dataType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-geo-data -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-geo-data --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/database-access/geo-data .agents/skills/bio-geo-data && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-geo-data" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/geo-data into .agents/skills/bio-geo-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-geo-data", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-geo-data -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-geo-data --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/database-access/geo-data .cursor/skills/bio-geo-data && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-geo-data" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/geo-data into .cursor/skills/bio-geo-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-geo-data", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path database-access/geo-data--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-geo-data -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-geo-data --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/database-access/geo-data .gemini/skills/bio-geo-data && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-geo-data" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/geo-data into .gemini/skills/bio-geo-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-geo-data", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-geo-dataInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-geo-data -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/database-access/geo-data .github/skills/bio-geo-data && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-geo-data" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/geo-data into .github/skills/bio-geo-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-geo-data", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-geo-data -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-geo-data --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/database-access/geo-data .opencode/skills/bio-geo-data && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-geo-data" agent skill from https://github.com/GPTomics/bioSkills/tree/main/database-access/geo-data into .opencode/skills/bio-geo-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-geo-data", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-geo-dataQuery and download from NCBI Gene Expression Omnibus (GEO) and EMBL-EBI's BioStudies/ArrayExpress mirror.
Bio Geo Data is an agent skill from GPTomics/bioSkills. Query and download from NCBI Gene Expression Omnibus (GEO) and EMBL-EBI's BioStudies/ArrayExpress mirror. Use when finding expression datasets, navigating SuperSeries vs SubSeries, choosing between series-matrix (submitter-normalized) and raw supplementary files, downloading via GEOparse (Python) or GEOquery (R/Bioconductor), linking GEO to SRA for raw reads, or distinguishing GSE/GSM/GPL/GDS record types. Encodes the SuperSeries trap, the series-matrix normalization-trust caveat, GEOmetadb deprecation…
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `examples/geo_from_pubmed.py`, `examples/geo_to_sra.py` and `examples/search_geo.py`).
It sits in Research & Science, covering Bioinformatics and Database schema design. It works with NCBI and Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipwgetFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
ebi.ac.ukftp.ncbi.nlm.nih.govAlso links to:
archs4.orggithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Geo Data loads about 4.4k tokens when it runs. Until then it costs about 149 tokens; SKILL.md has 1,415 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,415 words, ~4,423 tokens.
.claude/skills/bio-geo-data/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Reference examples tested with: BioPython 1.83+, GEOparse 2.0+, R Bioconductor GEOquery 2.70+, pandas 2.2+
Before using code patterns, verify installed versions match. If versions differ:
pip show biopython geoparse then introspect signaturespackageVersion('GEOquery')If the GSE structure doesn't match expectations (missing fields, malformed series matrix), re-fetch from FTP directly and inspect the SOFT or MINiML file as source of truth.
"Pull expression data from GEO accession GSE..." -> GEO stores Series (GSE), Samples (GSM), Platforms (GPL), and curated DataSets (GDS, frozen 2018). The single most consequential decision is processed (series matrix) vs raw (supplementary files / linked SRA) — the answer turns on how much trust the submitter's normalization deserves.
The single most-missed gotcha: SuperSeries. A GSE may be a meta-container (!Series_relation = SuperSeries of: GSExxxxx) holding multiple sub-studies on different platforms. Naively pulling samples from a SuperSeries gives mixed Affymetrix + Illumina + RNA-seq, mis-batched.
Entrez.esearch(db='gds'), GEOparse for full series downloadGEOquery::getGEO() (Bioconductor; more mature than GEOparse)wget from ftp.ncbi.nlm.nih.gov/geo/series/...pip install biopython GEOparse pandas
# OR for R-side:
# R: BiocManager::install('GEOquery')from Bio import Entrez
Entrez.email = 'researcher@institution.edu'
Entrez.api_key = 'optional'| Prefix | Type | Granularity | What's in it |
|---|---|---|---|
| GSE | Series | One study | Title, summary, design, links to GSMs, supplementary files |
| GSM | Sample | One biological/technical sample | Submitter metadata, per-sample processed data, link to raw SRA |
| GPL | Platform | One array / sequencer | Probe annotations or sequencer model |
| GDS | DataSet | Curated, normalized subset of one GSE | Re-normalized expression matrix (frozen 2018; new GDS no longer created) |
| GSEXXX SuperSeries | Series meta-container | Wraps multiple SubSeries | !Series_relation = SuperSeries of: ... |
GDS is dead-as-format: NCBI stopped creating new GDS records in 2018. Existing GDS still queryable but use GSE for anything current.
A SuperSeries (GSE) wraps multiple SubSeries, often with different platforms. Detection:
# Read the !Series_relation field from SOFT format
from Bio import Entrez
h = Entrez.esummary(db='gds', id='200122288') # example
r = Entrez.read(h)[0]; h.close()
print(r.get('summary')) # may or may not flag SuperSeries
# Definitive check: download SOFT and grep:
# curl ftp://ftp.ncbi.nlm.nih.gov/geo/series/GSE122nnn/GSE122288/soft/GSE122288_family.soft.gz | zgrep Series_relationA SuperSeries of: GSE12345 line means the SuperSeries' samples are the union of all SubSeries — almost certainly mixed-platform / mixed-batch. Process each SubSeries independently.
Symmetric trap: a paper may cite a SubSeries (SubSeries of: GSEsuper) where the wider context is essential — check both directions.
| Question | Source | Trust level |
|---|---|---|
| "I want expression values; submitter normalization is fine" | Series matrix (GSE_series_matrix.txt.gz) | Trust submitter's normalization |
| "I want raw Affymetrix CEL files and to do my own RMA" | Supplementary files (suppl/) | Re-normalize locally |
| "I want raw RNA-seq FASTQ" | pysradb gse_to_srp -> srp_to_srr (Entrez gds->sra ELink unreliable) | Always raw; processed at submitter is rarely re-usable |
| "I want submitter-provided counts (RNA-seq)" | Supplementary files (usually a *_counts.txt.gz) | Trust at risk; submitter pipelines vary |
| "I want a curated subset across many studies" | Use ArchS4 (https://archs4.org) or recount3 | Curated re-processing |
Default to raw whenever possible. For Affymetrix: CEL + locally-run RMA is far more reliable than the submitter's "normalized" matrix. For RNA-seq: SRA FASTQ + locally-run alignment/quantification is the only reproducible path; submitter counts often use a private pipeline.
A series matrix (GSE12345_series_matrix.txt.gz) is a header (sample metadata as !Sample_* lines) plus a sample-by-feature expression table. The format is fragile and the values' provenance is whatever the submitter chose. Critical caveats:
!Series_overall_design and !Sample_data_processing to know.!Sample_characteristics_ch1 rows that hold the metadata of interest — these are submitter-formatted strings, often inconsistent within one series.| Format | Content | Parser support |
|---|---|---|
SOFT (*_family.soft.gz) | Plain-text, key=value style | GEOparse (Python), GEOquery (R), Entrez Direct |
MINiML (*_family.xml.tgz) | XML-structured | GEOparse, GEOquery, custom XML |
Both contain the same content. SOFT is the legacy, MINiML the XML successor. GEOparse handles SOFT well; for very large series (1000+ samples) MINiML's XML structure is slower to parse.
| Aspect | GEOparse (Python) | GEOquery (R/Bioconductor) |
|---|---|---|
| Maturity | OK; some known supplementary-file fetch issues since ~2022 | Mature; Bioconductor-supported |
| Output | GEOparse.GSE object with gsms, gpls, metadata dicts | ExpressionSet or list per platform |
| Supplementary files | gse.download_supplementary_files() (sometimes flakey) | getGEOSuppFiles(gse) (more reliable) |
| Integration | Pandas DataFrames | Bioconductor ecosystem |
| When | Python-first pipelines | R-first / use ExpressionSet downstream |
For production GEO workflows in R, GEOquery is the stable choice. For Python, GEOparse is the only option but verify file counts after download.
GEOmetadb (Zhu 2008) was a SQLite mirror of GEO metadata enabling fast SQL queries. Unmaintained since 2020; downloads still work but data is stale. Modern replacement: pysradb (pysradb gse_to_srp, pysradb metadata) covers most of the GEO->SRA mapping; for full GEO queries fall back to Entrez gds.
ArrayExpress (EMBL-EBI's microarray archive, mirroring GEO) was migrated into BioStudies in 2020. Old E-MTAB-#### accessions still resolve but the API moved:
| Old (pre-2020) | New (BioStudies) |
|---|---|
https://www.ebi.ac.uk/arrayexpress/... | https://www.ebi.ac.uk/biostudies/... |
| ArrayExpress REST | BioStudies REST: https://www.ebi.ac.uk/biostudies/api/v1/... |
For new workflows, use BioStudies. For legacy ArrayExpress URLs in old papers, redirect via BioStudies.
Goal: Find GSE accessions matching keywords + organism + study type.
Approach: ESearch on gds db with field-qualified terms; filter to gse[Entry Type]; summarize with ESummary.
Reference (BioPython 1.83+):
from Bio import Entrez
import time
Entrez.email = 'researcher@institution.edu'
def search_geo(term, study_type='gse', organism=None, max_results=50):
full_term = f'{term} AND {study_type}[Entry Type]'
if organism:
full_term += f' AND {organism}[Organism]'
h = Entrez.esearch(db='gds', term=full_term, retmax=max_results)
s = Entrez.read(h); h.close()
if not s['IdList']:
return []
h = Entrez.esummary(db='gds', id=','.join(s['IdList']))
summaries = Entrez.read(h); h.close()
return summaries
for s in search_geo('breast cancer RNA-seq', organism='Homo sapiens', max_results=10):
# Surface SuperSeries
relation = s.get('summary', '')
is_super = 'SuperSeries' in str(relation)
print(f" {s['Accession']:12} {s['n_samples']:>4} samples {'[SuperSeries]' if is_super else '':12} {s['title'][:60]}")Goal: Avoid mixing platforms by detecting SuperSeries structure first.
Approach: Download SOFT family file and read !Series_relation keys.
import gzip
import urllib.request
def check_super_or_sub_series(gse):
prefix = gse[:-3] + 'nnn'
url = f'https://ftp.ncbi.nlm.nih.gov/geo/series/{prefix}/{gse}/soft/{gse}_family.soft.gz'
urllib.request.urlretrieve(url, f'{gse}.soft.gz')
super_of = []
sub_of = None
with gzip.open(f'{gse}.soft.gz', 'rt') as f:
for line in f:
if line.startswith('!Series_relation'):
if 'SuperSeries of' in line:
super_of.append(line.split('SuperSeries of: ')[1].strip())
elif 'SubSeries of' in line:
sub_of = line.split('SubSeries of: ')[1].strip()
if line.startswith('^SAMPLE'):
break # Speed: don't read past header
return {'super_of': super_of, 'sub_of': sub_of}
print(check_super_or_sub_series('GSE122288'))
# {'super_of': ['GSExxxxx', 'GSEyyyyy'], 'sub_of': None} -> SuperSeries; process subseries separatelyimport gzip
import pandas as pd
def download_series_matrix(gse):
prefix = gse[:-3] + 'nnn'
url = f'https://ftp.ncbi.nlm.nih.gov/geo/series/{prefix}/{gse}/matrix/{gse}_series_matrix.txt.gz'
urllib.request.urlretrieve(url, f'{gse}_matrix.txt.gz')
return f'{gse}_matrix.txt.gz'
def parse_series_matrix(path):
metadata = {}
with gzip.open(path, 'rt') as f:
for line in f:
if line.startswith('!series_matrix_table_begin'):
break
if line.startswith('!'):
key, *vals = line.rstrip('\n').split('\t')
metadata[key] = [v.strip('"') for v in vals]
expr = pd.read_csv(f, sep='\t', index_col=0, comment='!')
# Series matrix values are whatever submitter chose -- check metadata['!Sample_data_processing']
return metadata, expr
meta, expr = parse_series_matrix(download_series_matrix('GSE123456'))
print('Sample-level data processing notes:')
for note in set(meta.get('!Sample_data_processing', [])):
print(f' - {note}')from pysradb import SRAweb
def gse_to_srr(gse):
db = SRAweb()
srp_df = db.gse_to_srp(gse)
if srp_df.empty:
return []
srp = srp_df['study_accession'].iloc[0]
srr_df = db.srp_to_srr(srp)
return srr_df['run_accession'].tolist()
srrs = gse_to_srr('GSE123456')
print(f'GSE123456 -> {len(srrs)} SRR runs')import GEOparse
def get_gse(gse_id, dest='./geo_cache'):
gse = GEOparse.get_GEO(geo=gse_id, destdir=dest)
print(f'{gse_id}: {len(gse.gsms)} samples, {len(gse.gpls)} platforms')
for gsm_name, gsm in list(gse.gsms.items())[:3]:
print(f' {gsm_name}: {gsm.metadata.get("title", ["?"])[0]}')
return gse
# Supplementary files (raw data) -- verify file count manually after
gse = get_gse('GSE123456')
gse.download_supplementary_files(directory='./geo_cache')# Reference: Bioconductor GEOquery 2.70+ | Verify API if version differs
library(GEOquery)
gse <- getGEO('GSE123456', GSEMatrix = TRUE)
length(gse) # one ExpressionSet per platform
head(pData(gse[[1]])) # sample metadata
head(exprs(gse[[1]])) # expression matrix (submitter-normalized -- verify processing notes)
# Raw / supplementary files
supp_dir <- getGEOSuppFiles('GSE123456', baseDir = './geo_cache')
list.files(rownames(supp_dir))def geo_from_pubmed(pmid):
h = Entrez.elink(dbfrom='pubmed', db='gds', id=pmid)
r = Entrez.read(h); h.close()
if not r[0]['LinkSetDb']:
return []
gds_ids = [l['Id'] for l in r[0]['LinkSetDb'][0]['Link']]
h = Entrez.esummary(db='gds', id=','.join(gds_ids))
summaries = Entrez.read(h); h.close()
return summaries!Series_relation in SOFT before pulling; process SubSeries independently.!Sample_data_processing to know what's in the matrix; re-normalize from raw if in doubt.*_counts.txt.gz supplementary file as the count matrix.GSE_series_matrix.txt.gz is the merged one; per-platform are GSE-GPLxxx_series_matrix.txt.gz.!Series_platform_id count.gse.download_supplementary_files() silently misses files.wget -r on the suppl/ subdirectory.https://www.ebi.ac.uk/arrayexpress/experiments/E-MTAB-1234/.https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-1234.GEOmetadb.sqlite for fast queries.| Error / symptom | Cause | Solution |
|---|---|---|
Empty IdList for gse[entry_type] | Wrong field name | Use gse[Entry Type] (case-sensitive) |
| Matrix file has no expression data | SuperSeries with no aggregate matrix | Pull per-SubSeries matrices |
| Submitter "normalized" matrix gives different result than paper | Hidden submitter transforms | Re-process from raw |
| 404 on ArrayExpress URL | Migrated to BioStudies | Use new BioStudies URL |
| GEOparse missing CEL files | Known flake | Use R GEOquery or direct FTP |
| GEOmetadb-based pipeline missing recent series | DB unmaintained | Switch to pysradb / Entrez |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files in database-access/geo-data of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 9, 2026.
Bio Geo Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Geo Data this skillGPTomics/bioSkills | 1.2k | 2 repos | ~4.4k | Automated safety check: Pass | MIT | |
| ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates | 32k | 11 repos | ~4.5k | Automated safety check: Notes | MIT | |
| Biopythondavila7/claude-code-templates | 32k | 12 repos | ~3.4k | Automated safety check: Pass | MIT | |
| BiopythonK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~4.3k | Automated safety check: Notes | MIT | |
| Biopythonlamm-mit/scienceclaw | 244 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | |
| Bio Single Cell PreprocessingFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | 1 repos | ~2.4k | Automated safety check: Pass | None |
davila7/claude-code-templates
Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.
davila7/claude-code-templates
Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.
K-Dense-AI/scientific-agent-skills
Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).
lamm-mit/scienceclaw
Computational molecular biology library (sequence I/O, alignment, phylogenetics).
FreedomIntelligence/OpenClaw-Medical-Skills
Quality control, filtering, and normalization for single-cell RNA-seq using Seurat (R) and Scanpy (Python).
aipoch/medical-research-skills
Unified CLI/Python interface for querying genomic, proteomic, structure, and expression data across 20+ bioinformatics databases; use when you need fast, scriptable retrieval by gene/protein IDs or…
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Query and download from NCBI Gene Expression Omnibus (GEO) and EMBL-EBI's BioStudies/ArrayExpress mirror. Bio Geo Data is an agent skill from GPTomics/bioSkills. Query and download from NCBI Gene Expression Omnibus (GEO) and EMBL-EBI's BioStudies/ArrayExpress mirror.
Bio Geo Data fits situations like: finding expression datasets; navigating SuperSeries vs SubSeries; choosing between series-matrix (submitter-normalized) and raw supplementary files; downloading via GEOparse (Python).
Run `npx skills add GPTomics/bioSkills --skill bio-geo-data -a claude-code`. Or copy the skill folder (database-access/geo-data in GPTomics/bioSkills) into .claude/skills/bio-geo-data in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-geo-data -a codex`. Or copy the skill folder (database-access/geo-data in GPTomics/bioSkills) into .agents/skills/bio-geo-data in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-geo-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-geo-data, .gemini/skills/bio-geo-data, .github/skills/bio-geo-data and .opencode/skills/bio-geo-data in your project.
Going by SKILL.md and its folder, Bio Geo Data needs Python for the scripts in its folder and the command-line tools its instructions call (pip and wget). Our summary lists: Python 3.
SKILL.md names 4 domains. In commands or code: ebi.ac.uk and ftp.ncbi.nlm.nih.gov; the agent is likely to contact these when it follows the instructions. As links in the text: archs4.org and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Geo Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Geo Data: ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 32k stars), Biopython (davila7/claude-code-templates, 32k stars), Biopython (K-Dense-AI/scientific-agent-skills, 48k stars) and Biopython (lamm-mit/scienceclaw, 244 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.