Gget
davila7/claude-code-templates
CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.
Unified CLI/Python interface to 20+ genomic databases. An agent skill from jaechang-hits/SciAgent-Skills.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gget-genomic-databases --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gget-genomic-databases .claude/skills/gget-genomic-databases && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gget-genomic-databases" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gget-genomic-databases into .claude/skills/gget-genomic-databases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gget-genomic-databases", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gget-genomic-databasesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gget-genomic-databases --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gget-genomic-databases .agents/skills/gget-genomic-databases && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gget-genomic-databases" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gget-genomic-databases into .agents/skills/gget-genomic-databases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gget-genomic-databases", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gget-genomic-databases --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gget-genomic-databases .cursor/skills/gget-genomic-databases && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gget-genomic-databases" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gget-genomic-databases into .cursor/skills/gget-genomic-databases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gget-genomic-databases", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/databases/gget-genomic-databases--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gget-genomic-databases --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gget-genomic-databases .gemini/skills/gget-genomic-databases && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gget-genomic-databases" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gget-genomic-databases into .gemini/skills/gget-genomic-databases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gget-genomic-databases", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills gget-genomic-databasesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gget-genomic-databases .github/skills/gget-genomic-databases && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gget-genomic-databases" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gget-genomic-databases into .github/skills/gget-genomic-databases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gget-genomic-databases", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gget-genomic-databases --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gget-genomic-databases .opencode/skills/gget-genomic-databases && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gget-genomic-databases" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gget-genomic-databases into .opencode/skills/gget-genomic-databases/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gget-genomic-databases", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gget-genomic-databasesUnified CLI/Python interface to 20+ genomic databases. An agent skill from jaechang-hits/SciAgent-Skills.
Gget Genomic Databases is an agent skill from jaechang-hits/SciAgent-Skills. Unified CLI/Python interface to 20+ genomic databases. Gene lookups (Ensembl search/info/seq), BLAST/BLAT, AlphaFold, Enrichr enrichment, OpenTargets disease/drug, CELLxGENE single-cell, cBioPortal/COSMIC cancer, ARCHS4 expression. Spans genomics, proteomics, disease. For batch/advanced BLAST use biopython; for multi-DB Python SDK use bioservices.
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/databases_workflows.md` and `references/module_parameters.md`).
It sits in Research & Science, covering Bioinformatics. It works with Python, AlphaFold, Ensembl and Biopython. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-2-Clause.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
pachterlab.github.iogithub.comdoi.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gget Genomic Databases loads about 5.3k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 1,317 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-2-Clause licence (© jaechang-hits). 1,317 words, ~5,256 tokens.
.claude/skills/gget-genomic-databases/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.gget is a command-line and Python package providing unified access to 20+ genomic databases and analysis methods. Query gene information, sequences, protein structures, expression data, and disease associations through a consistent interface. All modules work as both CLI tools and Python functions, returning DataFrames (Python) or JSON/CSV (CLI).
biopython insteadbioservices insteadggetgget setup <module> before first use (alphafold, cellxgene, elm, gpt)time.sleep(). Databases update biweekly; keep gget updated. Max ~1000 Ensembl IDs per gget.info() callpip install gget
# Optional: setup modules that need additional dependencies
gget setup alphafold # ~4GB model parameters, requires OpenMM
gget setup cellxgene # cellxgene-census package
gget setup elm # local ELM databaseimport gget
# Search for genes by keyword
results = gget.search(["BRCA1", "tumor suppressor"], species="homo_sapiens")
print(f"Found {len(results)} genes")
# Get detailed gene information (Ensembl + UniProt + NCBI)
info = gget.info(["ENSG00000012048"])
print(f"Gene: {info.iloc[0]['primary_gene_name']}")
# Enrichment analysis on a gene list
enrichment = gget.enrichr(["ACE2", "AGT", "AGTR1"], database="ontology")
print(f"Enriched terms: {len(enrichment)}")Query Ensembl for gene references, search by keywords, retrieve gene metadata, and fetch sequences.
import gget
# Search for genes by keyword
results = gget.search(["BRCA1", "tumor suppressor"], species="homo_sapiens")
print(f"Found {len(results)} genes")
print(results[["ensembl_id", "gene_name", "biotype"]].head())
# Get detailed gene information (Ensembl + UniProt + NCBI)
info = gget.info(["ENSG00000012048", "ENSG00000139618"])
print(f"Gene info columns: {list(info.columns)}")import gget
# Retrieve sequences
nucleotide_seqs = gget.seq(["ENSG00000012048"])
protein_seqs = gget.seq(["ENSG00000012048"], translate=True, isoforms=True)
print(f"Retrieved {len(protein_seqs)} isoform sequences")
# Download reference genome files (specify release for reproducibility)
ref_links = gget.ref("homo_sapiens", which="gtf", release=112)
print(f"GTF download link: {ref_links}")BLAST/BLAT remote searches, multiple sequence alignment, and fast local alignment.
import gget
import time
# BLAST against SwissProt (remote API — add delay for batch queries)
blast_results = gget.blast(
"MKWMFKEDHSLEHRCVESAKIRAKYPDRVPVIVEKVSGSQIVDIDKRKYLVPSDITVAQFMWIIRKRIQLPSEKAIFLFVDKTVPQSR",
database="swissprot", limit=10
)
print(f"Top hit: {blast_results.iloc[0]['Description']}, E-value: {blast_results.iloc[0]['e-value']}")
time.sleep(2) # Rate-limit between BLAST queries
# BLAT — find genomic position (UCSC)
blat_results = gget.blat("ATCGATCGATCGATCGATCG", assembly="human")
print(f"Genomic location: chr{blat_results.iloc[0]['chromosome']}:{blat_results.iloc[0]['start']}")import gget
# Multiple sequence alignment with Muscle5
aligned = gget.muscle("sequences.fasta", save=True)
# Fast local alignment with DIAMOND (local, no rate limit needed)
diamond_results = gget.diamond(
"GGETISAWESQME",
reference="reference.fasta",
sensitivity="very-sensitive",
threads=4
)
print(f"Alignments found: {len(diamond_results)}")Download PDB structures, predict structures with AlphaFold2, find linear motifs.
import gget
# Download PDB structure
pdb_data = gget.pdb("7S7U", save=True)
# Predict structure with AlphaFold2 (requires gget setup alphafold)
structure = gget.alphafold(
"MKWMFKEDHSLEHRCVESAKIRAKYPDRVPVIVEKVSGSQIVDIDKRKYLVPSDITVAQFMWIIRKRIQLPSEKAIFLFVDKTVPQSR",
plot=True, show_sidechains=True
)
print("Structure prediction complete, PDB file saved")import gget
# Find Eukaryotic Linear Motifs (requires gget setup elm)
ortholog_df, regex_df = gget.elm("LIAQSIGQASFV")
print(f"Ortholog motifs: {len(ortholog_df)}, Regex motifs: {len(regex_df)}")Gene expression, tissue expression, correlated genes, single-cell data.
import gget
# Tissue expression from ARCHS4
tissue_expr = gget.archs4("ACE2", which="tissue")
print(f"Expression across {len(tissue_expr)} tissues")
# Correlated genes from ARCHS4
correlated = gget.archs4("ACE2", which="correlation")
print(f"Top correlated gene: {correlated.iloc[0]['gene_symbol']}")import gget
# Single-cell data from CELLxGENE (requires gget setup cellxgene)
adata = gget.cellxgene(
gene=["ACE2", "TMPRSS2"],
tissue="lung",
cell_type="epithelial cell",
census_version="2023-07-25" # pin version for reproducibility
)
print(f"Cells: {adata.n_obs}, Genes: {adata.n_vars}")
# Orthologs and expression from Bgee
orthologs = gget.bgee("ENSG00000169194", type="orthologs")
print(f"Orthologs in {len(orthologs)} species")Disease associations, drug targets, enrichment analysis.
import gget
# Disease associations from OpenTargets
diseases = gget.opentargets("ENSG00000169194", resource="diseases", limit=10)
print(f"Associated diseases: {len(diseases)}")
# Drug associations
drugs = gget.opentargets("ENSG00000169194", resource="drugs", limit=10)
print(f"Associated drugs: {len(drugs)}")
# OpenTargets resources: diseases, drugs, tractability, pharmacogenetics,
# expression, depmap, interactionsimport gget
# Enrichment analysis via Enrichr
# Database shortcuts: 'pathway' (KEGG), 'transcription' (ChEA),
# 'ontology' (GO_BP), 'diseases_drugs' (GWAS), 'celltypes' (PanglaoDB)
enrichment = gget.enrichr(
["ACE2", "AGT", "AGTR1", "TMPRSS2", "DPP4"],
database="ontology"
)
print(f"Enriched terms: {len(enrichment)}")
print(enrichment[["Term", "Adjusted P-value"]].head())Cancer mutations, copy number alterations, and somatic mutation databases.
import gget
# Search cBioPortal studies
studies = gget.cbio_search(["breast", "lung"])
print(f"Studies found: {len(studies)}")
# Plot cancer genomics heatmap
gget.cbio_plot(
["msk_impact_2017"],
["AKT1", "ALK", "BRAF"],
stratification="tissue",
variation_type="mutation_occurrences"
)import gget
# COSMIC: requires account + local database download
# First-time: gget.cosmic(searchterm="", download_cosmic=True,
# email="user@example.com", password="xxx", cosmic_project="cancer")
cosmic_results = gget.cosmic("EGFR", cosmic_tsv_path="cosmic_data.tsv", limit=10)
print(f"COSMIC mutations: {len(cosmic_results)}")Generate mutated sequences and manage module dependencies.
import gget
import pandas as pd
# Generate mutated sequences from mutation annotations
mutations_df = pd.DataFrame({
"seq_ID": ["seq1", "seq1"],
"mutation": ["c.4G>T", "c.10del"]
})
mutated = gget.mutate(["ATCGCTAAGCTGATCG"], mutations=mutations_df)
print(f"Generated {len(mutated)} mutated sequences")gget organizes 20+ modules by domain. Python interface uses gget.<module>():
| Domain | Modules | Primary Database |
|---|---|---|
| Gene reference | ref, search, info, seq | Ensembl, UniProt, NCBI |
| Sequence alignment | blast, blat, muscle, diamond | NCBI BLAST, UCSC, local |
| Protein structure | pdb, alphafold, elm | RCSB PDB, AlphaFold2, ELM |
| Expression | archs4, cellxgene, bgee | ARCHS4, CZ CELLxGENE, Bgee |
| Disease/drugs | opentargets, enrichr | OpenTargets, Enrichr |
| Cancer | cbio, cosmic | cBioPortal, COSMIC |
| Utilities | mutate, setup, gpt | local / OpenAI |
| Context | Default Format | Alternatives |
|---|---|---|
| Python | DataFrame or dict | json=True for JSON; save=True to file |
| CLI | JSON | -csv for CSV; -o file to save |
| Sequences | FASTA (seq, mutate) | -- |
| Structures | PDB file (pdb, alphafold) | JSON alignment error data |
| Single-cell | AnnData object (cellxgene) | meta_only=True for metadata only |
| Visualization | PNG (cbio plot) | show=True for interactive display |
| Shortcut | Full Database Name |
|---|---|
'pathway' | KEGG_2021_Human |
'transcription' | ChEA_2016 |
'ontology' | GO_Biological_Process_2021 |
'diseases_drugs' | GWAS_Catalog_2019 |
'celltypes' | PanglaoDB_Augmented_2021 |
Custom libraries: pass any Enrichr library name directly (e.g., "Jensen_TISSUES").
| Resource | Description |
|---|---|
diseases | Disease associations with evidence scores |
drugs | Drug associations and clinical trial data |
tractability | Target tractability assessment |
pharmacogenetics | Pharmacogenetic variants |
expression | Baseline tissue expression |
depmap | DepMap gene-disease effects |
interactions | Protein-protein interactions |
Pin database versions for consistent results across analyses:
import gget
# Pin Ensembl release
ref = gget.ref("homo_sapiens", release=112)
# Pin CELLxGENE Census version
adata = gget.cellxgene(gene=["ACE2"], census_version="2023-07-25")
# Always record gget version
print(f"gget version: {gget.__version__}")Goal: Find genes of interest, get their sequences, and perform enrichment analysis.
import gget
# 1. Search for genes
results = gget.search(["GABA", "receptor"], species="homo_sapiens")
gene_ids = results["ensembl_id"].tolist()[:10]
# 2. Get detailed information
info = gget.info(gene_ids)
print(f"Retrieved info for {len(info)} genes")
# 3. Get protein sequences
sequences = gget.seq(gene_ids, translate=True)
# 4. Find correlated genes
correlated = gget.archs4(info.index[0], which="correlation")
# 5. Enrichment analysis on correlated genes
gene_list = correlated["gene_symbol"].tolist()[:50]
enrichment = gget.enrichr(gene_list, database="ontology")
print(f"Top enriched term: {enrichment.iloc[0]['Term']}")Goal: Investigate a gene's disease associations, druggability, and cancer mutations.
import gget
gene_id = "ENSG00000169194" # ZBTB16
# 1. Disease associations
diseases = gget.opentargets(gene_id, resource="diseases", limit=20)
# 2. Drug associations
drugs = gget.opentargets(gene_id, resource="drugs")
# 3. Tractability assessment
tractability = gget.opentargets(gene_id, resource="tractability")
# 4. Protein interactions
interactions = gget.opentargets(gene_id, resource="interactions")
print(f"Diseases: {len(diseases)}, Drugs: {len(drugs)}, Interactions: {len(interactions)}")
# 5. Cancer genomics
gget.cbio_plot(["msk_impact_2017"], ["ZBTB16"], stratification="cancer_type")Goal: Compare a gene across species using orthologs and sequence alignment.
import gget
# 1. Find orthologs
orthologs = gget.bgee("ENSG00000169194", type="orthologs")
# 2. Get sequences for human and mouse
human_seq = gget.seq("ENSG00000169194", translate=True)
mouse_seq = gget.seq("ENSMUSG00000026091", translate=True)
# 3. Align sequences
alignment = gget.muscle([human_seq, mouse_seq])
# 4. Get human protein structure from PDB
pdb_structure = gget.pdb("7S7U")
print("Comparative analysis complete")| Parameter | Module(s) | Default | Range / Options | Effect |
|---|---|---|---|---|
species | search, archs4, cellxgene, enrichr | "homo_sapiens" | Any Ensembl species; shortcuts: 'human', 'mouse' | Target organism |
limit | blast, opentargets, cosmic | 50 / 100 | 1-1000 | Maximum results returned |
database | blast, enrichr | varies | blast: nt/nr/swissprot/pdbaa; enrichr: shortcuts or library names | Target database for query |
which | ref, archs4 | varies | ref: gtf,cdna,dna,cds,pep; archs4: correlation,tissue | Data type to retrieve |
translate | seq | False | True/False | Return amino acid instead of nucleotide sequences |
resource | opentargets | "diseases" | diseases, drugs, tractability, pharmacogenetics, expression, depmap, interactions | OpenTargets data type |
release | ref, search | latest | Integer Ensembl release number | Pin database version for reproducibility |
census_version | cellxgene | "stable" | "stable", "latest", date string | Pin CELLxGENE Census version |
sensitivity | diamond, elm | "very-sensitive" | fast to ultra-sensitive | Alignment sensitivity vs speed |
threads | diamond, elm | 1 | 1-N | CPU threads for alignment |
multimer_recycles | alphafold | 3 | 3-20 | Higher = more accurate multimer prediction |
Pin database versions for reproducibility: Use release=112 for Ensembl and census_version="2023-07-25" for CELLxGENE to ensure consistent results across analyses.
Rate-limit batch queries: gget queries remote APIs. Add time.sleep(2) between BLAST/BLAT queries in loops. For gget.info(), limit to ~1000 IDs per call.
Keep gget updated: Databases change their structure biweekly. Run pip install --upgrade gget regularly to avoid breakage from schema changes.
Use Python interface for pipelines, CLI for exploration: Python functions return DataFrames suitable for chaining. CLI with -csv is better for quick one-off lookups.
Check PDB before running AlphaFold: gget.pdb() is instant; AlphaFold prediction takes minutes to hours. Always check if the structure already exists in PDB.
Use database shortcuts in enrichr: The shortcuts ('pathway', 'ontology', etc.) map to curated Enrichr libraries. For custom analyses, pass any Enrichr library name directly.
Cache cBioPortal data for repeated analyses: Use data_dir="./cache" parameter to avoid re-downloading large cancer genomics datasets.
When to use: Need information for many genes at once (up to ~1000 IDs per call).
import gget
import time
gene_ids = ["ENSG00000012048", "ENSG00000139618", "ENSG00000141510"]
info = gget.info(gene_ids)
info.to_csv("gene_info_batch.csv")
print(f"Saved info for {len(info)} genes")
# For >1000 genes, batch with rate limiting
all_ids = [f"ENSG{i:011d}" for i in range(2000)]
results = []
for i in range(0, len(all_ids), 500):
batch = all_ids[i:i+500]
results.append(gget.info(batch))
time.sleep(1)When to use: Running enrichment against a custom background gene set.
import gget
# Use specific Enrichr library with background genes
enrichment = gget.enrichr(
["ACE2", "AGT", "AGTR1"],
database="Jensen_TISSUES",
background_list=["ACE2", "AGT", "AGTR1", "TP53", "BRCA1", "MYC"]
)
print(enrichment[["Term", "Adjusted P-value"]].head())When to use: Predicting and visualizing protein structures with confidence coloring.
import gget
# Predict with visualization (PAE + 3D structure)
result = gget.alphafold(
"MKWMFKEDHSLEHRCVESAKIRAKYPDRVPVIVEKVSGSQIVDIDKRKYLVPSDITVAQFMWIIRKRIQLPSEKAIFLFVDKTVPQSR",
plot=True,
show_sidechains=True,
relax=True # AMBER relaxation for final structure
)
# Output: PDB file + predicted aligned error (PAE) JSON
# PAE heatmap auto-generated with plot=TrueWhen to use: Setting up reference files for RNA-seq alignment pipelines.
# Download GTF and cDNA for human (specific release)
gget ref -w gtf -w cdna -d -r 112 homo_sapiens
# Download genome DNA
gget ref -w dna -d homo_sapiens| Problem | Cause | Solution |
|---|---|---|
ModuleNotFoundError: gget | Package not installed | pip install gget in clean virtual environment |
gget setup alphafold fails | Python version incompatibility | Use Python 3.8-3.10; check gget --version |
| Empty BLAST results | Sequence too short or no matches | Try longer sequence, different database, or megablast_off=True |
cellxgene gene not found | Case-sensitive gene symbols | Use 'ACE2' for human, 'Ace2' for mouse (exact capitalization required) |
gget info timeout | Too many IDs at once | Limit to ~1000 Ensembl IDs per call; batch with time.sleep() |
| Database structure changed | gget databases update biweekly | pip install --upgrade gget |
| COSMIC authentication error | Missing or expired credentials | Re-enter email/password; check COSMIC account status |
| AlphaFold out of memory | Protein too long for GPU memory | Use shorter sequences or split into domains |
| Different results on re-run | Database updated between runs | Pin versions: release=112 for Ensembl, census_version for CELLxGENE |
2 reference files provide extended coverage of capabilities from the original 3 reference files and 3 script files:
references/module_parameters.md — Consolidates module_reference.md (468 lines). Covers: detailed parameter tables for all 15+ modules with types, defaults, and return value descriptions; CLI vs Python interface differences; setup requirements per module. Relocated inline: most-used module parameters (Core API code blocks), output format summary (Key Concepts table). Omitted: gget gpt module details — trivial OpenAI wrapper, not genomics-specific.
references/databases_workflows.md — Consolidates database_info.md (301 lines) and workflows.md (815 lines). Covers: complete database directory with update frequencies and citation info, extended workflow examples (building reference indices, disease-drug pipeline, multi-species comparative analysis), data consistency and reproducibility guidance. Relocated inline: core database overview (Key Concepts table), top 3 workflows (Common Workflows), reproducibility patterns (Key Concepts). Omitted: scripts/ content (3 files, 590 lines total) — thin wrappers around gget API calls for CLI automation; core patterns absorbed into Core API and Common Workflows.
gget.cellxgene()© jaechang-hits, BSD-2-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/genomics-bioinformatics/databases/gget-genomic-databases of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Gget Genomic Databases next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gget Genomic Databases this skilljaechang-hits/SciAgent-Skills | 374 | 2 repos | ~5.3k | Automated safety check: Pass | BSD-2-Clause | |
| Ggetdavila7/claude-code-templates | 33k | 10 repos | ~6.3k | Automated safety check: Pass | MIT | |
| Ggetaipoch/medical-research-skills | 1.9k | — | ~816 | Automated safety check: Pass | MIT | |
| Biopythondavila7/claude-code-templates | 33k | 12 repos | ~3.4k | Automated safety check: Pass | MIT | |
| Scientific Pkg Ggetaffaan-m/ECC | 277k | 1 repos | ~1.3k | Automated safety check: Pass | MIT | |
| GgetK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~2.8k | Automated safety check: Notes | BSD-2-Clause |
davila7/claude-code-templates
CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.
aipoch/medical-research-skills
Unified CLI/Python interface for querying genomic, proteomic, structure, and expression data across 20+ bioinformatics databases; use when you need fast, scriptable retrieval by gene/protein IDs or…
davila7/claude-code-templates
Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.
affaan-m/ECC
gget CLI and Python workflow for quick genomic database queries, sequence lookup, BLAST-style searches, enrichment checks, and reproducible bioinformatics evidence logs.
K-Dense-AI/scientific-agent-skills
Queries 20+ bioinformatics resources through CLI/Python. An agent skill from K-Dense-AI/scientific-agent-skills.
K-Dense-AI/scientific-agent-skills
Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
Unified CLI/Python interface to 20+ genomic databases. An agent skill from jaechang-hits/SciAgent-Skills. Gget Genomic Databases is an agent skill from jaechang-hits/SciAgent-Skills. Unified CLI/Python interface to 20+ genomic databases.
Gget Genomic Databases fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/databases/gget-genomic-databases in jaechang-hits/SciAgent-Skills) into .claude/skills/gget-genomic-databases in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/databases/gget-genomic-databases in jaechang-hits/SciAgent-Skills) into .agents/skills/gget-genomic-databases in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gget-genomic-databases, .gemini/skills/gget-genomic-databases, .github/skills/gget-genomic-databases and .opencode/skills/gget-genomic-databases in your project.
Going by SKILL.md and its folder, Gget Genomic Databases needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 3 domains. As links in the text: pachterlab.github.io, github.com and doi.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gget Genomic Databases is published under the BSD-2-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Gget Genomic Databases: Gget (davila7/claude-code-templates, 33k stars), Gget (aipoch/medical-research-skills, 1.9k stars), Biopython (davila7/claude-code-templates, 33k stars) and Scientific Pkg Gget (affaan-m/ECC, 277k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.