Agent skill

Gget Genomic Databases

by jaechang-hits in jaechang-hits/SciAgent-Skills

Unified CLI/Python interface to 20+ genomic databases. An agent skill from jaechang-hits/SciAgent-Skills.

BSD-2-ClauseAuto-check passedResearch & Science

Install Gget Genomic Databases

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills gget-genomic-databases --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gget-genomic-databases .claude/skills/gget-genomic-databases && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gget-genomic-databases
GitHub stars
374
Used in
2 other repos
Token cost
~5.3k tokens
SKILL.md length
1,317 words
Files
3 (incl. references)
Skills in repo
169
Repo updated
First seen
Licence
BSD-2-Clause

At a glance

Unified CLI/Python interface to 20+ genomic databases. An agent skill from jaechang-hits/SciAgent-Skills.

  • Works in 7 steps: Pin database versions for… → Rate-limit batch queries: gget queries… → Keep gget updated: Databases change… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 6 more sections
  • Calls pip

What it does

Gget Genomic Databases is an agent skill from jaechang-hits/SciAgent-Skills. Unified CLI/Python interface to 20+ genomic databases. Gene lookups (Ensembl search/info/seq), BLAST/BLAT, AlphaFold, Enrichr enrichment, OpenTargets disease/drug, CELLxGENE single-cell, cBioPortal/COSMIC cancer, ARCHS4 expression. Spans genomics, proteomics, disease. For batch/advanced BLAST use biopython; for multi-DB Python SDK use bioservices.

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/databases_workflows.md` and `references/module_parameters.md`).

It sits in Research & Science, covering Bioinformatics. It works with Python, AlphaFold, Ensembl and Biopython. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-2-Clause.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/gget-genomic-databases”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Pin database versions for reproducibility: Use release=112 for Ensembl and census_version="2023-07-25" for CELLxGENE to ensure consistent…
  2. Rate-limit batch queries: gget queries remote APIs. Add time.sleep(2) between BLAST/BLAT queries in loops. For gget.info(), limit to ~1000…
  3. Keep gget updated: Databases change their structure biweekly. Run pip install --upgrade gget regularly to avoid breakage from schema…
  4. Use Python interface for pipelines, CLI for exploration: Python functions return DataFrames suitable for chaining. CLI with -csv is better…
  5. Check PDB before running AlphaFold: gget.pdb() is instant; AlphaFold prediction takes minutes to hours. Always check if the structure…
  6. Use database shortcuts in enrichr: The shortcuts ('pathway', 'ontology', etc.) map to curated Enrichr libraries. For custom analyses, pass…
  7. Cache cBioPortal data for repeated analyses: Use data_dir="./cache" parameter to avoid re-downloading large cancer genomics datasets.

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • pachterlab.github.io
    • github.com
    • doi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gget Genomic Databases loads about 5.3k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 1,317 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-2-Clause licence (© jaechang-hits). 1,317 words, ~5,256 tokens.

Download SKILL.mdSave it as .claude/skills/gget-genomic-databases/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
gget-genomic-databases
description
Unified CLI/Python interface to 20+ genomic databases. Gene lookups (Ensembl search/info/seq), BLAST/BLAT, AlphaFold, Enrichr enrichment, OpenTargets disease/drug, CELLxGENE single-cell, cBioPortal/COSMIC cancer, ARCHS4 expression. Spans genomics, proteomics, disease. For batch/advanced BLAST use biopython; for multi-DB Python SDK use bioservices.
license
BSD-2-Clause

gget — Unified Genomic Database Access

Overview

gget is a command-line and Python package providing unified access to 20+ genomic databases and analysis methods. Query gene information, sequences, protein structures, expression data, and disease associations through a consistent interface. All modules work as both CLI tools and Python functions, returning DataFrames (Python) or JSON/CSV (CLI).

When to Use

  • Looking up gene information (names, IDs, descriptions) across species from Ensembl
  • Retrieving nucleotide or protein sequences for Ensembl gene/transcript IDs
  • Running BLAST or BLAT searches against standard reference databases
  • Predicting protein 3D structures with AlphaFold2 from amino acid sequences
  • Performing gene set enrichment analysis (GO, KEGG, disease terms) via Enrichr
  • Querying single-cell RNA-seq datasets from CELLxGENE Census
  • Finding disease and drug associations for a gene target via OpenTargets
  • Downloading Ensembl reference genomes and annotations for a species
  • Finding cancer mutations and genomic alterations via cBioPortal or COSMIC
  • Getting tissue expression and correlated genes from ARCHS4
  • For batch processing or advanced BLAST parameters, use biopython instead
  • For programmatic multi-database workflows with rate limiting, use bioservices instead

Prerequisites

  • Python packages: gget
  • Optional setup: Some modules require gget setup <module> before first use (alphafold, cellxgene, elm, gpt)
  • Environment: Clean virtual environment recommended to avoid dependency conflicts
  • API notes: gget queries remote databases — rate-limit large batch queries with time.sleep(). Databases update biweekly; keep gget updated. Max ~1000 Ensembl IDs per gget.info() call
bash
pip install gget

# Optional: setup modules that need additional dependencies
gget setup alphafold   # ~4GB model parameters, requires OpenMM
gget setup cellxgene   # cellxgene-census package
gget setup elm         # local ELM database

Quick Start

python
import gget

# Search for genes by keyword
results = gget.search(["BRCA1", "tumor suppressor"], species="homo_sapiens")
print(f"Found {len(results)} genes")

# Get detailed gene information (Ensembl + UniProt + NCBI)
info = gget.info(["ENSG00000012048"])
print(f"Gene: {info.iloc[0]['primary_gene_name']}")

# Enrichment analysis on a gene list
enrichment = gget.enrichr(["ACE2", "AGT", "AGTR1"], database="ontology")
print(f"Enriched terms: {len(enrichment)}")

Core API

Module 1: Reference & Gene Search (ref, search, info, seq)

Query Ensembl for gene references, search by keywords, retrieve gene metadata, and fetch sequences.

python
import gget

# Search for genes by keyword
results = gget.search(["BRCA1", "tumor suppressor"], species="homo_sapiens")
print(f"Found {len(results)} genes")
print(results[["ensembl_id", "gene_name", "biotype"]].head())

# Get detailed gene information (Ensembl + UniProt + NCBI)
info = gget.info(["ENSG00000012048", "ENSG00000139618"])
print(f"Gene info columns: {list(info.columns)}")
python
import gget

# Retrieve sequences
nucleotide_seqs = gget.seq(["ENSG00000012048"])
protein_seqs = gget.seq(["ENSG00000012048"], translate=True, isoforms=True)
print(f"Retrieved {len(protein_seqs)} isoform sequences")

# Download reference genome files (specify release for reproducibility)
ref_links = gget.ref("homo_sapiens", which="gtf", release=112)
print(f"GTF download link: {ref_links}")
Module 2: Sequence Alignment (blast, blat, muscle, diamond)

BLAST/BLAT remote searches, multiple sequence alignment, and fast local alignment.

python
import gget
import time

# BLAST against SwissProt (remote API — add delay for batch queries)
blast_results = gget.blast(
    "MKWMFKEDHSLEHRCVESAKIRAKYPDRVPVIVEKVSGSQIVDIDKRKYLVPSDITVAQFMWIIRKRIQLPSEKAIFLFVDKTVPQSR",
    database="swissprot", limit=10
)
print(f"Top hit: {blast_results.iloc[0]['Description']}, E-value: {blast_results.iloc[0]['e-value']}")
time.sleep(2)  # Rate-limit between BLAST queries

# BLAT — find genomic position (UCSC)
blat_results = gget.blat("ATCGATCGATCGATCGATCG", assembly="human")
print(f"Genomic location: chr{blat_results.iloc[0]['chromosome']}:{blat_results.iloc[0]['start']}")
python
import gget

# Multiple sequence alignment with Muscle5
aligned = gget.muscle("sequences.fasta", save=True)

# Fast local alignment with DIAMOND (local, no rate limit needed)
diamond_results = gget.diamond(
    "GGETISAWESQME",
    reference="reference.fasta",
    sensitivity="very-sensitive",
    threads=4
)
print(f"Alignments found: {len(diamond_results)}")
Module 3: Protein Structure (pdb, alphafold, elm)

Download PDB structures, predict structures with AlphaFold2, find linear motifs.

python
import gget

# Download PDB structure
pdb_data = gget.pdb("7S7U", save=True)

# Predict structure with AlphaFold2 (requires gget setup alphafold)
structure = gget.alphafold(
    "MKWMFKEDHSLEHRCVESAKIRAKYPDRVPVIVEKVSGSQIVDIDKRKYLVPSDITVAQFMWIIRKRIQLPSEKAIFLFVDKTVPQSR",
    plot=True, show_sidechains=True
)
print("Structure prediction complete, PDB file saved")
python
import gget

# Find Eukaryotic Linear Motifs (requires gget setup elm)
ortholog_df, regex_df = gget.elm("LIAQSIGQASFV")
print(f"Ortholog motifs: {len(ortholog_df)}, Regex motifs: {len(regex_df)}")
Module 4: Expression & Correlation (archs4, cellxgene, bgee)

Gene expression, tissue expression, correlated genes, single-cell data.

python
import gget

# Tissue expression from ARCHS4
tissue_expr = gget.archs4("ACE2", which="tissue")
print(f"Expression across {len(tissue_expr)} tissues")

# Correlated genes from ARCHS4
correlated = gget.archs4("ACE2", which="correlation")
print(f"Top correlated gene: {correlated.iloc[0]['gene_symbol']}")
python
import gget

# Single-cell data from CELLxGENE (requires gget setup cellxgene)
adata = gget.cellxgene(
    gene=["ACE2", "TMPRSS2"],
    tissue="lung",
    cell_type="epithelial cell",
    census_version="2023-07-25"  # pin version for reproducibility
)
print(f"Cells: {adata.n_obs}, Genes: {adata.n_vars}")

# Orthologs and expression from Bgee
orthologs = gget.bgee("ENSG00000169194", type="orthologs")
print(f"Orthologs in {len(orthologs)} species")
Module 5: Disease & Drug Associations (opentargets, enrichr)

Disease associations, drug targets, enrichment analysis.

python
import gget

# Disease associations from OpenTargets
diseases = gget.opentargets("ENSG00000169194", resource="diseases", limit=10)
print(f"Associated diseases: {len(diseases)}")

# Drug associations
drugs = gget.opentargets("ENSG00000169194", resource="drugs", limit=10)
print(f"Associated drugs: {len(drugs)}")

# OpenTargets resources: diseases, drugs, tractability, pharmacogenetics,
#   expression, depmap, interactions
python
import gget

# Enrichment analysis via Enrichr
# Database shortcuts: 'pathway' (KEGG), 'transcription' (ChEA),
#   'ontology' (GO_BP), 'diseases_drugs' (GWAS), 'celltypes' (PanglaoDB)
enrichment = gget.enrichr(
    ["ACE2", "AGT", "AGTR1", "TMPRSS2", "DPP4"],
    database="ontology"
)
print(f"Enriched terms: {len(enrichment)}")
print(enrichment[["Term", "Adjusted P-value"]].head())
Module 6: Cancer Genomics (cbio, cosmic)

Cancer mutations, copy number alterations, and somatic mutation databases.

python
import gget

# Search cBioPortal studies
studies = gget.cbio_search(["breast", "lung"])
print(f"Studies found: {len(studies)}")

# Plot cancer genomics heatmap
gget.cbio_plot(
    ["msk_impact_2017"],
    ["AKT1", "ALK", "BRAF"],
    stratification="tissue",
    variation_type="mutation_occurrences"
)
python
import gget

# COSMIC: requires account + local database download
# First-time: gget.cosmic(searchterm="", download_cosmic=True,
#   email="user@example.com", password="xxx", cosmic_project="cancer")
cosmic_results = gget.cosmic("EGFR", cosmic_tsv_path="cosmic_data.tsv", limit=10)
print(f"COSMIC mutations: {len(cosmic_results)}")
Module 7: Mutation Generation & Utilities (mutate, setup)

Generate mutated sequences and manage module dependencies.

python
import gget
import pandas as pd

# Generate mutated sequences from mutation annotations
mutations_df = pd.DataFrame({
    "seq_ID": ["seq1", "seq1"],
    "mutation": ["c.4G>T", "c.10del"]
})
mutated = gget.mutate(["ATCGCTAAGCTGATCG"], mutations=mutations_df)
print(f"Generated {len(mutated)} mutated sequences")

Key Concepts

Module Overview

gget organizes 20+ modules by domain. Python interface uses gget.<module>():

DomainModulesPrimary Database
Gene referenceref, search, info, seqEnsembl, UniProt, NCBI
Sequence alignmentblast, blat, muscle, diamondNCBI BLAST, UCSC, local
Protein structurepdb, alphafold, elmRCSB PDB, AlphaFold2, ELM
Expressionarchs4, cellxgene, bgeeARCHS4, CZ CELLxGENE, Bgee
Disease/drugsopentargets, enrichrOpenTargets, Enrichr
Cancercbio, cosmiccBioPortal, COSMIC
Utilitiesmutate, setup, gptlocal / OpenAI
Output Formats
ContextDefault FormatAlternatives
PythonDataFrame or dictjson=True for JSON; save=True to file
CLIJSON-csv for CSV; -o file to save
SequencesFASTA (seq, mutate)--
StructuresPDB file (pdb, alphafold)JSON alignment error data
Single-cellAnnData object (cellxgene)meta_only=True for metadata only
VisualizationPNG (cbio plot)show=True for interactive display
Enrichr Database Shortcuts
ShortcutFull Database Name
'pathway'KEGG_2021_Human
'transcription'ChEA_2016
'ontology'GO_Biological_Process_2021
'diseases_drugs'GWAS_Catalog_2019
'celltypes'PanglaoDB_Augmented_2021

Custom libraries: pass any Enrichr library name directly (e.g., "Jensen_TISSUES").

OpenTargets Resources
ResourceDescription
diseasesDisease associations with evidence scores
drugsDrug associations and clinical trial data
tractabilityTarget tractability assessment
pharmacogeneticsPharmacogenetic variants
expressionBaseline tissue expression
depmapDepMap gene-disease effects
interactionsProtein-protein interactions
Reproducibility

Pin database versions for consistent results across analyses:

python
import gget
# Pin Ensembl release
ref = gget.ref("homo_sapiens", release=112)

# Pin CELLxGENE Census version
adata = gget.cellxgene(gene=["ACE2"], census_version="2023-07-25")

# Always record gget version
print(f"gget version: {gget.__version__}")

Common Workflows

Workflow 1: Gene Discovery to Functional Analysis

Goal: Find genes of interest, get their sequences, and perform enrichment analysis.

python
import gget

# 1. Search for genes
results = gget.search(["GABA", "receptor"], species="homo_sapiens")
gene_ids = results["ensembl_id"].tolist()[:10]

# 2. Get detailed information
info = gget.info(gene_ids)
print(f"Retrieved info for {len(info)} genes")

# 3. Get protein sequences
sequences = gget.seq(gene_ids, translate=True)

# 4. Find correlated genes
correlated = gget.archs4(info.index[0], which="correlation")

# 5. Enrichment analysis on correlated genes
gene_list = correlated["gene_symbol"].tolist()[:50]
enrichment = gget.enrichr(gene_list, database="ontology")
print(f"Top enriched term: {enrichment.iloc[0]['Term']}")
Workflow 2: Target Validation for Drug Discovery

Goal: Investigate a gene's disease associations, druggability, and cancer mutations.

python
import gget

gene_id = "ENSG00000169194"  # ZBTB16

# 1. Disease associations
diseases = gget.opentargets(gene_id, resource="diseases", limit=20)

# 2. Drug associations
drugs = gget.opentargets(gene_id, resource="drugs")

# 3. Tractability assessment
tractability = gget.opentargets(gene_id, resource="tractability")

# 4. Protein interactions
interactions = gget.opentargets(gene_id, resource="interactions")
print(f"Diseases: {len(diseases)}, Drugs: {len(drugs)}, Interactions: {len(interactions)}")

# 5. Cancer genomics
gget.cbio_plot(["msk_impact_2017"], ["ZBTB16"], stratification="cancer_type")
Workflow 3: Comparative Genomics

Goal: Compare a gene across species using orthologs and sequence alignment.

python
import gget

# 1. Find orthologs
orthologs = gget.bgee("ENSG00000169194", type="orthologs")

# 2. Get sequences for human and mouse
human_seq = gget.seq("ENSG00000169194", translate=True)
mouse_seq = gget.seq("ENSMUSG00000026091", translate=True)

# 3. Align sequences
alignment = gget.muscle([human_seq, mouse_seq])

# 4. Get human protein structure from PDB
pdb_structure = gget.pdb("7S7U")
print("Comparative analysis complete")

Key Parameters

ParameterModule(s)DefaultRange / OptionsEffect
speciessearch, archs4, cellxgene, enrichr"homo_sapiens"Any Ensembl species; shortcuts: 'human', 'mouse'Target organism
limitblast, opentargets, cosmic50 / 1001-1000Maximum results returned
databaseblast, enrichrvariesblast: nt/nr/swissprot/pdbaa; enrichr: shortcuts or library namesTarget database for query
whichref, archs4variesref: gtf,cdna,dna,cds,pep; archs4: correlation,tissueData type to retrieve
translateseqFalseTrue/FalseReturn amino acid instead of nucleotide sequences
resourceopentargets"diseases"diseases, drugs, tractability, pharmacogenetics, expression, depmap, interactionsOpenTargets data type
releaseref, searchlatestInteger Ensembl release numberPin database version for reproducibility
census_versioncellxgene"stable""stable", "latest", date stringPin CELLxGENE Census version
sensitivitydiamond, elm"very-sensitive"fast to ultra-sensitiveAlignment sensitivity vs speed
threadsdiamond, elm11-NCPU threads for alignment
multimer_recyclesalphafold33-20Higher = more accurate multimer prediction
Show full SKILL.md (589 more words)Show less

Best Practices

  1. Pin database versions for reproducibility: Use release=112 for Ensembl and census_version="2023-07-25" for CELLxGENE to ensure consistent results across analyses.

  2. Rate-limit batch queries: gget queries remote APIs. Add time.sleep(2) between BLAST/BLAT queries in loops. For gget.info(), limit to ~1000 IDs per call.

  3. Keep gget updated: Databases change their structure biweekly. Run pip install --upgrade gget regularly to avoid breakage from schema changes.

  4. Use Python interface for pipelines, CLI for exploration: Python functions return DataFrames suitable for chaining. CLI with -csv is better for quick one-off lookups.

  5. Check PDB before running AlphaFold: gget.pdb() is instant; AlphaFold prediction takes minutes to hours. Always check if the structure already exists in PDB.

  6. Use database shortcuts in enrichr: The shortcuts ('pathway', 'ontology', etc.) map to curated Enrichr libraries. For custom analyses, pass any Enrichr library name directly.

  7. Cache cBioPortal data for repeated analyses: Use data_dir="./cache" parameter to avoid re-downloading large cancer genomics datasets.

Common Recipes

Recipe: Batch Gene Information Retrieval

When to use: Need information for many genes at once (up to ~1000 IDs per call).

python
import gget
import time

gene_ids = ["ENSG00000012048", "ENSG00000139618", "ENSG00000141510"]
info = gget.info(gene_ids)
info.to_csv("gene_info_batch.csv")
print(f"Saved info for {len(info)} genes")

# For >1000 genes, batch with rate limiting
all_ids = [f"ENSG{i:011d}" for i in range(2000)]
results = []
for i in range(0, len(all_ids), 500):
    batch = all_ids[i:i+500]
    results.append(gget.info(batch))
    time.sleep(1)
Recipe: Custom Enrichment with Background

When to use: Running enrichment against a custom background gene set.

python
import gget

# Use specific Enrichr library with background genes
enrichment = gget.enrichr(
    ["ACE2", "AGT", "AGTR1"],
    database="Jensen_TISSUES",
    background_list=["ACE2", "AGT", "AGTR1", "TP53", "BRCA1", "MYC"]
)
print(enrichment[["Term", "Adjusted P-value"]].head())
Recipe: AlphaFold Structure Prediction with Visualization

When to use: Predicting and visualizing protein structures with confidence coloring.

python
import gget

# Predict with visualization (PAE + 3D structure)
result = gget.alphafold(
    "MKWMFKEDHSLEHRCVESAKIRAKYPDRVPVIVEKVSGSQIVDIDKRKYLVPSDITVAQFMWIIRKRIQLPSEKAIFLFVDKTVPQSR",
    plot=True,
    show_sidechains=True,
    relax=True  # AMBER relaxation for final structure
)
# Output: PDB file + predicted aligned error (PAE) JSON
# PAE heatmap auto-generated with plot=True
Recipe: Download Reference Genome for RNA-seq Pipeline

When to use: Setting up reference files for RNA-seq alignment pipelines.

bash
# Download GTF and cDNA for human (specific release)
gget ref -w gtf -w cdna -d -r 112 homo_sapiens

# Download genome DNA
gget ref -w dna -d homo_sapiens

Troubleshooting

ProblemCauseSolution
ModuleNotFoundError: ggetPackage not installedpip install gget in clean virtual environment
gget setup alphafold failsPython version incompatibilityUse Python 3.8-3.10; check gget --version
Empty BLAST resultsSequence too short or no matchesTry longer sequence, different database, or megablast_off=True
cellxgene gene not foundCase-sensitive gene symbolsUse 'ACE2' for human, 'Ace2' for mouse (exact capitalization required)
gget info timeoutToo many IDs at onceLimit to ~1000 Ensembl IDs per call; batch with time.sleep()
Database structure changedgget databases update biweeklypip install --upgrade gget
COSMIC authentication errorMissing or expired credentialsRe-enter email/password; check COSMIC account status
AlphaFold out of memoryProtein too long for GPU memoryUse shorter sequences or split into domains
Different results on re-runDatabase updated between runsPin versions: release=112 for Ensembl, census_version for CELLxGENE

Bundled Resources

2 reference files provide extended coverage of capabilities from the original 3 reference files and 3 script files:

  1. references/module_parameters.md — Consolidates module_reference.md (468 lines). Covers: detailed parameter tables for all 15+ modules with types, defaults, and return value descriptions; CLI vs Python interface differences; setup requirements per module. Relocated inline: most-used module parameters (Core API code blocks), output format summary (Key Concepts table). Omitted: gget gpt module details — trivial OpenAI wrapper, not genomics-specific.

  2. references/databases_workflows.md — Consolidates database_info.md (301 lines) and workflows.md (815 lines). Covers: complete database directory with update frequencies and citation info, extended workflow examples (building reference indices, disease-drug pipeline, multi-species comparative analysis), data consistency and reproducibility guidance. Relocated inline: core database overview (Key Concepts table), top 3 workflows (Common Workflows), reproducibility patterns (Key Concepts). Omitted: scripts/ content (3 files, 590 lines total) — thin wrappers around gget API calls for CLI automation; core patterns absorbed into Core API and Common Workflows.

  • biopython — advanced BLAST parameters, batch sequence processing, GenBank record parsing
  • bioservices — programmatic multi-database queries with built-in rate limiting (UniProt, KEGG, ChEMBL)
  • anndata-data-structure — working with AnnData objects returned by gget.cellxgene()
  • enrichr — deeper enrichment analysis with custom gene set libraries

References

© jaechang-hits, BSD-2-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/genomics-bioinformatics/databases/gget-genomic-databases of jaechang-hits/SciAgent-Skills.

  • SKILL.md
  • references/databases_workflows.md
  • references/module_parameters.md

Open the folder on GitHubat commit 82c862c

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Gget Genomic Databases next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gget Genomic Databases compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gget Genomic Databases this skilljaechang-hits/SciAgent-Skills3742 repos~5.3kAutomated safety check: PassBSD-2-Clause
Ggetdavila7/claude-code-templates33k10 repos~6.3kAutomated safety check: PassMIT
Ggetaipoch/medical-research-skills1.9k—~816Automated safety check: PassMIT
Biopythondavila7/claude-code-templates33k12 repos~3.4kAutomated safety check: PassMIT
Scientific Pkg Ggetaffaan-m/ECC277k1 repos~1.3kAutomated safety check: PassMIT
GgetK-Dense-AI/scientific-agent-skills48k1 repos~2.8kAutomated safety check: NotesBSD-2-Clause

Similar skills

  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 10 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • Gget

    aipoch/medical-research-skills

    Unified CLI/Python interface for querying genomic, proteomic, structure, and expression data across 20+ bioinformatics databases; use when you need fast, scriptable retrieval by gene/protein IDs or…

    1.9k GitHub stars~816 tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 12 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • gget CLI and Python workflow for quick genomic database queries, sequence lookup, BLAST-style searches, enrichment checks, and reproducible bioinformatics evidence logs.

    277k GitHub starsUsed in 1 repo~1.3k tokens
    Research & ScienceAuto-check passed
  • Gget

    K-Dense-AI/scientific-agent-skills

    Queries 20+ bioinformatics resources through CLI/Python. An agent skill from K-Dense-AI/scientific-agent-skills.

    48k GitHub starsUsed in 1 repo~2.8k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    K-Dense-AI/scientific-agent-skills

    Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).

    48k GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check: notes

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Questions about Gget Genomic Databases

What does Gget Genomic Databases do?

Unified CLI/Python interface to 20+ genomic databases. An agent skill from jaechang-hits/SciAgent-Skills. Gget Genomic Databases is an agent skill from jaechang-hits/SciAgent-Skills. Unified CLI/Python interface to 20+ genomic databases.

When should I use Gget Genomic Databases?

Gget Genomic Databases fits situations like: tasks that involve Bioinformatics.

How do I install Gget Genomic Databases in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/databases/gget-genomic-databases in jaechang-hits/SciAgent-Skills) into .claude/skills/gget-genomic-databases in your project. Claude Code loads it when a task matches its description.

How do I install Gget Genomic Databases in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/databases/gget-genomic-databases in jaechang-hits/SciAgent-Skills) into .agents/skills/gget-genomic-databases in your project. Codex loads it when a task matches its description.

Can I use Gget Genomic Databases in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill gget-genomic-databases -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gget-genomic-databases, .gemini/skills/gget-genomic-databases, .github/skills/gget-genomic-databases and .opencode/skills/gget-genomic-databases in your project.

What does Gget Genomic Databases need to run?

Going by SKILL.md and its folder, Gget Genomic Databases needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Gget Genomic Databases access the network?

SKILL.md names 3 domains. As links in the text: pachterlab.github.io, github.com and doi.org. This is read from the text; nothing was executed.

Is Gget Genomic Databases safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gget Genomic Databases use?

Gget Genomic Databases is published under the BSD-2-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gget Genomic Databases use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.8k tokens, read only when the agent opens those files.

What are the alternatives to Gget Genomic Databases?

Skills that share tags, products or a category with Gget Genomic Databases: Gget (davila7/claude-code-templates, 33k stars), Gget (aipoch/medical-research-skills, 1.9k stars), Biopython (davila7/claude-code-templates, 33k stars) and Scientific Pkg Gget (affaan-m/ECC, 277k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gget Genomic Databases?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.