Agent skill

Geo Database

by davila7 in davila7/claude-code-templates

Access NCBI GEO for gene expression/genomics data. An agent skill from davila7/claude-code-templates.

MITAuto-check passedResearch & Science

Install Geo Database

skills CLI
$ npx skills add davila7/claude-code-templates --skill geo-database -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates geo-database --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/scientific/geo-database .claude/skills/geo-database && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
geo-database
GitHub stars
32k
Used in
11 other repos
Token cost
~6.1k tokens
SKILL.md length
991 words
Files
2 (incl. references)
Skills in repo
478
Repo updated
First seen
Licence
MIT

At a glance

Access NCBI GEO for gene expression/genomics data. An agent skill from davila7/claude-code-templates.

  • Works in 7 steps: Understanding GEO Data Organization → Searching GEO Data → Retrieving GEO Data with GEOparse… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use This Skill, Core Capabilities and Installation and Setup, plus 7 more sections
  • Calls uv and wget; reaches ncbi.nlm.nih.gov

What it does

Geo Database is an agent skill from davila7/claude-code-templates. Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis.

Its SKILL.md is about 6.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/geo_reference.md`).

It sits in Research & Science, covering Bioinformatics. It works with NCBI. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/geo-database”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Understanding GEO Data Organization
  2. Searching GEO Data
  3. Retrieving GEO Data with GEOparse (Recommended)
  4. Using NCBI E-utilities for GEO Access
  5. Direct FTP Access for Data Files
  6. Analyzing GEO Data
  7. Batch Processing Multiple Datasets

What it can do on your machine

Read from SKILL.md and the folder at commit 46b4d8b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • wget

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ncbi.nlm.nih.gov

    Also links to:

    • geoparse.readthedocs.io
    • ncbiinsights.ncbi.nlm.nih.gov
    • biopython.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Geo Database loads about 6.1k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 50 tokens; SKILL.md has 991 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~6.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 46b4d8b, republished under its MIT licence (© davila7). 991 words, ~6,118 tokens.

Download SKILL.mdSave it as .claude/skills/geo-database/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
geo-database
description
Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis.

GEO Database

Overview

The Gene Expression Omnibus (GEO) is NCBI's public repository for high-throughput gene expression and functional genomics data. GEO contains over 264,000 studies with more than 8 million samples from both array-based and sequence-based experiments.

When to Use This Skill

This skill should be used when searching for gene expression datasets, retrieving experimental data, downloading raw and processed files, querying expression profiles, or integrating GEO data into computational analysis workflows.

Core Capabilities

1. Understanding GEO Data Organization

GEO organizes data hierarchically using different accession types:

Series (GSE): A complete experiment with a set of related samples

  • Example: GSE123456
  • Contains experimental design, samples, and overall study information
  • Largest organizational unit in GEO
  • Current count: 264,928+ series

Sample (GSM): A single experimental sample or biological replicate

  • Example: GSM987654
  • Contains individual sample data, protocols, and metadata
  • Linked to platforms and series
  • Current count: 8,068,632+ samples

Platform (GPL): The microarray or sequencing platform used

  • Example: GPL570 (Affymetrix Human Genome U133 Plus 2.0 Array)
  • Describes the technology and probe/feature annotations
  • Shared across multiple experiments
  • Current count: 27,739+ platforms

DataSet (GDS): Curated collections with consistent formatting

  • Example: GDS5678
  • Experimentally-comparable samples organized by study design
  • Processed for differential analysis
  • Subset of GEO data (4,348 curated datasets)
  • Ideal for quick comparative analyses

Profiles: Gene-specific expression data linked to sequence features

  • Queryable by gene name or annotation
  • Cross-references to Entrez Gene
  • Enables gene-centric searches across all studies
2. Searching GEO Data

GEO DataSets Search:

Search for studies by keywords, organism, or experimental conditions:

python
from Bio import Entrez

# Configure Entrez (required)
Entrez.email = "your.email@example.com"

# Search for datasets
def search_geo_datasets(query, retmax=20):
    """Search GEO DataSets database"""
    handle = Entrez.esearch(
        db="gds",
        term=query,
        retmax=retmax,
        usehistory="y"
    )
    results = Entrez.read(handle)
    handle.close()
    return results

# Example searches
results = search_geo_datasets("breast cancer[MeSH] AND Homo sapiens[Organism]")
print(f"Found {results['Count']} datasets")

# Search by specific platform
results = search_geo_datasets("GPL570[Accession]")

# Search by study type
results = search_geo_datasets("expression profiling by array[DataSet Type]")

GEO Profiles Search:

Find gene-specific expression patterns:

python
# Search for gene expression profiles
def search_geo_profiles(gene_name, organism="Homo sapiens", retmax=100):
    """Search GEO Profiles for a specific gene"""
    query = f"{gene_name}[Gene Name] AND {organism}[Organism]"
    handle = Entrez.esearch(
        db="geoprofiles",
        term=query,
        retmax=retmax
    )
    results = Entrez.read(handle)
    handle.close()
    return results

# Find TP53 expression across studies
tp53_results = search_geo_profiles("TP53", organism="Homo sapiens")
print(f"Found {tp53_results['Count']} expression profiles for TP53")

Advanced Search Patterns:

python
# Combine multiple search terms
def advanced_geo_search(terms, operator="AND"):
    """Build complex search queries"""
    query = f" {operator} ".join(terms)
    return search_geo_datasets(query)

# Find recent high-throughput studies
search_terms = [
    "RNA-seq[DataSet Type]",
    "Homo sapiens[Organism]",
    "2024[Publication Date]"
]
results = advanced_geo_search(search_terms)

# Search by author and condition
search_terms = [
    "Smith[Author]",
    "diabetes[Disease]"
]
results = advanced_geo_search(search_terms)

GEOparse is the primary Python library for accessing GEO data:

Installation:

bash
uv pip install GEOparse

Basic Usage:

python
import GEOparse

# Download and parse a GEO Series
gse = GEOparse.get_GEO(geo="GSE123456", destdir="./data")

# Access series metadata
print(gse.metadata['title'])
print(gse.metadata['summary'])
print(gse.metadata['overall_design'])

# Access sample information
for gsm_name, gsm in gse.gsms.items():
    print(f"Sample: {gsm_name}")
    print(f"  Title: {gsm.metadata['title'][0]}")
    print(f"  Source: {gsm.metadata['source_name_ch1'][0]}")
    print(f"  Characteristics: {gsm.metadata.get('characteristics_ch1', [])}")

# Access platform information
for gpl_name, gpl in gse.gpls.items():
    print(f"Platform: {gpl_name}")
    print(f"  Title: {gpl.metadata['title'][0]}")
    print(f"  Organism: {gpl.metadata['organism'][0]}")

Working with Expression Data:

python
import GEOparse
import pandas as pd

# Get expression data from series
gse = GEOparse.get_GEO(geo="GSE123456", destdir="./data")

# Extract expression matrix
# Method 1: From series matrix file (fastest)
if hasattr(gse, 'pivot_samples'):
    expression_df = gse.pivot_samples('VALUE')
    print(expression_df.shape)  # genes x samples

# Method 2: From individual samples
expression_data = {}
for gsm_name, gsm in gse.gsms.items():
    if hasattr(gsm, 'table'):
        expression_data[gsm_name] = gsm.table['VALUE']

expression_df = pd.DataFrame(expression_data)
print(f"Expression matrix: {expression_df.shape}")

Accessing Supplementary Files:

python
import GEOparse

gse = GEOparse.get_GEO(geo="GSE123456", destdir="./data")

# Download supplementary files
gse.download_supplementary_files(
    directory="./data/GSE123456_suppl",
    download_sra=False  # Set to True to download SRA files
)

# List available supplementary files
for gsm_name, gsm in gse.gsms.items():
    if hasattr(gsm, 'supplementary_files'):
        print(f"Sample {gsm_name}:")
        for file_url in gsm.metadata.get('supplementary_file', []):
            print(f"  {file_url}")

Filtering and Subsetting Data:

python
import GEOparse

gse = GEOparse.get_GEO(geo="GSE123456", destdir="./data")

# Filter samples by metadata
control_samples = [
    gsm_name for gsm_name, gsm in gse.gsms.items()
    if 'control' in gsm.metadata.get('title', [''])[0].lower()
]

treatment_samples = [
    gsm_name for gsm_name, gsm in gse.gsms.items()
    if 'treatment' in gsm.metadata.get('title', [''])[0].lower()
]

print(f"Control samples: {len(control_samples)}")
print(f"Treatment samples: {len(treatment_samples)}")

# Extract subset expression matrix
expression_df = gse.pivot_samples('VALUE')
control_expr = expression_df[control_samples]
treatment_expr = expression_df[treatment_samples]
4. Using NCBI E-utilities for GEO Access

E-utilities provide lower-level programmatic access to GEO metadata:

Basic E-utilities Workflow:

python
from Bio import Entrez
import time

Entrez.email = "your.email@example.com"

# Step 1: Search for GEO entries
def search_geo(query, db="gds", retmax=100):
    """Search GEO using E-utilities"""
    handle = Entrez.esearch(
        db=db,
        term=query,
        retmax=retmax,
        usehistory="y"
    )
    results = Entrez.read(handle)
    handle.close()
    return results

# Step 2: Fetch summaries
def fetch_geo_summaries(id_list, db="gds"):
    """Fetch document summaries for GEO entries"""
    ids = ",".join(id_list)
    handle = Entrez.esummary(db=db, id=ids)
    summaries = Entrez.read(handle)
    handle.close()
    return summaries

# Step 3: Fetch full records
def fetch_geo_records(id_list, db="gds"):
    """Fetch full GEO records"""
    ids = ",".join(id_list)
    handle = Entrez.efetch(db=db, id=ids, retmode="xml")
    records = Entrez.read(handle)
    handle.close()
    return records

# Example workflow
search_results = search_geo("breast cancer AND Homo sapiens")
id_list = search_results['IdList'][:5]

summaries = fetch_geo_summaries(id_list)
for summary in summaries:
    print(f"GDS: {summary.get('Accession', 'N/A')}")
    print(f"Title: {summary.get('title', 'N/A')}")
    print(f"Samples: {summary.get('n_samples', 'N/A')}")
    print()

Batch Processing with E-utilities:

python
from Bio import Entrez
import time

Entrez.email = "your.email@example.com"

def batch_fetch_geo_metadata(accessions, batch_size=100):
    """Fetch metadata for multiple GEO accessions"""
    results = {}

    for i in range(0, len(accessions), batch_size):
        batch = accessions[i:i + batch_size]

        # Search for each accession
        for accession in batch:
            try:
                query = f"{accession}[Accession]"
                search_handle = Entrez.esearch(db="gds", term=query)
                search_results = Entrez.read(search_handle)
                search_handle.close()

                if search_results['IdList']:
                    # Fetch summary
                    summary_handle = Entrez.esummary(
                        db="gds",
                        id=search_results['IdList'][0]
                    )
                    summary = Entrez.read(summary_handle)
                    summary_handle.close()
                    results[accession] = summary[0]

                # Be polite to NCBI servers
                time.sleep(0.34)  # Max 3 requests per second

            except Exception as e:
                print(f"Error fetching {accession}: {e}")

    return results

# Fetch metadata for multiple datasets
gse_list = ["GSE100001", "GSE100002", "GSE100003"]
metadata = batch_fetch_geo_metadata(gse_list)
5. Direct FTP Access for Data Files

FTP URLs for GEO Data:

GEO data can be downloaded directly via FTP:

python
import ftplib
import os

def download_geo_ftp(accession, file_type="matrix", dest_dir="./data"):
    """Download GEO files via FTP"""
    # Construct FTP path based on accession type
    if accession.startswith("GSE"):
        # Series files
        gse_num = accession[3:]
        base_num = gse_num[:-3] + "nnn"
        ftp_path = f"/geo/series/GSE{base_num}/{accession}/"

        if file_type == "matrix":
            filename = f"{accession}_series_matrix.txt.gz"
        elif file_type == "soft":
            filename = f"{accession}_family.soft.gz"
        elif file_type == "miniml":
            filename = f"{accession}_family.xml.tgz"

    # Connect to FTP server
    ftp = ftplib.FTP("ftp.ncbi.nlm.nih.gov")
    ftp.login()
    ftp.cwd(ftp_path)

    # Download file
    os.makedirs(dest_dir, exist_ok=True)
    local_file = os.path.join(dest_dir, filename)

    with open(local_file, 'wb') as f:
        ftp.retrbinary(f'RETR {filename}', f.write)

    ftp.quit()
    print(f"Downloaded: {local_file}")
    return local_file

# Download series matrix file
download_geo_ftp("GSE123456", file_type="matrix")

# Download SOFT format file
download_geo_ftp("GSE123456", file_type="soft")

Using wget or curl for Downloads:

bash
# Download series matrix file
wget ftp://ftp.ncbi.nlm.nih.gov/geo/series/GSE123nnn/GSE123456/matrix/GSE123456_series_matrix.txt.gz

# Download all supplementary files for a series
wget -r -np -nd ftp://ftp.ncbi.nlm.nih.gov/geo/series/GSE123nnn/GSE123456/suppl/

# Download SOFT format family file
wget ftp://ftp.ncbi.nlm.nih.gov/geo/series/GSE123nnn/GSE123456/soft/GSE123456_family.soft.gz
6. Analyzing GEO Data

Quality Control and Preprocessing:

python
import GEOparse
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# Load dataset
gse = GEOparse.get_GEO(geo="GSE123456", destdir="./data")
expression_df = gse.pivot_samples('VALUE')

# Check for missing values
print(f"Missing values: {expression_df.isnull().sum().sum()}")

# Log transformation (if needed)
if expression_df.min().min() > 0:  # Check if already log-transformed
    if expression_df.max().max() > 100:
        expression_df = np.log2(expression_df + 1)
        print("Applied log2 transformation")

# Distribution plots
plt.figure(figsize=(12, 5))

plt.subplot(1, 2, 1)
expression_df.plot.box(ax=plt.gca())
plt.title("Expression Distribution per Sample")
plt.xticks(rotation=90)

plt.subplot(1, 2, 2)
expression_df.mean(axis=1).hist(bins=50)
plt.title("Gene Expression Distribution")
plt.xlabel("Average Expression")

plt.tight_layout()
plt.savefig("geo_qc.png", dpi=300, bbox_inches='tight')

Differential Expression Analysis:

python
import GEOparse
import pandas as pd
import numpy as np
from scipy import stats

gse = GEOparse.get_GEO(geo="GSE123456", destdir="./data")
expression_df = gse.pivot_samples('VALUE')

# Define sample groups
control_samples = ["GSM1", "GSM2", "GSM3"]
treatment_samples = ["GSM4", "GSM5", "GSM6"]

# Calculate fold changes and p-values
results = []
for gene in expression_df.index:
    control_expr = expression_df.loc[gene, control_samples]
    treatment_expr = expression_df.loc[gene, treatment_samples]

    # Calculate statistics
    fold_change = treatment_expr.mean() - control_expr.mean()
    t_stat, p_value = stats.ttest_ind(treatment_expr, control_expr)

    results.append({
        'gene': gene,
        'log2_fold_change': fold_change,
        'p_value': p_value,
        'control_mean': control_expr.mean(),
        'treatment_mean': treatment_expr.mean()
    })

# Create results DataFrame
de_results = pd.DataFrame(results)

# Multiple testing correction (Benjamini-Hochberg)
from statsmodels.stats.multitest import multipletests
_, de_results['q_value'], _, _ = multipletests(
    de_results['p_value'],
    method='fdr_bh'
)

# Filter significant genes
significant_genes = de_results[
    (de_results['q_value'] < 0.05) &
    (abs(de_results['log2_fold_change']) > 1)
]

print(f"Significant genes: {len(significant_genes)}")
significant_genes.to_csv("de_results.csv", index=False)

Correlation and Clustering Analysis:

python
import GEOparse
import seaborn as sns
import matplotlib.pyplot as plt
from scipy.cluster import hierarchy
from scipy.spatial.distance import pdist

gse = GEOparse.get_GEO(geo="GSE123456", destdir="./data")
expression_df = gse.pivot_samples('VALUE')

# Sample correlation heatmap
sample_corr = expression_df.corr()

plt.figure(figsize=(10, 8))
sns.heatmap(sample_corr, cmap='coolwarm', center=0,
            square=True, linewidths=0.5)
plt.title("Sample Correlation Matrix")
plt.tight_layout()
plt.savefig("sample_correlation.png", dpi=300, bbox_inches='tight')

# Hierarchical clustering
distances = pdist(expression_df.T, metric='correlation')
linkage = hierarchy.linkage(distances, method='average')

plt.figure(figsize=(12, 6))
hierarchy.dendrogram(linkage, labels=expression_df.columns)
plt.title("Hierarchical Clustering of Samples")
plt.xlabel("Samples")
plt.ylabel("Distance")
plt.xticks(rotation=90)
plt.tight_layout()
plt.savefig("sample_clustering.png", dpi=300, bbox_inches='tight')
7. Batch Processing Multiple Datasets

Download and Process Multiple Series:

python
import GEOparse
import pandas as pd
import os

def batch_download_geo(gse_list, destdir="./geo_data"):
    """Download multiple GEO series"""
    results = {}

    for gse_id in gse_list:
        try:
            print(f"Processing {gse_id}...")
            gse = GEOparse.get_GEO(geo=gse_id, destdir=destdir)

            # Extract key information
            results[gse_id] = {
                'title': gse.metadata.get('title', ['N/A'])[0],
                'organism': gse.metadata.get('organism', ['N/A'])[0],
                'platform': list(gse.gpls.keys())[0] if gse.gpls else 'N/A',
                'num_samples': len(gse.gsms),
                'submission_date': gse.metadata.get('submission_date', ['N/A'])[0]
            }

            # Save expression data
            if hasattr(gse, 'pivot_samples'):
                expr_df = gse.pivot_samples('VALUE')
                expr_df.to_csv(f"{destdir}/{gse_id}_expression.csv")
                results[gse_id]['num_genes'] = len(expr_df)

        except Exception as e:
            print(f"Error processing {gse_id}: {e}")
            results[gse_id] = {'error': str(e)}

    # Save summary
    summary_df = pd.DataFrame(results).T
    summary_df.to_csv(f"{destdir}/batch_summary.csv")

    return results

# Process multiple datasets
gse_list = ["GSE100001", "GSE100002", "GSE100003"]
results = batch_download_geo(gse_list)

Meta-Analysis Across Studies:

python
import GEOparse
import pandas as pd
import numpy as np

def meta_analysis_geo(gse_list, gene_of_interest):
    """Perform meta-analysis of gene expression across studies"""
    results = []

    for gse_id in gse_list:
        try:
            gse = GEOparse.get_GEO(geo=gse_id, destdir="./data")

            # Get platform annotation
            gpl = list(gse.gpls.values())[0]

            # Find gene in platform
            if hasattr(gpl, 'table'):
                gene_probes = gpl.table[
                    gpl.table['Gene Symbol'].str.contains(
                        gene_of_interest,
                        case=False,
                        na=False
                    )
                ]

                if not gene_probes.empty:
                    expr_df = gse.pivot_samples('VALUE')

                    for probe_id in gene_probes['ID']:
                        if probe_id in expr_df.index:
                            expr_values = expr_df.loc[probe_id]

                            results.append({
                                'study': gse_id,
                                'probe': probe_id,
                                'mean_expression': expr_values.mean(),
                                'std_expression': expr_values.std(),
                                'num_samples': len(expr_values)
                            })

        except Exception as e:
            print(f"Error in {gse_id}: {e}")

    return pd.DataFrame(results)

# Meta-analysis for TP53
gse_studies = ["GSE100001", "GSE100002", "GSE100003"]
meta_results = meta_analysis_geo(gse_studies, "TP53")
print(meta_results)

Installation and Setup

Python Libraries
bash
# Primary GEO access library (recommended)
uv pip install GEOparse

# For E-utilities and programmatic NCBI access
uv pip install biopython

# For data analysis
uv pip install pandas numpy scipy

# For visualization
uv pip install matplotlib seaborn

# For statistical analysis
uv pip install statsmodels scikit-learn
Configuration

Set up NCBI E-utilities access:

python
from Bio import Entrez

# Always set your email (required by NCBI)
Entrez.email = "your.email@example.com"

# Optional: Set API key for increased rate limits
# Get your API key from: https://www.ncbi.nlm.nih.gov/account/
Entrez.api_key = "your_api_key_here"

# With API key: 10 requests/second
# Without API key: 3 requests/second

Common Use Cases

Transcriptomics Research
  • Download gene expression data for specific conditions
  • Compare expression profiles across studies
  • Identify differentially expressed genes
  • Perform meta-analyses across multiple datasets
Drug Response Studies
  • Analyze gene expression changes after drug treatment
  • Identify biomarkers for drug response
  • Compare drug effects across cell lines or patients
  • Build predictive models for drug sensitivity
Disease Biology
  • Study gene expression in disease vs. normal tissues
  • Identify disease-associated expression signatures
  • Compare patient subgroups and disease stages
  • Correlate expression with clinical outcomes
Biomarker Discovery
  • Screen for diagnostic or prognostic markers
  • Validate biomarkers across independent cohorts
  • Compare marker performance across platforms
  • Integrate expression with clinical data

Key Concepts

SOFT (Simple Omnibus Format in Text): GEO's primary text-based format containing metadata and data tables. Easily parsed by GEOparse.

MINiML (MIAME Notation in Markup Language): XML format for GEO data, used for programmatic access and data exchange.

Series Matrix: Tab-delimited expression matrix with samples as columns and genes/probes as rows. Fastest format for getting expression data.

MIAME Compliance: Minimum Information About a Microarray Experiment - standardized annotation that GEO enforces for all submissions.

Expression Value Types: Different types of expression measurements (raw signal, normalized, log-transformed). Always check platform and processing methods.

Platform Annotation: Maps probe/feature IDs to genes. Essential for biological interpretation of expression data.

Show full SKILL.md (408 more words)Show less

GEO2R Web Tool

For quick analysis without coding, use GEO2R:

  • Web-based statistical analysis tool integrated into GEO
  • Accessible at: https://www.ncbi.nlm.nih.gov/geo/geo2r/?acc=GSExxxxx
  • Performs differential expression analysis
  • Generates R scripts for reproducibility
  • Useful for exploratory analysis before downloading data

Rate Limiting and Best Practices

NCBI E-utilities Rate Limits:

  • Without API key: 3 requests per second
  • With API key: 10 requests per second
  • Implement delays between requests: time.sleep(0.34) (no API key) or time.sleep(0.1) (with API key)

FTP Access:

  • No rate limits for FTP downloads
  • Preferred method for bulk downloads
  • Can download entire directories with wget -r

GEOparse Caching:

  • GEOparse automatically caches downloaded files in destdir
  • Subsequent calls use cached data
  • Clean cache periodically to save disk space

Optimal Practices:

  • Use GEOparse for series-level access (easiest)
  • Use E-utilities for metadata searching and batch queries
  • Use FTP for direct file downloads and bulk operations
  • Cache data locally to avoid repeated downloads
  • Always set Entrez.email when using Biopython

Resources

references/geo_reference.md

Comprehensive reference documentation covering:

  • Detailed E-utilities API specifications and endpoints
  • Complete SOFT and MINiML file format documentation
  • Advanced GEOparse usage patterns and examples
  • FTP directory structure and file naming conventions
  • Data processing pipelines and normalization methods
  • Troubleshooting common issues and error handling
  • Platform-specific considerations and quirks

Consult this reference for in-depth technical details, complex query patterns, or when working with uncommon data formats.

Important Notes

Data Quality Considerations
  • GEO accepts user-submitted data with varying quality standards
  • Always check platform annotation and processing methods
  • Verify sample metadata and experimental design
  • Be cautious with batch effects across studies
  • Consider reprocessing raw data for consistency
File Size Warnings
  • Series matrix files can be large (>1 GB for large studies)
  • Supplementary files (e.g., CEL files) can be very large
  • Plan for adequate disk space before downloading
  • Consider downloading samples incrementally
Data Usage and Citation
  • GEO data is freely available for research use
  • Always cite original studies when using GEO data
  • Cite GEO database: Barrett et al. (2013) Nucleic Acids Research
  • Check individual dataset usage restrictions (if any)
  • Follow NCBI guidelines for programmatic access
Common Pitfalls
  • Different platforms use different probe IDs (requires annotation mapping)
  • Expression values may be raw, normalized, or log-transformed (check metadata)
  • Sample metadata can be inconsistently formatted across studies
  • Not all series have series matrix files (older submissions)
  • Platform annotations may be outdated (genes renamed, IDs deprecated)

Additional Resources

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in cli-tool/components/skills/scientific/geo-database of davila7/claude-code-templates.

  • SKILL.md
  • references/geo_reference.md

Open the folder on GitHubat commit 46b4d8b

Used in 11 other repositories

We found 17 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 11 other GitHub owners. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Geo Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Geo Database compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Geo Database this skilldavila7/claude-code-templates32k11 repos~6.1kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
Bio Write SequencesGPTomics/bioSkills1.2k3 repos~2.1kAutomated safety check: PassMIT
Ncbi DatasetsClawBio/ClawBio1.2k1 repos~2.8kAutomated safety check: PassMIT
EtetoolkitK-Dense-AI/scientific-agent-skills48k1 repos~3.3kAutomated safety check: NotesGPL-3.0-or-later

Similar skills

  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Research & ScienceAuto-check passed
  • Etetoolkit

    K-Dense-AI/scientific-agent-skills

    Analyzes, manipulates, compares, annotates, and visualizes phylogenetic or other hierarchical trees with ETE 4.

    48k GitHub starsUsed in 1 repo~3.3k tokens
    Research & ScienceAuto-check: notes
  • Bulkrna Geneid Mapping

    TianGzlab/OmicsClaw

    Load when converting Ensembl, Entrez or symbol IDs in a bulk RNA count matrix using an explicit mapping or a small human demo reference.

    161 GitHub stars~1.2k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed

More from davila7/claude-code-templates

All 478 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 11 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 9 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Works with

Questions about Geo Database

What does Geo Database do?

Access NCBI GEO for gene expression/genomics data. An agent skill from davila7/claude-code-templates. Geo Database is an agent skill from davila7/claude-code-templates. Access NCBI GEO for gene expression/genomics data.

When should I use Geo Database?

Geo Database fits situations like: tasks that involve Bioinformatics.

How do I install Geo Database in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill geo-database -a claude-code`. Or copy the skill folder (cli-tool/components/skills/scientific/geo-database in davila7/claude-code-templates) into .claude/skills/geo-database in your project. Claude Code loads it when a task matches its description.

How do I install Geo Database in Codex?

Run `npx skills add davila7/claude-code-templates --skill geo-database -a codex`. Or copy the skill folder (cli-tool/components/skills/scientific/geo-database in davila7/claude-code-templates) into .agents/skills/geo-database in your project. Codex loads it when a task matches its description.

Can I use Geo Database in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill geo-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/geo-database, .gemini/skills/geo-database, .github/skills/geo-database and .opencode/skills/geo-database in your project.

What does Geo Database need to run?

Going by SKILL.md and its folder, Geo Database needs the command-line tools its instructions call (uv and wget). Our summary lists: Python 3.

Does Geo Database access the network?

SKILL.md names 4 domains. In commands or code: ncbi.nlm.nih.gov; the agent is likely to contact it when it follows the instructions. As links in the text: geoparse.readthedocs.io, ncbiinsights.ncbi.nlm.nih.gov and biopython.org. This is read from the text; nothing was executed.

Is Geo Database safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Geo Database use?

Geo Database is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Geo Database use?

About 6.1k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.

What are the alternatives to Geo Database?

Skills that share tags, products or a category with Geo Database: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Bio Write Sequences (GPTomics/bioSkills, 1.2k stars) and Ncbi Datasets (ClawBio/ClawBio, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Geo Database?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,483 GitHub stars. The repository holds 478 skills in this directory. The repository was last updated on October 9, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.