Agent skill

Expression Data Retrieval

by lamm-mit in lamm-mit/scienceclaw

ToolUniverse workflow — Expression Data Retrieval. An agent skill from lamm-mit/scienceclaw.

Apache-2.0Auto-check passedResearch & Science

Install Expression Data Retrieval

skills CLI
$ npx skills add lamm-mit/scienceclaw --skill expression-data-retrieval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lamm-mit/scienceclaw expression-data-retrieval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lamm-mit/scienceclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/expression-data-retrieval .claude/skills/expression-data-retrieval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
expression-data-retrieval
GitHub stars
244
Token cost
~2.6k tokens
SKILL.md length
607 words
Files
3 (incl. scripts)
Skills in repo
85
Repo updated
First seen
Licence
Apache-2.0

At a glance

ToolUniverse workflow — Expression Data Retrieval. An agent skill from lamm-mit/scienceclaw.

  • Works in 4 steps: Clarification (When Needed) → Query Disambiguation → Data Retrieval (Internal) → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Workflow Overview, Phase 0: Clarification (When…, Phase 1: Query Disambiguation and Phase 2: Data Retrieval…, plus 7 more sections
  • Runs Python scripts from its folder; reaches ebi.ac.uk

What it does

Expression Data Retrieval is an agent skill from lamm-mit/scienceclaw. ToolUniverse workflow — Expression Data Retrieval

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/run.py`).

It sits in Research & Science, covering Bioinformatics. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/expression-data-retrieval”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Clarification (When Needed)
  2. Query Disambiguation
  3. Data Retrieval (Internal)
  4. Report Dataset Profile

What it can do on your machine

Read from SKILL.md and the folder at commit ab9aba1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ebi.ac.uk

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Expression Data Retrieval loads about 2.6k tokens when it runs. Until then it costs about 19 tokens; SKILL.md has 607 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from lamm-mit/scienceclaw at commit ab9aba1, republished under its Apache-2.0 licence (© lamm-mit). 607 words, ~2,612 tokens.

Download SKILL.mdSave it as .claude/skills/expression-data-retrieval/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
expression-data-retrieval
description
ToolUniverse workflow — Expression Data Retrieval
source
https://github.com/mims-harvard/ToolUniverse/tree/main/skills/tooluniverse-expression-data-retrieval

name: tooluniverse-expression-data-retrieval description: Retrieves gene expression and omics datasets from ArrayExpress and BioStudies with gene disambiguation, experiment quality assessment, and structured reports. Creates comprehensive dataset profiles with metadata, sample information, and download links. Use when users need expression data, omics datasets, or mention ArrayExpress (E-MTAB, E-GEOD) or BioStudies (S-BSST) accessions.

Gene Expression & Omics Data Retrieval

Retrieve gene expression experiments and multi-omics datasets with proper disambiguation and quality assessment.

IMPORTANT: Always use English terms in tool calls (gene names, tissue names, condition descriptions), even if the user writes in another language. Only try original-language terms as a fallback if English returns no results. Respond in the user's language.

Workflow Overview

Phase 0: Clarify Query (if ambiguous)
    ↓
Phase 1: Disambiguate Gene/Condition
    ↓
Phase 2: Search & Retrieve (Internal)
    ↓
Phase 3: Report Dataset Profile

Phase 0: Clarification (When Needed)

Ask the user ONLY if:

  • Gene name is ambiguous (e.g., "p53" → TP53 or MDM2 studies?)
  • Tissue/condition unclear for comparative studies
  • Organism not specified for non-human research

Skip clarification for:

  • Specific accession numbers (E-MTAB-, E-GEOD-, S-BSST*)
  • Clear disease/tissue + organism combinations
  • Explicit platform requests (RNA-seq, microarray)

Phase 1: Query Disambiguation

1.1 Gene Name Resolution

If searching by gene, first resolve official identifiers:

python
from tooluniverse import ToolUniverse
tu = ToolUniverse()
tu.load_tools()

# For gene-focused searches, resolve official symbol first
# This helps construct better search queries
# Example: "p53" → "TP53" (official HGNC symbol)

Gene Disambiguation Checklist:

  • Official gene symbol identified (HGNC for human, MGI for mouse)
  • Common aliases noted for search expansion
  • Species confirmed
1.2 Construct Search Strategy
User Query TypeSearch Strategy
Specific accessionDirect retrieval
Gene + condition"[gene] [condition]" + species filter
Disease only"[disease]" + species filter
Technology-specificAdd platform keywords (RNA-seq, microarray)

Phase 2: Data Retrieval (Internal)

Search silently. Do NOT narrate the process.

2.1 Search Experiments
python
# ArrayExpress search
result = tu.tools.arrayexpress_search_experiments(
    keywords="[gene/disease] [condition]",
    species="[species]",
    limit=20
)

# BioStudies for multi-omics
biostudies_result = tu.tools.biostudies_search_studies(
    query="[keywords]",
    limit=10
)
2.2 Get Experiment Details

For top results, retrieve full metadata:

python
# Get details for each relevant experiment
details = tu.tools.arrayexpress_get_experiment_details(
    accession=accession
)

# Get sample information
samples = tu.tools.arrayexpress_get_experiment_samples(
    accession=accession
)

# Get available files
files = tu.tools.arrayexpress_get_experiment_files(
    accession=accession
)
2.3 BioStudies Retrieval
python
# Multi-omics study details
study_details = tu.tools.biostudies_get_study_details(
    accession=study_accession
)

# Study structure
sections = tu.tools.biostudies_get_study_sections(
    accession=study_accession
)

# Available files
files = tu.tools.biostudies_get_study_files(
    accession=study_accession
)
Fallback Chains
PrimaryFallbackNotes
ArrayExpress searchBioStudies searchArrayExpress empty
arrayexpress_get_experiment_detailsbiostudies_get_study_detailsE-GEOD may have BioStudies mirror
arrayexpress_get_experiment_filesNote "Files unavailable"Some studies restrict downloads

Phase 3: Report Dataset Profile

Output Structure

Present as a Dataset Search Report. Hide search process.

markdown
# Expression Data: [Query Topic]

**Search Summary**
- Query: [gene/disease] in [species]
- Databases: ArrayExpress, BioStudies
- Results: [N] relevant experiments found

**Data Quality Overview**: [assessment based on criteria below]

---

## Top Experiments

### 1. [E-MTAB-XXXX]: [Title]

| Attribute | Value |
|-----------|-------|
| **Accession** | [accession with link] |
| **Organism** | [species] |
| **Experiment Type** | RNA-seq / Microarray |
| **Platform** | [specific platform] |
| **Samples** | [N] samples |
| **Release Date** | [date] |

**Description**: [Brief description from metadata]

**Experimental Design**:
- Conditions: [treatment vs control, etc.]
- Replicates: [N biological, M technical]
- Tissue/Cell type: [if specified]

**Sample Groups**:
| Group | Samples | Description |
|-------|---------|-------------|
| Control | [N] | [description] |
| Treatment | [N] | [description] |

**Data Files Available**:
| File | Type | Size |
|------|------|------|
| [filename] | Processed data | [size] |
| [filename] | Raw data | [size] |
| [filename] | Sample metadata | [size] |

**Quality Assessment**: ●●● High / ●●○ Medium / ●○○ Low
- Sample size: [adequate/limited]
- Replication: [yes/no]
- Metadata completeness: [complete/partial]

---

### 2. [E-GEOD-XXXXX]: [Title]
[Same structure as above]

---

## Multi-Omics Studies (from BioStudies)

### [S-BSST-XXXXX]: [Title]

| Attribute | Value |
|-----------|-------|
| **Accession** | [accession] |
| **Study Type** | [proteomics/metabolomics/integrated] |
| **Organism** | [species] |
| **Samples** | [N] |

**Data Types Included**:
- [ ] Transcriptomics
- [ ] Proteomics
- [ ] Metabolomics
- [ ] Other: [specify]

---

## Summary Table

| Accession | Type | Samples | Platform | Quality |
|-----------|------|---------|----------|---------|
| [E-MTAB-X] | RNA-seq | [N] | Illumina | ●●● |
| [E-GEOD-X] | Microarray | [N] | Affymetrix | ●●○ |

---

## Recommendations

**For [specific analysis type]**:
- Best experiment: [accession] - [reason]
- Alternative: [accession] - [reason]

**Data Integration Notes**:
- Platform compatibility: [notes on combining datasets]
- Batch considerations: [if applicable]

---

## Data Access

### Direct Download Links
- [E-MTAB-XXXX processed data](link)
- [E-MTAB-XXXX raw data](link)

### Database Links
- ArrayExpress: https://www.ebi.ac.uk/arrayexpress/experiments/[accession]
- BioStudies: https://www.ebi.ac.uk/biostudies/studies/[accession]

Retrieved: [date]

Data Quality Tiers

Assessment criteria for expression experiments:

TierSymbolCriteria
High Quality●●●≥3 bio replicates, complete metadata, processed data available
Medium Quality●●○2-3 replicates OR some metadata gaps, data accessible
Low Quality●○○No replicates, sparse metadata, or data access issues
Use with Caution○○○Single sample, no replication, outdated platform

Include assessment rationale:

markdown
**Quality**: ●●● High
- ✓ 4 biological replicates per condition
- ✓ Complete sample annotations
- ✓ Processed and raw data available
- ✓ Recent RNA-seq platform

Completeness Checklist

Every dataset report MUST include:

Per Experiment (Required)
  • Accession number with database link
  • Organism
  • Experiment type (RNA-seq/microarray/etc.)
  • Sample count
  • Brief description
  • Quality assessment
Show full SKILL.md (234 more words)Show less
Search Summary (Required)
  • Query parameters stated
  • Number of results
  • Databases searched
Recommendations (Required)
  • Best dataset for user's purpose (or "No suitable data found")
  • Data access notes
Include Even If Empty
  • Multi-omics studies section (or "No multi-omics studies found")
  • Data integration notes (or "Single-platform data, no integration needed")

Common Use Cases

Disease Gene Expression

User: "Find breast cancer RNA-seq data"

python
result = tu.tools.arrayexpress_search_experiments(
    keywords="breast cancer RNA-seq",
    species="Homo sapiens",
    limit=20
)

→ Report top experiments with quality assessment

Gene-Specific Studies

User: "Find TP53 expression experiments in mouse"

python
result = tu.tools.arrayexpress_search_experiments(
    keywords="TP53 p53",  # Include aliases
    species="Mus musculus",
    limit=15
)

→ Report experiments studying this gene

Specific Accession Lookup

User: "Get details for E-MTAB-5214" → Single experiment profile with all details and files

Multi-Omics Integration

User: "Find proteomics and transcriptomics studies for liver disease" → Search both ArrayExpress and BioStudies, note integration potential


Error Handling

ErrorResponse
"No experiments found"Broaden keywords, remove species filter, try synonyms
"Accession not found"Verify format (E-MTAB-, E-GEOD-, S-BSST*), check if withdrawn
"Files not available"Note in report: "Data files restricted by submitter"
"API timeout"Retry once, then note: "(metadata retrieval incomplete)"

Tool Reference

ArrayExpress (Gene Expression)

ToolPurpose
arrayexpress_search_experimentsKeyword/species search
arrayexpress_get_experiment_detailsFull metadata
arrayexpress_get_experiment_filesDownload links
arrayexpress_get_experiment_samplesSample annotations

BioStudies (Multi-Omics)

ToolPurpose
biostudies_search_studiesMulti-omics search
biostudies_get_study_detailsStudy metadata
biostudies_get_study_filesData files
biostudies_get_study_sectionsStudy structure

Search Parameters Reference

ArrayExpress

ParameterDescriptionExample
keywordsFree text search"breast cancer RNA-seq"
speciesScientific name"Homo sapiens"
arrayPlatform filter"Illumina"
limitMax results20

BioStudies

ParameterDescriptionExample
queryFree text"proteomics liver"
limitMax results10

© lamm-mit, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/expression-data-retrieval of lamm-mit/scienceclaw.

  • SKILL.md
  • scripts/__pycache__/run.cpython-313.pyc
  • scripts/run.py

Open the folder on GitHubat commit ab9aba1

Compare with similar skills

Expression Data Retrieval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Expression Data Retrieval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Expression Data Retrieval this skilllamm-mit/scienceclaw244—~2.6kAutomated safety check: PassApache-2.0
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from lamm-mit/scienceclaw

All 85 skills in this repo
  • Fred Economic Data

    lamm-mit/scienceclaw

    Query FRED (Federal Reserve Economic Data) API for 800,000+ economic time series from 100+ sources.

    244 GitHub starsUsed in 4 repos~3k tokens
    Auto-check passed
  • Drug Research

    lamm-mit/scienceclaw

    Generates comprehensive drug research reports with compound disambiguation, evidence grading, and mandatory completeness sections.

    244 GitHub starsUsed in 3 repos~1.7k tokens
    Auto-check passed
  • Imaging Data Commons

    lamm-mit/scienceclaw

    Query and download public cancer imaging data from NCI Imaging Data Commons using idc-index.

    244 GitHub starsUsed in 5 repos~11k tokens
    Auto-check passed
  • Rowan

    lamm-mit/scienceclaw

    Cloud-based quantum chemistry platform with Python API. An agent skill from lamm-mit/scienceclaw.

    244 GitHub starsUsed in 4 repos~3.1k tokens
    Auto-check: warnings
  • Infographics

    lamm-mit/scienceclaw

    Create professional infographics using Nano Banana Pro AI with smart iterative refinement.

    244 GitHub starsUsed in 6 repos~4.4k tokens
    Auto-check: notes
  • Disease Research

    lamm-mit/scienceclaw

    Generate comprehensive disease research reports using 100+ ToolUniverse tools.

    244 GitHub stars~946 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Expression Data Retrieval

What does Expression Data Retrieval do?

ToolUniverse workflow — Expression Data Retrieval. An agent skill from lamm-mit/scienceclaw. Expression Data Retrieval is an agent skill from lamm-mit/scienceclaw.

When should I use Expression Data Retrieval?

Expression Data Retrieval fits situations like: tasks that involve Bioinformatics.

How do I install Expression Data Retrieval in Claude Code?

Run `npx skills add lamm-mit/scienceclaw --skill expression-data-retrieval -a claude-code`. Or copy the skill folder (skills/expression-data-retrieval in lamm-mit/scienceclaw) into .claude/skills/expression-data-retrieval in your project. Claude Code loads it when a task matches its description.

How do I install Expression Data Retrieval in Codex?

Run `npx skills add lamm-mit/scienceclaw --skill expression-data-retrieval -a codex`. Or copy the skill folder (skills/expression-data-retrieval in lamm-mit/scienceclaw) into .agents/skills/expression-data-retrieval in your project. Codex loads it when a task matches its description.

Can I use Expression Data Retrieval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lamm-mit/scienceclaw --skill expression-data-retrieval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/expression-data-retrieval, .gemini/skills/expression-data-retrieval, .github/skills/expression-data-retrieval and .opencode/skills/expression-data-retrieval in your project.

What does Expression Data Retrieval need to run?

Going by SKILL.md and its folder, Expression Data Retrieval needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Expression Data Retrieval access the network?

SKILL.md names 1 domain. In commands or code: ebi.ac.uk; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Expression Data Retrieval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Expression Data Retrieval use?

Expression Data Retrieval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Expression Data Retrieval use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Expression Data Retrieval?

Skills that share tags, products or a category with Expression Data Retrieval: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Expression Data Retrieval?

lamm-mit (a GitHub user) maintains it in lamm-mit/scienceclaw, which has 244 GitHub stars. The repository holds 85 skills in this directory. The repository was last updated on August 21, 2026.

Source: lamm-mit/scienceclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.