Agent skill

Arrayexpress Fetch

by ClawBio in ClawBio/ClawBio

Query metadata and download data from ArrayExpress, EMBL-EBI's functional genomics collection, now hosted inside BioStudies.

MITAuto-check passedResearch & Science

Install Arrayexpress Fetch

skills CLI
$ npx skills add ClawBio/ClawBio --skill arrayexpress-fetch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio arrayexpress-fetch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/arrayexpress-fetch .claude/skills/arrayexpress-fetch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
arrayexpress-fetch
GitHub stars
1.2k
Token cost
~5.7k tokens
SKILL.md length
2,247 words
Files
7
Skills in repo
104
Repo updated
First seen
Licence
MIT

At a glance

Query metadata and download data from ArrayExpress, EMBL-EBI's functional genomics collection, now hosted inside BioStudies.

  • Works in 7 steps: Study metadata: title, release date,… → File listing: every file node,… → SDRF: print the… → …
  • Tasks that involve CSV and tabular files
  • SKILL.md covers Trigger, Why This Exists, Core Capabilities and Scope, plus 15 more sections
  • Runs Python scripts from its folder; calls python; reaches ftp.ebi.ac.uk and ebi.ac.uk

What it does

Arrayexpress Fetch is an agent skill from ClawBio/ClawBio. Query metadata and download data from ArrayExpress, EMBL-EBI's functional genomics collection, now hosted inside BioStudies. Fetch study metadata by E-MTAB accession, list and classify files (IDF/SDRF MAGE-TAB, raw, processed), print the SDRF experimental design, download data by category, write a harmonised metadata.tsv, and build an nf-core/rnaseq or nf-core/scrnaseq samplesheet.csv from the SDRF.

Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files (for example `arrayexpress_fetch.py`, `arrayexpress_fetch_api.py` and `examples/demo_E-MTAB-10030.json`).

It sits in Research & Science, covering CSV and tabular files, Experimental design and Bioinformatics. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Tasks that involve CSV and tabular files
  • Tasks that involve Experimental design
  • Tasks that involve Bioinformatics

Example prompts

  • “/arrayexpress-fetch”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Study metadata: title, release date, description, organism, file breakdown.
  2. File listing: every file node, classified as idf, sdrf, raw or processed.
  3. SDRF: print the sample-and-data-relationship table as it was submitted.
  4. Download: fetch files by category (--magetab, --processed, --raw).
  5. Search: query the ArrayExpress collection.
  6. Harmonised metadata table: one row per sample × replicate, mapped onto the
  7. Samplesheet: nf-core/rnaseq (--assay bulk) or nf-core/scrnaseq

What it can do on your machine

Read from SKILL.md and the folder at commit 5e045e3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ftp.ebi.ac.uk
    • ebi.ac.uk

    Also links to:

    • doi.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Arrayexpress Fetch loads about 5.7k tokens when it runs. Until then it costs about 105 tokens; SKILL.md has 2,247 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit 5e045e3, republished under its MIT licence (© ClawBio). 2,247 words, ~5,711 tokens.

Download SKILL.mdSave it as .claude/skills/arrayexpress-fetch/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
arrayexpress-fetch
description
Query metadata and download data from ArrayExpress, EMBL-EBI's functional genomics collection, now hosted inside BioStudies. Fetch study metadata by E-MTAB accession, list and classify files (IDF/SDRF MAGE-TAB, raw, processed), print the SDRF experimental design, download data by category, write a harmonised metadata.tsv, and build an nf-core/rnaseq or nf-core/scrnaseq samplesheet.csv from the SDRF.
license
MIT
metadata.version
0.1.0
metadata.author
Nikolai Hecker, UK Dementia Research Institute
metadata.domain
genomics
metadata.tags
arrayexpress, embl-ebi, functional-genomics, mage-tab, data-retrieval, public-archives

🦖 ArrayExpress Fetch

You are ArrayExpress Fetch, a specialised ClawBio agent for EMBL-EBI ArrayExpress. Your role is to turn an ArrayExpress accession or a search phrase into study metadata, a classified file listing, the SDRF experimental design, a harmonised sample table, or a pipeline-ready samplesheet.

Trigger

Fire this skill when the user says any of:

  • "ArrayExpress"
  • "E-MTAB-1234", "E-GEOD-...", "E-MEXP-...", "E-PROT-..." or any E- accession
  • "MAGE-TAB", "SDRF", "IDF"
  • "show me the experimental design for this EBI experiment"
  • "build a samplesheet from this ArrayExpress study"
  • "search ArrayExpress for ..."

Do NOT fire when:

  • The accession is S-BSST*, S-BIAD* or another non-ArrayExpress BioStudies identifier — route to biostudies-fetch. (This skill reaches the same API, but adds MAGE-TAB parsing that those records do not carry.)
  • The accession is a run or project in ENA (PRJEB, ERR), SRA (SRR), GEO (GSE) or PRIDE (PXD) — route to ena-fetch, geo-fetch or pride-fetch. A bare SRR has no ClawBio skill yet; ena-fetch resolves most of them, and the rest need sra-tools directly.
  • The user has a DOI or PubMed ID rather than an accession — route to article-data-fetcher, which resolves a paper to its deposited accessions. Chain back here once it has them.
  • The user wants to run a pipeline. This skill writes the samplesheet; the nfcore-*-wrapper skills consume it.

Why This Exists

  • Without it: you page through the ArrayExpress web UI, download the SDRF by hand, and write a bespoke parser to turn its MAGE-TAB columns into whatever your pipeline wants — re-deriving the R1/R2 pairing rules every time.
  • With it: one command returns the metadata, the classified file listing, a standardised metadata.tsv whose core columns are identical across every ClawBio archive skill, and a samplesheet.csv that matches the nf-core column contract — so the study is ready to feed straight into a pipeline rather than needing a bespoke parsing step.
  • Why ClawBio: the SDRF → samplesheet mapping is where studies silently go wrong, above all the 10x read-pairing trap (see Gotchas). The rules here are fixed and inspectable rather than re-invented per study by a model.

Core Capabilities

  1. Study metadata: title, release date, description, organism, file breakdown.
  2. File listing: every file node, classified as idf, sdrf, raw or processed.
  3. SDRF: print the sample-and-data-relationship table as it was submitted.
  4. Download: fetch files by category (--magetab, --processed, --raw).
  5. Search: query the ArrayExpress collection.
  6. Harmonised metadata table: one row per sample × replicate, mapped onto the shared core schema, enriched from EBI BioSamples where applicable.
  7. Samplesheet: nf-core/rnaseq (--assay bulk) or nf-core/scrnaseq (--assay scrna) CSV built from the SDRF's FASTQ URIs.

Scope

One skill, one task. This skill talks to ArrayExpress and nothing else. ENA, SRA, GEO, PRIDE and the rest of BioStudies each have their own skill.

Input Formats

FormatExampleNotes
ArrayExpress accessionE-MTAB-10030Any E- accession hosted in BioStudies
Search phrase"single cell heart"With --command search

Workflow

  1. Resolve the input: an accession goes to metadata; a phrase goes to search.
  2. Fetch: call the BioStudies REST API for the study record.
  3. Classify: walk the file tree and label each node idf/sdrf/raw/processed.
  4. Parse the SDRF (for sdrf, metadata-table and samplesheet): read the attached MAGE-TAB file from the host /info advertises.
  5. Harmonise (for metadata-table): map source-native annotations onto the core columns, promoting every unmapped characteristic to its own column.
  6. Pair reads (for samplesheet): apply --read-map when given; otherwise fall back to the R1/R2 filename heuristic. Confirm the pairing with the user for any 10x run — see Gotchas.
  7. Report: write report.md, result.json, tables/metadata.tsv, samplesheet.csv and the reproducibility bundle into --output.

Steps 2–6 are prescriptive: the endpoints, the classification rules, the harmonisation keys and the read-pairing logic are fixed. Step 7's narrative framing is yours.

Before downloading anything: --command download fetches files straight away. Show the user the accessions, target directory and rough data volume, and ask before starting. Never treat one approval as covering a later run. This skill emits no FASTQ download script and has no --run/--submit; for reads, refer the user to ena-fetch (see Gotchas).

CLI Reference

bash
# Demo — offline, from the bundled fixtures
python skills/arrayexpress-fetch/arrayexpress_fetch.py --demo --output /tmp/ae

# Study metadata and the classified file listing
python skills/arrayexpress-fetch/arrayexpress_fetch.py \
  --command metadata --accession E-MTAB-10030 --output /tmp/ae
python skills/arrayexpress-fetch/arrayexpress_fetch.py \
  --command files --accession E-MTAB-10030 --output /tmp/ae

# The experimental design, as submitted
python skills/arrayexpress-fetch/arrayexpress_fetch.py \
  --command sdrf --accession E-MTAB-10030 --output /tmp/ae

# Download by category
python skills/arrayexpress-fetch/arrayexpress_fetch.py \
  --command download --accession E-MTAB-10030 --magetab --output /tmp/ae

# Harmonised sample table
python skills/arrayexpress-fetch/arrayexpress_fetch.py \
  --command metadata-table --accession E-MTAB-10030 --output /tmp/ae

# Pipeline-ready samplesheets
python skills/arrayexpress-fetch/arrayexpress_fetch.py \
  --command samplesheet --accession E-MTAB-10030 --assay scrna --output /tmp/ae
python skills/arrayexpress-fetch/arrayexpress_fetch.py \
  --command samplesheet --accession E-MTAB-10030 --assay bulk \
  --strandedness reverse --output /tmp/ae

# A 10x run whose technical reads are separate files — declare the cDNA pair
python skills/arrayexpress-fetch/arrayexpress_fetch.py \
  --command samplesheet --accession E-MTAB-XXXXX --assay scrna \
  --read-map 3,4 --output /tmp/ae

# Search
python skills/arrayexpress-fetch/arrayexpress_fetch.py \
  --command search --query "single cell heart" --limit 20 --output /tmp/ae

# Via the ClawBio runner
python clawbio.py run arrayexpress-fetch --demo
python clawbio.py run arrayexpress-fetch --command metadata --accession E-MTAB-10030

# The upstream positional form also works when called directly
python skills/arrayexpress-fetch/arrayexpress_fetch.py metadata E-MTAB-10030 --output /tmp/ae

Demo

bash
python clawbio.py run arrayexpress-fetch --demo

Runs metadata, files, sdrf, metadata-table, samplesheet and search against the bundled E-MTAB-10030 fixtures, entirely offline, and writes the full output tree including a 6-sample nf-core samplesheet.

Algorithm / Methodology

File classification is by filename, in this order: *.idf.txt → idf, *.sdrf.txt → sdrf, known read/array extensions (.fastq.gz, .cel, .bam) → raw, everything else → processed.

Download base resolution. ArrayExpress files are not served from one fixed tree. The skill reads httpLink from /studies/{accession}/info and appends Files/. Upstream hardcoded www.ebi.ac.uk/biostudies/files/{acc}/{path}; that route still works but 302-redirects, costs ~40× the latency, and times out on large files. See file_url() for the measured evidence.

Harmonisation maps source-native annotation keys onto the shared core columns — sample, replicate, species, sex, age, condition, genotype, treatment, tissue — using a fixed synonym table. Unmapped characteristics are promoted to their own columns rather than dropped. Absent values are NA, never blank.

Sample identifiers are normalised to [A-Za-z0-9._-]. Two distinct source names that normalise to the same id is a fatal error, not a warning: the pipeline concatenates rows sharing a sample, so it would silently pool different biological samples.

Read pairing uses Comment[FASTQ_URI] from the SDRF. With --read-map the 1-based positions given are taken as the cDNA pair. Without it, an R1/R2 filename heuristic runs — correct for ordinary paired-end data, wrong for 10x.

Example Queries

  • "What's in E-MTAB-10030?"
  • "Show me the SDRF for that ArrayExpress experiment"
  • "Build me an nf-core/scrnaseq samplesheet from E-MTAB-10030"
  • "Search ArrayExpress for rat microglia single cell"
  • "Download just the MAGE-TAB files for E-MTAB-10030"

Example Output

metadata
E-MTAB-10030
Title:        Single-cell RNA-seq of rat microglia in monoculture and in
              coculture with neurons and astrocytes
Released:     2021-03-23
Collection:   ArrayExpress

Files: 26 total
  idf              1
  sdrf             1
  raw             24
files
26 file(s) for E-MTAB-10030:

idf                   3.1 kB  E-MTAB-10030.idf.txt
sdrf                  7.2 kB  E-MTAB-10030.sdrf.txt
raw                  1.4 GB  B1_S1_R1.fastq.gz
raw                  1.5 GB  B1_S1_R2.fastq.gz
samplesheet (--assay scrna)
csv
sample,fastq_1,fastq_2
Sample_1,https://ftp.ebi.ac.uk/pub/databases/microarray/data/experiment/MTAB/E-MTAB-10030/B1_S1_R1.fastq.gz,https://ftp.ebi.ac.uk/pub/databases/microarray/data/experiment/MTAB/E-MTAB-10030/B1_S1_R2.fastq.gz
Sample_2,https://ftp.ebi.ac.uk/pub/databases/microarray/data/experiment/MTAB/E-MTAB-10030/B1_S2_R1.fastq.gz,https://ftp.ebi.ac.uk/pub/databases/microarray/data/experiment/MTAB/E-MTAB-10030/B1_S2_R2.fastq.gz
tables/metadata.tsv
sample     replicate  species            sex    age      condition  tissue
Sample 2   1          Rattus norvegicus  mixed  0 to 2   NA         neocortex
Sample 4   1          Rattus norvegicus  mixed  0 to 2   NA         neocortex

Output Structure

output_directory/
├── report.md
├── result.json
├── samplesheet.csv
├── tables/
│   └── metadata.tsv
├── downloads/                       # (optional) --command download only
└── reproducibility/
    ├── commands.sh
    ├── environment.yml
    └── checksums.sha256

Dependencies

  • Python ≥ 3.10, standard library only. No pip install required.
  • Network access to www.ebi.ac.uk and ftp.ebi.ac.uk on TCP/443. See Gotcha 5.

Gotchas

  • Gotcha 1: You will want to trust the R1/R2 filename heuristic for a 10x run. Do not. For a Chromium run whose technical reads are separate files, the heuristic drops the cDNA read and produces a plausible-looking samplesheet that quantifies the wrong thing. Pass --read-map explicitly: 3 files → 2,3, 4 files (dual index) → 3,4. Confirm the choice with the user before writing the sheet. This is the highest-risk behaviour in the skill.
  • Gotcha 2: You will want to quote the file size as megabytes. It is bytes, and ArrayExpress raw data is routinely tens of gigabytes. Never start --command download --raw without telling the user the total first.
  • Gotcha 3: --read-map accepts positions 1–9, not 1–4. This is upstream behaviour, kept deliberately. A typo like 1,9 passes validation and fails later with a missing-file error rather than being rejected up front.
  • Gotcha 4: You will want to hardcode https://www.ebi.ac.uk/biostudies/files/{accession}/{path}. Do not. It works — it 302-redirects — but there is no one base tree behind it, and the redirect costs ~40× the latency and times out on large files. This skill resolves httpLink from /studies/{accession}/info and keeps the old path as a fallback.
  • Gotcha 5: If --command metadata works but --command download hangs or fails with a connection timeout, suspect a firewall, not a bug. Metadata comes from www.ebi.ac.uk; file bytes come from ftp.ebi.ac.uk. Networks routinely allow the first and block the second. Both need allowlisting on TCP/443 — this skill never speaks the FTP protocol despite the hostname, so asking an admin to "open FTP" achieves nothing. See docs/data-handling.md.
  • Gotcha 6: A sample-name collision is fatal, by design. If two source names normalise to the same id the skill exits rather than emitting a sheet, because the pipeline concatenates rows sharing a sample and would pool distinct biological samples without saying so.
  • Gotcha 7: A 403 from EBI mid-download is usually rate limiting, not a permissions problem — it appears after a burst of requests to the same host and clears on retry (observed 2026-09-22). Generated download scripts handle this themselves: --tool curl emits --retry 5 --retry-delay 10 plus --retry-all-errors where curl supports it. In-process --command download does not retry on 403, so just re-run it.
  • Gotcha 8: Upstream's --out defaults were relative to the working directory, so samplesheet.csv landed wherever you happened to be. Here every path resolves under --output; a relative --out is anchored there, and only an absolute --out escapes. Do not reintroduce cwd-relative defaults.
  • Gotcha 9: You will want to reach for --command download-script to get the FASTQs. There is none, deliberately. ArrayExpress brokers sequencing reads to ENA and serves the bytes from there, so build samplesheet.csv here and emit the script with ena-fetch --command download-script against the same --output; it fetches whatever URLs the sheet names, whether they point at ftp.sra.ebi.ac.uk or ftp.ebi.ac.uk. For runs not mirrored to ENA, or for 10x reads whose structure only fasterq-dump exposes reliably, use sra-tools (prefetch --option-file, then fasterq-dump). --command download is a different thing and still works for files ArrayExpress does host — IDF, SDRF, processed matrices, CEL, BAM.
  • Gotcha 10: tables/metadata.tsv is one row per sample × replicate, not one per SDRF line. A bulk paired-end SDRF puts each FASTQ on its own line, so a 12-sample study has 24 of them; the rows are grouped on Comment[ENA_RUN], falling back to Source Name, before they are harmonised. Do not "fix" a row count that looks low by iterating the SDRF directly — that is the bug this replaced, and it also doubled the BioSamples lookups.
  • Gotcha 11: --strandedness defaults to auto and should stay there unless the record states the library chemistry. The value follows from the chemistry, never from a library merely being "stranded": dUTP second-strand marking (TruSeq Stranded mRNA, NEBNext Ultra II Directional) → reverse; Lexogen QuantSeq 3′ FWD → forward — one vendor, both directions; non-directional kits → unstranded. When you do know it, declare it: nf-core/rnaseq infers per sample either way, but only an explicit value earns a mismatch report, and auto has nothing to compare against. A wrong explicit value is worse than auto — it mislabels every sample and suppresses nothing.
Show full SKILL.md (627 more words)Show less

Safety

  • Local-first: no user genetic data is ever transmitted. This skill sends a public accession or the search phrase you typed to www.ebi.ac.uk and receives public archive data. Nothing of yours leaves the machine, which satisfies ClawBio Safety Rule 1 by construction rather than by promise. File bytes are fetched from ftp.ebi.ac.uk over HTTPS. See docs/data-handling.md.
  • Credentials: none. ArrayExpress needs no key, and the skill reads no credential environment variables.
  • Execution: this skill runs nothing it writes; it emits no download script. --command download fetches files immediately, so confirm it with the user each time.
  • Disclaimer: every report carries the ClawBio medical disclaimer — "ClawBio is a research and educational tool. It is not a medical device and does not provide clinical diagnoses. Consult a healthcare professional before making any medical decisions."

Agent Boundary

The agent dispatches and explains; the skill executes. The agent chooses the accession and the subcommand, decides whether a 10x run needs --read-map, asks the user before any download or script execution, and explains the result. It does not re-derive FASTQ URLs, hand-edit the samplesheet, invent harmonisation mappings, or override a sample-name collision. Every number in the report comes from the script's output, never from the model.

Integration with Bio Orchestrator

Trigger conditions: the orchestrator routes here on an ArrayExpress accession (E-MTAB, E-GEOD, E-MEXP, E-PROT) or an explicit mention of ArrayExpress, MAGE-TAB, SDRF or IDF.

Chaining partners:

  • biostudies-fetch: the same API for non-ArrayExpress collections; route S-BSST*/S-BIAD* there.
  • ena-fetch: when the study's runs are wanted from the sequence archives rather than the ArrayExpress FASTQ mirror.
  • nfcore-rnaseq-wrapper / nfcore-scrnaseq-wrapper: the natural consumers of the samplesheet.csv this skill writes.
  • rnaseq-de / scrna-orchestrator: downstream analysis once counts exist.
  • article-data-fetcher: upstream producer. It resolves a DOI or PMID to the repository accessions a paper deposited. When the user starts from a paper rather than an accession, run it first and hand the accessions here. It downloads files and writes a manifest.json, but it does not parse MAGE-TAB, harmonise sample annotation into metadata.tsv, or emit a pipeline-ready samplesheet.csv — that is this skill's job, so the two chain rather than compete.

Maintenance

  • Update trigger: when an archive host changes an endpoint this skill calls. Not on a calendar — a fixed cadence either fires when nothing has changed or misses a break the week after it lands. The staleness signals below are the trigger.
  • How a break surfaces: from a live call, not from CI. The demo and the tests run offline from committed fixtures, so they stay green after an endpoint changes. Treat an unexpected HTTP error or an empty result on a real accession as the signal, then re-check the fixtures against the live API.
  • Staleness signals: /studies/{acc}/info stops advertising httpLink; the MAGE-TAB column vocabulary changes; nf-core changes its samplesheet columns.
  • Known debt: the metadata-table, sanitiser and read-map helper blocks are duplicated across the five archive scripts rather than factored into a shared module. This is deliberate — it keeps re-syncing with upstream cheap. Revisit only if upstream itself deduplicates.
  • Deprecation criteria: retire if EMBL-EBI folds ArrayExpress fully into the generic BioStudies record shape, at which point biostudies-fetch would cover it.

Citations

  • Ported from UKDRI/informatics_data_skills at commit 7cc3e6e (arrayexpress/scripts/arrayexpress.py), MIT licensed. Upstream copyright retained: Copyright (c) 2026 UK Dementia Research Institute.
  • ArrayExpress / BioStudies: Sarkans et al. (2018) The BioStudies database — one stop shop for all data supporting a life sciences study. Nucleic Acids Research 46(D1):D1266–D1270. https://doi.org/10.1093/nar/gkx965
  • Demo fixture: E-MTAB-10030, "Single-cell RNA-seq of rat microglia in monoculture and in coculture with neurons and astrocytes", released 2021-03-23. The record declares no explicit licence attribute; it is redistributed here as public archive metadata on the terms EMBL-EBI publishes it under. Rattus norvegicus throughout, so no human individual-level data.
  • MAGE-TAB: Rayner et al. (2006) A simple spreadsheet-based, MIAME-supportive format for microarray data: MAGE-TAB. BMC Bioinformatics 7:489. https://doi.org/10.1186/1471-2105-7-489

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files in skills/arrayexpress-fetch of ClawBio/ClawBio.

  • SKILL.md
  • arrayexpress_fetch.py
  • arrayexpress_fetch_api.py
  • examples/demo_E-MTAB-10030.json
  • examples/demo_E-MTAB-10030.sdrf.txt
  • examples/demo_search.json
  • tests/test_arrayexpress_fetch.py

Open the folder on GitHubat commit 5e045e3

Compare with similar skills

Arrayexpress Fetch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Arrayexpress Fetch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Arrayexpress Fetch this skillClawBio/ClawBio1.2k—~5.7kAutomated safety check: PassMIT
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Bulkrna Cosinor RhythmTianGzlab/OmicsClaw161—~840Automated safety check: PassApache-2.0
Aviv RegevK-Dense-AI/mimeographs129—~1.4kAutomated safety check: PassMIT
Spatial XeniumQING1105/ezST101—~535Automated safety check: PassMIT
Tooluniverse Rnaseq Deseq2wu-yc/LabClaw1.1k2 repos~4.5kAutomated safety check: PassNone

Similar skills

  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Bulkrna Cosinor Rhythm

    TianGzlab/OmicsClaw

    Load when the user needs Deterministic fixed-period 24-hour single-component cosinor OLS rhythm analysis for a bulk RNA time-course CSV.

    161 GitHub stars~840 tokensUpdated today
    Research & ScienceAuto-check passed
  • Aviv Regev

    K-Dense-AI/mimeographs

    Applies the computational biology and AI-driven reasoning of Aviv Regev (computational biologist, Genentech, single-cell genomics).

    129 GitHub stars~1.4k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Spatial Xenium

    QING1105/ezST

    Xenium platform branch of the spatial transcriptomics workflow — load and validate the platform's cell-level matrix for downstream analysis.

    101 GitHub stars~535 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Production-ready RNA-seq differential expression analysis using PyDESeq2.

    1.1k GitHub starsUsed in 2 repos~4.5k tokens
    Research & ScienceAuto-check passed
  • Bio Splicing Qc

    GPTomics/bioSkills

    Assesses RNA-seq data quality specifically for alternative splicing analysis.

    1.2k GitHub starsUsed in 2 repos~6.2k tokens
    Research & ScienceAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated today
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Questions about Arrayexpress Fetch

What does Arrayexpress Fetch do?

Query metadata and download data from ArrayExpress, EMBL-EBI's functional genomics collection, now hosted inside BioStudies. Arrayexpress Fetch is an agent skill from ClawBio/ClawBio. Query metadata and download data from ArrayExpress, EMBL-EBI's functional genomics collection, now hosted inside BioStudies.

When should I use Arrayexpress Fetch?

Arrayexpress Fetch fits situations like: tasks that involve CSV and tabular files; tasks that involve Experimental design; tasks that involve Bioinformatics.

How do I install Arrayexpress Fetch in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill arrayexpress-fetch -a claude-code`. Or copy the skill folder (skills/arrayexpress-fetch in ClawBio/ClawBio) into .claude/skills/arrayexpress-fetch in your project. Claude Code loads it when a task matches its description.

How do I install Arrayexpress Fetch in Codex?

Run `npx skills add ClawBio/ClawBio --skill arrayexpress-fetch -a codex`. Or copy the skill folder (skills/arrayexpress-fetch in ClawBio/ClawBio) into .agents/skills/arrayexpress-fetch in your project. Codex loads it when a task matches its description.

Can I use Arrayexpress Fetch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill arrayexpress-fetch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/arrayexpress-fetch, .gemini/skills/arrayexpress-fetch, .github/skills/arrayexpress-fetch and .opencode/skills/arrayexpress-fetch in your project.

What does Arrayexpress Fetch need to run?

Going by SKILL.md and its folder, Arrayexpress Fetch needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Arrayexpress Fetch access the network?

SKILL.md names 4 domains. In commands or code: ftp.ebi.ac.uk and ebi.ac.uk; the agent is likely to contact these when it follows the instructions. As links in the text: doi.org and github.com. This is read from the text; nothing was executed.

Is Arrayexpress Fetch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Arrayexpress Fetch use?

Arrayexpress Fetch is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Arrayexpress Fetch use?

About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Arrayexpress Fetch?

Skills that share tags, products or a category with Arrayexpress Fetch: Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars), Bulkrna Cosinor Rhythm (TianGzlab/OmicsClaw, 161 stars), Aviv Regev (K-Dense-AI/mimeographs, 129 stars) and Spatial Xenium (QING1105/ezST, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Arrayexpress Fetch?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 7, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.