Agent skill

Bio Sra Data

by GPTomics in GPTomics/bioSkills

Download raw sequencing reads from NCBI SRA using sra-tools (prefetch, fasterq-dump, vdb-validate) or the ENA mirror.

MITAuto-check passedSecurity

Install Bio Sra Data

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-sra-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-sra-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/database-access/sra-data .claude/skills/bio-sra-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-sra-data
GitHub stars
1.2k
Used in
2 other repos
Token cost
~4k tokens
SKILL.md length
1,339 words
Files
6
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Download raw sequencing reads from NCBI SRA using sra-tools (prefetch, fasterq-dump, vdb-validate) or the ENA mirror.

  • Pulling FASTQ for SRR/ERR/DRR accessions
  • SKILL.md covers Version Compatibility, Required Setup, Decision matrix: where to pull… and SRA accession hierarchy, plus 11 more sections
  • Runs Shell and Python scripts from its folder; calls aws, curl and pip; reaches ftp.sra.ebi.ac.uk and ebi.ac.uk
  • Deciding between SRA-direct

What it does

Bio Sra Data is an agent skill from GPTomics/bioSkills. Download raw sequencing reads from NCBI SRA using sra-tools (prefetch, fasterq-dump, vdb-validate) or the ENA mirror. Use when pulling FASTQ for SRR/ERR/DRR accessions, deciding between SRA-direct, ENA mirror, or AWS/GCP cloud mirror (STRIDES), handling --include-technical for 10x and other single-cell records, validating with MD5/vdb-validate, navigating SRR/SRX/SRS/SRP/PRJNA hierarchy, or finding accessions via pysradb. Encodes SRA cloud-egress economics, the fasterq-dump uncompressed-scratch trap, and the…

Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `examples/download_batch.sh`, `examples/download_single.sh` and `examples/find_sra_runs.py`).

It sits in Security, covering Threat modeling and Bioinformatics. It works with Amazon Web Services, Google Cloud and NCBI. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Pulling FASTQ for SRR/ERR/DRR accessions
  • Deciding between SRA-direct
  • AWS/GCP cloud mirror (STRIDES)
  • Handling --include-technical for 10x and other single-cell records

Example prompts

  • “/bio-sra-data”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • aws
    • curl
    • pip
    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ftp.sra.ebi.ac.uk
    • ebi.ac.uk

    Also links to:

    • github.com
    • datascience.nih.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Sra Data loads about 4k tokens when it runs. Until then it costs about 147 tokens; SKILL.md has 1,339 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,339 words, ~3,975 tokens.

Download SKILL.mdSave it as .claude/skills/bio-sra-data/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
bio-sra-data
description
Download raw sequencing reads from NCBI SRA using sra-tools (prefetch, fasterq-dump, vdb-validate) or the ENA mirror. Use when pulling FASTQ for SRR/ERR/DRR accessions, deciding between SRA-direct, ENA mirror, or AWS/GCP cloud mirror (STRIDES), handling --include-technical for 10x and other single-cell records, validating with MD5/vdb-validate, navigating SRR/SRX/SRS/SRP/PRJNA hierarchy, or finding accessions via pysradb. Encodes SRA cloud-egress economics, the fasterq-dump uncompressed-scratch trap, and the --max-size default that silently truncates large prefetches.
tool_type
cli
primary_tool
sra-tools

Version Compatibility

Reference examples tested with: sra-tools 3.0+ (fasterq-dump, prefetch, vdb-validate, vdb-config), pysradb 2.2+, ENA portal API 2.0+

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: fasterq-dump --version, prefetch --version
  • Python: pip show pysradb

If a flag is unrecognized or behavior changes, run <tool> --help and adapt.

SRA Data

"Download FASTQ from this SRA accession" -> Two paths exist in 2026: the SRA toolkit (NCBI's official, with prefetch + fasterq-dump) and the ENA mirror (EMBL-EBI's mirror with direct FASTQ download, often faster). For >1 TB workflows, a third path: AWS Open Data (STRIDES program) where same-region EC2 pulls SRA data with zero egress cost.

The single most impactful decision is where to pull from. SRA-direct is the default but ENA is faster more often than not, and AWS Open Data is the right answer for cloud-native analysis pipelines.

  • CLI: prefetch SRR..., fasterq-dump SRR..., vdb-validate SRR... (sra-tools)
  • CLI: curl https://ftp.sra.ebi.ac.uk/... (ENA mirror; direct FASTQ)
  • CLI: aws s3 cp s3://sra-pub-run-odp/sra/SRR.../SRR... ./SRR....sra ... (STRIDES; object is unsuffixed; same-region free)
  • Python: pysradb for metadata; subprocess for download

Required Setup

bash
# sra-tools (toolkit)
conda install -c bioconda sra-tools           # 3.0+
fasterq-dump --version                        # confirm

# Configure cache location (default ~/ncbi/ -- often too small)
vdb-config --cfg                              # show current config
vdb-config --set /repository/user/main/public/root=/data/sra_cache

# Optional: pysradb for metadata
pip install pysradb

For STRIDES cloud:

bash
# AWS CLI (no NCBI auth needed for public buckets)
aws s3 ls s3://sra-pub-run-odp/sra/SRR12345678/ --no-sign-request

Decision matrix: where to pull from

SourceWhen bestSpeedCost
ENA mirror (FTP/Aspera)Default for most workflowsOften fastest; direct FASTQ (no SRA->FASTQ conversion needed)Free; no rate limit observed
SRA toolkit + AWS STRIDESSame-region EC2/EKSFastest within AWS us-east-1Free egress within region; small storage cost
SRA toolkit + GCP STRIDESSame-region GCP Compute EngineFastest within GCP us-central1Free egress within region
SRA-direct (prefetch + fasterq-dump)On-prem; small downloads; need SRA-format accessVariable; can be slow off-peak failsFree; NCBI throttles by IP
Aspera (ascp)Institutional accounts onlyFaster than HTTPS on long linksNCBI public Aspera retired 2019; ENA public Aspera retired ~2023; institutional use still possible

Default recommendation: ENA mirror for off-cloud, STRIDES (AWS/GCP) for in-cloud analysis. SRA-direct only when neither is available or when SRA format itself is needed (e.g. for re-extraction of technical reads).

SRA accession hierarchy

PrefixTypeGranularity
SRR / ERR / DRRRunOne sequencing run (file-level)
SRX / ERX / DRXExperimentLibrary prep + sequencing strategy
SRS / ERS / DRSSampleBiological sample
SRP / ERP / DRPStudyProject (deprecated; superseded by BioProject)
PRJNA / PRJEB / PRJDBBioProjectTop-level project ID
SAMN / SAMEA / SAMDBioSampleBiological sample (cross-archive)

Conversion is via SRA metadata: pysradb metadata <ID> or efetch -db sra -id <UID> -rettype runinfo.

The actual download unit is SRR/ERR/DRR (runs). The BioProject (PRJNA...) is the convenient top-level handle for "pull all data for paper X".

fasterq-dump vs fastq-dump

fasterq-dump (sra-tools 2.10+) is the multi-threaded successor. Always prefer it, with two exceptions noted below.

Aspectfasterq-dumpfastq-dump
ThreadsMulti (-e N)Single
Speed~5-10x fasterBaseline
Disk overheadWrites uncompressed FASTQ to scratch (~3x final size)In-place; lower scratch
CompressionNOT built-in (post-process with pigz)--gzip flag built-in
Single-cell technical reads--include-technical worksSome 10x records need fastq-dump for full extraction
10x split semanticsSometimes incompleteSometimes the only way to get all reads

The uncompressed-scratch trap: fasterq-dump writes uncompressed FASTQ first, then leaves it uncompressed. A 100 GB compressed FASTQ needs ~300 GB of scratch space + 300 GB of final output. Either compress post-hoc with pigz or use --mem to control RAM/disk tradeoff.

prefetch and the --max-size trap

prefetch downloads .sra files to the configured cache before extraction. Default --max-size 20G silently skips runs larger than 20 GB.

bash
# Wrong: silently skips runs >20 GB
prefetch SRR12345678

# Right: set max-size explicitly to your largest expected size
prefetch SRR12345678 --max-size 100G -p

For unknown-size queues, set max-size to a generous upper bound (e.g. --max-size 200G) or query metadata first with pysradb metadata.

ENA mirror: direct FASTQ URLs

ENA stores FASTQ files directly (no SRA-format intermediate). Discover URLs via the ENA portal API:

bash
curl 'https://www.ebi.ac.uk/ena/portal/api/filereport?accession=SRR12345678&result=read_run&fields=fastq_ftp,fastq_md5,read_count&format=tsv'

Returns TSV with semicolon-separated paired-end URLs and md5 checksums.

Direct download:

bash
curl -O 'https://ftp.sra.ebi.ac.uk/vol1/fastq/SRR123/078/SRR12345678/SRR12345678_1.fastq.gz'

ENA's mirror is typically faster than SRA's because (a) it's hosted on Aspera-aware servers, (b) the FASTQ is pre-compressed (no SRA->FASTQ conversion needed), (c) EMBL-EBI's bandwidth is generous. For most downloads in 2026, ENA is the right default.

Single-cell / 10x quirks

10x Genomics records include "technical reads" (cell barcodes, UMIs) interleaved with biological reads. Default fasterq-dump (or fastq-dump) skips them. To get all reads:

bash
# fasterq-dump with technical reads
fasterq-dump SRR12345678 --include-technical --split-files -p -O ./fastq/

# Some 10x records require fastq-dump -- check sra-stat first
sra-stat --xml SRR12345678 | grep -E '(spotCount|baseCount|tag)'

For 10x v3, expect 3 files per run: R1 (barcode+UMI), R2 (cDNA), I1 (index). For 10x v2: R1 (barcode), R2 (UMI+cDNA), I1.

MD5 / vdb-validate

Always verify downloads.

bash
# vdb-validate for SRA-format files (toolkit path)
vdb-validate SRR12345678

# md5sum for ENA FASTQ files
md5sum -c <(echo "<expected_md5>  SRR12345678_1.fastq.gz")

ENA provides md5 in the portal API response. SRA-toolkit's vdb-validate is the equivalent for .sra files (different file format).

Cloud (STRIDES) access

NCBI's STRIDES initiative mirrored SRA data to AWS Open Data (us-east-1) and GCP (us-central1). Same-region pulls have zero egress cost.

bash
# List SRA cloud-hosted files (no NCBI auth needed)
aws s3 ls s3://sra-pub-run-odp/sra/SRR12345678/ --no-sign-request

# Direct copy to EC2 in us-east-1. The STRIDES object is named without a `.sra`
# suffix (just SRR12345678); rename on copy to keep fasterq-dump happy.
aws s3 cp s3://sra-pub-run-odp/sra/SRR12345678/SRR12345678 ./SRR12345678.sra --no-sign-request

# Then fasterq-dump locally
fasterq-dump ./SRR12345678.sra -p -e 8

For cloud-native analysis pipelines (Nextflow on AWS Batch, Cromwell, etc.), STRIDES is the right path.

Code patterns

Single SRR via ENA mirror (preferred default)

Goal: Download paired-end FASTQ for one SRR; verify md5; minimal dependencies.

Approach: Query ENA portal API for FASTQ URLs and md5; download with curl; verify with md5sum.

Reference (ENA portal API 2.0+, curl):

bash
#!/bin/bash
SRR="${1:-SRR12345678}"
OUT="${2:-./fastq}"
mkdir -p "${OUT}"

# Get FASTQ URLs + md5 from ENA portal API
META=$(curl -s "https://www.ebi.ac.uk/ena/portal/api/filereport?accession=${SRR}&result=read_run&fields=fastq_ftp,fastq_md5&format=tsv" | tail -1)
URLS=$(echo "${META}" | cut -f1 | tr ';' '\n')
MD5S=$(echo "${META}" | cut -f2 | tr ';' '\n')

i=0
while read url; do
    fname="${OUT}/$(basename ${url})"
    expected_md5=$(echo "${MD5S}" | sed -n "$((i+1))p")
    echo "Downloading ${fname}"
    curl -sL -o "${fname}" "https://${url}"
    actual_md5=$(md5sum "${fname}" | awk '{print $1}')
    if [ "${actual_md5}" != "${expected_md5}" ]; then
        echo "MD5 MISMATCH ${fname}: expected ${expected_md5}, got ${actual_md5}"
        exit 1
    fi
    echo "  md5 OK"
    i=$((i+1))
done <<< "${URLS}"
prefetch + fasterq-dump (SRA toolkit, classic)
bash
#!/bin/bash
SRR="${1:-SRR12345678}"
OUT="${2:-./fastq}"
THREADS="${3:-8}"
mkdir -p "${OUT}"

# prefetch with explicit max-size (default 20G silently skips larger)
prefetch "${SRR}" --max-size 100G -p

# Validate SRA file
vdb-validate "${SRR}" || { echo "Validation FAILED"; exit 1; }

# Extract FASTQ (multi-threaded; uncompressed scratch ~3x final size)
fasterq-dump "${SRR}" -O "${OUT}" -e "${THREADS}" -p --split-files

# Compress post-hoc (fasterq-dump does NOT compress)
pigz -p "${THREADS}" "${OUT}/${SRR}"_*.fastq

# Cleanup SRA cache if you don't need it
# rm -rf ~/ncbi/sra/${SRR}.sra
Show full SKILL.md (543 more words)Show less
Batch via pysradb metadata

Goal: Convert a list of GSE / BioProject / SRX IDs to SRR run accessions.

Approach: pysradb metadata returns a full hierarchy table; pull SRR column.

Reference (pysradb 2.2+):

python
from pysradb import SRAweb
import pandas as pd


def gse_to_srr(gse):
    db = SRAweb()
    df = db.gse_to_srp(gse)
    if df.empty:
        return []
    srp = df['study_accession'].iloc[0]
    runs = db.srp_to_srr(srp)
    return runs['run_accession'].tolist()


def bioproject_to_runs(prjna):
    db = SRAweb()
    return db.sra_metadata(prjna, detailed=True)


def batch_resolve(ids):
    db = SRAweb()
    rows = []
    for id in ids:
        try:
            meta = db.sra_metadata(id, detailed=True)
            rows.append(meta)
        except Exception as e:
            print(f'{id}: {e}')
    return pd.concat(rows, ignore_index=True) if rows else pd.DataFrame()


# Resolve a GSE to all its SRRs
srrs = gse_to_srr('GSE123456')
print(f'GSE123456 -> {len(srrs)} SRRs')
Cloud (STRIDES) via AWS
bash
#!/bin/bash
# Run from EC2 in us-east-1 for zero egress
SRR="${1:-SRR12345678}"

# Check if available on AWS Open Data
aws s3 ls "s3://sra-pub-run-odp/sra/${SRR}/" --no-sign-request

# Download .sra (then extract locally)
aws s3 cp "s3://sra-pub-run-odp/sra/${SRR}/${SRR}" "./${SRR}.sra" --no-sign-request

fasterq-dump "./${SRR}.sra" -p -e 8 --split-files
pigz -p 8 "${SRR}"_*.fastq
10x single-cell with technical reads
bash
#!/bin/bash
SRR="${1:-SRR_10x_run}"
OUT="${2:-./fastq_10x}"
mkdir -p "${OUT}"

# Get all reads including technical (barcode/UMI/index)
fasterq-dump "${SRR}" --include-technical --split-files -p -O "${OUT}" -e 8

# 10x v3 expects: R1 (28-bp barcode+UMI), R2 (cDNA), I1 (sample index)
ls -la "${OUT}/${SRR}"_*.fastq
pigz -p 8 "${OUT}/${SRR}"_*.fastq

Failure modes

prefetch --max-size silent skip
  • Trigger: Default 20 GB limit; run is 50 GB.
  • Mechanism: prefetch returns success but downloads nothing.
  • Symptom: vdb-validate or fasterq-dump fails because no file exists.
  • Fix: Always set --max-size explicitly to a generous upper bound (e.g. 200G).
fasterq-dump scratch space exhaustion
  • Trigger: Run is 100 GB compressed; scratch dir has 200 GB free.
  • Mechanism: fasterq-dump writes ~300 GB uncompressed, fills disk.
  • Symptom: "out of disk space" mid-extraction.
  • Fix: Use a scratch dir with 4-5x the compressed size; or use --mem to trade memory for disk; or stick with fastq-dump --gzip (slower but lower scratch).
10x technical reads missing
  • Trigger: Default fasterq-dump on a 10x record.
  • Mechanism: Technical reads (barcodes, UMIs) are skipped by default.
  • Symptom: Only the cDNA file (R2) appears; CellRanger / STARsolo errors.
  • Fix: Add --include-technical; verify with sra-stat --xml first.
SRA-direct slowness during US business hours
  • Trigger: Downloading from NCBI 9 AM-5 PM ET weekdays.
  • Mechanism: NCBI bandwidth contention; institutional users have priority.
  • Symptom: kbps-level download speeds.
  • Fix: Switch to ENA mirror or AWS STRIDES; run outside US business hours.
Aspera deprecation
  • Trigger: Old script using ascp against anonftp@ftp.ncbi.nlm.nih.gov.
  • Mechanism: NCBI retired public Aspera in 2019; ENA followed ~2023; only institutional accounts retain support.
  • Symptom: Connection refused or auth fails.
  • Fix: Switch to HTTPS (slower but works); for fastest cloud transfer use STRIDES (AWS/GCP).
Cloud egress costs surprise
  • Trigger: STRIDES pull from EC2 in us-west-2 against bucket in us-east-1.
  • Mechanism: Cross-region egress is charged.
  • Symptom: Unexpected AWS bill.
  • Fix: Match compute region to bucket region (us-east-1 for AWS, us-central1 for GCP).
vdb-config not persisted across containers
  • Trigger: Docker container without persisted ~/.ncbi/user-settings.mkfg.
  • Mechanism: Cache config is per-user, per-home; container rebuild loses it.
  • Symptom: Cache fills container's small layer; download fails.
  • Fix: Mount a host volume at ~/.ncbi/ and persist user-settings.mkfg; or set --temp and -O explicitly in commands.

Common errors

Error / symptomCauseSolution
"item not found"Invalid accession or not in current SRAVerify; check ENA mirror
Scratch disk full mid-extractionfasterq-dump uncompressed writeUse larger scratch or fastq-dump --gzip
Slow SRA-direct downloadBusiness-hours contentionENA or STRIDES
10x reads missing--include-technical not setAdd the flag
Container loses cache configvdb-config not persistedMount ~/.ncbi as volume
prefetch returns "success" but no file--max-size silent skipSet --max-size explicitly
AWS bill on STRIDESCross-region pullMatch compute region

References

  • NCBI. SRA Toolkit documentation. https://github.com/ncbi/sra-tools/wiki
  • NCBI. STRIDES program. https://datascience.nih.gov/strides
  • Leinonen R, Sugawara H, Shumway M; International Nucleotide Sequence Database Collaboration. (2011) The sequence read archive. Nucleic Acids Res 39:D19-D21.
  • Cochrane G, Karsch-Mizrachi I, Takagi T; International Nucleotide Sequence Database Collaboration. (2016) The International Nucleotide Sequence Database Collaboration. Nucleic Acids Res 44:D48-D50.
  • Choudhary S. (2019) pysradb: A Python package to query next-generation sequencing metadata and data from NCBI Sequence Read Archive. F1000Research 8:532.
  • entrez-search - Search the SRA db for accessions before downloading
  • geo-data - GEO Series often link to SRA; gds -> sra ELink
  • read-qc/quality-reports - QC the downloaded FASTQ
  • read-qc/fastp-workflow - Adapter trim downloaded FASTQ
  • ncbi-datasets-cli - Modern bulk path for genome data (NOT for SRA reads)

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in database-access/sra-data of GPTomics/bioSkills.

  • SKILL.md
  • examples/download_batch.sh
  • examples/download_single.sh
  • examples/find_sra_runs.py
  • examples/prefetch_large.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Sra Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Sra Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Sra Data this skillGPTomics/bioSkills1.2k2 repos~4kAutomated safety check: PassMIT
Cloud Auditbriiirussell/cybersecurity-skills413—~1.3kAutomated safety check: NotesMIT
Hunting For Living Off The Cloud Techniquesmukul975/Anthropic-Cybersecurity-Skills34k—~925Automated safety check: PassApache-2.0
Performing Cloud Incident Containment Proceduresmukul975/Anthropic-Cybersecurity-Skills34k—~2.8kAutomated safety check: PassApache-2.0
Auditing Cloud With Cis Benchmarksmukul975/Anthropic-Cybersecurity-Skills34k—~3kAutomated safety check: PassApache-2.0
Building Cloud Siem With Sentinelmukul975/Anthropic-Cybersecurity-Skills34k—~3.3kAutomated safety check: PassApache-2.0

Similar skills

  • Cloud Audit

    briiirussell/cybersecurity-skills

    Audit cloud infrastructure (AWS, GCP, Azure) for misconfigurations, excessive permissions, and security gaps.

    413 GitHub stars~1.3k tokensUpdated 4 mo ago
    SecurityAuto-check: notes
  • Hunting For Living Off The Cloud Techniques

    mukul975/Anthropic-Cybersecurity-Skills

    Hunts for adversary abuse of legitimate cloud services (Azure, AWS, GCP, and SaaS platforms) for command-and-control, data staging, and exfiltration, i.e.

    34k GitHub stars~925 tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Performing Cloud Incident Containment Procedures

    mukul975/Anthropic-Cybersecurity-Skills

    Execute cloud-native incident containment across AWS, Azure, and GCP using platform CLIs to revoke or disable compromised IAM credentials, isolate resources with security groups and network ACLs…

    34k GitHub stars~2.8k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Auditing Cloud With Cis Benchmarks

    mukul975/Anthropic-Cybersecurity-Skills

    Audit AWS, Azure, and GCP environments against the CIS Foundations Benchmarks by running automated scans with tools like Prowler and ScoutSuite, interpreting failed controls, and tracking…

    34k GitHub stars~3k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Building Cloud Siem With Sentinel

    mukul975/Anthropic-Cybersecurity-Skills

    Deploy Microsoft Sentinel as a cloud-native SIEM/SOAR by configuring multi-cloud data connectors (AWS, Azure, GCP), writing KQL detection and hunting queries, and building automated Logic Apps…

    34k GitHub stars~3.3k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Implementing Cloud Security Posture Management

    mukul975/Anthropic-Cybersecurity-Skills

    Continuously monitor multi-cloud environments (AWS, Azure, GCP) for misconfigurations, compliance violations, and security risks using Prowler, ScoutSuite, AWS Security Hub, Microsoft Defender for…

    34k GitHub stars~3k tokensUpdated 1 mo ago
    SecurityAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Sra Data

What does Bio Sra Data do?

Download raw sequencing reads from NCBI SRA using sra-tools (prefetch, fasterq-dump, vdb-validate) or the ENA mirror. Bio Sra Data is an agent skill from GPTomics/bioSkills. Download raw sequencing reads from NCBI SRA using sra-tools (prefetch, fasterq-dump, vdb-validate) or the ENA mirror.

When should I use Bio Sra Data?

Bio Sra Data fits situations like: pulling FASTQ for SRR/ERR/DRR accessions; deciding between SRA-direct; AWS/GCP cloud mirror (STRIDES); handling --include-technical for 10x and other single-cell records.

How do I install Bio Sra Data in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-sra-data -a claude-code`. Or copy the skill folder (database-access/sra-data in GPTomics/bioSkills) into .claude/skills/bio-sra-data in your project. Claude Code loads it when a task matches its description.

How do I install Bio Sra Data in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-sra-data -a codex`. Or copy the skill folder (database-access/sra-data in GPTomics/bioSkills) into .agents/skills/bio-sra-data in your project. Codex loads it when a task matches its description.

Can I use Bio Sra Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-sra-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-sra-data, .gemini/skills/bio-sra-data, .github/skills/bio-sra-data and .opencode/skills/bio-sra-data in your project.

What does Bio Sra Data need to run?

Going by SKILL.md and its folder, Bio Sra Data needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (aws, curl, pip and conda). Our summary lists: Python 3; A Bash shell.

Does Bio Sra Data access the network?

SKILL.md names 4 domains. In commands or code: ftp.sra.ebi.ac.uk and ebi.ac.uk; the agent is likely to contact these when it follows the instructions. As links in the text: github.com and datascience.nih.gov. This is read from the text; nothing was executed.

Is Bio Sra Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Sra Data use?

Bio Sra Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Sra Data use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Sra Data?

Skills that share tags, products or a category with Bio Sra Data: Cloud Audit (briiirussell/cybersecurity-skills, 413 stars), Hunting For Living Off The Cloud Techniques (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Performing Cloud Incident Containment Procedures (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Auditing Cloud With Cis Benchmarks (mukul975/Anthropic-Cybersecurity-Skills, 34k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Sra Data?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.