Agent skill

Bio Local Blast

by GPTomics in GPTomics/bioSkills

Build local BLAST databases and run searches using NCBI BLAST+ command-line tools.

MITAuto-check: notesResearch & Science

Install Bio Local Blast

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-local-blast -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-local-blast --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/database-access/local-blast .claude/skills/bio-local-blast && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-local-blast
GitHub stars
1.2k
Used in
2 other repos
Token cost
~4.1k tokens
SKILL.md length
1,430 words
Files
6
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Build local BLAST databases and run searches using NCBI BLAST+ command-line tools.

  • Running 50 queries
  • SKILL.md covers Version Compatibility, Installation, Database format: v5 vs v4 and makeblastdb flag taxonomy, plus 10 more sections
  • Runs Shell and Python scripts from its folder; calls conda, brew and apt
  • Building custom databases with -parseseqids and -taxid

What it does

Bio Local Blast is an agent skill from GPTomics/bioSkills. Build local BLAST databases and run searches using NCBI BLAST+ command-line tools. Use when running 50 queries, building custom databases with -parseseqids and -taxid, downloading prebuilt NCBI databases via updateblastdb.pl, choosing -task variants (megablast/dc-megablast/blastn/blastn-short), tuning soft/hard masking, scaling threads, or extracting hits with blastdbcmd. Encodes BLAST v5 vs v4 database format, taxonomy filtering, makeblastdb pitfalls.

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `examples/blast_wrapper.py`, `examples/create_database.sh` and `examples/reciprocal_best.sh`).

It sits in Research & Science. It works with NCBI. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Running 50 queries
  • Building custom databases with -parseseqids and -taxid
  • Downloading prebuilt NCBI databases via updateblastdb.pl
  • Choosing -task variants (megablast/dc-megablast/blastn/blastn-short)

Example prompts

  • “/bio-local-blast”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • conda
    • brew
    • apt

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Local Blast loads about 4.1k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 1,430 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:37
    sudo apt install ncbi-blast+

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,430 words, ~4,071 tokens.

Download SKILL.mdSave it as .claude/skills/bio-local-blast/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
bio-local-blast
description
Build local BLAST databases and run searches using NCBI BLAST+ command-line tools. Use when running >50 queries, building custom databases with -parse_seqids and -taxid, downloading prebuilt NCBI databases via update_blastdb.pl, choosing -task variants (megablast/dc-megablast/blastn/blastn-short), tuning soft/hard masking, scaling threads, or extracting hits with blastdbcmd. Encodes BLAST v5 vs v4 database format, taxonomy filtering, makeblastdb pitfalls.
tool_type
cli
primary_tool
BLAST+

Version Compatibility

Reference examples tested with: NCBI BLAST+ 2.15+

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: blastn -version then blastn -help to confirm flags
  • CLI: makeblastdb -help to confirm database build options

If a flag is unrecognized or behavior changes, introspect with -help and adapt the example to match the installed version rather than retrying.

Local BLAST

"Run BLAST locally for speed and control" -> Build or download a BLAST+ database, run the appropriate program with carefully chosen -task, masking, and thread settings, parse tabular output. Local BLAST is the right tool when remote is rate-limited or when the database must be reproducible (frozen).

The biggest mistakes are (a) using nt/nr without realizing they're >250 GB and grow weekly, (b) not building with -parse_seqids and then being unable to extract hit sequences with blastdbcmd, (c) using default blastn for cross-species when dc-megablast is correct, and (d) thinking -num_threads 32 will scale -- past ~16 threads BLAST is I/O bound.

  • CLI: makeblastdb, blastn/blastp, blastdbcmd, update_blastdb.pl (NCBI BLAST+)
  • Python: subprocess wrapper (preferred); Bio.Blast.Applications was deprecated and removed -- do not use

Installation

bash
# conda (preferred)
conda install -c bioconda blast

# macOS
brew install blast

# Ubuntu
sudo apt install ncbi-blast+

# Verify
blastn -version    # NCBI BLAST+ 2.15+ expected
update_blastdb.pl --showall pretty | head

Database format: v5 vs v4

NCBI introduced BLAST database v5 in BLAST+ 2.10 (2020). v5 includes taxonomy indexing directly in the database files, enabling -taxids and -taxidlist filtering without a companion file. v4 databases require taxonomy4blast.sqlite3 to be present and discoverable.

Featurev4v5
Default for prebuilt NCBI dbsNo (legacy)Yes (since 2020)
-taxids, -taxidlist supportNoYes
blastdbcmd -taxidsNoYes
New -info output fieldsNoYes

update_blastdb.pl downloads v5 by default. When building a database manually with makeblastdb, v5 format requires -blastdb_version 5. Always pass -blastdb_version 5 and -parse_seqids when building from scratch.

makeblastdb flag taxonomy

FlagEffectWhen
-dbtype nucl or -dbtype protRequiredAlways
-parse_seqidsIndexes accessions so blastdbcmd -entry <acc> worksAlmost always (downstream extraction)
-hash_indexSpeeds up extraction by accessionLarge dbs
-blastdb_version 5Use v5 formatAlways
-taxid 9606Single taxid for all seqsSingle-species DB
-taxid_map file.tsvPer-sequence taxid mapping (seqid<TAB>taxid)Multi-species DB
-mask_data masking.asnbApply precomputed soft-maskingProduction pipelines
-title "..."Free-text labelCosmetic
-out path/prefixDB file path prefixAlways
bash
makeblastdb -in reference.fasta -dbtype nucl \
            -blastdb_version 5 \
            -parse_seqids \
            -hash_index \
            -title "Custom reference 2026-05" \
            -out custom_db

-task taxonomy (the most-misused BLAST setting)

For blastn, the -task flag picks among heuristics with different word sizes and gap parameters.

-taskWordGappedUse caseMistake to avoid
megablast (default)28linear>=95% identity, intra-species, primer hits, contamination checkUsed for cross-species and misses everything
dc-megablast11 (discontiguous)yesCross-species mRNA homologyUnderused -- this is what blastn "should" be for cross-species
blastn11yesGeneral sensitive DNASlower than dc-megablast for same job
blastn-short7yesQueries <50 nt (primers, small RNAs)Default megablast can't seed at length 7
rmblastn11yesRepeat masking; bundled with RepeatModelerSpecialized

For blastp:

-taskWordUse case
blastp (default)3General protein similarity
blastp-fast6Faster, less sensitive
blastp-short2Peptides <30 aa, with PAM30 + word_size=2 typical

Soft vs hard masking

SettingEffect on seedEffect on extensionEffect on score
-soft_masking true (default for several tasks)Skip masked positions when seedingAllow extension through maskedScore includes masked positions
-soft_masking false + -dust yes / -seg yesSkip masked positions when seedingSkip masked positions in extensionScore excludes masked positions
Hard-mask in input FASTA (N or X)Hard exclusion everywhereHard exclusionTreated as mismatches

Soft masking is correct for almost all cases. Hard masking creates artificial mismatches at masked boundaries and can split true alignments. The exception: searching against a database of repeats explicitly, where hard masking on the query is the right choice.

Thread scaling

BLAST+ parallelizes per-query (with -num_threads) but is I/O bound past ~16 threads on most hardware. For >100,000 query batches the better answer is splitting the input FASTA into N chunks and running N parallel blastn invocations -- this saturates CPUs better than -num_threads 64.

ThreadsTypical speedup vs singleNotes
1-8Near-linearDefault sweet spot
8-16Sub-linear (1.5-2x over 8)Useful on big SMP boxes
16-32Diminishing returnsI/O bound for most DBs
32+Often slowerCache thrash + I/O contention

For massive workflows, prefer DIAMOND (Buchfink et al. 2021 Nat Methods 18:366) or MMseqs2 (Steinegger & Soding 2017 Nat Biotechnol 35:1026) -- 100-10,000x faster than BLASTP at comparable sensitivity. See remote-homology skill.

Output format reference (-outfmt)

-outfmtDescriptionUse
0Pairwise (default; human-readable)Debugging, inspection
5XMLProgrammatic parsing (Bio.SearchIO)
6Tabular (no header)Most pipelines
7Tabular with comment headersSelf-documenting
11ASN.1 binaryRe-parse with later versions

Custom tabular fields:

bash
blastn -query q.fa -db db -outfmt "6 qseqid sseqid pident length qcovs qcovhsp evalue bitscore staxids sscinames stitle"

Field key fields for analysis:

  • pident = percent identity over the HSP (NOT the query); for query-level, use qcovhsp
  • qcovs = total query coverage by all HSPs of this subject (the "coverage" most users want)
  • qcovhsp = query coverage by best HSP alone (use when there's only one HSP per hit)
  • staxids = taxonomy IDs (v5 only); critical for any "what species" workflow

Prebuilt NCBI databases via update_blastdb.pl

bash
# List available
update_blastdb.pl --showall pretty | grep -E 'refseq|swissprot|nt|nr'

# Download (with decompress)
update_blastdb.pl --decompress refseq_select_rna

# Download specific volume of split database
update_blastdb.pl --decompress refseq_protein

# Download with parallelism
update_blastdb.pl --decompress --num_threads 4 refseq_select_rna

Sizes (approximate, 2026):

  • refseq_select_rna: ~5 GB
  • refseq_protein: ~30 GB
  • swissprot: <1 GB
  • nt: ~250 GB
  • nr: ~300 GB

For most use cases, refseq_select_* is the right starting point. nt/nr are storage-heavy and reproducibility-hostile.

Code patterns

Build and search a custom protein database

Goal: Build a BLAST+ protein database from a custom FASTA and search against it.

Approach: makeblastdb with v5 + parse_seqids + hash_index; blastp with explicit outfmt.

Reference (NCBI BLAST+ 2.15+):

bash
#!/bin/bash
# Reference: NCBI BLAST+ 2.15+ | Verify API if version differs

REF=reference_proteins.fasta
DB=ref_prot_db
QUERY=query.fasta
OUT=hits.tsv

makeblastdb -in "$REF" -dbtype prot \
            -blastdb_version 5 -parse_seqids -hash_index \
            -title "$REF $(date +%Y-%m-%d)" \
            -out "$DB"

blastp -query "$QUERY" -db "$DB" \
       -evalue 1e-10 \
       -num_threads 8 \
       -max_target_seqs 500 \
       -outfmt "6 qseqid sseqid pident length qcovs evalue bitscore stitle" \
       -out "$OUT"

# Top hit per query by bit-score (column 7)
sort -k1,1 -k7,7gr "$OUT" | awk '!seen[$1]++' > top_hit_per_query.tsv
Show full SKILL.md (560 more words)Show less
Cross-species DNA with dc-megablast
bash
blastn -query mouse_cdna.fa -db human_refseq_rna \
       -task dc-megablast \
       -word_size 11 \
       -evalue 1e-10 \
       -outfmt "6 qseqid sseqid pident length qcovs evalue bitscore" \
       -num_threads 8 \
       -out cross_species.tsv
bash
blastn -query primers.fa -db genome_db \
       -task blastn-short \
       -word_size 7 \
       -evalue 1000 \
       -outfmt 6 \
       -out primer_hits.tsv
Taxonomy-filtered search (BLAST v5 only)
bash
# Restrict to specific taxids
blastp -query query.fa -db nr \
       -taxids 9606,10090,10116 \
       -outfmt "6 qseqid sseqid staxids sscinames evalue bitscore" \
       -out mammalian_hits.tsv

# Or to a taxid subtree (NCBI BLAST+ 2.13+)
echo 9606 > human_only.txt
blastp -query query.fa -db nr -taxidlist human_only.txt -outfmt 6 -out human_hits.tsv
Extract subject sequences for top hits
bash
# Requires database built with -parse_seqids
cut -f2 top_hit_per_query.tsv | sort -u > hit_accessions.txt
blastdbcmd -db ref_prot_db -entry_batch hit_accessions.txt -out hits.fasta

# Pull a range of a sequence
blastdbcmd -db genome_db -entry NC_000001.11 -range 1000000-1001000 -out region.fa
Reciprocal best hit (RBH) for ortholog candidates

See ortholog-inference skill for the principled treatment. Quick version:

bash
blastp -query A.fa -db B_db -outfmt 6 -evalue 1e-5 -num_threads 8 \
       -max_target_seqs 5 -out A_vs_B.tsv
blastp -query B.fa -db A_db -outfmt 6 -evalue 1e-5 -num_threads 8 \
       -max_target_seqs 5 -out B_vs_A.tsv

# Best forward + reverse, intersect
awk '!seen[$1]++ {print $1"\t"$2}' A_vs_B.tsv | sort > A_best
awk '!seen[$1]++ {print $1"\t"$2}' B_vs_A.tsv | sort > B_best
awk 'NR==FNR{a[$1]=$2; next} a[$2]==$1' A_best B_best > rbh.tsv

This works but does NOT handle paralog mis-pairs from gene duplication; for that use OrthoFinder or OMA (in ortholog-inference).

Python wrapper with version pinning
python
import subprocess
import shutil


def require_tool(name, min_version=None):
    if not shutil.which(name):
        raise RuntimeError(f'{name} not on PATH')
    out = subprocess.run([name, '-version'], capture_output=True, text=True)
    print(f'  {out.stdout.strip().splitlines()[0]}')


def run_blast(query, db, out, program='blastp', evalue=1e-10, threads=8, hitlist=500):
    require_tool(program)
    cmd = [program, '-query', query, '-db', db, '-out', out,
           '-evalue', str(evalue),
           '-num_threads', str(threads),
           '-max_target_seqs', str(hitlist),
           '-outfmt', '6 qseqid sseqid pident length qcovs qcovhsp evalue bitscore stitle']
    subprocess.run(cmd, check=True)


def parse_tabular(path):
    cols = ['qseqid', 'sseqid', 'pident', 'length', 'qcovs', 'qcovhsp', 'evalue', 'bitscore', 'stitle']
    rows = []
    with open(path) as f:
        for line in f:
            vals = line.rstrip('\n').split('\t')
            d = dict(zip(cols, vals))
            for k in ('pident', 'qcovs', 'qcovhsp', 'evalue', 'bitscore'):
                d[k] = float(d[k])
            d['length'] = int(d['length'])
            rows.append(d)
    return rows

Failure modes

nt/nr size shock
  • Trigger: update_blastdb.pl --decompress nt without realizing the size.
  • Mechanism: nt is ~250 GB compressed, ~1 TB indexed.
  • Symptom: Disk fills mid-download; partial DB unusable.
  • Fix: Use refseq_select for most workflows; only pull nt/nr with intent and >1 TB free.
Missing -parse_seqids
  • Trigger: Built DB without -parse_seqids; later try blastdbcmd -entry.
  • Mechanism: Without the parsed index, blastdbcmd can't look up by accession.
  • Symptom: Error: ... not found in database.
  • Fix: Rebuild with -parse_seqids (cheap if FASTA still on disk).
Wrong -task for the question
  • Trigger: Default blastn for cross-species mRNA (word=11 but ungapped seeding).
  • Mechanism: Discontiguous seed (dc-megablast) is much more sensitive across species.
  • Symptom: Far fewer hits than the question warrants.
  • Fix: Use -task dc-megablast for cross-species; -task megablast only for >=95% identity.
Thread saturation
  • Trigger: -num_threads 64 on a 32-core box.
  • Mechanism: I/O bound past ~16 threads; cache thrash hurts past CPU count.
  • Symptom: No speedup or slowdown.
  • Fix: Cap at 8-16; split FASTA and run parallel processes instead for very large batches.
v4 database, expecting v5 features
  • Trigger: Old prebuilt DB; -taxids flag returns "Taxonomy database not available".
  • Mechanism: v4 needs taxonomy4blast.sqlite3 companion; v5 has taxonomy indexed in DB.
  • Symptom: Taxonomy filtering silently no-ops or errors.
  • Fix: Re-download with update_blastdb.pl --decompress (gets v5); or use v5 explicitly when building.
Soft-masking confusion
  • Trigger: Hard-masking input (replacing repeats with N or X) instead of using -dust/-seg.
  • Mechanism: Hard-mask creates artificial mismatches at boundaries.
  • Symptom: True alignments split into multiple short HSPs.
  • Fix: Pass unmasked FASTA + soft-mask via -soft_masking true + -dust yes/-seg yes.
max_target_seqs truncation
  • Trigger: -max_target_seqs 10 (Shah et al. 2019 Bioinformatics 35:1613).
  • Mechanism: Early termination, not top-N filter.
  • Symptom: Different top-10 than -max_target_seqs 500 + post-filter.
  • Fix: Set -max_target_seqs large (500+); filter top N in awk/Python.

Common errors

Error / symptomCauseSolution
BLAST Database errorDB path wrong, or alias missingblastdbcmd -db <db> -info to confirm
Error: entry not foundBuilt without -parse_seqidsRebuild
Taxonomy filter no-opv4 DBUpgrade to v5
Threads >16 not fasterI/O boundSplit input + parallel invocations
nt download fills diskDatabase is hugeUse refseq_select
Sequence too shortQuery < word_sizeUse -task blastn-short (word=7)
Out of memorySingle large queryReduce -num_threads, split query

References

  • Camacho C, Coulouris G, Avagyan V, Ma N, Papadopoulos J, Bealer K, Madden TL. (2009) BLAST+: architecture and applications. BMC Bioinformatics 10:421.
  • Altschul SF, Madden TL, Schaffer AA, Zhang J, Zhang Z, Miller W, Lipman DJ. (1997) Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res 25:3389-3402.
  • Shah N, Nute MG, Warnow T, Pop M. (2019) Misunderstood parameter of NCBI BLAST impacts the correctness of bioinformatics workflows. Bioinformatics 35:1613-1614.
  • Boratyn GM, Camacho C, Cooper PS, et al. (2013) BLAST: a more efficient report with usability improvements. Nucleic Acids Res 41:W29-W33.
  • blast-searches - Remote BLAST against NCBI servers
  • remote-homology - PSI-BLAST, jackhmmer, HHblits, MMseqs2, DIAMOND, Foldseek for distant homology
  • ortholog-inference - Reciprocal best hit, OrthoFinder, OMA for ortholog calls
  • sequence-io/read-sequences - Load query/reference FASTA files
  • batch-downloads - Download large reference FASTA sets before makeblastdb

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in database-access/local-blast of GPTomics/bioSkills.

  • SKILL.md
  • examples/blast_wrapper.py
  • examples/create_database.sh
  • examples/reciprocal_best.sh
  • examples/run_blast.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Local Blast next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Local Blast compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Local Blast this skillGPTomics/bioSkills1.2k2 repos~4.1kAutomated safety check: NotesMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
Mako Loreliebaojun/MakoCode155—~692Automated safety check: PassCustom licence
PubMed REST API Searchdavila7/claude-code-templates32k14 repos~3.9kAutomated safety check: PassMIT
ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates32k11 repos~4.5kAutomated safety check: NotesMIT

Similar skills

  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Mako Lore

    liebaojun/MakoCode

    穗织世界观、神话与诅咒、身边人物、API速查表——常陆茉子的背景知识库,自动加载. An agent skill from liebaojun/MakoCode.

    155 GitHub stars~692 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • PubMed REST API Search

    davila7/claude-code-templates

    Searches PubMed directly through its E-utilities REST API, with guidance on Boolean and MeSH query syntax, batch retrieval and citation data.

    32k GitHub starsUsed in 14 repos~3.9k tokens
    Research & ScienceAuto-check passed
  • ETE Toolkit for Phylogenetic Trees

    davila7/claude-code-templates

    Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.

    32k GitHub starsUsed in 11 repos~4.5k tokens
    Research & ScienceAuto-check: notes
  • Gene Database

    davila7/claude-code-templates

    Query NCBI Gene via E-utilities/Datasets API. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~1.6k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Local Blast

What does Bio Local Blast do?

Build local BLAST databases and run searches using NCBI BLAST+ command-line tools. Bio Local Blast is an agent skill from GPTomics/bioSkills. Build local BLAST databases and run searches using NCBI BLAST+ command-line tools.

When should I use Bio Local Blast?

Bio Local Blast fits situations like: running 50 queries; building custom databases with -parseseqids and -taxid; downloading prebuilt NCBI databases via updateblastdb.pl; choosing -task variants (megablast/dc-megablast/blastn/blastn-short).

How do I install Bio Local Blast in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-local-blast -a claude-code`. Or copy the skill folder (database-access/local-blast in GPTomics/bioSkills) into .claude/skills/bio-local-blast in your project. Claude Code loads it when a task matches its description.

How do I install Bio Local Blast in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-local-blast -a codex`. Or copy the skill folder (database-access/local-blast in GPTomics/bioSkills) into .agents/skills/bio-local-blast in your project. Codex loads it when a task matches its description.

Can I use Bio Local Blast in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-local-blast -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-local-blast, .gemini/skills/bio-local-blast, .github/skills/bio-local-blast and .opencode/skills/bio-local-blast in your project.

What does Bio Local Blast need to run?

Going by SKILL.md and its folder, Bio Local Blast needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (conda, brew and apt). Our summary lists: Python 3; A Bash shell.

Does Bio Local Blast access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Local Blast safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Bio Local Blast use?

Bio Local Blast is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Local Blast use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Local Blast?

Skills that share tags, products or a category with Bio Local Blast: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Mako Lore (liebaojun/MakoCode, 155 stars) and PubMed REST API Search (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Local Blast?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.