Agent skill

Bio Filter Sequences

by GPTomics in GPTomics/bioSkills

Filter and select sequences by criteria (length, ID, GC content, N content, motifs, patterns, description) using Biopython, streaming so large files never load into RAM.

MITAuto-check passedResearch & Science

Install Bio Filter Sequences

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-filter-sequences -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-filter-sequences --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/sequence-io/filter-sequences .claude/skills/bio-filter-sequences && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-filter-sequences
GitHub stars
1.2k
Used in
1 other repo
Token cost
~3.3k tokens
SKILL.md length
1,063 words
Files
3
Skills in repo
553
Repo updated
First seen
Licence
MIT

At a glance

Filter and select sequences by criteria (length, ID, GC content, N content, motifs, patterns, description) using Biopython, streaming so large files never load into RAM.

  • Works in 2 steps: Never filter one mate of a paired-end… → Filter by STREAMING, not by loading.…
  • Subsetting a FASTA/FASTQ file
  • SKILL.md covers Version Compatibility, The governing principle, Required Imports and Core Pattern, plus 8 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Bio Filter Sequences is an agent skill from GPTomics/bioSkills. Filter and select sequences by criteria (length, ID, GC content, N content, motifs, patterns, description) using Biopython, streaming so large files never load into RAM. Use when subsetting a FASTA/FASTQ file, removing unwanted or low-quality records, or selecting records by specific criteria. Use the paired-end-fastq skill instead whenever the input is paired R1/R2 reads.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/filter_seqs.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. It works with Biopython. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Subsetting a FASTA/FASTQ file
  • Removing unwanted
  • Low-quality records
  • Selecting records by specific criteria

Example prompts

  • “/bio-filter-sequences”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Never filter one mate of a paired-end set independently. Aligners (bwa, bowtie2) read R1 and R2 as two parallel streams and pair the i-th…
  2. Filter by STREAMING, not by loading. SeqIO.parse() yields one record at a time; a generator expression into SeqIO.write() holds a single…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Filter Sequences loads about 3.3k tokens when it runs. Until then it costs about 99 tokens; SKILL.md has 1,063 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,063 words, ~3,308 tokens.

Download SKILL.mdSave it as .claude/skills/bio-filter-sequences/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-filter-sequences
description
Filter and select sequences by criteria (length, ID, GC content, N content, motifs, patterns, description) using Biopython, streaming so large files never load into RAM. Use when subsetting a FASTA/FASTQ file, removing unwanted or low-quality records, or selecting records by specific criteria. Use the paired-end-fastq skill instead whenever the input is paired R1/R2 reads.
tool_type
python
primary_tool
Bio.SeqIO

Version Compatibility

Reference examples tested with: BioPython 1.83+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Filter Sequences

"Filter sequences by length, quality, or content" -> Apply boolean criteria to a stream of sequence records and write survivors to output.

  • Python: generator expression with SeqIO.parse() + SeqIO.write() (BioPython)
  • CLI: seqkit seq -m 200 (SeqKit) or awk on FASTA

Filter and select sequences based on length, ID, GC content, N content, motifs, regex patterns, and description.

The governing principle

Two traps cause silent, downstream-corrupting errors. Neither raises an exception, so the agent must guard against both up front.

  1. Never filter one mate of a paired-end set independently. Aligners (bwa, bowtie2) read R1 and R2 as two parallel streams and pair the i-th record of each, assuming SAME ORDER and SAME COUNT. Dropping a read from R1 alone DESYNCS the files: best case the aligner crashes on a name mismatch, worst case it silently pairs the wrong R1 with the wrong R2, producing mismapped reads and garbage insert sizes with no error. If the input is paired, route to the paired-end-fastq skill (synchronized filtering that writes matched output plus separate orphan/singleton files). Do NOT apply the single-file patterns below to one mate.
  2. Filter by STREAMING, not by loading. SeqIO.parse() yields one record at a time; a generator expression into SeqIO.write() holds a single record in RAM regardless of file size. list(SeqIO.parse(...)) materializes every record and OOMs on large FASTQ. Only the patterns that genuinely need all records at once (random sampling, splitting into multiple files) load the file; they say so explicitly.

Required Imports

python
from Bio import SeqIO
from Bio.SeqUtils import gc_fraction

Core Pattern

Stream records through a generator expression so memory stays flat:

python
records = SeqIO.parse('input.fasta', 'fasta')
filtered = (rec for rec in records if len(rec.seq) >= 100)
SeqIO.write(filtered, 'output.fasta', 'fasta')

SeqIO.write() consumes the generator lazily and returns the count written.

Filter by Length

Minimum Length
python
records = SeqIO.parse('input.fasta', 'fasta')
long_seqs = (rec for rec in records if len(rec.seq) >= 500)
SeqIO.write(long_seqs, 'long.fasta', 'fasta')
Length Range
python
records = SeqIO.parse('input.fasta', 'fasta')
sized = (rec for rec in records if 100 <= len(rec.seq) <= 1000)
SeqIO.write(sized, 'sized.fasta', 'fasta')
Remove Short Sequences
python
min_length = 200
records = SeqIO.parse('input.fasta', 'fasta')
filtered = (rec for rec in records if len(rec.seq) >= min_length)
count = SeqIO.write(filtered, 'filtered.fasta', 'fasta')

len(rec.seq) counts every base, including soft-masked lowercase (see Case Sensitivity below).

Filter by ID

Select Specific IDs
python
wanted_ids = {'seq1', 'seq2', 'seq3'}
records = SeqIO.parse('input.fasta', 'fasta')
selected = (rec for rec in records if rec.id in wanted_ids)
SeqIO.write(selected, 'selected.fasta', 'fasta')

rec.id is the first whitespace-delimited token of the header. For CASAVA 1.8+ paired FASTQ, R1 and R2 share this id (the mate number lives after the space), so an id set matches both mates equally - another reason mate-aware filtering belongs in paired-end-fastq.

Select from ID File

Goal: Extract sequences whose IDs appear in an external list file.

Approach: Load IDs into a set for O(1) lookup, then stream-filter and write matches.

Reference (BioPython 1.83+):

python
with open('ids.txt') as f:
    wanted_ids = {line.strip() for line in f}

records = SeqIO.parse('input.fasta', 'fasta')
selected = (rec for rec in records if rec.id in wanted_ids)
SeqIO.write(selected, 'selected.fasta', 'fasta')
Exclude Specific IDs
python
exclude_ids = {'bad_seq1', 'bad_seq2'}
records = SeqIO.parse('input.fasta', 'fasta')
kept = (rec for rec in records if rec.id not in exclude_ids)
SeqIO.write(kept, 'kept.fasta', 'fasta')
Filter by ID Pattern
python
import re

pattern = re.compile(r'^chr\d+$')  # matches chr1, chr2, etc.
records = SeqIO.parse('input.fasta', 'fasta')
chromosomes = (rec for rec in records if pattern.match(rec.id))
SeqIO.write(chromosomes, 'chromosomes.fasta', 'fasta')

Filter by GC Content

Goal: Keep records whose GC fraction falls in a target band.

Approach: Use gc_fraction(), which returns a FRACTION (0-1), NOT a percentage - thresholds must be 0.4, not 40. The ambiguous= mode decides how N and other IUPAC ambiguity codes are counted, and the same sequence yields a different GC value per mode, so set it explicitly rather than relying on the default.

Reference (BioPython 1.83+):

python
from Bio.SeqUtils import gc_fraction

records = SeqIO.parse('input.fasta', 'fasta')
moderate_gc = (rec for rec in records if 0.4 <= gc_fraction(rec.seq, ambiguous='ignore') <= 0.6)
SeqIO.write(moderate_gc, 'moderate_gc.fasta', 'fasta')
Choosing the ambiguous= mode

gc_fraction(seq, ambiguous='remove') is the default. For the same sequence the three modes give different answers - an N-containing read can pass or fail purely because of the mode:

ModeDenominatorgc_fraction('GCGCNNNN')When to use
'remove' (default)only unambiguous A,T,G,C,S,W,U1.0GC of the called bases only; ignores how many N's are present
'ignore'full len(seq) (N's dilute GC)0.5GC over the whole read; matches a naive (G+C)/len
'weighted'full length, ambiguous codes add expected GC0.75each IUPAC code contributes its mean GC (S=1.0, W=0.0, N=0.5, V/B=0.667, H/D=0.333)

A naive (G+C)/len silently equals 'ignore' mode and under-reports GC whenever N's are present. The default 'remove' ignores N's entirely, so a heavily-N read can post a misleadingly extreme GC. Pick the mode that matches the intent and pass it explicitly.

High / Low GC bands
python
records = SeqIO.parse('input.fasta', 'fasta')
high_gc = (rec for rec in records if gc_fraction(rec.seq, ambiguous='ignore') >= 0.6)
SeqIO.write(high_gc, 'high_gc.fasta', 'fasta')
Show full SKILL.md (430 more words)Show less

Case Sensitivity (soft-masking)

Seq is CASE-PRESERVING: lowercase soft-masked bases (from RepeatMasker, Ensembl, dustmasker) survive parse and round-trip unchanged. Length, GC, motif, and regex filters are CASE-SENSITIVE - a naive uppercase test silently misses masked bases. Always .upper() the sequence before content matching when the masking should not affect the decision:

python
seq_upper = str(rec.seq).upper()
has_site = 'GAATTC' in seq_upper            # matches gaattc and GAATTC

gc_fraction() itself is case-insensitive, but a hand-rolled .count('G') is not - count on the uppercased string.

Filter by Sequence Content

Remove Sequences with N's
python
records = SeqIO.parse('input.fasta', 'fasta')
clean = (rec for rec in records if 'N' not in str(rec.seq).upper())
SeqIO.write(clean, 'clean.fasta', 'fasta')
Limit N Content
python
def n_fraction(seq):
    upper = str(seq).upper()
    return upper.count('N') / len(seq)

records = SeqIO.parse('input.fasta', 'fasta')
low_n = (rec for rec in records if n_fraction(rec.seq) < 0.05)  # under 5% ambiguous bases
SeqIO.write(low_n, 'low_n.fasta', 'fasta')
Contains Specific Motif
python
motif = 'GAATTC'  # EcoRI site
records = SeqIO.parse('input.fasta', 'fasta')
with_motif = (rec for rec in records if motif in str(rec.seq).upper())
SeqIO.write(with_motif, 'with_ecori.fasta', 'fasta')
Regex Pattern in Sequence
python
import re

pattern = re.compile(r'ATG.{30,100}T(AA|AG|GA)')  # ORF-like pattern
records = SeqIO.parse('input.fasta', 'fasta')
matches = (rec for rec in records if pattern.search(str(rec.seq).upper()))
SeqIO.write(matches, 'orf_like.fasta', 'fasta')

Filter by Description

Description Contains Keyword
python
records = SeqIO.parse('input.fasta', 'fasta')
kinases = (rec for rec in records if 'kinase' in rec.description.lower())
SeqIO.write(kinases, 'kinases.fasta', 'fasta')
Multiple Keywords (OR)
python
keywords = ['kinase', 'phosphatase', 'transferase']
records = SeqIO.parse('input.fasta', 'fasta')
enzymes = (rec for rec in records if any(k in rec.description.lower() for k in keywords))
SeqIO.write(enzymes, 'enzymes.fasta', 'fasta')

Combine Multiple Filters

Goal: Remove sequences that fail any of several length/content thresholds.

Approach: Define a predicate that checks all criteria against the uppercased sequence once, set the GC ambiguous= mode explicitly, apply the predicate as a generator filter, and stream survivors to output.

Reference (BioPython 1.83+):

python
from Bio.SeqUtils import gc_fraction

def passes_filters(record):
    if len(record.seq) < 100:
        return False
    gc = gc_fraction(record.seq, ambiguous='ignore')
    if gc < 0.3 or gc > 0.7:
        return False
    if 'N' in str(record.seq).upper():
        return False
    return True

records = SeqIO.parse('input.fasta', 'fasta')
filtered = (rec for rec in records if passes_filters(rec))
SeqIO.write(filtered, 'filtered.fasta', 'fasta')

Sample Sequences

Random Sample (requires loading all)
python
import random

records = list(SeqIO.parse('input.fasta', 'fasta'))  # loads file - needs all records up front
sample = random.sample(records, min(100, len(records)))
SeqIO.write(sample, 'sample.fasta', 'fasta')
First N Sequences (streams)
python
from itertools import islice

records = SeqIO.parse('input.fasta', 'fasta')
first_100 = islice(records, 100)
SeqIO.write(first_100, 'first100.fasta', 'fasta')
Every Nth Sequence (streams)
python
records = SeqIO.parse('input.fasta', 'fasta')
every_10th = (rec for i, rec in enumerate(records) if i % 10 == 0)
SeqIO.write(every_10th, 'sampled.fasta', 'fasta')

Split by Criteria

Split by Length

Goal: Partition sequences into separate files based on a length threshold.

Approach: Load all records once, partition with list comprehensions, and write each partition. Loading is acceptable here because both partitions are needed in a single pass; for very large files, run two streaming passes instead.

Reference (BioPython 1.83+):

python
records = list(SeqIO.parse('input.fasta', 'fasta'))
short = [r for r in records if len(r.seq) < 500]
long = [r for r in records if len(r.seq) >= 500]
SeqIO.write(short, 'short.fasta', 'fasta')
SeqIO.write(long, 'long.fasta', 'fasta')

Common Errors

SymptomCauseFix
Downstream mismapping, wrong insert sizes, no errorFiltered one mate of a paired-end set independently, desyncing R1/R2Never filter one mate alone; use paired-end-fastq for synchronized filtering with orphan output
GC filter keeps/drops the wrong readsgc_fraction returns a fraction 0-1 but threshold written as a percent (40 instead of 0.4)Use 0-1 thresholds; multiply by 100 only for display
N-containing read unexpectedly passes or fails GC bandWrong ambiguous= mode (default 'remove' drops N's; 'ignore' dilutes GC)Set ambiguous= explicitly to match intent
Soft-masked read fails a motif/regex/uppercase testSeq is case-preserving; lowercase masked bases do not match an uppercase pattern.upper() the sequence before content matching
Generator yields nothing on second useSeqIO.parse() is one-pass and exhausts silentlyRe-create the generator, or list() it if it must be reused
MemoryError on large FASTQlist(SeqIO.parse(...)) materialized every recordUse a generator expression; only load for sampling/splitting
Empty output fileFilter too strict, or matched against the wrong case/fieldLoosen thresholds; confirm id vs description and case
  • read-sequences - Parse sequences before filtering
  • write-sequences - Write filtered sequences to output
  • fastq-quality - Filter FASTQ by per-base quality scores and encoding
  • paired-end-fastq - Synchronized filtering of R1/R2 with orphan handling
  • sequence-manipulation/sequence-properties - Per-sequence GC, length, and composition
  • sequence-manipulation/motif-search - Filter by complex motif patterns
  • alignment-files/alignment-filtering - Filter aligned reads with samtools view -f/-F

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in sequence-io/filter-sequences of GPTomics/bioSkills.

  • SKILL.md
  • examples/filter_seqs.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Filter Sequences next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Filter Sequences compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Filter Sequences this skillGPTomics/bioSkills1.2k1 repos~3.3kAutomated safety check: PassMIT
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
Biopythondavila7/claude-code-templates32k13 repos~3.4kAutomated safety check: PassMIT
Ggetdavila7/claude-code-templates32k11 repos~6.3kAutomated safety check: PassMIT
Bio Alignment Pairwisemajiayu000/claude-skill-registry6664 repos~1.7kAutomated safety check: PassMIT
GgetK-Dense-AI/scientific-agent-skills48k1 repos~2.8kAutomated safety check: NotesBSD-2-Clause

Similar skills

  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 13 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Research & ScienceAuto-check passed
  • Gget

    K-Dense-AI/scientific-agent-skills

    Queries 20+ bioinformatics resources through CLI/Python. An agent skill from K-Dense-AI/scientific-agent-skills.

    48k GitHub starsUsed in 1 repo~2.8k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    K-Dense-AI/scientific-agent-skills

    Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).

    48k GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check: notes

More from GPTomics/bioSkills

All 553 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Works with

Questions about Bio Filter Sequences

What does Bio Filter Sequences do?

Filter and select sequences by criteria (length, ID, GC content, N content, motifs, patterns, description) using Biopython, streaming so large files never load into RAM. Bio Filter Sequences is an agent skill from GPTomics/bioSkills. Filter and select sequences by criteria (length, ID, GC content, N content, motifs, patterns, description) using Biopython, streaming so large files never load into RAM.

When should I use Bio Filter Sequences?

Bio Filter Sequences fits situations like: subsetting a FASTA/FASTQ file; removing unwanted; low-quality records; selecting records by specific criteria.

How do I install Bio Filter Sequences in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-filter-sequences -a claude-code`. Or copy the skill folder (sequence-io/filter-sequences in GPTomics/bioSkills) into .claude/skills/bio-filter-sequences in your project. Claude Code loads it when a task matches its description.

How do I install Bio Filter Sequences in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-filter-sequences -a codex`. Or copy the skill folder (sequence-io/filter-sequences in GPTomics/bioSkills) into .agents/skills/bio-filter-sequences in your project. Codex loads it when a task matches its description.

Can I use Bio Filter Sequences in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-filter-sequences -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-filter-sequences, .gemini/skills/bio-filter-sequences, .github/skills/bio-filter-sequences and .opencode/skills/bio-filter-sequences in your project.

What does Bio Filter Sequences need to run?

Going by SKILL.md and its folder, Bio Filter Sequences needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Filter Sequences access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Filter Sequences safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Filter Sequences use?

Bio Filter Sequences is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Filter Sequences use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Filter Sequences?

Skills that share tags, products or a category with Bio Filter Sequences: Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Biopython (davila7/claude-code-templates, 32k stars), Gget (davila7/claude-code-templates, 32k stars) and Bio Alignment Pairwise (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Filter Sequences?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.