Agent skill

Bio Sequence Slicing

by GPTomics in GPTomics/bioSkills

Slice, extract, and concatenate biological sequences and annotated records using Biopython.

MITAuto-check passedResearch & Science

Install Bio Sequence Slicing

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-sequence-slicing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-sequence-slicing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/sequence-manipulation/sequence-slicing .claude/skills/bio-sequence-slicing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-sequence-slicing
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.7k tokens
SKILL.md length
984 words
Files
6
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Slice, extract, and concatenate biological sequences and annotated records using Biopython.

  • Extracting subsequences by position
  • SKILL.md covers Version Compatibility, The Governing Principle, Required Imports and Coordinate Systems: the…, plus 5 more sections
  • Runs Python scripts from its folder; calls pip
  • Splicing exons into a transcript

What it does

Bio Sequence Slicing is an agent skill from GPTomics/bioSkills. Slice, extract, and concatenate biological sequences and annotated records using Biopython. Use when extracting subsequences by position, splicing exons into a transcript, joining sequences, or carrying a sub-region of an annotated record (with quality scores and features) into a new record.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `examples/basic_slicing.py`, `examples/chunking.py` and `examples/concatenation.py`).

It sits in Research & Science, covering Bioinformatics. It works with Biopython and Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Extracting subsequences by position
  • Splicing exons into a transcript
  • Joining sequences
  • Carrying a sub-region of an annotated record (with quality scores and features) into a new record

Example prompts

  • “/bio-sequence-slicing”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Sequence Slicing loads about 2.7k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 984 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 984 words, ~2,697 tokens.

Download SKILL.mdSave it as .claude/skills/bio-sequence-slicing/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
bio-sequence-slicing
description
Slice, extract, and concatenate biological sequences and annotated records using Biopython. Use when extracting subsequences by position, splicing exons into a transcript, joining sequences, or carrying a sub-region of an annotated record (with quality scores and features) into a new record.
tool_type
python
primary_tool
Bio.Seq

Version Compatibility

Reference examples tested with: BioPython 1.83+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Sequence Slicing

Extract sub-regions, splice non-contiguous regions, and concatenate sequences and annotated records.

"Extract a subsequence" -> Slice a Seq with 0-based half-open coordinates.

  • Python: seq[start:end] (Bio.Seq)

"Pull out a sub-region but keep its quality scores and features" -> Slice the SeqRecord, not the bare Seq.

  • Python: record[start:end] (Bio.SeqRecord)

"Splice exons into a transcript" -> Extract each region and concatenate.

  • Python: sum((seq[s:e] for s, e in coords), Seq('')) or the + operator

The Governing Principle

Slicing a bare Seq is pure string math: seq[start:end] returns a new Seq, half-open, with no metadata to lose. Slicing a SeqRecord carries metadata, and the rule for WHAT survives is the single most error-prone part of this skill:

record[start:end] (verified against Bio/SeqRecord.__getitem__):

  • PRESERVES id, name, description, and molecule_type.
  • AUTO-SLICES letter_annotations (per-letter data such as PHRED phred_quality) to match the new coordinates -- this is why a FASTQ slice keeps the right per-base qualities for free.
  • KEEPS only features FULLY CONTAINED in [start:end]; their locations are recalculated relative to the new start.
  • SILENTLY DROPS the annotations dict (organism, taxonomy, references, comments), the dbxrefs list, and any feature that STRADDLES the slice boundary (dropped whole, never truncated). A non-trivial stride (record[::2]) drops features entirely.

Nothing warns when annotations vanish. The GenBank source feature spans the whole record, so it straddles almost any slice and disappears along with organism/taxonomy. To carry that metadata across, copy it explicitly:

python
sub = record[start:end]
sub.annotations = record.annotations.copy()

and re-add any boundary-straddling feature manually (with a clamped, recalculated location) if a truncated copy is needed.

Required Imports

python
from Bio.Seq import Seq
from Bio.SeqRecord import SeqRecord
from Bio import SeqIO

Coordinate Systems: the 0-based vs 1-based trap

Python and Biopython slicing is 0-based and half-open: seq[start:end] includes start, excludes end, and returns end - start letters. File formats disagree, and mixing them is a SILENT off-by-one (no error, just the wrong bases):

SourceConventionPosition 1234..5678 means
Python / Bio.Seq slice0-based, half-openseq[1234:5678]
GenBank / EMBL / GFF / VCF feature line1-based, INCLUSIVEseq[1233:5678] (subtract 1 from start only)
BED file0-based, half-openseq[1234:5678] (already matches Python)

The asymmetry is the catch: convert a 1-based inclusive interval by subtracting 1 from the START only; the end already lands correctly because Python's exclusive end cancels the inclusive end. Reading a coordinate straight off a GFF and slicing seq[start:end] without the -1 silently shifts everything one base left.

Bio.SeqFeature locations sidestep this entirely: they store a 0-based start and a Python-style end, so int(feature.location.start):int(feature.location.end) slices the parent directly, and feature.extract(record.seq) does the same automatically (handling strand and compound/joined locations).

python
def extract_1based(seq, start, end):
    '''Extract a 1-based inclusive interval (GenBank/GFF style).'''
    return seq[start - 1:end]

Slicing a Bare Seq

Slicing returns a Seq (not a string); negative indices and strides behave exactly like str (Seq has behaved like str since BioPython 1.78).

python
seq = Seq('ATGCGATCGATCG')
seq[0]       # 'A'  single base, 0-indexed -> returns a str
seq[-1]      # 'G'  last base
seq[0:3]     # Seq('ATG')   first 3 bases
seq[-5:]     # Seq('GATCG')  last 5
seq[::2]     # Seq('AGGTGTG')  every 2nd base (stride)
seq[::-1]    # Seq('GCTAGCTAGCGTA')  reversed (not the reverse complement)

str(record.seq) returns the raw string, but raises UndefinedSequenceError when the record's sequence content is undefined (e.g. Seq(None, length=n) from a header-only FASTA or a pysam-backed record). Guard with len() (always defined) before forcing the content to a string.

Code Patterns

Splice Non-Contiguous Regions (Exons -> Transcript)

Goal: Join several separated regions of a genomic sequence into one continuous sequence.

Approach: Extract each region with half-open coordinates and concatenate. sum() needs an explicit Seq('') start value because the default 0 cannot be added to a Seq.

python
def extract_regions(seq, regions):
    '''Concatenate multiple [start, end) regions in order.'''
    return sum((seq[start:end] for start, end in regions), Seq(''))

exon_coords = [(0, 50), (100, 150), (200, 250)]
mrna = extract_regions(genomic_seq, exon_coords)

For a real annotated transcript, let the feature do the work -- feature.extract honors strand and joined exon locations:

python
for feature in record.features:
    if feature.type == 'mRNA':
        transcript = feature.extract(record.seq)
Show full SKILL.md (406 more words)Show less
Carry a Sub-Region into a New Annotated Record

Goal: Keep id, per-base quality, and contained features when extracting a window, and decide deliberately what metadata to carry.

Approach: Slice the SeqRecord (qualities and contained features ride along automatically), then explicitly copy the annotations dict, which slicing always drops.

python
sub = record[100:400]                      # qualities + contained features auto-sliced
sub.annotations = record.annotations.copy()  # organism/taxonomy/refs would be lost otherwise
sub.id = f'{record.id}:101-400'            # 1-based label for humans

To build a fresh record from a bare Seq slice instead (no source metadata to carry):

python
sub = SeqRecord(record.seq[100:400], id=f'{record.id}_sub', description='positions 101-400')
Extract a Feature by Type
python
for record in SeqIO.parse('sequence.gb', 'genbank'):
    for feature in record.features:
        if feature.type == 'CDS':
            cds = feature.extract(record.seq)      # strand-aware
            gene = feature.qualifiers.get('gene', ['?'])[0]
Concatenate Sequences and Records
python
seq1 + seq2                      # Seq + Seq -> Seq
seq1 + 'NNNN'                    # Seq + str -> Seq
Seq('NNN').join([s1, s2, s3])   # linker between each -> Seq

Adding SeqRecord objects works (rec1 + rec2 concatenates sequences and per-letter annotations), but follows the same rule as slicing: the result keeps id/name/description only when both share them, and the annotations dict is reset. Set metadata on the result explicitly.

Split into Codons or Fixed Chunks
python
def split_codons(seq):
    '''Whole codons only; trailing 1-2 nt remainder is dropped.'''
    return [seq[i:i + 3] for i in range(0, len(seq) - len(seq) % 3, 3)]

def chunk_sequence(seq, size):
    '''Fixed-size chunks; final chunk may be shorter.'''
    return [seq[i:i + size] for i in range(0, len(seq), size)]
Tile Overlapping Windows
python
def sliding_windows(seq, window_size, step=1):
    for i in range(0, len(seq) - window_size + 1, step):
        yield i, seq[i:i + window_size]
Flanking Region Around a Position
python
def get_flanking(seq, position, flank):
    '''Clamp to sequence ends so the slice never runs past the edges.'''
    start = max(0, position - flank)
    end = min(len(seq), position + flank + 1)
    return seq[start:end]

Common Errors

SymptomCauseFix
Organism/taxonomy/references gone from a sub-recordrecord[start:end] silently drops the annotations dict and dbxrefssub.annotations = record.annotations.copy() after slicing
A feature spanning the cut is missing from the sliceFeatures straddling the boundary are dropped whole, not truncatedRe-add manually with a clamped, recalculated location
All features gone after record[::2]A non-trivial stride drops features entirelySlice without a stride, or rebuild features by hand
Everything shifted one base leftGFF/GenBank 1-based start sliced as if 0-basedSubtract 1 from the START only: seq[start-1:end]
UndefinedSequenceError on str(record.seq)Sequence content undefined (Seq(None, length=n))Use len(record); do not force undefined content to a string
TypeError from sum(slices)Default start 0 cannot add to a SeqPass a start: sum(slices, Seq(''))
Reversed but wrong strandseq[::-1] reverses only; it does not complementUse seq.reverse_complement() (see reverse-complement)
IndexError on single-base indexPosition past the endCheck len(seq) first; slices clamp but seq[i] does not

Decision Guide

  • Bare sequence, no metadata to keep -> slice the Seq: seq[start:end].
  • Need per-base quality or contained features to ride along -> slice the SeqRecord: record[start:end], then copy annotations.
  • Coordinates came from a GFF/GenBank/EMBL/VCF line -> subtract 1 from the start before slicing.
  • Coordinates came from a BED file -> use as-is (already 0-based half-open).
  • Strand-aware or joined/compound location -> feature.extract(record.seq), never a manual slice.
  • Joining separated regions -> sum((seq[s:e] for s, e in coords), Seq('')).
  • seq-objects - Create Seq/SeqRecord objects and handle undefined sequence content
  • reverse-complement - Reverse-complement an extracted region (slicing reverses but does not complement)
  • transcription-translation - Translate an extracted CDS or spliced transcript
  • sequence-io/read-sequences - Parse GenBank/FASTQ records (with features and qualities) to slice
  • genome-intervals/gtf-gff-handling - Read 1-based GFF/GTF feature coordinates before slicing
  • alignment-files/sam-bam-basics - Extract sequences from BAM regions with samtools

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in sequence-manipulation/sequence-slicing of GPTomics/bioSkills.

  • SKILL.md
  • examples/basic_slicing.py
  • examples/chunking.py
  • examples/concatenation.py
  • examples/seqrecord_slicing.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Sequence Slicing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Sequence Slicing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Sequence Slicing this skillGPTomics/bioSkills1.2k1 repos~2.7kAutomated safety check: PassMIT
Biopythondavila7/claude-code-templates32k12 repos~3.4kAutomated safety check: PassMIT
Ggetdavila7/claude-code-templates32k10 repos~6.3kAutomated safety check: PassMIT
GgetK-Dense-AI/scientific-agent-skills48k1 repos~2.8kAutomated safety check: NotesBSD-2-Clause
BiopythonK-Dense-AI/scientific-agent-skills48k1 repos~4.3kAutomated safety check: NotesMIT
Biopythonlamm-mit/scienceclaw244—~3.9kAutomated safety check: PassApache-2.0

Similar skills

  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 12 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • Gget

    K-Dense-AI/scientific-agent-skills

    Queries 20+ bioinformatics resources through CLI/Python. An agent skill from K-Dense-AI/scientific-agent-skills.

    48k GitHub starsUsed in 1 repo~2.8k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    K-Dense-AI/scientific-agent-skills

    Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).

    48k GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    lamm-mit/scienceclaw

    Computational molecular biology library (sequence I/O, alignment, phylogenetics).

    244 GitHub stars~3.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Bio Compressed Files

    FreedomIntelligence/OpenClaw-Medical-Skills

    Read and write compressed sequence files (gzip, bzip2, BGZF) using Biopython.

    3.1k GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Sequence Slicing

What does Bio Sequence Slicing do?

Slice, extract, and concatenate biological sequences and annotated records using Biopython. Bio Sequence Slicing is an agent skill from GPTomics/bioSkills. Slice, extract, and concatenate biological sequences and annotated records using Biopython.

When should I use Bio Sequence Slicing?

Bio Sequence Slicing fits situations like: extracting subsequences by position; splicing exons into a transcript; joining sequences; carrying a sub-region of an annotated record (with quality scores and features) into a new record.

How do I install Bio Sequence Slicing in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-sequence-slicing -a claude-code`. Or copy the skill folder (sequence-manipulation/sequence-slicing in GPTomics/bioSkills) into .claude/skills/bio-sequence-slicing in your project. Claude Code loads it when a task matches its description.

How do I install Bio Sequence Slicing in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-sequence-slicing -a codex`. Or copy the skill folder (sequence-manipulation/sequence-slicing in GPTomics/bioSkills) into .agents/skills/bio-sequence-slicing in your project. Codex loads it when a task matches its description.

Can I use Bio Sequence Slicing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-sequence-slicing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-sequence-slicing, .gemini/skills/bio-sequence-slicing, .github/skills/bio-sequence-slicing and .opencode/skills/bio-sequence-slicing in your project.

What does Bio Sequence Slicing need to run?

Going by SKILL.md and its folder, Bio Sequence Slicing needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Sequence Slicing access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Sequence Slicing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Sequence Slicing use?

Bio Sequence Slicing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Sequence Slicing use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Sequence Slicing?

Skills that share tags, products or a category with Bio Sequence Slicing: Biopython (davila7/claude-code-templates, 32k stars), Gget (davila7/claude-code-templates, 32k stars), Gget (K-Dense-AI/scientific-agent-skills, 48k stars) and Biopython (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Sequence Slicing?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.