Agent skill

Bio Write Sequences

by GPTomics in GPTomics/bioSkills

Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

MITAuto-check passedResearch & Science

Install Bio Write Sequences

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-write-sequences -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-write-sequences --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/sequence-io/write-sequences .claude/skills/bio-write-sequences && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-write-sequences
GitHub stars
1.2k
Used in
3 other repos
Token cost
~2.1k tokens
SKILL.md length
845 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

  • Saving sequences
  • SKILL.md covers Version Compatibility, Governing Principle, Required Import and Core Functions, plus 6 more sections
  • Runs Python scripts from its folder; calls pip
  • Creating new sequence files

What it does

Bio Write Sequences is an agent skill from GPTomics/bioSkills. Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO. Use when saving sequences, creating new sequence files, or outputting modified records.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/basic_writing.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. It works with NCBI and Biopython. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Saving sequences
  • Creating new sequence files
  • Outputting modified records

Example prompts

  • “/bio-write-sequences”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Write Sequences loads about 2.1k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 845 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 845 words, ~2,133 tokens.

Download SKILL.mdSave it as .claude/skills/bio-write-sequences/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-write-sequences
description
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO. Use when saving sequences, creating new sequence files, or outputting modified records.
tool_type
python
primary_tool
Bio.SeqIO

Version Compatibility

Reference examples tested with: BioPython 1.83+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Write Sequences

"Write sequences to a file" -> Serialize SeqRecord objects into a formatted sequence file.

  • Python: SeqIO.write() (BioPython)
  • R: writeXStringSet() (Biostrings)

Governing Principle

The write is only as complete as the SeqRecord. Each format reads specific record fields and silently ignores the rest, so what survives a write is decided by which fields are populated before the call, not by the format string. FASTA serializes only id/description+seq; FASTQ additionally requires letter_annotations['phred_quality']; GenBank/EMBL additionally require annotations['molecule_type']. Populate the fields a format needs, or the write either drops data quietly (FASTA) or raises (FASTQ/GenBank).

Required Import

python
from Bio import SeqIO
from Bio.Seq import Seq
from Bio.SeqRecord import SeqRecord

Core Functions

SeqIO.write() - Write Records to File
python
SeqIO.write(records, 'output.fasta', 'fasta')
  • records - Single SeqRecord, list, or iterator of SeqRecords
  • handle - Filename (string) or open file handle
  • format - Lowercase output format string
  • Returns the number of records written (integer)
record.format() - Get Formatted String
python
formatted = record.format('fasta')

The FASTA Header Trap (id vs description)

FASTA output is built from record.description, NOT record.id. The writer compares the first whitespace token of description to id: if they match it writes description as-is; otherwise it prepends id + a space. When a record is parsed from FASTA, description already leads with the id token, so a round trip is faithful. But when id and description are set independently, a stale id-like token inside the description gets duplicated, and older BioPython releases dropped the id entirely instead of prepending it.

record.idrecord.descriptionHeader written
seq1seq1 kinase domain>seq1 kinase domain (clean: description leads with id)
seq1kinase domain>seq1 kinase domain (id auto-prepended)
seq1`` (empty)>seq1 (id used as fallback)
seq1gene7 kinase domain>seq1 gene7 kinase domain (stale id duplicated)

To control the header exactly and stay robust across versions, make description begin with id + a space: description=f'{rec_id} kinase domain'. The FASTA writer wraps the sequence at 60 characters per line by default.

Format Field Requirements

FormatStringRecord fields readHard requirement
FASTA'fasta'id/description, seqnone (header trap above)
FASTQ'fastq'seq, letter_annotationsphred quality scores
GenBank'genbank' / 'gb'seq, annotations, featuresmolecule_type
EMBL'embl'seq, annotations, featuresmolecule_type
Tab'tab'id, seqnone

Creating SeqRecord Objects

Goal: Construct in-memory records that carry the fields the target format requires.

Approach: Build a SeqRecord from a Seq plus id; add letter_annotations['phred_quality'] for FASTQ and annotations['molecule_type'] for GenBank/EMBL.

"Create a sequence record from scratch" -> Wrap a Seq in a SeqRecord with metadata.

  • Python: SeqRecord(Seq(...), id=...) (BioPython)
python
record = SeqRecord(Seq('ATGCGATCGATCG'), id='seq1', description='seq1 example sequence')

Code Patterns

Write Single or Multiple Records
python
records = [SeqRecord(Seq('ATGC'), id='seq1'), SeqRecord(Seq('GCTA'), id='seq2')]
count = SeqIO.write(records, 'output.fasta', 'fasta')
Write to a File Handle (and Append)
python
with open('output.fasta', 'w') as handle:
    SeqIO.write(records, handle, 'fasta')

with open('output.fasta', 'a') as handle:
    SeqIO.write(new_records, handle, 'fasta')
Write Modified Records via Generator

Goal: Transform sequences in memory and write the modified versions to a new file.

Approach: Parse input, map a transform over a generator, write the generator. Streaming avoids loading every record into RAM.

"Modify sequences and save" -> Parse records, transform each, write with SeqIO.write().

python
def uppercase_record(rec):
    return SeqRecord(rec.seq.upper(), id=rec.id, description=rec.description)

records = SeqIO.parse('input.fasta', 'fasta')
modified = (uppercase_record(rec) for rec in records)
SeqIO.write(modified, 'output.fasta', 'fasta')
Show full SKILL.md (343 more words)Show less
Write FASTQ with Quality Scores

FASTQ requires letter_annotations['phred_quality'] as a list of ints. letter_annotations is length-locked to len(seq): assigning a list whose length differs from the sequence raises. Set the sequence first, then the quality list of matching length.

python
record = SeqRecord(Seq('ATGCGATCG'), id='read1')
record.letter_annotations['phred_quality'] = [40] * len(record.seq)
SeqIO.write(record, 'output.fastq', 'fastq')
Quality Encoding on Write (Phred vs Solexa)

When both phred_quality and solexa_quality keys are present, the writer uses Phred. Writing 'fastq-solexa' from a Phred-only record forces an on-the-fly lossy conversion (the scales diverge in the low-quality region) and emits a BiopythonWarning once any score reaches the high end (max quality >= ~62). For modern data, write plain 'fastq' (Sanger/Phred+33); only use 'fastq-solexa'/'fastq-illumina' when a tool explicitly demands that legacy encoding.

Write GenBank Format

GenBank and EMBL writing requires annotations['molecule_type'] (the alphabet that once carried this was removed in BioPython 1.78). Missing it raises on write.

python
record = SeqRecord(Seq('ATGCGATCGATCG'), id='SEQ001', name='example')
record.annotations['molecule_type'] = 'DNA'
record.annotations['topology'] = 'linear'
record.annotations['organism'] = 'Example organism'
SeqIO.write(record, 'output.gb', 'genbank')

Common Errors

SymptomCauseFix
Header has a duplicated or mangled idFASTA builds the header from description; it does not lead with id + spaceSet description=f'{rec.id} ...' or leave description empty to fall back to id
ValueError: No suitable quality scores found in letter_annotations of SeqRecord (id=...) on FASTQ writeRecord has no letter_annotations['phred_quality']Assign record.letter_annotations['phred_quality'] = [q]*len(seq)
TypeError: Any per-letter annotation should be a Python sequence ... of the same lengthQuality list length != len(seq) (annotations are length-locked)Set seq first, then a quality list of matching length
ValueError: missing molecule_type ... on GenBank/EMBL writeNo annotations['molecule_type'] since the 1.78 alphabet removalAdd record.annotations['molecule_type'] = 'DNA' (or 'RNA'/'protein')
BiopythonWarning: Data loss - max Solexa quality ...Writing 'fastq-solexa' from a high Phred-only record forces lossy conversionWrite plain 'fastq' unless a tool requires the Solexa encoding
TypeError passing a raw str/Seq to writeSeqIO.write expects SeqRecord(s)Wrap the sequence in a SeqRecord first
ValueError: Sequences must all be the same lengthPHYLIP/alignment format with unequal lengthsAlign, pad, or trim to equal length first
  • read-sequences - Read sequences before modifying and writing
  • format-conversion - Direct format conversion without intermediate processing
  • filter-sequences - Filter sequences before writing a subset
  • fastq-quality - Phred/Solexa encodings and quality-score handling
  • sequence-manipulation/seq-objects - Create SeqRecord objects to write
  • alignment-files/sam-bam-basics - For SAM/BAM output, use samtools/pysam

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in sequence-io/write-sequences of GPTomics/bioSkills.

  • SKILL.md
  • examples/basic_writing.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Write Sequences next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Write Sequences compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Write Sequences this skillGPTomics/bioSkills1.2k3 repos~2.1kAutomated safety check: PassMIT
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
Biopythondavila7/claude-code-templates32k12 repos~3.4kAutomated safety check: PassMIT
BiopythonK-Dense-AI/scientific-agent-skills48k1 repos~4.3kAutomated safety check: NotesMIT
Biopythonlamm-mit/scienceclaw244—~3.9kAutomated safety check: PassApache-2.0
Tooluniverse Phylogeneticswu-yc/LabClaw1.1k2 repos~4.2kAutomated safety check: PassNone

Similar skills

  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 12 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Biopython

    K-Dense-AI/scientific-agent-skills

    Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).

    48k GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    lamm-mit/scienceclaw

    Computational molecular biology library (sequence I/O, alignment, phylogenetics).

    244 GitHub stars~3.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Production-ready phylogenetics and sequence analysis skill for alignment processing, tree analysis, and evolutionary metrics.

    1.1k GitHub starsUsed in 2 repos~4.2k tokens
    Research & ScienceAuto-check passed
  • Bio Format Conversion

    FreedomIntelligence/OpenClaw-Medical-Skills

    Convert between sequence file formats (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    3.1k GitHub starsUsed in 1 repo~1.5k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Works with

Questions about Bio Write Sequences

What does Bio Write Sequences do?

Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO. Bio Write Sequences is an agent skill from GPTomics/bioSkills.SeqIO.

When should I use Bio Write Sequences?

Bio Write Sequences fits situations like: saving sequences; creating new sequence files; outputting modified records.

How do I install Bio Write Sequences in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-write-sequences -a claude-code`. Or copy the skill folder (sequence-io/write-sequences in GPTomics/bioSkills) into .claude/skills/bio-write-sequences in your project. Claude Code loads it when a task matches its description.

How do I install Bio Write Sequences in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-write-sequences -a codex`. Or copy the skill folder (sequence-io/write-sequences in GPTomics/bioSkills) into .agents/skills/bio-write-sequences in your project. Codex loads it when a task matches its description.

Can I use Bio Write Sequences in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-write-sequences -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-write-sequences, .gemini/skills/bio-write-sequences, .github/skills/bio-write-sequences and .opencode/skills/bio-write-sequences in your project.

What does Bio Write Sequences need to run?

Going by SKILL.md and its folder, Bio Write Sequences needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Write Sequences access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Write Sequences safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Write Sequences use?

Bio Write Sequences is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Write Sequences use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Write Sequences?

Skills that share tags, products or a category with Bio Write Sequences: Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Biopython (davila7/claude-code-templates, 32k stars), Biopython (K-Dense-AI/scientific-agent-skills, 48k stars) and Biopython (lamm-mit/scienceclaw, 244 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Write Sequences?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.