Agent skill

Bio Compressed Files

by GPTomics in GPTomics/bioSkills

Read, write, and index compressed sequence files (gzip, bzip2, xz, BGZF) with Biopython and bgzip/samtools.

MITAuto-check passedResearch & Science

Install Bio Compressed Files

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-compressed-files -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-compressed-files --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/sequence-io/compressed-files .claude/skills/bio-compressed-files && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-compressed-files
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.8k tokens
SKILL.md length
1,217 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Read, write, and index compressed sequence files (gzip, bzip2, xz, BGZF) with Biopython and bgzip/samtools.

  • Working with .gz
  • SKILL.md covers Version Compatibility, The Governing Principle: BGZF…, Required Imports and Reading Compressed Files, plus 9 more sections
  • Runs Python scripts from its folder; calls pip
  • .bgz sequence files

What it does

Bio Compressed Files is an agent skill from GPTomics/bioSkills. Read, write, and index compressed sequence files (gzip, bzip2, xz, BGZF) with Biopython and bgzip/samtools. Use when working with .gz, .bz2, or .bgz sequence files, when random access into a compressed FASTA/FASTQ is needed, or when SeqIO.index/faidx/tabix rejects a plain .gz. Covers the BGZF-vs-gzip seekability asymmetry, the 'rt'-not-'rb' handle trap, virtual offsets, and gzip-to-BGZF conversion.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/compressed_io.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. It works with Biopython and Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Working with .gz
  • .bgz sequence files
  • Random access into a compressed FASTA/FASTQ is needed
  • SeqIO.index/faidx/tabix rejects a plain .gz

Example prompts

  • “/bio-compressed-files”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Compressed Files loads about 2.8k tokens when it runs. Until then it costs about 106 tokens; SKILL.md has 1,217 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,217 words, ~2,799 tokens.

Download SKILL.mdSave it as .claude/skills/bio-compressed-files/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-compressed-files
description
Read, write, and index compressed sequence files (gzip, bzip2, xz, BGZF) with Biopython and bgzip/samtools. Use when working with .gz, .bz2, or .bgz sequence files, when random access into a compressed FASTA/FASTQ is needed, or when SeqIO.index/faidx/tabix rejects a plain .gz. Covers the BGZF-vs-gzip seekability asymmetry, the 'rt'-not-'rb' handle trap, virtual offsets, and gzip-to-BGZF conversion.
tool_type
mixed
primary_tool
Bio.bgzf

Version Compatibility

Reference examples tested with: BioPython 1.83+, htslib/bgzip 1.19+, samtools 1.19+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Compressed Files

Read, write, and randomly access gzip, bzip2, xz, and BGZF compressed sequence files.

"Read a compressed sequence file" -> Open a decompression handle in TEXT mode, then parse with the standard SeqIO interface.

  • gzip: gzip.open(path, 'rt') (Python stdlib)
  • bzip2: bz2.open(path, 'rt') (Python stdlib)
  • xz/LZMA: lzma.open(path, 'rt') (Python stdlib)
  • BGZF: bgzf.open(path, 'rt') (BioPython) - BGZF input ONLY

"Make a compressed file randomly accessible" -> Re-compress as BGZF, then index. Only BGZF supports SeqIO.index(), samtools faidx, and tabix on compressed data.

The Governing Principle: BGZF vs plain gzip is asymmetric

A BGZF (Blocked GNU Zip) file IS a valid gzip file - gunzip and zcat read it transparently. The reverse is FALSE: a plain .gz is NOT BGZF, so faidx/tabix/SeqIO.index reject it, and bgzf.open refuses to read it.

The reason is structural. BGZF is a series of concatenated gzip blocks, each <=64 KiB and independently decodable, so any record can be reached by seeking to its block. Plain gzip is one continuous DEFLATE stream with no block boundaries: reaching byte N means decompressing every byte before it (O(n)). Random access therefore REQUIRES BGZF; on a plain .gz, SeqIO.index() would re-decompress huge prefixes on every lookup, which is why Biopython forbids it outright (it raises rather than running slowly).

Consequences the agent must respect:

  • Reading sequentially: gzip, bzip2, xz, and BGZF all work via the matching *.open(path, 'rt') handle.
  • Random access / indexing: BGZF only. Convert plain gzip to BGZF first.
  • bgzf.open() reads BGZF input only. Pointing it at a plain .gz raises ValueError: A BGZF block should start with b'\x1f\x8b\x08\x04'.... To read plain gzip use gzip.open().
  • BAM and tabix-indexed files use BGZF natively; bzip2/xz are archive-only (no seekable index).

Required Imports

python
import gzip
import bz2
import lzma
from Bio import SeqIO
from Bio import bgzf

Reading Compressed Files

Goal: Parse sequence records from a compressed file without decompressing to disk.

Approach: Open a decompression handle in TEXT mode ('rt'), then pass the handle to SeqIO.parse(). The parser is format-agnostic about the underlying compression.

python
with gzip.open('reads.fastq.gz', 'rt') as handle:
    for record in SeqIO.parse(handle, 'fastq'):
        print(record.id, len(record.seq))

Swap gzip.open for bz2.open (.bz2), lzma.open (.xz), or bgzf.open (.bgz) - the parse loop is identical.

The 'rt' vs 'rb' trap

SeqIO.parse() in Python 3 needs a TEXT handle that yields str. A binary 'rb' handle yields bytes and raises TypeError: a bytes-like object is required (or a decode error). Always use 'rt' for reading and 'wt' for writing through SeqIO. The low-level SimpleFastaParser/FastqGeneralIterator also require text handles.

Writing Compressed Files

Goal: Save records straight to a compressed file with no intermediate plain copy.

Approach: Open a compression handle in TEXT mode ('wt'), then pass it to SeqIO.write().

python
with gzip.open('output.fasta.gz', 'wt') as handle:
    SeqIO.write(records, handle, 'fasta')

For an indexable result write BGZF instead:

python
with bgzf.open('output.fasta.bgz', 'wt') as handle:
    SeqIO.write(records, handle, 'fasta')

BgzfWriter.close() (and the with block exit) automatically appends the 28-byte empty-block EOF marker that htslib tools check for; let the context manager close the handle.

Random Access: index a BGZF file

Goal: Pull individual records by id from a large compressed file without a linear scan.

Approach: Compress as BGZF, then build a Biopython offset index. SeqIO.index() keeps virtual offsets in RAM; SeqIO.index_db() stores them in an on-disk SQLite index that persists across sessions and spans multiple files.

python
records = SeqIO.index('sequences.fasta.bgz', 'fasta')
target = records['gene_042'].seq
records.close()

# Persistent, multi-file, scales beyond RAM:
db = SeqIO.index_db('idx.sqlite', ['a.fasta.bgz', 'b.fasta.bgz'], 'fasta')

SeqIO.index() on a plain .gz raises ValueError: Gzipped files are not suitable for indexing, please use BGZF (blocked gzip format) instead. Convert first (below).

Convert plain gzip to BGZF

"Convert gzip to an indexable format" -> Decompress the gzip stream and re-compress it as BGZF.

CLI (fastest, htslib-native):

bash
# Either decompress then bgzip in place...
gzip -d sequences.fasta.gz && bgzip sequences.fasta      # -> sequences.fasta.gz (now BGZF)
# ...or stream without touching disk:
zcat sequences.fasta.gz | bgzip -@ 4 > sequences.fasta.bgz

# Index a BGZF FASTA for region extraction:
samtools faidx sequences.fasta.bgz                        # writes BOTH .fai and .gzi
samtools faidx sequences.fasta.bgz gene_042:1-200

samtools faidx on a BGZF FASTA writes TWO index files: .fai (record offsets in uncompressed coordinates) AND .gzi (the compressed-to-uncompressed block map). Deleting .gzi breaks region extraction even though .fai survives. Note that bgzip keeps the .gz extension, so a .gz may be EITHER plain gzip or BGZF - check with bgzip -t file.gz (tests for a valid BGZF stream) rather than trusting the suffix.

Pure-Python equivalent (no external tools):

python
import gzip
from Bio import SeqIO, bgzf

with gzip.open('input.fasta.gz', 'rt') as in_handle:
    with bgzf.open('output.fasta.bgz', 'wt') as out_handle:
        SeqIO.write(SeqIO.parse(in_handle, 'fasta'), out_handle, 'fasta')
Show full SKILL.md (541 more words)Show less

Bio.bgzf API and virtual offsets

Bio.bgzf exports open, BgzfReader, BgzfWriter, make_virtual_offset, split_virtual_offset.

A virtual offset packs two coordinates into one 64-bit integer: voffset = coffset << 16 | uoffset, where coffset is the byte position of the block start in the compressed file (top 48 bits) and uoffset is the offset within that block's decompressed data (low 16 bits - 16 bits suffices because a block holds at most 64 KiB).

python
vo = bgzf.make_virtual_offset(100, 7)        # 6553607
coffset, uoffset = bgzf.split_virtual_offset(vo)   # (100, 7)

with bgzf.open('sequences.fasta.bgz', 'rt') as handle:
    handle.readline()
    saved = handle.tell()        # a VIRTUAL offset, not a byte position
    handle.seek(saved)           # jumps back to the same record

Critical caveat: virtual offsets may be COMPARED for ordering but NEVER SUBTRACTED to get a byte length - they live in two coordinate spaces (compressed position and within-block position), so vo2 - vo1 is meaningless. BgzfReader.tell() returns a virtual offset; seek() consumes one. Text mode forces latin1 and does no newline translation.

Compression Format Comparison

FormatExtensionRandom accessSpeedRatioStdlib handle
gzip.gzNo (O(n) seek)FastGoodgzip.open
BGZF.bgz / .gzYes (block-seekable)Fast, threadableGoodbgzf.open (BioPython)
bzip2.bz2NoSlowBetterbz2.open
xz / LZMA.xzNoSlowestBestlzma.open

When to Use Each Format

Use caseFormatWhy
Sequential read/write, sharinggzipUniversal, fast, every tool reads it
Need faidx/tabix/SeqIO.indexBGZFOnly seekable compressed format
BAM, tabix-indexed VCF/GFF/BEDBGZFRequired natively
Cold archive, max shrink, no random accessxz then bzip2Highest ratios, slowest
Random access into an existing plain .gz without re-bgzippingpyfastxAdds a seek-point index over the gzip stream

pyfastx is the exception that gives random access into a PLAIN gzip FASTA/FASTQ: it builds a seek-point index (via zran from indexed_gzip) plus a SQLite .fxi/.fqi index alongside the file, a different strategy from faidx (which requires the stream itself to be BGZF). Use it when re-compressing a large gzipped genome to BGZF is not an option.

Common Errors

SymptomCauseFix
TypeError: a bytes-like object is requiredHandle opened 'rb' instead of 'rt'Open compressed handles with 'rt'/'wt' for SeqIO
ValueError: Gzipped files are not suitable for indexing, please use BGZF...SeqIO.index() on a plain .gzRe-compress as BGZF (`zcat ...
ValueError: A BGZF block should start with b'\x1f\x8b\x08\x04'...bgzf.open() pointed at a plain gzip fileRead plain gzip with gzip.open(); reserve bgzf.open for BGZF
[bgzf] file ... not BGZF / not compressed with bgzip (htslib)faidx/tabix given a plain .gzConvert to BGZF first
[faidx] Failed to read ... / could not load .gzi.gzi deleted next to a BGZF FASTARe-run samtools faidx to regenerate .fai + .gzi
gzip.BadGzipFile / OSError: Not a gzipped fileFile is not gzip (wrong suffix / corrupt)Verify with bgzip -t or file; match handle to real format
UnicodeDecodeErrorNon-UTF8 bytes in a text handlegzip.open(path, 'rt', encoding='latin-1')
  • read-sequences - parse vs index vs index_db trade-offs for compressed handles
  • write-sequences - write records through a compression handle
  • batch-processing - stream many compressed files without loading them into RAM
  • filter-sequences - keep R1/R2 in sync when filtering gzipped paired reads
  • alignment-files/sam-bam-basics - BAM is BGZF natively; samtools manages the compression

References

  • Li H, Handsaker B, Wysoker A, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25(16):2078-2079. (Defines BGZF in the SAM/BAM specification.)
  • Bonfield JK. CRAM 3.1: advances in the CRAM file format. Bioinformatics. 2022;38(6):1497. (Per-column block compression beyond BGZF.)
  • Du L, Liu Q, Fan Z, et al. Pyfastx: a robust Python package for fast random access to sequences from plain and gzipped FASTA/Q files. Briefings in Bioinformatics. 2021;22(4):bbaa368.

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in sequence-io/compressed-files of GPTomics/bioSkills.

  • SKILL.md
  • examples/compressed_io.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Compressed Files next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Compressed Files compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Compressed Files this skillGPTomics/bioSkills1.2k1 repos~2.8kAutomated safety check: PassMIT
Biopythondavila7/claude-code-templates33k12 repos~3.4kAutomated safety check: PassMIT
Ggetdavila7/claude-code-templates33k10 repos~6.3kAutomated safety check: PassMIT
GgetK-Dense-AI/scientific-agent-skills48k1 repos~2.8kAutomated safety check: NotesBSD-2-Clause
BiopythonK-Dense-AI/scientific-agent-skills48k1 repos~4.3kAutomated safety check: NotesMIT
Biopythonlamm-mit/scienceclaw246—~3.9kAutomated safety check: PassApache-2.0

Similar skills

  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 12 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 10 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • Gget

    K-Dense-AI/scientific-agent-skills

    Queries 20+ bioinformatics resources through CLI/Python. An agent skill from K-Dense-AI/scientific-agent-skills.

    48k GitHub starsUsed in 1 repo~2.8k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    K-Dense-AI/scientific-agent-skills

    Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).

    48k GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    lamm-mit/scienceclaw

    Computational molecular biology library (sequence I/O, alignment, phylogenetics).

    246 GitHub stars~3.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Bio Compressed Files

    FreedomIntelligence/OpenClaw-Medical-Skills

    Read and write compressed sequence files (gzip, bzip2, BGZF) using Biopython.

    3.1k GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Compressed Files

What does Bio Compressed Files do?

Read, write, and index compressed sequence files (gzip, bzip2, xz, BGZF) with Biopython and bgzip/samtools. Bio Compressed Files is an agent skill from GPTomics/bioSkills. Read, write, and index compressed sequence files (gzip, bzip2, xz, BGZF) with Biopython and bgzip/samtools.

When should I use Bio Compressed Files?

Bio Compressed Files fits situations like: working with .gz; .bgz sequence files; random access into a compressed FASTA/FASTQ is needed; seqIO.index/faidx/tabix rejects a plain .gz.

How do I install Bio Compressed Files in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-compressed-files -a claude-code`. Or copy the skill folder (sequence-io/compressed-files in GPTomics/bioSkills) into .claude/skills/bio-compressed-files in your project. Claude Code loads it when a task matches its description.

How do I install Bio Compressed Files in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-compressed-files -a codex`. Or copy the skill folder (sequence-io/compressed-files in GPTomics/bioSkills) into .agents/skills/bio-compressed-files in your project. Codex loads it when a task matches its description.

Can I use Bio Compressed Files in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-compressed-files -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-compressed-files, .gemini/skills/bio-compressed-files, .github/skills/bio-compressed-files and .opencode/skills/bio-compressed-files in your project.

What does Bio Compressed Files need to run?

Going by SKILL.md and its folder, Bio Compressed Files needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Compressed Files access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Compressed Files safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Compressed Files use?

Bio Compressed Files is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Compressed Files use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Compressed Files?

Skills that share tags, products or a category with Bio Compressed Files: Biopython (davila7/claude-code-templates, 33k stars), Gget (davila7/claude-code-templates, 33k stars), Gget (K-Dense-AI/scientific-agent-skills, 48k stars) and Biopython (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Compressed Files?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.