Agent skill

Bio Alignment Indexing

by GPTomics in GPTomics/bioSkills

Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

MITAuto-check passedResearch & Science

Install Bio Alignment Indexing

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-alignment-indexing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/alignment-files/alignment-indexing .claude/skills/bio-alignment-indexing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-alignment-indexing
GitHub stars
1.2k
Used in
2 other repos
Token cost
~2.4k tokens
SKILL.md length
715 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

  • Enabling random access to alignment files
  • SKILL.md covers Version Compatibility, Index Types, samtools index and Index Requirements, plus 5 more sections
  • Runs Python scripts from its folder; calls pip
  • Fetching specific genomic regions

What it does

Bio Alignment Indexing is an agent skill from GPTomics/bioSkills. Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam. Use when enabling random access to alignment files or fetching specific genomic regions.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/fetch_regions.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. It works with pysam and Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Enabling random access to alignment files
  • Fetching specific genomic regions

Example prompts

  • “/bio-alignment-indexing”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Alignment Indexing loads about 2.4k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 715 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 715 words, ~2,431 tokens.

Download SKILL.mdSave it as .claude/skills/bio-alignment-indexing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-alignment-indexing
description
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam. Use when enabling random access to alignment files or fetching specific genomic regions.
tool_type
cli
primary_tool
samtools

Version Compatibility

Reference examples tested with: pysam 0.22+, samtools 1.19+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Alignment Indexing

Create indices for random access to alignment files using samtools and pysam.

"Index a BAM file" -> Create a .bai/.csi index enabling random access to genomic regions.

  • CLI: samtools index file.bam
  • Python: pysam.index('file.bam')

Index Types

IndexExtensionMax contigBin shiftWhen required
BAI.bai / .bam.bai2^29 bp ≈ 537 Mbpfixed (16 kb)Default for human, mouse, fly, fish
CSI.csi / .bam.csi2^(min_shift + depth*3)configurable via -mRequired for any contig >537 Mbp
CRAI.crai / .cram.craichunk-basedn/aCRAM only
TBI.tbi2^29-1fixedtabix VCF/BED -- same limit as BAI
Which Index for Which Genome
GenomeLargest contigIndex
GRCh38 / GRCh37 (human)248 MbpBAI
GRCm39 (mouse)195 MbpBAI
GRCz11 (zebrafish), TAIR10 (Arabidopsis)78 Mbp / 30 MbpBAI
Wheat IWGSC (Triticum aestivum)~830 Mbp (chr3B)CSI
Pine, fir, axolotl, sugar pinemulti-GbpCSI with larger -m
Long-read assembly with very large contigsvariescheck cut -f2 ref.fa.fai | sort -nr | head -1

For polyploid plants and salamander-scale genomes, increase the bin shift:

bash
# Default CSI matches BAI bin layout: 2^(14 + 5*3) = 2^29 bp ≈ 537 Mbp per contig
samtools index -c file.bam

# Larger min_shift for contigs >537 Mbp (wheat, axolotl, sugar pine)
samtools index -c -m 18 file.bam   # 2^(18+15) = 2^33 = ~8.5 Gbp per contig

Index file precedence: htslib (and therefore the samtools CLI) tries .csi before .bai on auto-load, so when both exist the .csi is used -- a stale .bai left over after re-indexing to CSI is generally ignored. (Note: htsjdk/Java prefers .bai, the opposite order.) Removing the obsolete .bai still avoids confusion.

samtools index

Create BAI Index
bash
samtools index input.bam
# Creates input.bam.bai
Create CSI Index
bash
samtools index -c input.bam
# Creates input.bam.csi
Specify Output Name
bash
samtools index input.bam output.bai
Multi-threaded Indexing
bash
samtools index -@ 4 input.bam
Index CRAM
bash
samtools index input.cram
# Creates input.cram.crai

Index Requirements

Indexing requires coordinate-sorted files:

bash
# Check sort order
samtools view -H input.bam | grep "^@HD"
# Should show SO:coordinate

# Sort if needed, then index
samtools sort -o sorted.bam input.bam
samtools index sorted.bam

Using Indices for Region Access

Goal: Extract reads overlapping specific genomic coordinates from an indexed BAM.

Approach: With the index present, samtools view or pysam.fetch() can jump directly to the relevant file offset instead of scanning the entire file.

samtools view with Region
bash
# Requires index file present
samtools view input.bam chr1:1000000-2000000
Multiple Regions
bash
samtools view input.bam chr1:1000-2000 chr2:3000-4000
Regions from BED File
bash
samtools view -L regions.bed input.bam

pysam Python Alternative

Create Index
python
import pysam

pysam.index('input.bam')
# Creates input.bam.bai
Create CSI Index
python
# pysam.index passes through to samtools index; pass the -c flag for CSI.
pysam.index('-c', 'input.bam')
# Produces input.bam.csi.
Fetch with Index
python
with pysam.AlignmentFile('input.bam', 'rb') as bam:
    # fetch() requires index
    for read in bam.fetch('chr1', 1000000, 2000000):
        print(read.query_name)
Check if Indexed
python
import pysam
from pathlib import Path

def is_indexed(bam_path):
    bam_path = Path(bam_path)
    return (bam_path.with_suffix('.bam.bai').exists() or
            Path(str(bam_path) + '.bai').exists() or
            bam_path.with_suffix('.bam.csi').exists())

if not is_indexed('input.bam'):
    pysam.index('input.bam')
Fetch Multiple Regions
python
regions = [('chr1', 1000, 2000), ('chr1', 5000, 6000), ('chr2', 1000, 2000)]

with pysam.AlignmentFile('input.bam', 'rb') as bam:
    for chrom, start, end in regions:
        count = sum(1 for _ in bam.fetch(chrom, start, end))
        print(f'{chrom}:{start}-{end}: {count} reads')
Count Reads in Region
python
with pysam.AlignmentFile('input.bam', 'rb') as bam:
    count = bam.count('chr1', 1000000, 2000000)
    print(f'Reads in region: {count}')
Get Reads Covering Position
python
with pysam.AlignmentFile('input.bam', 'rb') as bam:
    for read in bam.fetch('chr1', 1000000, 1000001):
        if read.reference_start <= 1000000 < read.reference_end:
            print(f'{read.query_name} covers position 1000000')

Index File Locations

samtools looks for indices in two locations:

input.bam.bai   # Standard location
input.bai       # Alternative location

For CRAM:

input.cram.crai

idxstats - Index Statistics

Get Per-Chromosome Counts
bash
samtools idxstats input.bam

Output format:

chr1    248956422    5000000    0
chr2    242193529    4500000    0
*       0            0          10000

Columns: reference name, length, mapped reads, unmapped reads.

What "mapped" Actually Counts (Caveat)

The mapped column counts every alignment record with that RNAME, including secondary AND supplementary. For long-read minimap2 output, where a single read can produce many supplementary chimeric alignments, idxstats overcounts input reads -- typically 1.5-3x.

For unique read counts, use primary-only:

bash
samtools view -c -F 2304 input.bam chr1   # primary only

Cross-check unmapped consistency (a senior sanity check):

bash
samtools idxstats file.bam | awk '{sum+=$4} END {print sum}'   # idxstats unmapped (sum across all rows; PE orphans get a contig RNAME)
samtools view -c -f 4 -F 2304 file.bam                         # primary unmapped (should match)
Show full SKILL.md (266 more words)Show less
Sum Total Mapped Reads
bash
samtools idxstats input.bam | awk '{sum += $3} END {print sum}'
pysam idxstats
python
with pysam.AlignmentFile('input.bam', 'rb') as bam:
    for stat in bam.get_index_statistics():
        print(f'{stat.contig}: {stat.mapped} mapped, {stat.unmapped} unmapped')

FASTA Index (faidx)

Related but different - index reference FASTA for random access:

bash
samtools faidx reference.fa
# Creates reference.fa.fai

# Fetch region from indexed FASTA
samtools faidx reference.fa chr1:1000-2000
pysam FastaFile
python
with pysam.FastaFile('reference.fa') as ref:
    seq = ref.fetch('chr1', 1000, 2000)
    print(seq)

Quick Reference

Tasksamtoolspysam
Create BAIsamtools index file.bampysam.index('file.bam')
Create CSIsamtools index -c file.bampysam.index('-c', 'file.bam')
Fetch regionsamtools view file.bam chr1:1-1000bam.fetch('chr1', 0, 1000)
Count in regionsamtools view -c file.bam chr1:1-1000bam.count('chr1', 0, 1000)
Index statssamtools idxstats file.bambam.get_index_statistics()
Index FASTAsamtools faidx ref.faAutomatic with FastaFile

Index Staleness

If the BAM was modified after indexing, the index points to wrong file offsets and region queries return wrong (or zero) reads. Quick check:

bash
if [ input.bam -nt input.bam.bai ]; then
    echo "Index older than BAM; re-indexing"
    samtools index input.bam
fi

Contig-Naming Sanity Check

A leading cause of "my variant calling produced empty VCFs" tickets: querying chrM against a BAM that uses MT (or chr1 vs 1). Always inspect contig conventions before region queries:

bash
samtools view -H input.bam | grep '^@SQ' | head -3
# Compare with reference dict:
samtools dict ref.fa | head -3

UCSC convention uses chr1/chrM; Ensembl/NCBI uses 1/MT. The two are not interchangeable; tools fail with "contig not found" or silently return zero reads.

Common Errors

ErrorCauseSolution
random alignment retrieval only works for indexed BAMMissing indexRun samtools index file.bam
file is not sortedUnsorted BAMSort first with samtools sort
chromosome not foundWrong chromosome nameCheck names with samtools view -H
Region query returns zero reads on a known-populated locusStale BAI / chr vs no-chr mismatchRe-index; verify naming convention
BAI silently truncates reads on contigs >537 MbpPlant / amphibian / amplified genomeUse CSI: samtools index -c file.bam
  • sam-bam-basics - View and convert alignment files
  • alignment-sorting - Sort BAM files (required before indexing)
  • alignment-filtering - Filter by regions using index
  • bam-statistics - Use idxstats for quick counts
  • sequence-io/read-sequences - Index FASTA with SeqIO.index_db()

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in alignment-files/alignment-indexing of GPTomics/bioSkills.

  • SKILL.md
  • examples/fetch_regions.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Alignment Indexing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Alignment Indexing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Alignment Indexing this skillGPTomics/bioSkills1.2k2 repos~2.4kAutomated safety check: PassMIT
PysamK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: NotesMIT
Tooluniverse Epigenomicswu-yc/LabClaw1.1k2 repos~14kAutomated safety check: PassNone
Samtools Bam Processingjaechang-hits/SciAgent-Skills3711 repos~4.1kAutomated safety check: PassMIT
Pysam Genomic Filesjaechang-hits/SciAgent-Skills3711 repos~5.2kAutomated safety check: PassMIT
Bio Splicing QcFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~1.6kAutomated safety check: PassNone

Similar skills

  • Pysam

    K-Dense-AI/scientific-agent-skills

    Provides Python/HTSlib workflows for genomic files. An agent skill from K-Dense-AI/scientific-agent-skills.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Research & ScienceAuto-check: notes
  • Production-ready genomics and epigenomics data processing for BixBench questions.

    1.1k GitHub starsUsed in 2 repos~14k tokens
    Research & ScienceAuto-check passed
  • Samtools Bam Processing

    jaechang-hits/SciAgent-Skills

    CLI toolkit for SAM/BAM/CRAM: sort, index, convert, filter, QC alignments.

    371 GitHub starsUsed in 1 repo~4.1k tokens
    Research & ScienceAuto-check passed
  • Pysam Genomic Files

    jaechang-hits/SciAgent-Skills

    Read/write SAM/BAM/CRAM, VCF/BCF, FASTA/FASTQ. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub starsUsed in 1 repo~5.2k tokens
    Research & ScienceAuto-check passed
  • Bio Splicing Qc

    FreedomIntelligence/OpenClaw-Medical-Skills

    Assesses RNA-seq data quality for splicing analysis including junction saturation curves, splice site strength scoring, and junction coverage metrics using RSeQC.

    3.1k GitHub stars~1.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Works with

Questions about Bio Alignment Indexing

What does Bio Alignment Indexing do?

Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam. Bio Alignment Indexing is an agent skill from GPTomics/bioSkills. Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

When should I use Bio Alignment Indexing?

Bio Alignment Indexing fits situations like: enabling random access to alignment files; fetching specific genomic regions.

How do I install Bio Alignment Indexing in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a claude-code`. Or copy the skill folder (alignment-files/alignment-indexing in GPTomics/bioSkills) into .claude/skills/bio-alignment-indexing in your project. Claude Code loads it when a task matches its description.

How do I install Bio Alignment Indexing in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a codex`. Or copy the skill folder (alignment-files/alignment-indexing in GPTomics/bioSkills) into .agents/skills/bio-alignment-indexing in your project. Codex loads it when a task matches its description.

Can I use Bio Alignment Indexing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-alignment-indexing, .gemini/skills/bio-alignment-indexing, .github/skills/bio-alignment-indexing and .opencode/skills/bio-alignment-indexing in your project.

What does Bio Alignment Indexing need to run?

Going by SKILL.md and its folder, Bio Alignment Indexing needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Alignment Indexing access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Alignment Indexing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Alignment Indexing use?

Bio Alignment Indexing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Alignment Indexing use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Alignment Indexing?

Skills that share tags, products or a category with Bio Alignment Indexing: Pysam (K-Dense-AI/scientific-agent-skills, 48k stars), Tooluniverse Epigenomics (wu-yc/LabClaw, 1.1k stars), Samtools Bam Processing (jaechang-hits/SciAgent-Skills, 371 stars) and Pysam Genomic Files (jaechang-hits/SciAgent-Skills, 371 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Alignment Indexing?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.