Pysam
K-Dense-AI/scientific-agent-skills
Provides Python/HTSlib workflows for genomic files. An agent skill from K-Dense-AI/scientific-agent-skills.
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
$ npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-alignment-indexing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/alignment-files/alignment-indexing .claude/skills/bio-alignment-indexing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-alignment-indexing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/alignment-files/alignment-indexing into .claude/skills/bio-alignment-indexing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-alignment-indexing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/alignment-files/alignment-indexingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-alignment-indexing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/alignment-files/alignment-indexing .agents/skills/bio-alignment-indexing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-alignment-indexing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/alignment-files/alignment-indexing into .agents/skills/bio-alignment-indexing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-alignment-indexing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-alignment-indexing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/alignment-files/alignment-indexing .cursor/skills/bio-alignment-indexing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-alignment-indexing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/alignment-files/alignment-indexing into .cursor/skills/bio-alignment-indexing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-alignment-indexing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path alignment-files/alignment-indexing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-alignment-indexing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/alignment-files/alignment-indexing .gemini/skills/bio-alignment-indexing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-alignment-indexing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/alignment-files/alignment-indexing into .gemini/skills/bio-alignment-indexing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-alignment-indexing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-alignment-indexingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/alignment-files/alignment-indexing .github/skills/bio-alignment-indexing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-alignment-indexing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/alignment-files/alignment-indexing into .github/skills/bio-alignment-indexing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-alignment-indexing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-alignment-indexing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/alignment-files/alignment-indexing .opencode/skills/bio-alignment-indexing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-alignment-indexing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/alignment-files/alignment-indexing into .opencode/skills/bio-alignment-indexing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-alignment-indexing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-alignment-indexingCreate and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Bio Alignment Indexing is an agent skill from GPTomics/bioSkills. Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam. Use when enabling random access to alignment files or fetching specific genomic regions.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/fetch_regions.py` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. It works with pysam and Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Alignment Indexing loads about 2.4k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 715 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 715 words, ~2,431 tokens.
.claude/skills/bio-alignment-indexing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: pysam 0.22+, samtools 1.19+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signatures<tool> --version then <tool> --help to confirm flagsIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Create indices for random access to alignment files using samtools and pysam.
"Index a BAM file" -> Create a .bai/.csi index enabling random access to genomic regions.
samtools index file.bampysam.index('file.bam')| Index | Extension | Max contig | Bin shift | When required |
|---|---|---|---|---|
| BAI | .bai / .bam.bai | 2^29 bp ≈ 537 Mbp | fixed (16 kb) | Default for human, mouse, fly, fish |
| CSI | .csi / .bam.csi | 2^(min_shift + depth*3) | configurable via -m | Required for any contig >537 Mbp |
| CRAI | .crai / .cram.crai | chunk-based | n/a | CRAM only |
| TBI | .tbi | 2^29-1 | fixed | tabix VCF/BED -- same limit as BAI |
| Genome | Largest contig | Index |
|---|---|---|
| GRCh38 / GRCh37 (human) | 248 Mbp | BAI |
| GRCm39 (mouse) | 195 Mbp | BAI |
| GRCz11 (zebrafish), TAIR10 (Arabidopsis) | 78 Mbp / 30 Mbp | BAI |
| Wheat IWGSC (Triticum aestivum) | ~830 Mbp (chr3B) | CSI |
| Pine, fir, axolotl, sugar pine | multi-Gbp | CSI with larger -m |
| Long-read assembly with very large contigs | varies | check cut -f2 ref.fa.fai | sort -nr | head -1 |
For polyploid plants and salamander-scale genomes, increase the bin shift:
# Default CSI matches BAI bin layout: 2^(14 + 5*3) = 2^29 bp ≈ 537 Mbp per contig
samtools index -c file.bam
# Larger min_shift for contigs >537 Mbp (wheat, axolotl, sugar pine)
samtools index -c -m 18 file.bam # 2^(18+15) = 2^33 = ~8.5 Gbp per contigIndex file precedence: htslib (and therefore the samtools CLI) tries .csi before .bai on auto-load, so when both exist the .csi is used -- a stale .bai left over after re-indexing to CSI is generally ignored. (Note: htsjdk/Java prefers .bai, the opposite order.) Removing the obsolete .bai still avoids confusion.
samtools index input.bam
# Creates input.bam.baisamtools index -c input.bam
# Creates input.bam.csisamtools index input.bam output.baisamtools index -@ 4 input.bamsamtools index input.cram
# Creates input.cram.craiIndexing requires coordinate-sorted files:
# Check sort order
samtools view -H input.bam | grep "^@HD"
# Should show SO:coordinate
# Sort if needed, then index
samtools sort -o sorted.bam input.bam
samtools index sorted.bamGoal: Extract reads overlapping specific genomic coordinates from an indexed BAM.
Approach: With the index present, samtools view or pysam.fetch() can jump directly to the relevant file offset instead of scanning the entire file.
# Requires index file present
samtools view input.bam chr1:1000000-2000000samtools view input.bam chr1:1000-2000 chr2:3000-4000samtools view -L regions.bed input.bamimport pysam
pysam.index('input.bam')
# Creates input.bam.bai# pysam.index passes through to samtools index; pass the -c flag for CSI.
pysam.index('-c', 'input.bam')
# Produces input.bam.csi.with pysam.AlignmentFile('input.bam', 'rb') as bam:
# fetch() requires index
for read in bam.fetch('chr1', 1000000, 2000000):
print(read.query_name)import pysam
from pathlib import Path
def is_indexed(bam_path):
bam_path = Path(bam_path)
return (bam_path.with_suffix('.bam.bai').exists() or
Path(str(bam_path) + '.bai').exists() or
bam_path.with_suffix('.bam.csi').exists())
if not is_indexed('input.bam'):
pysam.index('input.bam')regions = [('chr1', 1000, 2000), ('chr1', 5000, 6000), ('chr2', 1000, 2000)]
with pysam.AlignmentFile('input.bam', 'rb') as bam:
for chrom, start, end in regions:
count = sum(1 for _ in bam.fetch(chrom, start, end))
print(f'{chrom}:{start}-{end}: {count} reads')with pysam.AlignmentFile('input.bam', 'rb') as bam:
count = bam.count('chr1', 1000000, 2000000)
print(f'Reads in region: {count}')with pysam.AlignmentFile('input.bam', 'rb') as bam:
for read in bam.fetch('chr1', 1000000, 1000001):
if read.reference_start <= 1000000 < read.reference_end:
print(f'{read.query_name} covers position 1000000')samtools looks for indices in two locations:
input.bam.bai # Standard location
input.bai # Alternative locationFor CRAM:
input.cram.craisamtools idxstats input.bamOutput format:
chr1 248956422 5000000 0
chr2 242193529 4500000 0
* 0 0 10000Columns: reference name, length, mapped reads, unmapped reads.
The mapped column counts every alignment record with that RNAME, including secondary AND supplementary. For long-read minimap2 output, where a single read can produce many supplementary chimeric alignments, idxstats overcounts input reads -- typically 1.5-3x.
For unique read counts, use primary-only:
samtools view -c -F 2304 input.bam chr1 # primary onlyCross-check unmapped consistency (a senior sanity check):
samtools idxstats file.bam | awk '{sum+=$4} END {print sum}' # idxstats unmapped (sum across all rows; PE orphans get a contig RNAME)
samtools view -c -f 4 -F 2304 file.bam # primary unmapped (should match)samtools idxstats input.bam | awk '{sum += $3} END {print sum}'with pysam.AlignmentFile('input.bam', 'rb') as bam:
for stat in bam.get_index_statistics():
print(f'{stat.contig}: {stat.mapped} mapped, {stat.unmapped} unmapped')Related but different - index reference FASTA for random access:
samtools faidx reference.fa
# Creates reference.fa.fai
# Fetch region from indexed FASTA
samtools faidx reference.fa chr1:1000-2000with pysam.FastaFile('reference.fa') as ref:
seq = ref.fetch('chr1', 1000, 2000)
print(seq)| Task | samtools | pysam |
|---|---|---|
| Create BAI | samtools index file.bam | pysam.index('file.bam') |
| Create CSI | samtools index -c file.bam | pysam.index('-c', 'file.bam') |
| Fetch region | samtools view file.bam chr1:1-1000 | bam.fetch('chr1', 0, 1000) |
| Count in region | samtools view -c file.bam chr1:1-1000 | bam.count('chr1', 0, 1000) |
| Index stats | samtools idxstats file.bam | bam.get_index_statistics() |
| Index FASTA | samtools faidx ref.fa | Automatic with FastaFile |
If the BAM was modified after indexing, the index points to wrong file offsets and region queries return wrong (or zero) reads. Quick check:
if [ input.bam -nt input.bam.bai ]; then
echo "Index older than BAM; re-indexing"
samtools index input.bam
fiA leading cause of "my variant calling produced empty VCFs" tickets: querying chrM against a BAM that uses MT (or chr1 vs 1). Always inspect contig conventions before region queries:
samtools view -H input.bam | grep '^@SQ' | head -3
# Compare with reference dict:
samtools dict ref.fa | head -3UCSC convention uses chr1/chrM; Ensembl/NCBI uses 1/MT. The two are not interchangeable; tools fail with "contig not found" or silently return zero reads.
| Error | Cause | Solution |
|---|---|---|
random alignment retrieval only works for indexed BAM | Missing index | Run samtools index file.bam |
file is not sorted | Unsorted BAM | Sort first with samtools sort |
chromosome not found | Wrong chromosome name | Check names with samtools view -H |
| Region query returns zero reads on a known-populated locus | Stale BAI / chr vs no-chr mismatch | Re-index; verify naming convention |
| BAI silently truncates reads on contigs >537 Mbp | Plant / amphibian / amplified genome | Use CSI: samtools index -c file.bam |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in alignment-files/alignment-indexing of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Alignment Indexing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Alignment Indexing this skillGPTomics/bioSkills | 1.2k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | |
| PysamK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.4k | Automated safety check: Notes | MIT | |
| Tooluniverse Epigenomicswu-yc/LabClaw | 1.1k | 2 repos | ~14k | Automated safety check: Pass | None | |
| Samtools Bam Processingjaechang-hits/SciAgent-Skills | 371 | 1 repos | ~4.1k | Automated safety check: Pass | MIT | |
| Pysam Genomic Filesjaechang-hits/SciAgent-Skills | 371 | 1 repos | ~5.2k | Automated safety check: Pass | MIT | |
| Bio Splicing QcFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | — | ~1.6k | Automated safety check: Pass | None |
K-Dense-AI/scientific-agent-skills
Provides Python/HTSlib workflows for genomic files. An agent skill from K-Dense-AI/scientific-agent-skills.
wu-yc/LabClaw
Production-ready genomics and epigenomics data processing for BixBench questions.
jaechang-hits/SciAgent-Skills
CLI toolkit for SAM/BAM/CRAM: sort, index, convert, filter, QC alignments.
jaechang-hits/SciAgent-Skills
Read/write SAM/BAM/CRAM, VCF/BCF, FASTA/FASTQ. An agent skill from jaechang-hits/SciAgent-Skills.
FreedomIntelligence/OpenClaw-Medical-Skills
Assesses RNA-seq data quality for splicing analysis including junction saturation curves, splice site strength scoring, and junction coverage metrics using RSeQC.
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Categories
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam. Bio Alignment Indexing is an agent skill from GPTomics/bioSkills. Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Bio Alignment Indexing fits situations like: enabling random access to alignment files; fetching specific genomic regions.
Run `npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a claude-code`. Or copy the skill folder (alignment-files/alignment-indexing in GPTomics/bioSkills) into .claude/skills/bio-alignment-indexing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a codex`. Or copy the skill folder (alignment-files/alignment-indexing in GPTomics/bioSkills) into .agents/skills/bio-alignment-indexing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-alignment-indexing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-alignment-indexing, .gemini/skills/bio-alignment-indexing, .github/skills/bio-alignment-indexing and .opencode/skills/bio-alignment-indexing in your project.
Going by SKILL.md and its folder, Bio Alignment Indexing needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Alignment Indexing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Alignment Indexing: Pysam (K-Dense-AI/scientific-agent-skills, 48k stars), Tooluniverse Epigenomics (wu-yc/LabClaw, 1.1k stars), Samtools Bam Processing (jaechang-hits/SciAgent-Skills, 371 stars) and Pysam Genomic Files (jaechang-hits/SciAgent-Skills, 371 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.