Agent skill

Bio Reference Operations

by GPTomics in GPTomics/bioSkills

Generate consensus sequences and manage reference files using samtools.

MITAuto-check passedResearch & Science

Install Bio Reference Operations

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-reference-operations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-reference-operations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/alignment-files/reference-operations .claude/skills/bio-reference-operations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-reference-operations
GitHub stars
1.2k
Used in
2 other repos
Token cost
~3.1k tokens
SKILL.md length
735 words
Files
3
Skills in repo
553
Repo updated
First seen
Licence
MIT

At a glance

Generate consensus sequences and manage reference files using samtools.

  • Creating consensus from alignments
  • SKILL.md covers Version Compatibility, samtools faidx - Index…, samtools dict - Create… and samtools consensus - Generate…, plus 1 more section
  • Runs Shell scripts from its folder; calls pip
  • Indexing references

What it does

Bio Reference Operations is an agent skill from GPTomics/bioSkills. Generate consensus sequences and manage reference files using samtools. Use when creating consensus from alignments, indexing references, or creating sequence dictionaries.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/prepare_reference.sh` and `usage-guide.md`).

It sits in Research & Science. It works with Python and pysam. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Creating consensus from alignments
  • Indexing references
  • Creating sequence dictionaries

Example prompts

  • “/bio-reference-operations”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Reference Operations loads about 3.1k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 735 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 735 words, ~3,095 tokens.

Download SKILL.mdSave it as .claude/skills/bio-reference-operations/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-reference-operations
description
Generate consensus sequences and manage reference files using samtools. Use when creating consensus from alignments, indexing references, or creating sequence dictionaries.
tool_type
cli
primary_tool
samtools

Version Compatibility

Reference examples tested with: GATK 4.5+, bcftools 1.19+, pysam 0.22+, samtools 1.19+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Reference Operations

Generate consensus sequences and manage reference files using samtools.

"Prepare a reference genome" -> Index the FASTA and create a sequence dictionary for downstream tools.

  • CLI: samtools faidx ref.fa + samtools dict ref.fa -o ref.dict
  • Python: pysam.FastaFile('ref.fa') (auto-uses .fai index)

"Build a consensus from BAM" -> Derive the most-supported base at each position from aligned reads.

  • CLI: samtools consensus input.bam -o consensus.fa
  • Python: iterate pileup columns and take majority base (pysam)

samtools faidx - Index Reference FASTA

Create index for random access to reference sequences.

Create Index
bash
samtools faidx reference.fa
# Creates reference.fa.fai
Fetch Region from Reference
bash
samtools faidx reference.fa chr1:1000-2000
Fetch Multiple Regions
bash
samtools faidx reference.fa chr1:1000-2000 chr2:3000-4000
Fetch Entire Chromosome
bash
samtools faidx reference.fa chr1
Output to File
bash
samtools faidx reference.fa chr1:1000-2000 > region.fa
Reverse Complement
bash
samtools faidx -i reference.fa chr1:1000-2000
FAI File Format
chr1    248956422    6    60    61
chr2    242193529    253105708    60    61

Columns: name, length, offset, line bases, line width

samtools dict - Create Sequence Dictionary

Create SAM header dictionary for reference (used by GATK, Picard).

Create Dictionary
bash
samtools dict reference.fa -o reference.dict
With Assembly Info
bash
samtools dict -a GRCh38 -s "Homo sapiens" reference.fa -o reference.dict
Dictionary Format
@HD VN:1.0 SO:unsorted
@SQ SN:chr1 LN:248956422 M5:6aef897c3d6ff0c78aff06ac189178dd UR:file:reference.fa
@SQ SN:chr2 LN:242193529 M5:f98db672eb0993dcfdabafe2a882905c UR:file:reference.fa

The M5: (MD5) tag is the only definitive reference-identity check -- two references named "GRCh38" with different decoy/alt content have different M5s. CRAM enforces M5 match on read-back. See alignment-validation for BAM-vs-reference M5 cross-check.

GRCh38 Is Not One Reference
Reference flavorALTDecoyEBVHLAUse case
GRCh38 no-altnonononoConservative analyses
GRCh38 + decoy + EBV (1000G analysis set)noyesyesnoCohort projects
GRCh38 ALT + decoy + EBV + HLA (Broad / hs38DH)yesyesyesyesGATK Best Practices
T2T-CHM13 v2.0n/an/an/an/aDistinct coordinates -- NOT interchangeable

Mixing no-alt and ALT-aware BAMs in one cohort produces inconsistent multi-mapping behavior at HLA, KIR, and segmental-duplication regions. Standardize before joint calling.

Contig Naming: The Silent Killer
ConventionSourcechr1mitochondrion
UCSC (hg19, hg38)UCSC Genome Browserchr1chrM
Ensembl (GRCh37, GRCh38)Ensembl, ENA1MT
NCBI RefSeq (recent)NCBIchr1chrM
1000G analysis sets1000G GRCh38 analysis setchr1chrM

A BAM with @SQ SN:chr1 cannot be analyzed against a 1-named reference (and vice versa). Detect:

bash
samtools view -H sample.bam | grep '^@SQ' | head -3
samtools dict ref.fa | head -3

Convert: bcftools annotate --rename-chrs for VCF; for BAM there is no clean conversion -- re-align.

samtools consensus - Generate Consensus

Create consensus sequence from alignments.

Basic Consensus
bash
samtools consensus input.bam -o consensus.fa
From Specific Region
bash
samtools consensus -r chr1:1000-2000 input.bam -o region_consensus.fa
Output Formats
bash
# FASTA (default)
samtools consensus -f fasta input.bam -o consensus.fa

# FASTQ (includes quality)
samtools consensus -f fastq input.bam -o consensus.fq
Quality Options
bash
# Minimum depth to call base
samtools consensus -d 5 input.bam -o consensus.fa

# Call all positions (including low coverage)
samtools consensus -a input.bam -o consensus.fa
IUPAC Ambiguity for Heterozygotes
bash
# Emit IUPAC codes (R, Y, S, W, K, M, B, D, H, V, N) for heterozygous columns
# --ambig is REQUIRED -- without it, output is restricted to A,C,G,T,N,*
samtools consensus --ambig --het-fract 0.2 --call-fract 0.5 input.bam -o consensus.fa

--het-fract controls the fraction of the second-most-common base relative to the most common required to call a heterozygote (verify the default for the installed version with samtools consensus --help; the manpage documents none). Without --ambig, columns where the second base passes --het-fract resolve to N rather than the IUPAC code. --show-ins / --show-del control insertion / deletion display, not ambiguity.

Show full SKILL.md (275 more words)Show less
Platform-Aware Consensus
bash
# Default: Bayesian algorithm (no --config needed)
samtools consensus -f fasta input.bam -o consensus.fa

# Platform-specific profiles (samtools 1.17+; verify via samtools consensus --help for installed version)
samtools consensus --config hifi       input.bam -o consensus.fa   # PacBio HiFi
samtools consensus --config r10.4_sup  input.bam -o consensus.fa   # ONT R10.4+ (r10.4_dup for duplex)
samtools consensus --config ultima     input.bam -o consensus.fa   # Ultima Genomics
samtools consensus --config hiseq      input.bam -o consensus.fa   # Illumina

# Report ref base where consensus unavailable (low coverage; -T added in samtools 1.22)
samtools consensus -T ref.fa input.bam -o consensus.fa
samtools consensus vs bcftools consensus

Different operations -- conflating them produces nonsense:

ToolInputOutputUse case
samtools consensusBAMConsensus FASTA derived from reads (Bayesian)Viral, de novo / amplicon, low-coverage species
bcftools consensusreference + VCFReference with VCF variants appliedApply called variants (haplotype reconstruction, custom ref for re-mapping)

For viral consensus from BAM:

bash
# Modern: samtools consensus
samtools consensus --config hiseq -d 10 --het-fract 0.5 \
    --show-ins yes --show-del yes input.bam -o consensus.fa

# Apply called variants to reference (different question)
bcftools consensus -f reference.fa variants.vcf.gz -o sample_consensus.fa
bcftools consensus -f reference.fa -H 1 phased.vcf.gz -o haplotype1.fa   # phased haplotype 1

For bacterial / phage assembly polishing, prefer Pilon (short-read) or medaka (ONT); samtools consensus is not iterative.

pysam Python Alternative

Fetch from Indexed FASTA
python
import pysam

with pysam.FastaFile('reference.fa') as ref:
    seq = ref.fetch('chr1', 999, 2000)  # 0-based
    print(seq)
Get Reference Lengths
python
with pysam.FastaFile('reference.fa') as ref:
    for name in ref.references:
        length = ref.get_reference_length(name)
        print(f'{name}: {length:,} bp')
Fetch All Chromosomes
python
with pysam.FastaFile('reference.fa') as ref:
    for chrom in ref.references:
        seq = ref.fetch(chrom)
        print(f'>{chrom}')
        print(seq[:100] + '...')
Generate Simple Consensus
python
import pysam
from collections import Counter

def consensus_at_position(bam, chrom, pos):
    bases = Counter()
    for pileup in bam.pileup(chrom, pos, pos + 1, truncate=True):
        if pileup.pos == pos:
            for read in pileup.pileups:
                if not read.is_del and not read.is_refskip:
                    bases[read.alignment.query_sequence[read.query_position]] += 1
    if bases:
        return bases.most_common(1)[0][0]
    return 'N'

with pysam.AlignmentFile('input.bam', 'rb') as bam:
    consensus = consensus_at_position(bam, 'chr1', 1000000)
    print(f'Consensus at chr1:1000000 = {consensus}')
Build Consensus Sequence (Pedagogical Only)

The Python majority-vote consensus below is illustrative, NOT production. samtools consensus is Bayesian, quality-aware, and platform-aware; majority vote ignores base qualities and produces wrong calls on low-coverage / low-quality regions. Use for teaching pileup iteration mechanics; use samtools consensus for any real consensus.

python
import pysam
from collections import Counter

def build_consensus(bam_path, chrom, start, end, min_depth=3):
    consensus = []

    with pysam.AlignmentFile(bam_path, 'rb') as bam:
        for pileup in bam.pileup(chrom, start, end, truncate=True):
            bases = Counter()
            for read in pileup.pileups:
                if not read.is_del and not read.is_refskip:
                    base = read.alignment.query_sequence[read.query_position]
                    bases[base] += 1

            if sum(bases.values()) >= min_depth:
                consensus.append(bases.most_common(1)[0][0])
            else:
                consensus.append('N')

    return ''.join(consensus)
Create Dictionary Header
python
import pysam

def create_dict_header(fasta_path):
    header = {'HD': {'VN': '1.6', 'SO': 'unsorted'}, 'SQ': []}

    with pysam.FastaFile(fasta_path) as ref:
        for name in ref.references:
            length = ref.get_reference_length(name)
            header['SQ'].append({'SN': name, 'LN': length})

    return header

header = create_dict_header('reference.fa')
for sq in header['SQ'][:5]:
    print(f'{sq["SN"]}: {sq["LN"]:,} bp')

Reference Preparation Workflow

Goal: Set up a reference genome with all indices needed by common analysis tools.

Approach: Create FASTA index (.fai), sequence dictionary (.dict), and aligner-specific indices in sequence.

Prepare Reference for Analysis
bash
# 1. Index FASTA for samtools/pysam
samtools faidx reference.fa

# 2. Create sequence dictionary for GATK/Picard
samtools dict reference.fa -o reference.dict

# 3. Pre-populate CRAM REF_CACHE (for offline HPC nodes)
seq_cache_populate.pl -root $REF_CACHE_DIR reference.fa

For aligner-specific indices (BWA, Bowtie2, STAR, minimap2, Salmon), see read-alignment.

Check Reference Setup
bash
# Verify FAI exists
ls -la reference.fa.fai

# Verify dict exists
head reference.dict

# Test fetch
samtools faidx reference.fa chr1:1-100

Common Operations

Extract Chromosome
bash
samtools faidx reference.fa chr1 > chr1.fa
samtools faidx chr1.fa  # Index the subset
Get Chromosome Sizes
bash
cut -f1,2 reference.fa.fai > chrom.sizes
Subset Reference
bash
samtools faidx reference.fa chr1 chr2 chr3 > subset.fa
samtools faidx subset.fa
Compare Consensus to Reference
bash
# Generate consensus
samtools consensus input.bam -o consensus.fa

# Align consensus back to reference
minimap2 -a reference.fa consensus.fa > comparison.sam

Quick Reference

TaskCommand
Index FASTAsamtools faidx ref.fa
Fetch regionsamtools faidx ref.fa chr1:1-1000
Create dictsamtools dict ref.fa -o ref.dict
Build consensussamtools consensus in.bam -o out.fa
Chrom sizescut -f1,2 ref.fa.fai
  • sam-bam-basics - CRAM reference resolution (REF_PATH, REF_CACHE)
  • alignment-indexing - faidx for reference access
  • alignment-validation - BAM-vs-reference M5 cross-validation
  • pileup-generation - Pileup for consensus building
  • variant-calling/vcf-basics - VCF I/O for bcftools consensus
  • variant-calling/consensus-sequences - Consensus from VCF (different operation)
  • read-alignment/bwa-alignment - BWA index preparation
  • sequence-io/read-sequences - Parse FASTA with Biopython

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in alignment-files/reference-operations of GPTomics/bioSkills.

  • SKILL.md
  • examples/prepare_reference.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Reference Operations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Reference Operations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Reference Operations this skillGPTomics/bioSkills1.2k2 repos~3.1kAutomated safety check: PassMIT
PysamK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: NotesMIT
Tooluniverse Epigenomicswu-yc/LabClaw1.1k2 repos~14kAutomated safety check: PassNone
Samtools Bam Processingjaechang-hits/SciAgent-Skills3701 repos~4.1kAutomated safety check: PassMIT
Pysam Genomic Filesjaechang-hits/SciAgent-Skills3701 repos~5.2kAutomated safety check: PassMIT
Bio Splicing QcFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~1.6kAutomated safety check: PassNone

Similar skills

  • Pysam

    K-Dense-AI/scientific-agent-skills

    Provides Python/HTSlib workflows for genomic files. An agent skill from K-Dense-AI/scientific-agent-skills.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Research & ScienceAuto-check: notes
  • Production-ready genomics and epigenomics data processing for BixBench questions.

    1.1k GitHub starsUsed in 2 repos~14k tokens
    Research & ScienceAuto-check passed
  • Samtools Bam Processing

    jaechang-hits/SciAgent-Skills

    CLI toolkit for SAM/BAM/CRAM: sort, index, convert, filter, QC alignments.

    370 GitHub starsUsed in 1 repo~4.1k tokens
    Research & ScienceAuto-check passed
  • Pysam Genomic Files

    jaechang-hits/SciAgent-Skills

    Read/write SAM/BAM/CRAM, VCF/BCF, FASTA/FASTQ. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 1 repo~5.2k tokens
    Research & ScienceAuto-check passed
  • Bio Splicing Qc

    FreedomIntelligence/OpenClaw-Medical-Skills

    Assesses RNA-seq data quality for splicing analysis including junction saturation curves, splice site strength scoring, and junction coverage metrics using RSeQC.

    3.1k GitHub stars~1.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 553 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Works with

Questions about Bio Reference Operations

What does Bio Reference Operations do?

Generate consensus sequences and manage reference files using samtools. Bio Reference Operations is an agent skill from GPTomics/bioSkills. Generate consensus sequences and manage reference files using samtools.

When should I use Bio Reference Operations?

Bio Reference Operations fits situations like: creating consensus from alignments; indexing references; creating sequence dictionaries.

How do I install Bio Reference Operations in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-reference-operations -a claude-code`. Or copy the skill folder (alignment-files/reference-operations in GPTomics/bioSkills) into .claude/skills/bio-reference-operations in your project. Claude Code loads it when a task matches its description.

How do I install Bio Reference Operations in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-reference-operations -a codex`. Or copy the skill folder (alignment-files/reference-operations in GPTomics/bioSkills) into .agents/skills/bio-reference-operations in your project. Codex loads it when a task matches its description.

Can I use Bio Reference Operations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-reference-operations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-reference-operations, .gemini/skills/bio-reference-operations, .github/skills/bio-reference-operations and .opencode/skills/bio-reference-operations in your project.

What does Bio Reference Operations need to run?

Going by SKILL.md and its folder, Bio Reference Operations needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Reference Operations access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Reference Operations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Reference Operations use?

Bio Reference Operations is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Reference Operations use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Reference Operations?

Skills that share tags, products or a category with Bio Reference Operations: Pysam (K-Dense-AI/scientific-agent-skills, 48k stars), Tooluniverse Epigenomics (wu-yc/LabClaw, 1.1k stars), Samtools Bam Processing (jaechang-hits/SciAgent-Skills, 370 stars) and Pysam Genomic Files (jaechang-hits/SciAgent-Skills, 370 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Reference Operations?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.