Agent skill

Bio Consensus Sequences

by GPTomics in GPTomics/bioSkills

Generate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar.

MITAuto-check passedResearch & Science

Install Bio Consensus Sequences

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-consensus-sequences --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/variant-calling/consensus-sequences .claude/skills/bio-consensus-sequences && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-consensus-sequences
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4k tokens
SKILL.md length
1,676 words
Files
3
Skills in repo
553
Repo updated
First seen
Licence
MIT

At a glance

Generate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar.

  • Works in 3 steps: A consensus silently emits REFERENCE… → H 1 on an UNPHASED VCF yields a chimeric… → A single FASTA cannot faithfully…
  • Reconstructing a sample-specific reference
  • SKILL.md covers Version Compatibility, The governing principle, Basic Usage and Haplotype Selection and the…, plus 11 more sections
  • Runs Shell scripts from its folder; calls pip

What it does

Bio Consensus Sequences is an agent skill from GPTomics/bioSkills. Generate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar. Use when reconstructing a sample-specific reference or haplotype, deciding -H haplotype vs IUPAC vs all-ALT projection, masking no-coverage sites so a consensus does not manufacture false reference calls, or setting iVar min-depth/min-frequency policy for surveillance genomes.

Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/generate_consensus.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Reconstructing a sample-specific reference
  • Deciding -H haplotype vs IUPAC vs all-ALT projection
  • Masking no-coverage sites so a consensus does not manufacture false reference calls
  • Setting iVar min-depth/min-frequency policy for surveillance genomes

Example prompts

  • “/bio-consensus-sequences”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. A consensus silently emits REFERENCE wherever the VCF is silent -- including positions with zero coverage. No data and…
  2. H 1 on an UNPHASED VCF yields a chimeric pseudo-haplotype. Haplotype selection is only meaningful when genotypes are phased; on unphased…
  3. A single FASTA cannot faithfully represent a diploid genome. Every projection (-H 1, -I, -H A) is lossy in a different way; for…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Consensus Sequences loads about 4k tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 1,676 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,676 words, ~4,009 tokens.

Download SKILL.mdSave it as .claude/skills/bio-consensus-sequences/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-consensus-sequences
description
Generate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar. Use when reconstructing a sample-specific reference or haplotype, deciding -H haplotype vs IUPAC vs all-ALT projection, masking no-coverage sites so a consensus does not manufacture false reference calls, or setting iVar min-depth/min-frequency policy for surveillance genomes.
tool_type
cli
primary_tool
bcftools

Version Compatibility

Reference examples tested with: bcftools 1.19+, samtools 1.19+, bedtools 2.31+, iVar 1.4+, minimap2 2.26+, BioPython 1.83+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Note: the -H argument vocabulary (N/R/A/I/LR/LA/SR/SA/NpIu, where I = IUPAC code for all genotypes) has grown across bcftools 1.x releases; IUPAC output is available both as the -H I code and the standalone -I/--iupac-codes flag. Always confirm the accepted letters with bcftools consensus on the installed version before scripting a selection.

Consensus Sequences

"Generate a consensus sequence from my VCF" -> Apply called variants onto a reference FASTA, producing a sample-specific sequence, with a deliberate choice of haplotype projection and no-coverage masking.

  • CLI (from a VCF): bcftools consensus -f reference.fa input.vcf.gz
  • CLI (viral/amplicon from a BAM): samtools mpileup ... | ivar consensus -p out
  • Python: cyvcf2 + Bio.SeqIO for SNP-only prototypes

The governing principle

bcftools consensus walks the reference and substitutes ALT alleles at the positions present in the VCF. Everything else is copied from the reference verbatim -- which drives three traps that ruin more consensus analyses than any tool bug:

  1. A consensus silently emits REFERENCE wherever the VCF is silent -- including positions with zero coverage. No data and confidently-reference look identical in the output. An unmasked consensus therefore manufactures false confidence at exactly the sites where the sample was never observed. Mask no-coverage sites (below) or the FASTA lies.
  2. -H 1 on an UNPHASED VCF yields a chimeric pseudo-haplotype. Haplotype selection is only meaningful when genotypes are phased; on unphased data it mixes alleles from different real chromosomes into a sequence that exists in no cell. Verify | phasing before selecting a haplotype.
  3. A single FASTA cannot faithfully represent a diploid genome. Every projection (-H 1, -I, -H A) is lossy in a different way; for phase-sensitive work keep the VCF, not the consensus.

The input VCF must be bgzipped and indexed (bgzip + bcftools index/tabix); plain-gzip or unindexed input errors out. The REF bases in the VCF must match the FASTA exactly or bcftools warns and skips those records. Normalize first (see Normalization).

Basic Usage

bcftools consensus reads variants from a bgzipped, indexed VCF and writes FASTA:

bash
bcftools index input.vcf.gz                                    # .csi index (or tabix -p vcf)
bcftools consensus -f reference.fa input.vcf.gz > consensus.fa
bcftools consensus -f reference.fa -o consensus.fa input.vcf.gz  # -o instead of redirect

For a multi-sample VCF, always pass -s -- without it, the applied genotypes are undefined:

bash
bcftools query -l input.vcf.gz                                 # list samples
bcftools consensus -f reference.fa -s sample1 input.vcf.gz > sample1.fa

Restrict to a region with -r (the FASTA header is then >chr:from-to):

bash
bcftools consensus -f reference.fa -r chr1:1000000-1010000 -s sample1 input.vcf.gz > gene.fa

Haplotype Selection and the Phasing Trap

-H chooses which allele to apply from FORMAT/GT. The codes are case-insensitive:

OptionAppliesUse when
-H 1 / -H 2Allele at GT index 1 or 2Emitting one true chromosome -- only valid on PHASED genotypes
-H AALT allele in every genotypeMaximum divergence from reference; a chimera of both chromosomes
-H RREF allele at heterozygous sitesConservative consensus; discards het ALT alleles
-H I (or the standalone -I / --iupac-codes flag)IUPAC ambiguity codeRetain heterozygosity in one sequence (see caveat below)
-H LA/LR/SA/SRLonger/shorter allele, tie broken by ALT/REFLength-driven selection; confirm the letter set on the installed version

The chimeric-haplotype footgun. -H 1/-H 2 are only meaningful when genotypes are phased (0|1, pipe separator). With unphased genotypes (0/1, slash), the assignment of "which allele is haplotype 1" is arbitrary per site, so -H 1 across many heterozygous sites produces a switch-error mosaic that corresponds to no real chromosome -- while looking like a clean haplotype FASTA. This is the single most dangerous consensus mistake. Verify phasing before any -H 1/-H 2:

bash
bcftools query -f '%CHROM\t%POS[\t%GT]\n' input.vcf.gz | head   # phased: 0|1 ; unphased: 0/1

If genotypes are unphased, phase first (read-backed WhatsHap/HapCUT2, trio, statistical SHAPEIT/Eagle -- accurate for common variants, poor for rare/singletons -- or native long-read phasing). See phasing-imputation/haplotype-phasing and variant-calling/vcf-basics for GT interpretation.

What a Consensus Cannot Represent

A single consensus FASTA is a lossy projection of a diploid genome; the right projection depends on the downstream use, and some tasks need the VCF instead:

StrategyFlagBest forLoses
Two haplotype sequences-H 1 + -H 2 (phased)Allele-specific expression, compound-het, HLA, cis-regulatory haplotypesNothing (if correctly phased)
IUPAC ambiguity codes-IRetaining het signal in one sequencePhase/linkage; many tree/alignment tools read IUPAC as N
All ALT alleles-H AMax divergence, quick draftReality -- exists in no cell
REF at het sites-H RConservative single sequenceEvery heterozygous ALT allele

Two hard boundaries:

  • For phase-sensitive work, keep the VCF (or two phased haplotype FASTAs), not a single consensus. Collapsing hets to IUPAC or picking one allele discards linkage that the analysis needs -- treating a consensus FASTA as "the sample's genome" for compound-het or allele-specific analysis is a category error.
  • bcftools consensus cannot apply symbolic SV alleles (<DEL>, <INS>, <DUP>, <INV>): those carry no ALT sequence, only INFO fields, so consensus has nothing to substitute. Short-read SV VCFs (Manta/DELLY) are mostly symbolic and are NOT directly consensus-able. Folding SVs into a consensus needs sequence-resolved records (long-read/assembly callers emit these) or an assembly-based approach -- see variant-calling/structural-variant-calling.

For phylogenetics specifically, prefer one clean phased haplotype or a homozygous-ALT-only sequence over IUPAC, because ambiguity codes are silently dropped by many tree builders:

bash
bcftools view -i 'GT="AA"' input.vcf.gz | bcftools consensus -f reference.fa > hom_alt.fa

Masking No-Coverage Sites (the load-bearing footgun)

Because unobserved positions are emitted as reference (trap 1), a consensus must mask sites with insufficient data. -m mask.bed replaces the listed regions (default char N via --mask-with N). The mask must be built from callable depth, and the depth step hides a silent bug:

samtools depth WITHOUT -a OMITS zero-coverage positions from its output -- so those positions never enter the low-depth BED, never get masked, and stay as reference: the exact false-confidence failure the mask was meant to prevent. Always use -a (report all positions) so no-coverage sites are captured:

bash
# Build a mask of every position below the callable-depth threshold. -a is mandatory:
# without it, zero-coverage positions are absent from the output and escape masking.
samtools depth -a aligned.bam | awk '$3 < 10 {print $1"\t"$2-1"\t"$2}' | bedtools merge > lowcov.bed

bcftools consensus -f reference.fa -m lowcov.bed input.vcf.gz > consensus.fa

The < 10 threshold is a minimum-callable-depth policy (10x is a common floor for confident base calls); set it to the depth below which the calls are not trusted. bedtools genomecov -bga -ibam aligned.bam is an equivalent zero-coverage-aware alternative that also emits 0-depth intervals.

Do NOT rely on -M/-a for this: -M N outputs N only for missing ./. genotypes already present in the VCF, and -a N replaces every position absent from the VCF (which N-outs the entire non-variant genome). Neither distinguishes no-coverage from confident-reference -- only a depth-derived mask does.

Show full SKILL.md (632 more words)Show less

Normalization Before Consensus

Goal: Apply indels at the correct reference position and sequence.

Approach: Left-align and split multiallelics with bcftools norm so each record matches the reference context; consensus applies records positionally and mis-represented indels corrupt the output.

bash
bcftools norm -f reference.fa input.vcf.gz -Oz -o norm.vcf.gz
bcftools index norm.vcf.gz
bcftools consensus -f reference.fa norm.vcf.gz > consensus.fa

Un-normalized or overlapping indels produce wrong sequence, and bcftools consensus only warns to stderr while still emitting output -- so the corruption is silent unless the stderr is inspected. Even after norm, two records whose REF spans collide remain a hazard; grep the run for warnings and inspect the region. See variant-calling/variant-normalization.

bash
bcftools consensus -f reference.fa norm.vcf.gz 2>&1 >consensus.fa | grep -i 'overlap\|warn'

Viral / Amplicon Consensus with iVar

For amplicon surveillance (SARS-CoV-2 and similar), ivar consensus builds a per-sample consensus directly from a pileup. Its two key thresholds are epidemiological policy decisions, not defaults to accept blindly -- they propagate into lineage assignment and transmission-cluster inference:

bash
# Trim PCR primers FIRST -- primer-derived bases are not sample sequence and, at
# primer-binding-site mutations, cause reference-biased miscalls if left in.
ivar trim -b primers.bed -p trimmed -i aligned.bam
samtools sort -o trimmed.sorted.bam trimmed.bam

# -aa keeps all positions (so no-coverage becomes N), -A keeps orphan mates, -d 0 lifts the depth cap.
samtools mpileup -aa -A -d 0 -B -Q 0 trimmed.sorted.bam | ivar consensus -p sample -q 20 -t 0.5 -m 10 -n N
FlagDefaultDecision
-m min depth10Below this, iVar emits N. Too low -> single-read sequencing errors become "mutations" that corrupt outbreak phylogenies. Too high -> excessive Ns, an unusably fragmented genome.
-t min frequency to call a base0 (majority)0 calls the most common base. For a strict majority consensus use 0.5. Too low bakes minority/within-host variants and contamination into the "genome", inflating diversity and creating phantom transmission links. Raise (e.g. 0.03) only deliberately for intrahost variant work, not for a reference consensus.
-q min base quality20Bases below this are not counted toward depth/frequency.
-n no-coverage charNCharacter emitted where depth < -m.

Always report -m and -t alongside a surveillance consensus -- the genome is only as trustworthy as those two numbers. Alternatives: bcftools consensus from a called VCF, or ViralConsensus (Moshiri 2023) which calls consensus directly from the alignment without an intermediate VCF, faster and lower-memory for large batches.

Filtering Before Consensus

Apply only trusted calls; pipe filtered VCF straight into consensus:

bash
bcftools view -f PASS input.vcf.gz -Oz -o pass.vcf.gz && bcftools index pass.vcf.gz
bcftools consensus -f reference.fa pass.vcf.gz > consensus.fa

bcftools view -v snps input.vcf.gz -Oz -o snps.vcf.gz && bcftools index snps.vcf.gz  # SNPs only

Filtered VCFs must be re-bgzipped and re-indexed before bcftools consensus reads them.

Chain Files and Naming

-c chain.txt writes a liftover chain mapping reference coordinates to consensus coordinates -- needed when indels shift positions and annotations must be lifted. -p PREFIX prepends a string to output sequence names (>sample1_chr1).

bash
bcftools consensus -f reference.fa -c chain.txt -p "sample1_" input.vcf.gz > consensus.fa

cyvcf2 Consensus (SNP-only prototypes)

For a quick SNP-only substitution in Python (production work should use bcftools consensus, which handles indels, phasing, and masking):

python
from cyvcf2 import VCF
from Bio import SeqIO

ref = {rec.id: list(str(rec.seq)) for rec in SeqIO.parse('reference.fa', 'fasta')}
for v in VCF('input.vcf.gz'):
    if v.is_snp and len(v.ALT) == 1:
        ref[v.CHROM][v.POS - 1] = v.ALT[0]   # POS is 1-based; list index is 0-based
with open('consensus.fa', 'w') as fh:
    for chrom, seq in ref.items():
        fh.write(f'>{chrom}\n{"".join(seq)}\n')

Verify the Consensus

bash
minimap2 -a reference.fa consensus.fa | samtools view -b -o aln.bam   # inspect where it diverges
bcftools view -H input.vcf.gz | wc -l                                 # variants available to apply

Common Errors

Error / SymptomCauseFix
the VCF file is not indexedPlain-gzip or missing indexbgzip then bcftools index (or tabix -p vcf)
sequence "chr1" not foundChromosome names differ between FASTA and VCFbcftools annotate --rename-chrs map.txt
REF does not matchDifferent reference than the caller usedUse the exact FASTA used for calling; normalize
Clean haplotype looks wrong-H 1 on an unphased VCF -> chimeraVerify `
Consensus reference-identical over gapsNo-coverage sites emitted as referenceMask with samtools depth -a derived BED and -m
Garbled indels, stderr overlap warningsUn-normalized/overlapping recordsbcftools norm -f ref.fa first; inspect warnings
<DEL>/<INS> not appliedSymbolic SV alleles carry no ALT sequenceUse sequence-resolved SV records; see structural-variant-calling
  • variant-calling/variant-calling - Generate the VCF consensus is built from
  • variant-calling/vcf-basics - Interpret GT and phasing (| vs /) before -H
  • variant-calling/variant-normalization - Left-align indels before consensus
  • variant-calling/filtering-best-practices - Restrict to trusted calls first
  • variant-calling/structural-variant-calling - Sequence-resolved SVs for SV-aware consensus
  • phasing-imputation/haplotype-phasing - Produce phased genotypes for true haplotypes
  • phylogenetics/modern-tree-inference - Build trees from a consensus alignment

References

  • Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, et al. Twelve years of SAMtools and BCFtools. GigaScience. 2021;10(2):giab008. doi:10.1093/gigascience/giab008. (bcftools consensus / norm / mpileup.)
  • Grubaugh ND, Gangavarapu K, Quick J, Matteson NL, De Jesus JG, Main BJ, et al. An amplicon-based sequencing framework for accurately measuring intrahost virus diversity using PrimalSeq and iVar. Genome Biology. 2019;20(1):8. doi:10.1186/s13059-018-1618-7. (iVar consensus/trim; depth -m and frequency -t thresholds.)
  • Moshiri N. ViralConsensus: a fast and memory-efficient tool for calling viral consensus genome sequences directly from read alignment data. Bioinformatics. 2023;39(5):btad317. doi:10.1093/bioinformatics/btad317.

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in variant-calling/consensus-sequences of GPTomics/bioSkills.

  • SKILL.md
  • examples/generate_consensus.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Consensus Sequences next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Consensus Sequences compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Consensus Sequences this skillGPTomics/bioSkills1.2k1 repos~4kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 553 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Consensus Sequences

What does Bio Consensus Sequences do?

Generate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar. Bio Consensus Sequences is an agent skill from GPTomics/bioSkills. Generate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar.

When should I use Bio Consensus Sequences?

Bio Consensus Sequences fits situations like: reconstructing a sample-specific reference; deciding -H haplotype vs IUPAC vs all-ALT projection; masking no-coverage sites so a consensus does not manufacture false reference calls; setting iVar min-depth/min-frequency policy for surveillance genomes.

How do I install Bio Consensus Sequences in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a claude-code`. Or copy the skill folder (variant-calling/consensus-sequences in GPTomics/bioSkills) into .claude/skills/bio-consensus-sequences in your project. Claude Code loads it when a task matches its description.

How do I install Bio Consensus Sequences in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a codex`. Or copy the skill folder (variant-calling/consensus-sequences in GPTomics/bioSkills) into .agents/skills/bio-consensus-sequences in your project. Codex loads it when a task matches its description.

Can I use Bio Consensus Sequences in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-consensus-sequences, .gemini/skills/bio-consensus-sequences, .github/skills/bio-consensus-sequences and .opencode/skills/bio-consensus-sequences in your project.

What does Bio Consensus Sequences need to run?

Going by SKILL.md and its folder, Bio Consensus Sequences needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: A Bash shell.

Does Bio Consensus Sequences access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Consensus Sequences safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Consensus Sequences use?

Bio Consensus Sequences is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Consensus Sequences use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Consensus Sequences?

Skills that share tags, products or a category with Bio Consensus Sequences: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Consensus Sequences?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.