Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Generate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar.
$ npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-consensus-sequences --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/variant-calling/consensus-sequences .claude/skills/bio-consensus-sequences && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-consensus-sequences" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/consensus-sequences into .claude/skills/bio-consensus-sequences/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-consensus-sequences", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/variant-calling/consensus-sequencesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-consensus-sequences --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/variant-calling/consensus-sequences .agents/skills/bio-consensus-sequences && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-consensus-sequences" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/consensus-sequences into .agents/skills/bio-consensus-sequences/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-consensus-sequences", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-consensus-sequences --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/variant-calling/consensus-sequences .cursor/skills/bio-consensus-sequences && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-consensus-sequences" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/consensus-sequences into .cursor/skills/bio-consensus-sequences/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-consensus-sequences", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path variant-calling/consensus-sequences--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-consensus-sequences --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/variant-calling/consensus-sequences .gemini/skills/bio-consensus-sequences && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-consensus-sequences" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/consensus-sequences into .gemini/skills/bio-consensus-sequences/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-consensus-sequences", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-consensus-sequencesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/variant-calling/consensus-sequences .github/skills/bio-consensus-sequences && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-consensus-sequences" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/consensus-sequences into .github/skills/bio-consensus-sequences/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-consensus-sequences", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-consensus-sequences --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/variant-calling/consensus-sequences .opencode/skills/bio-consensus-sequences && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-consensus-sequences" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/consensus-sequences into .opencode/skills/bio-consensus-sequences/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-consensus-sequences", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-consensus-sequencesGenerate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar.
Bio Consensus Sequences is an agent skill from GPTomics/bioSkills. Generate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar. Use when reconstructing a sample-specific reference or haplotype, deciding -H haplotype vs IUPAC vs all-ALT projection, masking no-coverage sites so a consensus does not manufacture false reference calls, or setting iVar min-depth/min-frequency policy for surveillance genomes.
Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/generate_consensus.sh` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Consensus Sequences loads about 4k tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 1,676 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,676 words, ~4,009 tokens.
.claude/skills/bio-consensus-sequences/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: bcftools 1.19+, samtools 1.19+, bedtools 2.31+, iVar 1.4+, minimap2 2.26+, BioPython 1.83+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signatures<tool> --version then <tool> --help to confirm flagsIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Note: the -H argument vocabulary (N/R/A/I/LR/LA/SR/SA/NpIu, where I = IUPAC code for all genotypes) has grown across bcftools 1.x releases; IUPAC output is available both as the -H I code and the standalone -I/--iupac-codes flag. Always confirm the accepted letters with bcftools consensus on the installed version before scripting a selection.
"Generate a consensus sequence from my VCF" -> Apply called variants onto a reference FASTA, producing a sample-specific sequence, with a deliberate choice of haplotype projection and no-coverage masking.
bcftools consensus -f reference.fa input.vcf.gzsamtools mpileup ... | ivar consensus -p outcyvcf2 + Bio.SeqIO for SNP-only prototypesbcftools consensus walks the reference and substitutes ALT alleles at the positions present in the VCF. Everything else is copied from the reference verbatim -- which drives three traps that ruin more consensus analyses than any tool bug:
-H 1 on an UNPHASED VCF yields a chimeric pseudo-haplotype. Haplotype selection is only meaningful when genotypes are phased; on unphased data it mixes alleles from different real chromosomes into a sequence that exists in no cell. Verify | phasing before selecting a haplotype.-H 1, -I, -H A) is lossy in a different way; for phase-sensitive work keep the VCF, not the consensus.The input VCF must be bgzipped and indexed (bgzip + bcftools index/tabix); plain-gzip or unindexed input errors out. The REF bases in the VCF must match the FASTA exactly or bcftools warns and skips those records. Normalize first (see Normalization).
bcftools consensus reads variants from a bgzipped, indexed VCF and writes FASTA:
bcftools index input.vcf.gz # .csi index (or tabix -p vcf)
bcftools consensus -f reference.fa input.vcf.gz > consensus.fa
bcftools consensus -f reference.fa -o consensus.fa input.vcf.gz # -o instead of redirectFor a multi-sample VCF, always pass -s -- without it, the applied genotypes are undefined:
bcftools query -l input.vcf.gz # list samples
bcftools consensus -f reference.fa -s sample1 input.vcf.gz > sample1.faRestrict to a region with -r (the FASTA header is then >chr:from-to):
bcftools consensus -f reference.fa -r chr1:1000000-1010000 -s sample1 input.vcf.gz > gene.fa-H chooses which allele to apply from FORMAT/GT. The codes are case-insensitive:
| Option | Applies | Use when |
|---|---|---|
-H 1 / -H 2 | Allele at GT index 1 or 2 | Emitting one true chromosome -- only valid on PHASED genotypes |
-H A | ALT allele in every genotype | Maximum divergence from reference; a chimera of both chromosomes |
-H R | REF allele at heterozygous sites | Conservative consensus; discards het ALT alleles |
-H I (or the standalone -I / --iupac-codes flag) | IUPAC ambiguity code | Retain heterozygosity in one sequence (see caveat below) |
-H LA/LR/SA/SR | Longer/shorter allele, tie broken by ALT/REF | Length-driven selection; confirm the letter set on the installed version |
The chimeric-haplotype footgun. -H 1/-H 2 are only meaningful when genotypes are phased (0|1, pipe separator). With unphased genotypes (0/1, slash), the assignment of "which allele is haplotype 1" is arbitrary per site, so -H 1 across many heterozygous sites produces a switch-error mosaic that corresponds to no real chromosome -- while looking like a clean haplotype FASTA. This is the single most dangerous consensus mistake. Verify phasing before any -H 1/-H 2:
bcftools query -f '%CHROM\t%POS[\t%GT]\n' input.vcf.gz | head # phased: 0|1 ; unphased: 0/1If genotypes are unphased, phase first (read-backed WhatsHap/HapCUT2, trio, statistical SHAPEIT/Eagle -- accurate for common variants, poor for rare/singletons -- or native long-read phasing). See phasing-imputation/haplotype-phasing and variant-calling/vcf-basics for GT interpretation.
A single consensus FASTA is a lossy projection of a diploid genome; the right projection depends on the downstream use, and some tasks need the VCF instead:
| Strategy | Flag | Best for | Loses |
|---|---|---|---|
| Two haplotype sequences | -H 1 + -H 2 (phased) | Allele-specific expression, compound-het, HLA, cis-regulatory haplotypes | Nothing (if correctly phased) |
| IUPAC ambiguity codes | -I | Retaining het signal in one sequence | Phase/linkage; many tree/alignment tools read IUPAC as N |
| All ALT alleles | -H A | Max divergence, quick draft | Reality -- exists in no cell |
| REF at het sites | -H R | Conservative single sequence | Every heterozygous ALT allele |
Two hard boundaries:
bcftools consensus cannot apply symbolic SV alleles (<DEL>, <INS>, <DUP>, <INV>): those carry no ALT sequence, only INFO fields, so consensus has nothing to substitute. Short-read SV VCFs (Manta/DELLY) are mostly symbolic and are NOT directly consensus-able. Folding SVs into a consensus needs sequence-resolved records (long-read/assembly callers emit these) or an assembly-based approach -- see variant-calling/structural-variant-calling.For phylogenetics specifically, prefer one clean phased haplotype or a homozygous-ALT-only sequence over IUPAC, because ambiguity codes are silently dropped by many tree builders:
bcftools view -i 'GT="AA"' input.vcf.gz | bcftools consensus -f reference.fa > hom_alt.faBecause unobserved positions are emitted as reference (trap 1), a consensus must mask sites with insufficient data. -m mask.bed replaces the listed regions (default char N via --mask-with N). The mask must be built from callable depth, and the depth step hides a silent bug:
samtools depth WITHOUT -a OMITS zero-coverage positions from its output -- so those positions never enter the low-depth BED, never get masked, and stay as reference: the exact false-confidence failure the mask was meant to prevent. Always use -a (report all positions) so no-coverage sites are captured:
# Build a mask of every position below the callable-depth threshold. -a is mandatory:
# without it, zero-coverage positions are absent from the output and escape masking.
samtools depth -a aligned.bam | awk '$3 < 10 {print $1"\t"$2-1"\t"$2}' | bedtools merge > lowcov.bed
bcftools consensus -f reference.fa -m lowcov.bed input.vcf.gz > consensus.faThe < 10 threshold is a minimum-callable-depth policy (10x is a common floor for confident base calls); set it to the depth below which the calls are not trusted. bedtools genomecov -bga -ibam aligned.bam is an equivalent zero-coverage-aware alternative that also emits 0-depth intervals.
Do NOT rely on -M/-a for this: -M N outputs N only for missing ./. genotypes already present in the VCF, and -a N replaces every position absent from the VCF (which N-outs the entire non-variant genome). Neither distinguishes no-coverage from confident-reference -- only a depth-derived mask does.
Goal: Apply indels at the correct reference position and sequence.
Approach: Left-align and split multiallelics with bcftools norm so each record matches the reference context; consensus applies records positionally and mis-represented indels corrupt the output.
bcftools norm -f reference.fa input.vcf.gz -Oz -o norm.vcf.gz
bcftools index norm.vcf.gz
bcftools consensus -f reference.fa norm.vcf.gz > consensus.faUn-normalized or overlapping indels produce wrong sequence, and bcftools consensus only warns to stderr while still emitting output -- so the corruption is silent unless the stderr is inspected. Even after norm, two records whose REF spans collide remain a hazard; grep the run for warnings and inspect the region. See variant-calling/variant-normalization.
bcftools consensus -f reference.fa norm.vcf.gz 2>&1 >consensus.fa | grep -i 'overlap\|warn'For amplicon surveillance (SARS-CoV-2 and similar), ivar consensus builds a per-sample consensus directly from a pileup. Its two key thresholds are epidemiological policy decisions, not defaults to accept blindly -- they propagate into lineage assignment and transmission-cluster inference:
# Trim PCR primers FIRST -- primer-derived bases are not sample sequence and, at
# primer-binding-site mutations, cause reference-biased miscalls if left in.
ivar trim -b primers.bed -p trimmed -i aligned.bam
samtools sort -o trimmed.sorted.bam trimmed.bam
# -aa keeps all positions (so no-coverage becomes N), -A keeps orphan mates, -d 0 lifts the depth cap.
samtools mpileup -aa -A -d 0 -B -Q 0 trimmed.sorted.bam | ivar consensus -p sample -q 20 -t 0.5 -m 10 -n N| Flag | Default | Decision |
|---|---|---|
-m min depth | 10 | Below this, iVar emits N. Too low -> single-read sequencing errors become "mutations" that corrupt outbreak phylogenies. Too high -> excessive Ns, an unusably fragmented genome. |
-t min frequency to call a base | 0 (majority) | 0 calls the most common base. For a strict majority consensus use 0.5. Too low bakes minority/within-host variants and contamination into the "genome", inflating diversity and creating phantom transmission links. Raise (e.g. 0.03) only deliberately for intrahost variant work, not for a reference consensus. |
-q min base quality | 20 | Bases below this are not counted toward depth/frequency. |
-n no-coverage char | N | Character emitted where depth < -m. |
Always report -m and -t alongside a surveillance consensus -- the genome is only as trustworthy as those two numbers. Alternatives: bcftools consensus from a called VCF, or ViralConsensus (Moshiri 2023) which calls consensus directly from the alignment without an intermediate VCF, faster and lower-memory for large batches.
Apply only trusted calls; pipe filtered VCF straight into consensus:
bcftools view -f PASS input.vcf.gz -Oz -o pass.vcf.gz && bcftools index pass.vcf.gz
bcftools consensus -f reference.fa pass.vcf.gz > consensus.fa
bcftools view -v snps input.vcf.gz -Oz -o snps.vcf.gz && bcftools index snps.vcf.gz # SNPs onlyFiltered VCFs must be re-bgzipped and re-indexed before bcftools consensus reads them.
-c chain.txt writes a liftover chain mapping reference coordinates to consensus coordinates -- needed when indels shift positions and annotations must be lifted. -p PREFIX prepends a string to output sequence names (>sample1_chr1).
bcftools consensus -f reference.fa -c chain.txt -p "sample1_" input.vcf.gz > consensus.faFor a quick SNP-only substitution in Python (production work should use bcftools consensus, which handles indels, phasing, and masking):
from cyvcf2 import VCF
from Bio import SeqIO
ref = {rec.id: list(str(rec.seq)) for rec in SeqIO.parse('reference.fa', 'fasta')}
for v in VCF('input.vcf.gz'):
if v.is_snp and len(v.ALT) == 1:
ref[v.CHROM][v.POS - 1] = v.ALT[0] # POS is 1-based; list index is 0-based
with open('consensus.fa', 'w') as fh:
for chrom, seq in ref.items():
fh.write(f'>{chrom}\n{"".join(seq)}\n')minimap2 -a reference.fa consensus.fa | samtools view -b -o aln.bam # inspect where it diverges
bcftools view -H input.vcf.gz | wc -l # variants available to apply| Error / Symptom | Cause | Fix |
|---|---|---|
the VCF file is not indexed | Plain-gzip or missing index | bgzip then bcftools index (or tabix -p vcf) |
sequence "chr1" not found | Chromosome names differ between FASTA and VCF | bcftools annotate --rename-chrs map.txt |
REF does not match | Different reference than the caller used | Use the exact FASTA used for calling; normalize |
| Clean haplotype looks wrong | -H 1 on an unphased VCF -> chimera | Verify ` |
| Consensus reference-identical over gaps | No-coverage sites emitted as reference | Mask with samtools depth -a derived BED and -m |
| Garbled indels, stderr overlap warnings | Un-normalized/overlapping records | bcftools norm -f ref.fa first; inspect warnings |
<DEL>/<INS> not applied | Symbolic SV alleles carry no ALT sequence | Use sequence-resolved SV records; see structural-variant-calling |
| vs /) before -H-m and frequency -t thresholds.)© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in variant-calling/consensus-sequences of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Consensus Sequences next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Consensus Sequences this skillGPTomics/bioSkills | 1.2k | 1 repos | ~4k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Categories
Generate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar. Bio Consensus Sequences is an agent skill from GPTomics/bioSkills. Generate consensus FASTA sequences by applying VCF variants onto a reference with bcftools consensus, or build viral/amplicon consensus with iVar.
Bio Consensus Sequences fits situations like: reconstructing a sample-specific reference; deciding -H haplotype vs IUPAC vs all-ALT projection; masking no-coverage sites so a consensus does not manufacture false reference calls; setting iVar min-depth/min-frequency policy for surveillance genomes.
Run `npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a claude-code`. Or copy the skill folder (variant-calling/consensus-sequences in GPTomics/bioSkills) into .claude/skills/bio-consensus-sequences in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a codex`. Or copy the skill folder (variant-calling/consensus-sequences in GPTomics/bioSkills) into .agents/skills/bio-consensus-sequences in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-consensus-sequences -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-consensus-sequences, .gemini/skills/bio-consensus-sequences, .github/skills/bio-consensus-sequences and .opencode/skills/bio-consensus-sequences in your project.
Going by SKILL.md and its folder, Bio Consensus Sequences needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: A Bash shell.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Consensus Sequences is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Consensus Sequences: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.