Agent skill

Bio Comparative Genomics Comparative Annotation Projection

by GPTomics in GPTomics/bioSkills

Project gene annotations across genomes using TOGA (Kirilenko 2023 whole-genome-alignment chain-based projection with intactness classification), CESAR 2.0 (Sharma, Schwede & Hiller 2017 codon-aware…

MITAuto-check passedResearch & Science

Install Bio Comparative Genomics Comparative Annotation Projection

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-comparative-annotation-projection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-comparative-genomics-comparative-annotation-projection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/comparative-genomics/comparative-annotation-projection .claude/skills/bio-comparative-genomics-comparative-annotation-projection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-comparative-genomics-comparative-annotation-projection
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6.9k tokens
SKILL.md length
2,738 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Project gene annotations across genomes using TOGA (Kirilenko 2023 whole-genome-alignment chain-based projection with intactness classification), CESAR 2.0 (Sharma, Schwede & Hiller 2017 codon-aware…

  • Transferring annotations from a well-annotated reference to query genome(s)
  • SKILL.md covers Version Compatibility, Algorithmic Taxonomy, Decision Tree by Experimental… and Per-Tool Failure Modes, plus 11 more sections
  • Runs Shell scripts from its folder; calls pip, conda and git; reaches github.com and raw.githubusercontent.com
  • Classifying gene-loss vs gene-intact across many genomes at scale

What it does

Bio Comparative Genomics Comparative Annotation Projection is an agent skill from GPTomics/bioSkills. Project gene annotations across genomes using TOGA (Kirilenko 2023 whole-genome-alignment chain-based projection with intactness classification), CESAR 2.0 (Sharma, Schwede & Hiller 2017 codon-aware exon projection), LiftOff (Shumate & Salzberg 2021 reference-based annotation transfer), Liftover (UCSC), GeMoMa (Keilwagen 2019 evidence-based projection), and Comparative Annotation Toolkit (CAT). Use when transferring annotations from a well-annotated reference to query genome(s), classifying gene-loss vs…

Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/toga_annotation_projection.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Transferring annotations from a well-annotated reference to query genome(s)
  • Classifying gene-loss vs gene-intact across many genomes at scale
  • Building Zoonomia-style comparative annotations across hundreds of mammals
  • Birds (Kirilenko 2023)

Example prompts

  • “/bio-comparative-genomics-comparative-annotation-projection”

Requirements

  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • conda
    • git
    • python
    • wget

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • raw.githubusercontent.com
    • gemoma.de

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Comparative Genomics Comparative Annotation Projection loads about 6.9k tokens when it runs. Until then it costs about 218 tokens; SKILL.md has 2,738 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~218
When it runs · the whole SKILL.md, loaded when a task matches
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,738 words, ~6,948 tokens.

Download SKILL.mdSave it as .claude/skills/bio-comparative-genomics-comparative-annotation-projection/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-comparative-genomics-comparative-annotation-projection
description
Project gene annotations across genomes using TOGA (Kirilenko 2023 whole-genome-alignment chain-based projection with intactness classification), CESAR 2.0 (Sharma, Schwede & Hiller 2017 codon-aware exon projection), LiftOff (Shumate & Salzberg 2021 reference-based annotation transfer), Liftover (UCSC), GeMoMa (Keilwagen 2019 evidence-based projection), and Comparative Annotation Toolkit (CAT). Use when transferring annotations from a well-annotated reference to query genome(s), classifying gene-loss vs gene-intact across many genomes at scale, building Zoonomia-style comparative annotations across hundreds of mammals or birds (Kirilenko 2023), detecting pseudogenization, projecting alternative isoforms, or selecting between WGA-anchored (TOGA) vs ortholog-based (LiftOff) annotation transfer strategies.
tool_type
cli
primary_tool
TOGA

Version Compatibility

Reference examples tested with: TOGA 1.1.7+ (hillerlab/TOGA; Kirilenko 2023 Science 380:eabn3107), CESAR 2.0 (Sharma, Schwede & Hiller 2017 Bioinformatics 33:3985), LiftOff 1.6.3+ (Shumate & Salzberg 2021 Bioinformatics 37(12):1639-1643), Comparative Annotation Toolkit (CAT) 2.4+, GeMoMa 1.9+ (Keilwagen 2019 Methods Mol Biol 1962:161), UCSC liftOver 2024+, Cactus 2.9.1+ (for HAL input), HAL toolkit 2.3+, NextFlow 24+ for TOGA pipeline, BUSCO 5.7+ / Compleasm 0.2.7+ for QC, Luigi + Toil for CAT, R 4.4+. The current TOGA expects HAL from Cactus 2.5+; older HAL formats may fail.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: toga.py --help; cesar --help; liftoff --version
  • Python: pip show liftoff
  • Java: gemoma --help (Java 11+)

If code throws TOGA chain file missing, CESAR fragment not found, LiftOff annotation not parsed, the toolchain expects specific input formats: TOGA needs HAL or chain/net files from Cactus / LASTZ; CESAR needs exon-level GFF; LiftOff needs reference GFF and aligned FASTA. Pre-process with the appropriate format conversion.

Comparative Annotation Projection

"Annotate this new genome using my well-annotated reference" -> Annotation projection from a reference is the modern alternative to de novo gene prediction; it produces high-quality, comparable annotations across genomes by leveraging evolutionary conservation. The 2023-era standard is TOGA + CESAR 2.0 (Kirilenko 2023 Science 380:eabn3107), which uses whole-genome alignment chains + ML classification + codon-aware exon projection to scale to hundreds of genomes (Zoonomia: 488 mammals; Bird10000 Genomes: 501 birds). For ortholog-based projection (no WGA needed), LiftOff (Shumate & Salzberg 2021 Bioinformatics 37(12):1639) is the standard. The critical decision is WGA-anchored (TOGA) vs ortholog-anchored (LiftOff): TOGA explicitly classifies gene intactness vs loss using the alignment chains, LiftOff relies on reciprocal-best-hit equivalents.

  • CLI: toga.py --chain chain.bb --bed ref.bed --tDB target.2bit --qDB query.2bit --pn project_name -- WGA-based projection
  • CLI: cesar -i exons.fa -d 4 -o output.aln -- codon-aware exon alignment (used internally by TOGA)
  • CLI: liftoff -g ref.gff query.fa ref.fa -o query.gff -- ortholog-based projection
  • CLI: gemoma -- evidence-based comparative annotation (Java)
  • CLI: cat (Comparative Annotation Toolkit) -- multi-species annotation projection

Algorithmic Taxonomy

ToolApproachOutputStrengthFails when
TOGA (Kirilenko 2023 Science 380:eabn3107)Cactus HAL or LASTZ chains -> ML projection + intactness classifierPer-gene I/PI/UL/L/M/PM codes; orthology classification; coding annotation via CESAR 2.0Modern paradigm; explicit gene-loss detection at scale; Zoonomia / Bird10000 standardRequires Cactus WGA; not for prokaryotes
CESAR 2.0 (Sharma, Schwede & Hiller 2017 Bioinformatics 33:3985)HMM-based codon-aware exon projectionAligned exons + frame preservationMost accurate exon projection from WGA; preserves frame across indelsUsed internally by TOGA; standalone use more rare
LiftOff (Shumate & Salzberg 2021 Bioinformatics 37(12):1639-1643)Read-mapping-style ortholog detection + GFF transferLifted GFFFast; no WGA required; standard for query-vs-reference pairsTandem duplicates ambiguous; not for gene loss detection
UCSC liftOverCoordinate-based lift using chain filesCoordinate-lifted regionsStandard for coordinate transfers; not for gene annotationsDoesn't handle gene structure changes
Comparative Annotation Toolkit (CAT)Luigi + Toil workflow integrating TransMap + AUGUSTUS (TM/TMR/CGP/PB) + homGeneMappingPer-species comparative annotationIntegrates de novo + projectionRequires Cactus HAL input; Toil/Luigi setup complex
GeMoMa (Keilwagen 2019 Methods Mol Biol 1962:161)Reference protein homology + evidence integrationComparative gene annotationCombines multiple reference species evidenceSlower; less popular than TOGA / LiftOff
AUGUSTUS (Stanke 2008)De novo prediction; not strictly projectionPer-genome annotationAugments projection with de novoStandalone de novo; lower comparative accuracy
BRAKER3 (Gabriel 2024)Augustus + GeneMark-ETP + RNA-Seq + proteinComparative-aware de novoModern de novo with evidenceNot strictly projection
Funannotate (palmer lab)Multi-evidence annotation including LiftOffFunannotate annotationsIntegrates evidenceSetup complex
Comparative Annotation Pipeline (CAP)Earlier WGA-based annotationPer-species per-geneHistorical; replaced by TOGAUse TOGA
TransMapUCSC genome browser annotation lifterPer-locus liftTool for UCSC tracksTool-specific
Maker (Cantarel 2008)Evidence-based de novo + projectionPer-genome annotationCombines evidenceMaker is for novel genomes; LiftOff for transfer

Methodology evolves; the Kirilenko 2023 TOGA paradigm (WGA-anchored + intactness classification) is the gold standard for vertebrate-scale comparative annotation. For pairwise transfers, LiftOff is the modern standard. Verify the current TOGA documentation (hillerlab/TOGA) before locking on a single approach.

Decision Tree by Experimental Scenario

ScenarioRecommended approachWhy
Annotate hundreds of mammal / bird genomesTOGA with Cactus HALScales to Zoonomia / Bird10000
Annotate single new genome from referenceLiftOffFast; no WGA required
Detect gene loss across mammalsTOGA intactness classificationExplicit I/PI/UL/L/M/PM codes
Project alternative isoformsTOGA (preserves multiple transcripts)Standard
Project annotations to assembly with high N50 + chromosome-levelTOGARequires good assembly
Project to fragmented draft assemblyLiftOff (more tolerant)LiftOff works on draft assemblies
Multi-species annotation pipelineCAT (Luigi + Toil)Integrated workflow
Annotate plant genome from ArabidopsisLiftOff with plant-specific optionsStandard for plant work
Pseudogenization detection at scaleTOGA + intactness analysisDesigned for this
Reference-free gene predictionBRAKER3 or AUGUSTUSDe novo; not projection
Comparative annotation of multiple referencesGeMoMaMulti-reference evidence integration
UCSC genome browser coordinate transferliftOver toolCoordinate-specific
Annotation transfer to closely related strain (>95% ANI)LiftOffHigh accuracy at close divergence
Annotation transfer to deep divergence (mammal to fish)TOGA + manual reviewRequires WGA; expect lower coverage
Project annotations with WGD-aware handlingAnchorWave + TOGA-like or custom workflowWGD-aware tools
Annotate non-coding RNAsSpecialized tools (Rfam, ncRNA-specific)RNA detection different problem
Annotate immune / repetitive genes (MHC, OR)PGR-TK MAP graph or manualRepetitive regions; use [[pangenome-analysis]]
Annotate transposable elementsRepeatMasker / RepeatModelerTE annotation different problem
Validate projected annotationsRNA-Seq alignment to projectedRNA-Seq evidence is gold standard

Per-Tool Failure Modes

TOGA chain file missing or incompatible

Trigger: Running TOGA on Cactus HAL without proper chain file extraction.

Mechanism: TOGA requires UCSC-style chain files derived from Cactus HAL or LASTZ chains/nets pipeline. Cactus HAL doesn't directly produce chain files; conversion via halSynteny + chainNet + axtChain is required.

Symptom: TOGA fails with "chain file not found" or "no syntenic blocks for query."

Fix: Use halSynteny (HAL toolkit) to extract syntenic blocks; convert to chain format via axtChain and chainNet. The TOGA Nextflow wrapper handles this automatically; for manual runs, see UCSC kentUtils chain documentation.

CESAR exon-fragment misalignment in highly divergent species

Trigger: Projecting from mouse to fish (~400 Myr divergence); many exons fail CESAR projection.

Mechanism: CESAR's HMM model is calibrated for vertebrate divergence (< 100 Myr typical). At deep divergence, exon boundaries shift; CESAR may misalign or fail to project.

Symptom: Many genes in mouse have TOGA "M" (missing) or "PI" (partial-intact) classification in fish; coverage of expected genes is low.

Fix: TOGA documentation recommends < 300 Myr divergence for reliable projection. For deeper divergence, manual review of failed exons; consider GeMoMa with multiple reference species. Some genes won't project because they're truly absent (orphan genes); others fail due to alignment limitations.

LiftOff tandem duplicate ambiguity

Trigger: LiftOff on genomes with extensive tandem duplications (e.g., NLR clusters in plants, olfactory receptors in mammals).

Mechanism: LiftOff uses ortholog detection similar to OrthoFinder; tandem duplicates create many similar sequences, making reciprocal-best-hit identification ambiguous.

Symptom: LiftOff reports many "multimapped" genes; tandem clusters have one-to-many or many-to-one orthology calls.

Fix: Pre-collapse tandem duplicates manually; or use LiftOff with -mismatch 5 and -flank 0.5 for more relaxed mapping; or use TOGA which has tandem-aware classification.

TOGA intactness classification false negatives

Trigger: TOGA classifies a gene as "Lost" when it is actually intact.

Mechanism: TOGA uses ML classifier on chain features + frame preservation; assembly gaps, short alignment fractions, or CESAR projection failures can cause false-loss calls.

Symptom: TOGA "Lost" gene actually present in independent validation (RNA-Seq, manual inspection); known biology contradicts loss.

Fix: Manual review of TOGA "Lost" calls in the loss_summ_data.tsv; cross-validate with RNA-Seq mapping; use ID + biology to verify. TOGA's classifier is calibrated for mammals/birds; non-canonical genome architectures may produce false losses.

Reference choice bias

Trigger: Projecting from one reference (e.g., mouse) produces different annotation than from another (e.g., human).

Mechanism: Each reference's annotation has its own biases (gene structures, splice variants, missing genes). Projection inherits these biases; different references produce somewhat different annotations.

Symptom: Mouse-reference TOGA annotation has gene X missing in query; human-reference TOGA has it present; or different exon structures.

Fix: Project from multiple references; consensus annotation. CAT integrates multi-reference projection; manual review of inconsistencies. Document reference choice impact.

Pseudogenization vs gene loss distinction

Trigger: Reporting a "lost" gene that retains coding sequence (frame may be conserved but expression lost).

Mechanism: TOGA detects loss of coding capacity (intactness), but doesn't directly identify pseudogenization (loss of expression). A pseudogene with intact reading frame may be classified "Intact" by TOGA.

Symptom: TOGA "Intact" annotation but the gene is pseudogene per RNA-Seq + Ribo-Seq evidence.

Fix: Combine TOGA intactness with expression data (RNA-Seq from species of interest); apply PseudoPipe (Zhang 2006 Bioinformatics 22:1437) or RetroFinder (Baertsch 2008 BMC Genomics 9:466) for systematic pseudogene detection.

Splice variant inconsistency across projections

Trigger: Projecting genes with extensive alternative splicing.

Mechanism: Reference may have alternative splice variants; projection of alternative isoforms requires per-isoform alignment which may fail for some variants.

Symptom: Query species has fewer projected isoforms than reference; canonical isoform present but alternatives missing.

Fix: TOGA projects each transcript independently; manual review of dropped isoforms. RNA-Seq from species of interest for novel splice variants.

Annotation pipeline reference quality affecting projection

Trigger: Projecting from an outdated or buggy reference annotation.

Mechanism: Errors in reference annotation propagate to all projections. A wrong exon boundary in mouse is propagated to all mammals.

Symptom: Same exon boundary error appears across many projected annotations.

Fix: Verify reference annotation quality via BUSCO + manual gene model review; use updated Ensembl / NCBI releases (current 2024-Q4).

Polyploid query genome handling

Trigger: Projecting annotations onto a polyploid query without subgenome consideration.

Mechanism: Projection sees multiple homeologous regions; ortholog detection ambiguous between subgenomes.

Symptom: Polyploid query annotation has ~2x the genes expected; many "redundant" projections from homeologs.

Fix: Assign subgenomes before projection (see [[whole-genome-duplication]]); project to each subgenome separately. AnchorWave proali handles ploidy; LiftOff doesn't natively.

Chromosome-level vs scaffold-level reference

Trigger: Projecting from chromosome-level reference to scaffold-level query.

Mechanism: Scaffold-level query has gaps and ambiguous gene assignments; projection mostly succeeds but some genes split across scaffolds.

Symptom: Projected GFF has fragmented gene models; some genes have 2-3 entries across scaffolds.

Fix: Pre-scaffold query (Hi-C scaffolding if possible); or accept partial annotations; document fragmentation rate. TOGA reports "Partial Intact" for these cases.

Show full SKILL.md (1,080 more words)Show less

Quantitative Thresholds

QuantityThresholdSource / Rationale
TOGA intactness classes (loss_summ_data.tsv)I (intact), PI (partial intact), UL (uncertain loss), L (lost), M (missing/assembly gap), PM (partial missing)Kirilenko 2023 + TOGA repo
TOGA orthology relationships (orthology_classification.tsv)one2one, one2many, many2one, many2many, PG (paralogous projection / no orthologous chain)Kirilenko 2023
TOGA "Intact" classification confidenceML classifier posterior > 0.9Kirilenko 2023 supp
LiftOff coverage>=80% of reference gene length alignedDefault
LiftOff identity>=70% nucleotide identity (default)Default
Maximum divergence for TOGA~300 Myr (vertebrate); validate per cladeKirilenko 2023
Maximum divergence for LiftOff~80% nucleotide identityEmpirical
Maximum divergence for CESAR~150 Myr (vertebrate)Sharma, Schwede & Hiller 2017
Reference annotation BUSCO completeness>= 95%Standard QC
Assembly N50 for projection>= 1 Mb; chromosome-level preferredStandard
Annotation transfer success rate90-95% for closely related (< 50 Myr); 60-80% for moderately divergedEmpirical
Multi-reference consensus>= 2 references agreeingManual standard
Pseudogene classification thresholdTOGA "Lost" + no RNA-Seq evidenceOperational
Tandem cluster window for LiftOff50 kb defaultDefault
GeMoMa minimum protein identity60%Default
CAT pipeline runtime per genome1-5 hours on 16 coresEmpirical
TOGA per-genome runtime30 min - 5 hr on 16 coresEmpirical
Nextflow scalingscales with coresStandard
Reference annotation versionEnsembl / NCBI release 2024-Q4 minimumStandard
Splice variant countreport; per-transcript projectionVariable

TOGA Standard Workflow

Goal: Project annotations from reference to query genome(s), classifying gene-loss / intactness.

Approach: Cactus WGA -> halSynteny + chainNet -> TOGA Nextflow pipeline.

bash
# Prerequisites: Cactus HAL file from [[whole-genome-alignment]]
# OR LASTZ chain/net pipeline output

# 1. Extract syntenic blocks from HAL
halSynteny output.hal reference query --queryGenome query > query.synteny.psl

# 2. Convert PSL to UCSC chain format
axtChain -psl -linearGap=loose query.synteny.psl reference.2bit query.2bit chains/query.chain.gz
# Note: `-psl` is correct here only if the input is PSL. When the input comes from
# `lastz --format=axt`, drop the `-psl` flag (or emit PSL from LASTZ first).

# 3. Run TOGA. The canonical invocation is `python toga.py` from the TOGA checkout;
# the Nextflow-style command shown below mirrors the same arguments but may not be
# the standard entry point in your release -- verify against the hillerlab/TOGA README.
python toga.py \
    chains/query.chain.gz \
    reference_annotation.bed \
    reference.2bit \
    query.2bit \
    --pn project_name \
    --cpus 32

# Output:
#   project_name/loss_summ_data.tsv           Per-gene intactness call
#   project_name/orthology_classification.tsv One-to-one / one-to-many / many-to-many
#   project_name/query_annotation.bed         Lifted gene annotation
#   project_name/query_annotation.gff         GFF format
#   project_name/cesar_alignment/             CESAR exon alignments
python
'''Parse TOGA loss summary to identify gene loss vs intact.'''
import pandas as pd


def load_toga_loss(loss_summary_path):
    '''loss_summary_data.tsv columns: TRANSCRIPT, STATUS, IS_INTACT, ...'''
    df = pd.read_csv(loss_summary_path, sep='\t')
    return df


def classify_genes(df):
    '''Standard TOGA classification: I/PI/UL/L/M/PM'''
    return df.groupby('STATUS')['TRANSCRIPT'].count()


def filter_high_confidence_intact(df):
    '''Filter to high-confidence intact (I) only. PI is partial-intact; L is Lost.'''
    return df[df['STATUS'] == 'I']

LiftOff for Pairwise Annotation Transfer

Goal: Transfer reference annotation to query genome via ortholog mapping.

Approach: Minimap2-based ortholog detection -> per-gene transfer.

bash
# Standard LiftOff (verify flags with `liftoff --help`)
liftoff -g reference.gff \
    query.fa reference.fa \
    -o query.gff \
    -u unmapped.txt \
    -copies \
    -overlap 0.5 \
    -mismatch 2 \
    -gap 5 \
    -threads 16

# Output:
#   query.gff             Lifted annotations
#   unmapped.txt          Genes failed to lift

For closely related species, default settings suffice. For divergent (75-90% identity), use -mismatch 5 -gap 10. For tandem-rich regions, -copies allows multiple projections.

Comparative Annotation Toolkit (CAT)

bash
# Setup
git clone https://github.com/ComparativeGenomicsToolkit/Comparative-Annotation-Toolkit
cd Comparative-Annotation-Toolkit && pip install .

# Edit the CAT config file with reference + query genomes; input is a Cactus HAL alignment
# CAT is orchestrated by Luigi (task graph) on top of Toil (execution); launch the RunCat module
luigi --module cat RunCat --hal=alignment.hal --ref-genome=mm10 --config=cat.config \
      --work-dir work --out-dir out --workers=10 --local-scheduler \
      --augustus --augustus-cgp --augustus-pb --assembly-hub > log.txt

CAT projects the reference annotation across the Cactus alignment with TransMap, then integrates AUGUSTUS (TM/TMR/CGP/PB) and homGeneMapping; output is multi-species comparative annotation.

Reconciliation: When Methods Disagree

PatternLikely causeAction
TOGA "Lost" vs LiftOff "Mapped"LiftOff more permissive; doesn't check frame preservationTrust TOGA for loss claims; LiftOff for pairwise transfer
TOGA "Partial Intact" vs LiftOff "Mapped"Frame disruptionCross-validate; consider biology
Mouse-reference annotation vs human-referenceReference choice biasMulti-reference consensus
TOGA "Intact" but no RNA-Seq evidencePossible pseudogene with intact ORFAdd expression evidence; reclassify
CESAR fails exonDeep divergence or assembly gapManual review; consider GeMoMa alternative
GeMoMa vs LiftOff disagreeEvidence integration vs ortholog-basedGeMoMa for evidence-rich; LiftOff for fast pairwise
TOGA chain doesn't cover geneCactus alignment quality issueRe-run Cactus with adjusted parameters; verify assembly
LiftOff produces duplicate transfersTandem clusterUse -copies or restrict; manual review
CAT pipeline integrates de novo + projectionCombined evidenceTrust integrated; better than single tool
BUSCO of projected annotation lowMissing core genesLikely tool failure; re-run with relaxed parameters

Operational rule for publication: TOGA + Cactus HAL for clade-level annotation (Zoonomia-style); LiftOff for pairwise transfer to closely related (<100 Myr); BRAKER3 / Funannotate for de novo where projection fails; manual review of TOGA "Lost" / "PI" calls.

Cohort Gotchas

  • Plant comparative annotation: GENESPACE handles synteny-aware ortholog detection; project via plant-specific tools
  • Bacterial annotation: different problem; use [[pangenome-analysis]] with Bakta consistent annotation
  • Single-cell expression data: RNA-Seq evidence for projected genes essential
  • Repetitive genes (MHC, OR): projection unreliable; use [[pangenome-analysis]] with PGR-TK
  • Recently diverged strains: LiftOff with strict parameters; high accuracy
  • Polyploid query: assign subgenomes first ([[whole-genome-duplication]])
  • Distantly related to reference (>300 Myr): projection rate drops; consider de novo with comparative evidence
  • Fragmented draft genomes: projection works on contigs but gene splits possible
  • Reference annotation quality: verify BUSCO before propagating to all projections
  • Non-canonical genomes (B chromosomes, supernumerary): typically excluded from projection

Anticipated Reviewer Pushback

PushbackStandard response
"Reference annotation quality?"BUSCO completeness reported; updated to current Ensembl / NCBI release
"Maximum divergence for TOGA?"< 300 Myr for vertebrate; documented and respected
"Tandem duplicate handling?"LiftOff -copies; or pre-collapse for TOGA
"Gene loss detection?"TOGA intactness classification (I/PI/UL/L/M/PM); cross-validated with RNA-Seq
"Pseudogenization?"TOGA "Intact" classification verified against RNA-Seq + Ribo-Seq
"Cactus WGA quality?"Pre-filtered repeats; BUSCO on reference and query; Toil reproducibility
"Multi-reference?"Consensus annotation from 2-3 reference species; documented disagreements
"Polyploid?"Subgenomes assigned via [[whole-genome-duplication]]; per-subgenome projection
"Annotation transfer rate?"Per-species coverage reported (e.g. 85% projection success)
"Splice variants?"Per-transcript projection; missing isoforms documented
"BUSCO of projected annotation?">= 90% complete; reported

Common Errors

Error / symptomCauseSolution
TOGA "chain file not found"Cactus HAL not converted to chainUse halSynteny + axtChain
TOGA stops with "no syntenic blocks"Genome unrelated to referenceVerify Cactus alignment is non-empty
LiftOff produces empty GFFReference / query genome mismatchVerify both are FASTA; check identifier consistency
GeMoMa OOMJava heap insufficientIncrease via -Xmx32g
CESAR fragment-not-foundExon GFF malformedVerify GFF format; convert if needed
TOGA Nextflow hangsCluster resource issueRestart with -resume; check Nextflow config
BUSCO of projected annotation lowMissing core genesLikely projection failure; manual review
Many "PI" classificationsAssembly fragmentationImprove assembly; or accept partial intactness
CAT Luigi/Toil step failsConda env or Toil jobstore issueRe-create conda envs; inspect the Toil jobstore and Luigi task logs
LiftOff "no overlap"Reference / query coordinate systems differVerify same genome version
TOGA intactness disagrees with biologyEdge case; manual review neededInspect CESAR alignment; cross-validate with RNA-Seq

Tool Installation Notes

bash
# TOGA
conda env create -f https://raw.githubusercontent.com/hillerlab/TOGA/master/toga_env.yml
# Requires Nextflow and nf-core

# CESAR 2.0 (bundled with TOGA)
git clone https://github.com/hillerlab/CESAR2.0

# LiftOff
pip install liftoff

# UCSC liftOver
conda install -c bioconda ucsc-liftover

# Comparative Annotation Toolkit (CAT)
git clone https://github.com/ComparativeGenomicsToolkit/Comparative-Annotation-Toolkit
cd Comparative-Annotation-Toolkit && pip install .

# GeMoMa
wget https://gemoma.de/jcag/gemoma.zip && unzip gemoma.zip

# Comparative tools
conda install -c bioconda cactus busco compleasm

# De novo annotation (alternative)
conda install -c bioconda braker3 funannotate maker

For TOGA pipeline at vertebrate-scale, use Nextflow with proper HPC config (Slurm / Kubernetes); allocate >= 32 cores per genome.

References

  • Kirilenko BM et al 2023 Science 380:eabn3107 (TOGA; Zoonomia + Bird10000 standard)
  • Sharma V, Schwede P & Hiller M 2017 Bioinformatics 33:3985 (CESAR 2.0)
  • Shumate A & Salzberg SL 2021 Bioinformatics 37(12):1639-1643 (LiftOff)
  • Hickey G et al 2013 Bioinformatics 29:1341 (HAL toolkit)
  • Keilwagen J et al 2019 Methods Mol Biol 1962:161 (GeMoMa)
  • Gabriel L, Brůna T et al 2024 Genome Res 34:769 (BRAKER3)
  • Cantarel BL et al 2008 Genome Res 18:188 (MAKER)
  • Stanke M et al 2008 Bioinformatics 24:637 (AUGUSTUS)
  • Zhang Z et al 2006 Bioinformatics 22:1437 (PseudoPipe)
  • Armstrong J et al 2020 Nature 587:246 (Progressive Cactus)
  • Baertsch R et al 2008 BMC Genomics 9:466 (RetroFinder)
  • Liao W-W et al 2023 Nature 617:312 (HPRC draft pangenome; relevant context)
  • Fiddes IT et al 2018 Genome Res 28:1029 (Comparative Annotation Toolkit)
  • Salzberg SL 2019 Genome Biol 20:92 (next-gen annotation pipelines)
  • Comparative Genomics Toolkit (CGT) GitHub
  • TimeTree (database) for divergence dates
  • comparative-genomics/whole-genome-alignment - Cactus WGA precedes TOGA
  • comparative-genomics/synteny-analysis - Synteny detection from WGA
  • comparative-genomics/ortholog-inference - TOGA orthology classification
  • comparative-genomics/pangenome-analysis - PGR-TK for repetitive / clinical genes
  • comparative-genomics/whole-genome-duplication - Subgenome assignment for polyploid query
  • genome-annotation/eukaryotic-gene-prediction - BRAKER3 / Funannotate de novo alternative
  • genome-annotation/functional-annotation - Function assignment downstream
  • genome-annotation/annotation-transfer - Related skill on annotation transfer mechanisms
  • genome-annotation/prokaryotic-annotation - Bakta for prokaryote annotation
  • read-qc/rnaseq-qc - RNA-Seq evidence to validate projected annotations
  • read-alignment/star-alignment - RNA-Seq alignment for validation

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in comparative-genomics/comparative-annotation-projection of GPTomics/bioSkills.

  • SKILL.md
  • examples/toga_annotation_projection.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Comparative Genomics Comparative Annotation Projection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Comparative Genomics Comparative Annotation Projection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Comparative Genomics Comparative Annotation Projection this skillGPTomics/bioSkills1.2k2 repos~6.9kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Comparative Genomics Comparative Annotation Projection

What does Bio Comparative Genomics Comparative Annotation Projection do?

Project gene annotations across genomes using TOGA (Kirilenko 2023 whole-genome-alignment chain-based projection with intactness classification), CESAR 2.0 (Sharma, Schwede & Hiller 2017 codon-aware…. Bio Comparative Genomics Comparative Annotation Projection is an agent skill from GPTomics/bioSkills.0 (Sharma, Schwede & Hiller 2017 codon-aware exon projection), LiftOff (Shumate & Salzberg 2021 reference-based annotation transfer), Liftover (UCSC), GeMoMa (Keilwagen 2019 evidence-based projection), and Comparative Annotation Toolkit (CAT).

When should I use Bio Comparative Genomics Comparative Annotation Projection?

Bio Comparative Genomics Comparative Annotation Projection fits situations like: transferring annotations from a well-annotated reference to query genome(s); classifying gene-loss vs gene-intact across many genomes at scale; building Zoonomia-style comparative annotations across hundreds of mammals; birds (Kirilenko 2023).

How do I install Bio Comparative Genomics Comparative Annotation Projection in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-comparative-annotation-projection -a claude-code`. Or copy the skill folder (comparative-genomics/comparative-annotation-projection in GPTomics/bioSkills) into .claude/skills/bio-comparative-genomics-comparative-annotation-projection in your project. Claude Code loads it when a task matches its description.

How do I install Bio Comparative Genomics Comparative Annotation Projection in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-comparative-annotation-projection -a codex`. Or copy the skill folder (comparative-genomics/comparative-annotation-projection in GPTomics/bioSkills) into .agents/skills/bio-comparative-genomics-comparative-annotation-projection in your project. Codex loads it when a task matches its description.

Can I use Bio Comparative Genomics Comparative Annotation Projection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-comparative-annotation-projection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-comparative-genomics-comparative-annotation-projection, .gemini/skills/bio-comparative-genomics-comparative-annotation-projection, .github/skills/bio-comparative-genomics-comparative-annotation-projection and .opencode/skills/bio-comparative-genomics-comparative-annotation-projection in your project.

What does Bio Comparative Genomics Comparative Annotation Projection need to run?

Going by SKILL.md and its folder, Bio Comparative Genomics Comparative Annotation Projection needs a shell for the scripts in its folder and the command-line tools its instructions call (pip, conda, git, python and wget). Our summary lists: A Bash shell.

Does Bio Comparative Genomics Comparative Annotation Projection access the network?

SKILL.md names 3 domains. In commands or code: github.com, raw.githubusercontent.com and gemoma.de; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Comparative Genomics Comparative Annotation Projection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Comparative Genomics Comparative Annotation Projection use?

Bio Comparative Genomics Comparative Annotation Projection is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Comparative Genomics Comparative Annotation Projection use?

About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Comparative Genomics Comparative Annotation Projection?

Skills that share tags, products or a category with Bio Comparative Genomics Comparative Annotation Projection: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Comparative Genomics Comparative Annotation Projection?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.