Agent skill

Bio Comparative Genomics Synteny Analysis

by GPTomics in GPTomics/bioSkills

Detect syntenic blocks and structural rearrangements between genomes using MCScanX (Wang 2012), JCVI/MCScan (Tang 2008 Python), GENESPACE (Lovell 2022) for orthology-anchored riparian visualization…

MITAuto-check passedResearch & Science

Install Bio Comparative Genomics Synteny Analysis

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-synteny-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-comparative-genomics-synteny-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/comparative-genomics/synteny-analysis .claude/skills/bio-comparative-genomics-synteny-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-comparative-genomics-synteny-analysis
GitHub stars
1.2k
Used in
2 other repos
Token cost
~8.3k tokens
SKILL.md length
3,199 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Detect syntenic blocks and structural rearrangements between genomes using MCScanX (Wang 2012), JCVI/MCScan (Tang 2008 Python), GENESPACE (Lovell 2022) for orthology-anchored riparian visualization…

  • Identifying collinear gene blocks across species
  • SKILL.md covers Version Compatibility, Algorithmic Taxonomy, Decision Tree by Experimental… and Per-Tool Failure Modes, plus 12 more sections
  • Runs Python scripts from its folder; calls python, conda and pip; reaches github.com
  • Distinguishing macrosynteny from microsynteny

What it does

Bio Comparative Genomics Synteny Analysis is an agent skill from GPTomics/bioSkills. Detect syntenic blocks and structural rearrangements between genomes using MCScanX (Wang 2012), JCVI/MCScan (Tang 2008 Python), GENESPACE (Lovell 2022) for orthology-anchored riparian visualization, SyRI for structural variation, AnchorWave for sequence-level synteny, i-ADHoRe 3.0 for highly diverged species, SynNet for synteny networks, and ntSynt for multi-genome macrosynteny. Use when identifying collinear gene blocks across species, distinguishing macrosynteny from microsynteny, detecting…

Its SKILL.md is about 8.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/synteny_analysis.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. It works with Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Identifying collinear gene blocks across species
  • Distinguishing macrosynteny from microsynteny
  • Detecting inversions/translocations/duplications
  • Anchoring orthology in WGD lineages

Example prompts

  • “/bio-comparative-genomics-synteny-analysis”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • conda
    • pip
    • git
    • make
    • cmake

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Comparative Genomics Synteny Analysis loads about 8.3k tokens when it runs. Until then it costs about 198 tokens; SKILL.md has 3,199 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~198
When it runs · the whole SKILL.md, loaded when a task matches
~8.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 3,199 words, ~8,326 tokens.

Download SKILL.mdSave it as .claude/skills/bio-comparative-genomics-synteny-analysis/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-comparative-genomics-synteny-analysis
description
Detect syntenic blocks and structural rearrangements between genomes using MCScanX (Wang 2012), JCVI/MCScan (Tang 2008 Python), GENESPACE (Lovell 2022) for orthology-anchored riparian visualization, SyRI for structural variation, AnchorWave for sequence-level synteny, i-ADHoRe 3.0 for highly diverged species, SynNet for synteny networks, and ntSynt for multi-genome macrosynteny. Use when identifying collinear gene blocks across species, distinguishing macrosynteny from microsynteny, detecting inversions/translocations/duplications, anchoring orthology in WGD lineages, producing publication riparian plots, computing synteny block age via Ks (cross-references whole-genome-duplication), or running synteny-aware ortholog inference in polyploids.
tool_type
mixed
primary_tool
MCScanX

Version Compatibility

Reference examples tested with: MCScanX 1.0+ (wyp1125/MCScanX commit 2020+), JCVI 1.4.21+ (Python port of MCScan), GENESPACE 1.4.0+ (Lovell 2022 eLife 11:e78526), SyRI 1.7.1+ (Goel 2019 Genome Biol 20:277), plotsr 1.1.1+, AnchorWave 1.2.5+ (Song 2022 PNAS 119:e2113075119), i-ADHoRe 3.0.01+, SynNet (Zhao 2017 Plant Cell 29:1278), ntSynt 1.0.4+ (2024), minimap2 2.28+, MUMmer 4.0.0+, OrthoFinder 3.0+, R 4.4+. plotsr requires pysam 0.22+ and seaborn 0.13+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: MCScanX -h; syri --version; python -m jcvi.compara.catalog ortholog --help
  • R: packageVersion('GENESPACE'); ?run_genespace
  • Python: pip show jcvi

If code throws MCScanX: argument bad format, syri: input alignment file missing required columns, or GENESPACE: GFF parse error, these tools have brittle input parsing: MCScanX requires 4-column species_chr gene start end BED (non-standard), JCVI expects 4-column simple BED, GENESPACE requires GFF3 with gene feature type. Pre-process with jcvi.formats.gff bed or custom AWK.

Synteny Analysis

"Compare genome architecture between these species" -> Detect conserved gene order (synteny) and infer rearrangement history. Synteny is NOT the same as collinearity: synteny is "genes on same chromosome", collinearity is "same order on same chromosome" (modern usage often conflates them). The choice of tool depends on whether the question is gene-level co-linearity (MCScanX, JCVI), whole-genome structural rearrangements (SyRI, AnchorWave), multi-genome macrosynteny (GENESPACE, ntSynt), or synteny-aware orthology (GENESPACE, ProteinOrtho-synteny). Repeat-masking quality is the dominant determinant of result reliability -- unmasked TEs produce ~100x more false anchor pairs than real syntenic anchors.

  • CLI: MCScanX for collinear gene blocks via dynamic programming
  • CLI: python -m jcvi.compara.catalog ortholog A B for JCVI/MCScan Python pipeline
  • R: run_genespace() (Lovell 2022) for orthology-anchored riparian plots + pan-gene tracks
  • CLI: syri for inversion / translocation / duplication detection
  • CLI: anchorwave proali for sequence-level WGD-aware synteny

Algorithmic Taxonomy

ToolApproachOutputStrengthFails when
MCScanX (Wang 2012 NAR 40:e49)Dynamic programming on BLAST hits with collinearity scoring.collinearity blocks; tandem duplications collapsedMost widely used; 14 downstream tools; well-benchmarkedRepeat-derived false BLAST hits; brittle 4-column input
MCScanX-hWithin-genome variant for WGD detectionSelf-comparison blocksStandard for paranome construction; WGD-awareSame input fragility
JCVI / MCScan Python (Tang 2008 GR 18:1944)Re-implementation of MCScanX with anchor stringency controlPublication-grade dotplots, karyotype, riparianBest plotting; --cscore reciprocal-hit controlSlower than MCScanX C version on huge genomes
GENESPACE (Lovell 2022 eLife 11:e78526)OrthoFinder + MCScanX orthogroup-constrained syntenyRiparian plots; pan-gene tracks; CNV across genomesModern standard for plant comparative; integrates orthologySlow at > 30 genomes; less rigorous for non-plant clades
i-ADHoRe 3.0 (Proost 2012 NAR 40:e11)Iterative ordered-gene-list detectionGhost gene families; deeply diverged syntenyCatches ancient synteny obscured by rearrangementsComputationally heavy at genome scale
AnchorWave (Song 2022 PNAS 119:e2113075119)CDS/exon anchors -> wave-front alignmentWhole-genome alignment with WGD awareness; SVsBest for plant WGD analyses with known ploidyCDS-anchored only; intergenic resolution limited
SyRI (Goel 2019 GB 20:277)Pairwise WGA -> syntenic-path identification -> SV callsINV / TRANS / DUP / SYN / INS / DEL annotatedComprehensive SV detection; works on chromosome-level assembliesRequires chromosome-level assemblies; pairwise only
plotsr (Goel 2022 Bioinformatics 38:2922)Multi-genome SyRI visualizationStacked synteny + SV maps across N genomesBest for visualizing 3-10 genome rearrangement historiesInherits SyRI's pairwise input limitation
ntSynt (2024)Minimizer-based alignment-free syntenyMulti-genome macrosynteny blocksAlignment-free; handles > 15% divergenceMacrosynteny only; misses microsynteny
SynNet (Zhao 2017 Plant Cell 29:1278)Synteny block adjacency graphsSynteny networks across many genomesPhylogenetic network from synteny; detects deep ancestryLess standard than block-based methods
Satsuma / progressive CactusReference-free WGAWhole-genome alignment (HAL format)Underlies large-scale orthology; sequence-level syntenySee [[whole-genome-alignment]]
LASTZ chain/net (UCSC; Kent 2003 PNAS 100:11484 chains/nets; Schwartz 2003 GR 13:103 BLASTZ)Pairwise WGA with chains and netsChains + nets in UCSC genome browserReference-anchored synteny; standard for UCSC tracksSee [[whole-genome-alignment]]
nucmer + dnadiff (MUMmer4; Marçais 2018 PLoS Comp Biol 14:e1005944)MUM-anchored pairwise alignmentWhole-genome alignment with SV summaryFast pairwise WGA for closely relatedSensitive only above ~70% identity
MashMap (Jain 2018 Bioinformatics 34:i748)Approximate mapping for fragment-fragment syntenyPairwise mappings with identityScales to thousands of genomesCoarse (window-based); no SV inference

Methodology evolves; GENESPACE has emerged as the de facto standard for plant comparative genomics (2022-2026). For non-plant clades, MCScanX or JCVI remain the workhorses. Whole-genome alignment for synteny is increasingly delegated to Cactus / Minigraph-Cactus and HAL toolkit (see [[whole-genome-alignment]]).

Decision Tree by Experimental Scenario

ScenarioRecommended approachWhy
Plant comparative genomics, 2-20 speciesGENESPACEIntegrated orthology + synteny + visualization; modern standard
Animal / fungal, 2 species comparisonJCVI/MCScanBest plots; flexible; widely cited
Animal / fungal, 10+ speciesOrthoFinder + MCScanX + custom plot OR JCVI multi-genomeAvoid plant-specific GENESPACE assumptions
Bacterial / prokaryote syntenyprogressiveMauve OR MCScanX-bacterialTools designed for compact genomes
Chromosome-level rearrangement inventorySyRI + plotsrComprehensive SV (INV/TRANS/DUP) with publication viz
Polyploid genome analysisAnchorWave proali mode with ploidyWGD-aware synteny; subgenome-aware
Closely related strains (>= 95% identity)nucmer + dnadiffFaster than chromosome-level WGA; appropriate sensitivity
Distantly related species (< 70% identity)i-ADHoRe 3.0 or LASTZ chains/netsSequence-level approaches lose power; ordered-list methods retain it
Ancient WGD detectionSee [[whole-genome-duplication]]wgd v2, KsRates; Ks-based dating outside synteny scope
WGA at clade-level (10+ vertebrates)Cactus / Minigraph-Cactus -> halSyntenySee [[whole-genome-alignment]]; multi-genome reference-free
Pan-genome of bacterial strainsSee [[pangenome-analysis]]Panaroo / PPanGGOLiN / PEPPAN; different problem
Synteny-aware ortholog disambiguationProteinOrtho -synteny or GENESPACETandem duplicates collapsed; co-orthologs assigned to syntenic position
Reference-guided assembly bias checkDO NOT use synteny across reference-guided scaffoldsReference-guided scaffolds propagate scaffold assumptions; circular reasoning
Microsynteny network analysisSynNetNetwork from many genomes; detects ancient micro-conserved clusters
Synteny across > 15% divergent genomesntSynt or i-ADHoReAlignment-free or ordered-list approaches
Identifying syntelogs (syntenic orthologs only)GENESPACE riparian() output; or MCScanX filtered by chr-pairSyntelog tracks only same-chromosome-pair orthologs

Per-Tool Failure Modes

Repeat-derived false synteny

Trigger: Running BLAST input for MCScanX/JCVI without softmasking repeats; or weak softmasking.

Mechanism: Transposable elements occupy 30-80% of plant and animal genomes; unmasked TEs produce millions of paralogous BLAST hits forming "synteny" blocks that are TE-driven, not real ancestry. The dynamic programming sees long apparent collinear blocks of TE-encoded proteins.

Symptom: MCScanX/JCVI output dominated by short (5-10 anchor) blocks; blocks cluster in TE-rich regions (heterochromatin, centromeric); BLAST hit count > 10x what is expected for the divergence.

Fix: Softmask genomes with RepeatModeler2 (de novo TE library) + RepeatMasker before BLAST; report softmasking statistics. For MCScanX, set -e 1e-10 (stricter) for close species, -e 1e-5 for distant. JCVI --cscore 0.95 enforces near-reciprocal-best-hit, dramatically reducing TE-driven blocks. Always check that block size distribution drops at 5-10 anchors (TE-driven) vs proper > 20 anchor blocks (real synteny).

Reference-guided assembly creating circular synteny

Trigger: Comparing a de novo assembly with a reference-guided scaffold of the same / related species.

Mechanism: Reference-guided scaffolding orders contigs based on synteny to the reference. Comparing the resulting scaffold to that same reference produces "synteny" that is actually the imposed reference order, not biological gene order. The comparison is circular.

Symptom: Suspicious uniformity in synteny across most chromosomes; visualization shows reference and query as perfectly parallel; SyRI reports few SVs where divergence-time-appropriate counts would predict many.

Fix: Verify assembly method via metadata. For reference-guided assemblies, treat synteny inferences with the reference as suspect; use a sister species reference or pure de novo assembly. Hi-C-scaffolded assemblies are not "reference-guided" in this sense and are safe.

Fragmented assemblies underestimating macrosynteny

Trigger: Running synteny on draft assemblies with N50 < 1 Mb.

Mechanism: Macrosynteny detection requires that syntenic blocks be on single contigs; if contigs are short, blocks are artificially truncated. JCVI reports show artifactual "synteny loss" between species when one assembly is fragmented.

Symptom: Synteny "loss" correlates with assembly N50; per-chromosome dotplots show dotted lines where solid diagonals expected; SyRI reports excessive INVs / TRANS where the divergence is actually low.

Fix: Require minimum N50 of 1 Mb for cross-species synteny; chromosome-level assemblies preferred for SyRI. For fragmented assemblies, restrict analysis to long-contig regions or use SyRI's contig-level mode with caveats.

Tandem duplicate inflation

Trigger: Including tandemly duplicated genes in synteny anchor input.

Mechanism: Tandem duplicates produce many same-region BLAST hits between species; the dynamic programming includes them as multiple anchors per gene position, artificially inflating block size and confidence.

Symptom: "High-confidence" synteny blocks in regions known for tandem expansions (NLR clusters in plants, OR genes in mammals); blocks anchored largely on adjacent tandem duplicates of the same gene.

Fix: MCScanX automatically collapses tandem duplicates (within 5 genes by default; -w parameter). Verify in the output the tandem-collapse step ran. JCVI --tandem_Nmax 10 is more aggressive. For NLR / OR genes specifically, manual tandem filtering before BLAST is preferred.

Cross-genome chromosome name mismatch

Trigger: GFF and FASTA from different sources with different chromosome naming (chr1 vs Chr1 vs 1 vs NC_001234.5).

Mechanism: MCScanX requires chromosomes to be named consistently; mismatch causes empty output or partial blocks.

Symptom: MCScanX completes but .collinearity file is small or empty; warnings about unknown chromosome IDs.

Fix: Normalize chromosome names with awk or samtools faidx --regions to unify naming convention; verify GFF + FASTA share the same names. For NCBI assemblies, use datasets to download with consistent labels.

SyRI complaining about inversions in absence of true inversions

Trigger: Running SyRI on alignments with many small-scale rearrangements; or when one genome has been re-ordered (different convention for which strand is "+").

Mechanism: SyRI's syntenic-path algorithm penalizes any rearrangement; high-divergence pairs accumulate small-scale rearrangements that are not "real" inversions but inheritance differences.

Symptom: SyRI reports thousands of small "inversions" (< 1 kb); blocks dominate the SV count; total INV length unrealistic.

Fix: Filter SyRI output by size: real biologically-relevant INVs are typically > 5 kb. Tighten minimap2 sensitivity (-x asm5 for close species; -x asm10 or -x asm20 for more divergent). Reverse-complement one assembly's chromosomes if the strand convention differs.

GENESPACE OrthoFinder version mismatch

Trigger: Running GENESPACE with an outdated bundled OrthoFinder.

Mechanism: GENESPACE bundles a specific OrthoFinder version; if user's installed OrthoFinder is newer (v3 vs v2.5), the HOG output layout differs and GENESPACE parsing fails.

Symptom: GENESPACE reports "Phylogenetic_Hierarchical_Orthogroups not found" or similar; orthogroup table is empty.

Fix: Use the OrthoFinder version GENESPACE expects (currently OrthoFinder 2.5.4 for GENESPACE 1.4.x). Pin via conda env. Future GENESPACE versions will adapt to OrthoFinder 3 layout (Lovell update expected).

Microsynteny vs macrosynteny conflation

Trigger: Reporting "synteny" without specifying scale.

Mechanism: Macrosynteny (same chromosome) is detectable across hundreds of millions of years but decays as rearrangements accumulate; microsynteny (same order within local region) can be deeply conserved even when macrosynteny is lost. Conflating them produces misleading "synteny loss" claims.

Symptom: Reports of "synteny lost between species X and Y" when both share microsyntenic gene clusters; or vice versa.

Fix: Always specify scale. SynNet (Zhao 2017) explicitly separates the two. MCScanX -s parameter (minimum block anchors) at 5 = microsynteny; at 20+ = macrosynteny. Report block-size distribution.

Polyploid / WGD-affected genome ambiguity

Trigger: Running MCScanX on a polyploid (e.g. allohexaploid wheat) without subgenome assignment.

Mechanism: WGD doubles all genes; synteny is between subgenomes within the polyploid AND across to outgroup. Without subgenome assignment, all paralogs and orthologs collapse into the same blocks.

Symptom: Many "1:many" or "many:many" synteny relationships; per-chromosome synteny counts indicate doubled or tripled blocks; Ks distribution multimodal.

Fix: Use AnchorWave proali with explicit ploidy specification. Alternatively, run synteny twice: within polyploid (subgenome-vs-subgenome) and across to outgroup. See [[whole-genome-duplication]] for subgenome assignment workflow.

Show full SKILL.md (1,368 more words)Show less

Quantitative Thresholds

QuantityThresholdSource / Rationale
Assembly N50 for synteny>= 1 MbBelow this, results tool-dependent and unreliable; chromosome-level strongly preferred
MCScanX -s (minimum anchors per block)5 default; 10 stringent; 3 sensitiveStandard configuration; 5 is balance
MCScanX -m (maximum gaps between anchors)25 defaultHigher for distant species; lower for recent
MCScanX -k (match score per anchor)50 defaultHigher rewards longer blocks
MCScanX -e (BLAST e-value)1e-5 default; 1e-10 for close speciesStricter for recent radiations
BLAST evalue threshold1e-5 to 1e-10 typicalDepends on divergence
JCVI --cscore for reciprocal best0.7 default; 0.95 near-RBH; 0.99 RBH-onlyHigher = fewer false positives, fewer hits
Tandem duplicate window5-10 genes (MCScanX default 5; JCVI 10)Wang 2012; species-specific tuning
SyRI INV minimum size for biological significance>= 5 kbBelow this, alignment noise dominates
SyRI TRANS minimum size>= 1 kbStandard convention
GENESPACE minimum syntenic block5 orthogroupsLovell 2022 default
Synteny block decay (macrosynteny half-life)~150 Myr in vertebratesApproximate convention
Microsynteny conservationup to 1 Gyr for metabolic gene clustersStated convention; verify per-clade
minimap2 preset for synteny-x asm5 for < 5% divergence; asm10 for ~10%; asm20 for ~20%minimap2 docs
MUMmer nucmer maxmatch--maxmatch for SyRI; --mum defaultMUMmer4 manual
Ks for syntenic block age (cross-references WGD)Ks 0.1-0.5 recent; 0.5-1.5 older; > 1.5 saturatedSee [[whole-genome-duplication]] for Ks plot interpretation
Repeat masking minimum>= 90% of known TE families masked (RepeatMasker .tbl)Below this, expect spurious synteny

MCScanX Standard Pipeline

Goal: Detect collinear gene blocks between two genomes.

Approach: Prepare 4-column BED with species prefix -> all-vs-all BLASTP -> run MCScanX -> parse .collinearity -> classify blocks.

bash
# 1. Prepare MCScanX-format BED (species_chr  gene  start  end)
python -m jcvi.formats.gff bed --type=gene --key=ID species_A.gff > A.gff.tmp
python -m jcvi.formats.gff bed --type=gene --key=ID species_B.gff > B.gff.tmp
awk 'BEGIN{OFS="\t"}{print "A"$1, $4, $2, $3}' A.gff.tmp > work/A.gff
awk 'BEGIN{OFS="\t"}{print "B"$1, $4, $2, $3}' B.gff.tmp > work/B.gff
cat work/A.gff work/B.gff > work/A_B.gff

# 2. All-vs-all DIAMOND BLAST (faster than BLAST+)
diamond makedb --in species_A.faa --db work/A.dmnd
diamond makedb --in species_B.faa --db work/B.dmnd
diamond blastp --db work/A.dmnd --query species_A.faa --threads 16 \
    --outfmt 6 --evalue 1e-10 --max-target-seqs 5 --out work/AA.blast
diamond blastp --db work/B.dmnd --query species_B.faa --threads 16 \
    --outfmt 6 --evalue 1e-10 --max-target-seqs 5 --out work/BB.blast
diamond blastp --db work/B.dmnd --query species_A.faa --threads 16 \
    --outfmt 6 --evalue 1e-10 --max-target-seqs 5 --out work/AB.blast
diamond blastp --db work/A.dmnd --query species_B.faa --threads 16 \
    --outfmt 6 --evalue 1e-10 --max-target-seqs 5 --out work/BA.blast
cat work/AA.blast work/BB.blast work/AB.blast work/BA.blast > work/A_B.blast

# 3. Run MCScanX
cd work && MCScanX -s 5 -m 25 -k 50 -e 1e-10 A_B
# Output: A_B.collinearity (block table), A_B.tandem (tandem clusters), A_B.html (browsing)
python
'''Parse MCScanX .collinearity and classify synteny relationships.'''
import re
from collections import defaultdict


def parse_collinearity(path):
    '''Returns list of dicts: {block_id, n_anchors, e_value, score, gene_pairs: [(g1, g2), ...]}'''
    blocks = []
    current = None
    with open(path) as fh:
        for line in fh:
            if line.startswith('## Alignment'):
                if current:
                    blocks.append(current)
                m = re.match(r'## Alignment (\d+): score=([0-9.]+) e_value=([0-9.e\-]+) N=(\d+)', line)
                if m:
                    current = {
                        'block_id': int(m.group(1)),
                        'score': float(m.group(2)),
                        'e_value': float(m.group(3)),
                        'n_anchors': int(m.group(4)),
                        'gene_pairs': []
                    }
            elif current and ':' in line and '\t' in line:
                parts = line.strip().split()
                if len(parts) >= 3:
                    current['gene_pairs'].append((parts[1], parts[2]))
    if current:
        blocks.append(current)
    return blocks


def classify_chromosome_synteny(blocks, gene_to_chr):
    '''Classify syntenic chromosome relationships: 1-1, 1-many, many-many.'''
    a_partners = defaultdict(set)
    for blk in blocks:
        for g1, g2 in blk['gene_pairs']:
            c1, c2 = gene_to_chr.get(g1), gene_to_chr.get(g2)
            if c1 and c2:
                a_partners[c1].add(c2)
    result = {}
    for chr_a, partners in a_partners.items():
        n = len(partners)
        result[chr_a] = '1-1' if n == 1 else ('1-many' if n <= 3 else 'many-many')
    return result

GENESPACE Plant-Focused Pipeline

Goal: Build pan-gene tracks across N plant genomes with orthology-anchored synteny visualization.

Approach: Format input: per-species GFF + protein FASTA -> initialize -> run -> riparian plot.

r
library(GENESPACE)

dir.create('work_dir')
parsedPaths <- parse_annotations(
    rawGenomeRepo = 'raw_genomes/',
    genomeDirs = list.dirs('raw_genomes/', recursive = FALSE),
    headerEntryIndex = 1,
    gffString = 'gff',
    faString = 'protein.faa',
    genespaceWd = 'work_dir/'
)

gpar <- init_genespace(
    wd = 'work_dir/',
    nCores = 8,
    rawOrthofinderDir = NULL,  # GENESPACE bundles OrthoFinder 2.5.x
    onewayBlast = TRUE
)

out <- run_genespace(gsParam = gpar, overwrite = TRUE)

# Riparian plot for 5 genomes vs reference
ripDat <- plot_riparian(
    gsParam = out,
    refGenome = 'Reference_species',
    useOrder = FALSE,
    backgroundColor = 'white'
)

GENESPACE outputs include results/syntenicHits.txt (anchor pairs), results/pangenes.txt (pan-gene presence/absence matrix), and results/riparian.pdf (riparian visualization). Pan-gene tracks across species are particularly useful for identifying lineage-specific genes (cf. pangenome-analysis for bacterial scope).

SyRI for Structural Rearrangement Inventory

Goal: Identify INVs, TRANSs, DUPs, INS, DEL between two chromosome-level genome assemblies.

Approach: Align with minimap2 -x asm5 -> filter to chromosome-level mappings -> run SyRI -> visualize with plotsr.

bash
# 1. Whole-genome alignment with minimap2
minimap2 -ax asm5 --eqx -t 16 reference.fa query.fa | \
    samtools sort -@ 8 -O bam - > work/aln.bam
samtools index work/aln.bam

# 2. SyRI on the alignment
syri -c work/aln.bam -r reference.fa -q query.fa -F B --prefix work/syri_out -k

# Outputs:
#   syri_out.syri.out      structural variant table (SYN/INV/TRANS/DUP/INS/DEL)
#   syri_out.syri.vcf      VCF-format variants
#   syri_out.syri.summary  block-level summary

# 3. Visualize with plotsr
echo "ref_species  query_species" > work/genome_pairs.txt
plotsr --sr work/syri_out.syri.out --genomes work/genome_pairs.txt \
    --output work/syri_plot.png --markers gene_markers.bed

For multi-genome rearrangement: run SyRI pairwise for each adjacent species in a phylogeny, then plotsr stacks the panel.

JCVI / MCScan Python for Publication Plots

Goal: Produce publication-grade dotplots, karyotype, and synteny visualization.

Approach: Use JCVI's catalog ortholog workflow to derive anchors -> generate dotplot or karyotype directly.

bash
# 1. Detect orthologs and synteny
python -m jcvi.formats.gff bed --type=mRNA --key=Name species_A.gff > A.bed
python -m jcvi.formats.gff bed --type=mRNA --key=Name species_B.gff > B.bed
python -m jcvi.formats.fasta format species_A.faa A.fasta
python -m jcvi.formats.fasta format species_B.faa B.fasta

python -m jcvi.compara.catalog ortholog --no_strip_names A B
# Produces A.B.lifted.anchors and other files

# 2. Dotplot
python -m jcvi.graphics.dotplot A.B.anchors --dpi 300 --output dotplot.png

# 3. Karyotype (chromosome-level synteny)
# Layout file: each row = species  chr  start  end  reverse
cat > layout.csv << 'EOF'
A.chr1 species_A 1 50000000 1
B.chr2 species_B 1 60000000 1
EOF
python -m jcvi.graphics.karyotype seqids.txt layout.csv

# 4. Microsynteny block plot
python -m jcvi.graphics.synteny blocks.layout

Reconciliation: When Methods Disagree

PatternLikely causeAction
MCScanX many blocks, GENESPACE fewGENESPACE uses orthogroup-constrained anchors (more conservative)GENESPACE excludes non-orthogroup hits; MCScanX includes any high BLAST hit; trust GENESPACE for orthology-anchored synteny
GENESPACE 1:1 syntelog, MCScanX 1:manyTandem duplicates in MCScanX collapsed at different granularityTrust GENESPACE for clear syntelogs; MCScanX 1:many often reflects tandem expansions
SyRI reports many INVs, plotsr shows mostly SYNSmall INVs below biological-relevance thresholdFilter SyRI INVs < 5 kb; SyRI was correct, just noisy
MCScanX block on chromosomes that JCVI doesn't pairJCVI --cscore 0.7 more restrictiveLower JCVI cscore to 0.5 OR raise MCScanX -e to 1e-10
AnchorWave finds collinearity in WGD region, MCScanX doesn'tAnchorWave WGD-aware; MCScanX treats all anchors equallyAnchorWave correct for WGD lineages
ntSynt macrosynteny across genus, MCScanX microsynteny onlyDifferent scalesBoth correct at their scale; report both with explicit scale labels
SyRI calls "translocation" but it's a known chromosome splitCentromere break / Robertsonian translocationManual cytogenetic context; SyRI's mechanical detection is correct but may not be biologically novel
Recent species pair has > 1000 SVs in SyRIEither real chromosome instability OR assembly errorsCheck assembly QC (telomere completeness; BUSCO); cross-check with PacBio long-read SVs
Synteny on reference-guided scaffoldCircular reasoningDiscard synteny calls between scaffold and its reference

Operational rule for publication: Synteny block annotation requires (1) BUSCO/Compleasm > 90% complete on both genomes; (2) softmasked repeats verified; (3) at least one cross-validation tool (MCScanX vs JCVI or MCScanX vs GENESPACE); (4) microsynteny vs macrosynteny scale explicit; (5) for SV calls, plotsr / SyRI followed by manual review of large rearrangements with read-coverage support.

Cohort Gotchas

  • Polyploid plants: WGD events confound 1:1 synteny; use AnchorWave or subgenome-assigned analysis
  • Salmonids / fish 4R: Ts3R + Ss4R WGD events; ohnologs (WGD paralogs) appear as syntenic but not orthologous; cross-reference [[whole-genome-duplication]]
  • Mammalian X chromosome: Massive recombination suppression; "synteny" sometimes inverted in opposite sex; verify strand convention
  • Centromeric regions: generally unalignable; synteny tools may produce false breaks near centromeres
  • Telomeres: assembly often incomplete; synteny calls near telomere boundaries unreliable
  • Sex chromosomes: rapidly evolving; lower synteny signal than autosomes
  • Plant B chromosomes: supernumerary; exclude from synteny analysis

Anticipated Reviewer Pushback

PushbackStandard response
"Assembly completeness?"BUSCO/Compleasm > 90%; N50 reported per assembly; chromosome-level (or scaffold N50 > 1 Mb)
"Repeat masking?"RepeatModeler2 de novo TE library + RepeatMasker; > 90% of known TE families masked; softmasked, not hardmasked
"Reference-guided scaffolding?"Verified de novo OR Hi-C scaffolded; no reference-guided steps
"Tandem duplicates?"Collapsed by MCScanX automatic detection (window 5); JCVI tandem_Nmax 10
"Cross-validation?"MCScanX results validated against JCVI (or GENESPACE) at consistent stringency
"Microsynteny vs macrosynteny?"Reported both scales; minimum block size 5 anchors for microsynteny, 20+ for macrosynteny
"WGD-aware?"AnchorWave proali with ploidy specification (for polyploids); or restricted to non-WGD lineages
"SV calling sensitivity?"SyRI INVs > 5 kb; cross-validated with PacBio long-read SV calls where available
"Synteny block age (Ks)?"Ks plotted per block; saturation threshold (Ks > 2) applied; see [[whole-genome-duplication]]
"Bacterial pangenome reference?"Not applicable here; see [[pangenome-analysis]] for prokaryotes

Common Errors

Error / symptomCauseSolution
MCScanX produces empty .collinearity4-column BED format wrong (extra columns, wrong delimiter, missing prefix)Check exact 4-column format: species_prefix_chr gene_id start end; tabs only
MCScanX dumps "no matches" messageBLAST e-value too strict; or gene IDs don't match BEDVerify gene IDs in BED match BLAST input; relax e-value
GENESPACE "OrthoFinder version mismatch"Bundled OrthoFinder version conflictPin OrthoFinder 2.5.x for GENESPACE 1.4.x; or update GENESPACE
JCVI dotplot empty--cscore too strict; few anchors surviveLower cscore to 0.5; or check BLAST hit count
SyRI "no chromosomes match"Reference and query use different chromosome IDsNormalize names: samtools faidx --regions to extract by name
SyRI INVs appear in all assembliesStrand convention issue OR repeat-drivenReverse-complement one assembly's chromosomes if convention differs; mask repeats
AnchorWave proali times outGenome > 1 Gb with high TE densityPre-mask repeats; consider GENESPACE / MCScanX alternative
GENESPACE riparian plot too crowded> 10 genomesSubset to representative species or use multi-page riparian
Per-chromosome BLAST hit count > 100,000Unmasked repeatsSoftmask before BLAST; verify by counting TE-encoded protein hits
Synteny blocks on Y chromosome appear absentY not assembled fully; or sex-specificDocument assembly limitation; exclude Y from synteny analysis

Tool Installation Notes

bash
# MCScanX
git clone https://github.com/wyp1125/MCScanX && cd MCScanX && make
# JCVI / MCScan Python
pip install jcvi
# GENESPACE
remotes::install_github('jtlovell/GENESPACE')
# SyRI + plotsr
conda install -c bioconda syri plotsr
# AnchorWave
conda install -c bioconda anchorwave
# i-ADHoRe
git clone https://github.com/VIB-PSB/i-ADHoRe && cd i-ADHoRe && mkdir build && cd build && cmake .. && make
# ntSynt
pip install ntsynt
# minimap2 + MUMmer4
conda install -c bioconda minimap2 mummer4

# Repeat masking
conda install -c bioconda repeatmodeler repeatmasker

For GENESPACE, OrthoFinder 2.5.x must be pinned; install via conda install -c bioconda orthofinder=2.5.5. JCVI is the most generally useful Python toolkit; install per-project alongside specific synteny tools.

References

  • Wang Y et al 2012 NAR 40:e49 (MCScanX)
  • Tang H et al 2008 GR 18:1944 (synteny / MCScan Python)
  • Lovell JT et al 2022 eLife 11:e78526 (GENESPACE)
  • Proost S et al 2012 NAR 40:e11 (i-ADHoRe 3.0)
  • Song B et al 2022 PNAS 119:e2113075119 (AnchorWave)
  • Goel M et al 2019 Genome Biol 20:277 (SyRI)
  • Goel M et al 2022 Bioinformatics 38:2922 (plotsr)
  • Zhao T et al 2017 Plant Cell 29:1278 (SynNet synteny network)
  • Marçais G et al 2018 PLoS Comp Biol 14:e1005944 (MUMmer4)
  • Kent WJ et al 2003 PNAS 100:11484 (chains and nets)
  • Schwartz S et al 2003 GR 13:103 (BLASTZ pairwise aligner)
  • Jain C et al 2018 Bioinformatics 34:i748 (MashMap)
  • Li H 2018 Bioinformatics 34:3094 (minimap2)
  • Holland PWH et al 1994 Development Suppl:125 (2R hypothesis)
  • Vanneste K et al 2013 MBE 30:177 (Ks saturation)
  • Birchler JA & Veitia RA 2007 Plant Cell 19:395 (gene balance)
  • Force A et al 1999 Genetics 151:1531 (subfunctionalization)
  • Zhao T & Schranz ME 2017 Curr Opin Plant Biol 36:129 (synteny network for phylogeny)
  • comparative-genomics/whole-genome-duplication - Ks-based WGD detection, paranome construction, KsRates
  • comparative-genomics/whole-genome-alignment - Cactus / Minigraph-Cactus / LASTZ chains-and-nets for sequence-level synteny
  • comparative-genomics/ortholog-inference - GENESPACE depends on OrthoFinder; synteny verifies orthology in WGD lineages
  • comparative-genomics/pangenome-analysis - Bacterial pangenome (different problem; for prokaryotes)
  • comparative-genomics/comparative-annotation-projection - TOGA + CESAR project genes through WGA-derived synteny
  • comparative-genomics/positive-selection - Selection on syntenic gene pairs (1:1 syntelogs)
  • phylogenetics/modern-tree-inference - Phylogenetic context for synteny block dating
  • genome-assembly/assembly-qc - BUSCO / Compleasm assembly completeness check before synteny
  • alignment/structural-alignment - Sequence-level alignment underlying SyRI
  • variant-calling/structural-variant-calling - SV detection from short reads (orthogonal to SyRI WGA-based)

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in comparative-genomics/synteny-analysis of GPTomics/bioSkills.

  • SKILL.md
  • examples/synteny_analysis.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Comparative Genomics Synteny Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Comparative Genomics Synteny Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Comparative Genomics Synteny Analysis this skillGPTomics/bioSkills1.2k2 repos~8.3kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Singlecell Qcxuzhougeng/wisp-science1k—~1.6kAutomated safety check: PassAGPL-3.0
Trackplotygidtu/trackplot109—~1.9kAutomated safety check: PassBSD-3-Clause
UniProt Database Accessdavila7/claude-code-templates33k14 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Singlecell Qc

    xuzhougeng/wisp-science

    A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.

    1k GitHub stars~1.6k tokensUpdated today
    Research & ScienceAuto-check passed
  • Trackplot

    ygidtu/trackplot

    Generate sashimi-style genome visualization plots (coverage, line, heatmap, IGV read-by-read, HiC, circRNA, motif) from BAM/bigWig/depth/HiC inputs.

    109 GitHub stars~1.9k tokensUpdated 14 days ago
    Research & ScienceAuto-check passed
  • UniProt Database Access

    davila7/claude-code-templates

    Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.

    33k GitHub starsUsed in 14 repos~1.7k tokens
    Research & ScienceAuto-check passed
  • End-to-end 10x Visium spatial transcriptomics analysis workflow with staged execution and human review gates.

    101 GitHub stars~1.4k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Comparative Genomics Synteny Analysis

What does Bio Comparative Genomics Synteny Analysis do?

Detect syntenic blocks and structural rearrangements between genomes using MCScanX (Wang 2012), JCVI/MCScan (Tang 2008 Python), GENESPACE (Lovell 2022) for orthology-anchored riparian visualization…. Bio Comparative Genomics Synteny Analysis is an agent skill from GPTomics/bioSkills.0 for highly diverged species, SynNet for synteny networks, and ntSynt for multi-genome macrosynteny.

When should I use Bio Comparative Genomics Synteny Analysis?

Bio Comparative Genomics Synteny Analysis fits situations like: identifying collinear gene blocks across species; distinguishing macrosynteny from microsynteny; detecting inversions/translocations/duplications; anchoring orthology in WGD lineages.

How do I install Bio Comparative Genomics Synteny Analysis in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-synteny-analysis -a claude-code`. Or copy the skill folder (comparative-genomics/synteny-analysis in GPTomics/bioSkills) into .claude/skills/bio-comparative-genomics-synteny-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Bio Comparative Genomics Synteny Analysis in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-synteny-analysis -a codex`. Or copy the skill folder (comparative-genomics/synteny-analysis in GPTomics/bioSkills) into .agents/skills/bio-comparative-genomics-synteny-analysis in your project. Codex loads it when a task matches its description.

Can I use Bio Comparative Genomics Synteny Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-comparative-genomics-synteny-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-comparative-genomics-synteny-analysis, .gemini/skills/bio-comparative-genomics-synteny-analysis, .github/skills/bio-comparative-genomics-synteny-analysis and .opencode/skills/bio-comparative-genomics-synteny-analysis in your project.

What does Bio Comparative Genomics Synteny Analysis need to run?

Going by SKILL.md and its folder, Bio Comparative Genomics Synteny Analysis needs Python for the scripts in its folder and the command-line tools its instructions call (python, conda, pip, git, make and cmake). Our summary lists: Python 3.

Does Bio Comparative Genomics Synteny Analysis access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Comparative Genomics Synteny Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Comparative Genomics Synteny Analysis use?

Bio Comparative Genomics Synteny Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Comparative Genomics Synteny Analysis use?

About 8.3k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Comparative Genomics Synteny Analysis?

Skills that share tags, products or a category with Bio Comparative Genomics Synteny Analysis: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Singlecell Qc (xuzhougeng/wisp-science, 1k stars) and Trackplot (ygidtu/trackplot, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Comparative Genomics Synteny Analysis?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.