Agent skill

Bio Epidemiological Genomics Pathogen Typing

by GPTomics in GPTomics/bioSkills

Assigns isolate identity at the right resolution for the question -- ANI/Mash species triage, 7-locus MLST historical comparability, cgMLST/wgMLST outbreak resolution (chewBBACA, BIGSdb, Ridom…

MITAuto-check passedResearch & Science

Install Bio Epidemiological Genomics Pathogen Typing

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-epidemiological-genomics-pathogen-typing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-epidemiological-genomics-pathogen-typing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/epidemiological-genomics/pathogen-typing .claude/skills/bio-epidemiological-genomics-pathogen-typing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-epidemiological-genomics-pathogen-typing
GitHub stars
1.2k
Used in
1 other repo
Token cost
~8.7k tokens
SKILL.md length
3,988 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Assigns isolate identity at the right resolution for the question -- ANI/Mash species triage, 7-locus MLST historical comparability, cgMLST/wgMLST outbreak resolution (chewBBACA, BIGSdb, Ridom…

  • Typing bacterial isolates for surveillance
  • SKILL.md covers Version Compatibility, The Single Most Important…, Algorithmic Taxonomy and Decision Tree by Scenario, plus 10 more sections
  • Runs Python scripts from its folder; calls jq and pip
  • Outbreak investigation

What it does

Bio Epidemiological Genomics Pathogen Typing is an agent skill from GPTomics/bioSkills. Assigns isolate identity at the right resolution for the question -- ANI/Mash species triage, 7-locus MLST historical comparability, cgMLST/wgMLST outbreak resolution (chewBBACA, BIGSdb, Ridom SeqSphere, EnteroBase HierCC), in-silico serotyping (SISTR/SeqSero2 Salmonella, SerotypeFinder E. coli, Kaptive Klebsiella, SeroBA pneumococcus, spa+SCCmec S. aureus), and lineage callers (TB-Profiler/Mykrobe barcode for MTBC, Pangolin + Nextclade for SARS-CoV-2, PopPUNK GPSC for S. pneumoniae). Use when typing bacterial…

Its SKILL.md is about 8.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/mlst_typing.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Performance optimization. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Typing bacterial isolates for surveillance
  • Outbreak investigation
  • Choosing between cgMLST allele distance and core-SNP distance for cluster definition
  • Harmonising calls across schemas/database versions

Example prompts

  • “Use the bio-epidemiological-genomics-pathogen-typing skill to assign isolate identity at the right resolution for the question -- ANI/Mash species…”
  • “/bio-epidemiological-genomics-pathogen-typing”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • jq
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Epidemiological Genomics Pathogen Typing loads about 8.7k tokens when it runs. Until then it costs about 254 tokens; SKILL.md has 3,988 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~254
When it runs · the whole SKILL.md, loaded when a task matches
~8.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 3,988 words, ~8,725 tokens.

Download SKILL.mdSave it as .claude/skills/bio-epidemiological-genomics-pathogen-typing/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-epidemiological-genomics-pathogen-typing
description
Assigns isolate identity at the right resolution for the question -- ANI/Mash species triage, 7-locus MLST historical comparability, cgMLST/wgMLST outbreak resolution (chewBBACA, BIGSdb, Ridom SeqSphere, EnteroBase HierCC), in-silico serotyping (SISTR/SeqSero2 Salmonella, SerotypeFinder E. coli, Kaptive Klebsiella, SeroBA pneumococcus, spa+SCCmec S. aureus), and lineage callers (TB-Profiler/Mykrobe barcode for MTBC, Pangolin + Nextclade for SARS-CoV-2, PopPUNK GPSC for S. pneumoniae). Use when typing bacterial isolates for surveillance or outbreak investigation, choosing between cgMLST allele distance and core-SNP distance for cluster definition, harmonising calls across schemas/database versions, assigning MTBC lineage with the Napier 90-SNP barcode, calling Salmonella serovar via SISTR with monophasic Typhimurium awareness, running Pangolin UShER mode with pangolin-data version pinning, or selecting a typing resolution to match the surveillance question.
tool_type
mixed
primary_tool
chewBBACA

Version Compatibility

Reference examples tested with: mlst 2.23+, chewBBACA 3.3+, SISTR 1.1+, SeqSero2 1.3+, SerotypeFinder 2.0+, Kleborate 3.0+, Kaptive 3.0+, SeroBA 1.0+, PopPUNK 2.7+, pangolin 4.3+ (pangolin-data 1.30+), nextclade 3.8+, tb-profiler 6.2+, mykrobe 0.13+, mash 2.3+, skani 0.2+, snippy 4.6+, snp-dists 0.8+, pandas 2.2+, BioPython 1.84+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name
  • CLI: <tool> --version then <tool> --help to confirm flags
  • Pangolin: pangolin --all-versions records pangolin + pangolin-data + scorpio + constellations
  • Nextclade dataset: nextclade dataset list --tag latest sars-cov-2
  • chewBBACA schema: schemas are versioned independently; record the schema source AND the schema-fetch date for any published call
  • TB-Profiler bundled DB: tb-profiler list_db for the WHO catalogue edition currently bundled

If a tool reports an unexpected serovar / lineage / ST, the schema or barcode version is the first thing to check; introspect the installed package and database date before re-running.

Pathogen Typing

"What is this isolate, and is it the same as that one?" -> Pick a typing resolution that matches the epidemiological question, then run the appropriate caller and report the call with explicit schema / database / lineage-barcode versions. Resolution mismatch is the single most common typing error -- using 7-locus MLST to investigate a 20-isolate outbreak (multiple unrelated isolates share an ST) and using cgMLST to track decade-long lineage trends (allele drift obscures the lineage signal) are equally wrong, in opposite directions.

  • CLI: mlst assembly.fa -- 7-locus MLST via Seemann's mlst with PubMLST schemas
  • CLI: chewBBACA.py AlleleCall -i assemblies/ -g schema/ -o alleles/ --cpu 8 -- cgMLST allele profile
  • CLI: pangolin sequences.fasta --analysis-mode usher --all-versions -- SARS-CoV-2 Pango lineage with version provenance
  • CLI: tb-profiler profile -1 r1.fq.gz -2 r2.fq.gz -p sample -- MTBC lineage (Coll/Napier barcode) plus DR call
  • Python: pandas + snp-dists for cluster definition with pathogen-specific thresholds

The Single Most Important Modern Insight -- cgMLST and core-SNP distance answer different questions

cgMLST counts allele changes (one wobble within a 1-kb locus = 1 allelic difference, regardless of how many SNPs sit inside that allele) and is robust to small assembly artifacts; core-SNP counts every position and is sensitive to alignment / mapping artifacts but recovers higher resolution. The EFSA harmonised foodborne approach uses cgMLST for Salmonella / Listeria / E. coli (Salm cgMLST <=5 alleles = cluster; PulseNet Listeria <=4 alleles); UK / EU TB outbreak literature uses core-SNP with recombination masking (Walker 2013: <=12 SNPs = likely transmission; <=5 = recent). Mixing the two yields incompatible cluster definitions: the same 50-isolate outbreak under cgMLST may cluster as 1 group at threshold 5, under core-SNP as 3 groups at threshold 12. The output of pathogen-typing is not a single distance -- it is a distance metric, a threshold derived from a specific population, and a schema or reference version. All three must travel with the call.

Algorithmic Taxonomy

MethodMechanismResolutionStrengthFails when
ANI / fastANI / skaniPairwise nucleotide identity over orthologous regionsSpecies (>=95% ANI = same species)Species-level QC; cross-genus is non-metric for MashSpecies below ~80% ANI; Mash distance violates triangle inequality
Mash (Ondov 2016 Genome Biol 17:132)MinHash sketches; fast pairwise distancesSpecies triageSeconds per pair; the rapid screening standardDistance non-metric at low identity; ANI <80% unreliable
7-locus MLST (Maiden 2013 Nat Rev Microbiol 11:728; Seemann mlst tool)PubMLST allele lookup at 7 housekeeping lociSequence typeHistorical comparability across decadesInsufficient resolution for outbreaks; multiple unrelated isolates may share ST
cgMLST (chewBBACA; BIGSdb; Ridom SeqSphere)Allele lookup across ~1000-3000 core lociOutbreak / multi-country surveillanceAllele-distance is robust to small mapping errorsSchema-version-dependent; cross-schema NON-comparable; missing-locus handling matters
wgMLSTAllele lookup across all genes including accessoryHighest typing resolutionMaximum resolution within schemaEven more schema-dependent than cgMLST
Core-SNP typing (snippy + snp-dists; Parsnp; Lyve-SET)Reference-based SNP calling on core genomeSingle-SNP resolutionHighest resolution; underlies most Mtb workReference-dependent; recombination must be masked for bacteria
HierCC (Zhou 2021 Bioinformatics 37:3645)Hierarchical clustering on cgMLST at multiple thresholds (HC5, HC10, HC50, etc.)Stable nomenclature across schema updatesCross-schema-update stability for EnteroBase pathogensLimited to EnteroBase organisms
PopPUNK GPSC (Lees 2019 Genome Res 29:304)k-mer-based clustering with Gaussian mixture on core+accessory distancePopulation clusterStable IDs across additions; scales to >100k genomesCluster membership of an individual isolate can shift as model is refined
Pangolin (O'Toole 2021 Virus Evol 7:veab064)Phylogenetic placement (UShER) or ML classifier (pangoLEARN)SARS-CoV-2 Pango lineageCurated dynamic nomenclature; recombinant X-prefix designationspangoLEARN deprecated mid-2023; lab-to-lab version skew silently flips lineage calls
Nextclade (Aksamentov 2021 JOSS 6:3773)Reference-tree placement + clade calling + QCClade nomenclature + mutations + QCMutation reports + QC integrated; multi-pathogen datasetsDataset version drift changes which mutations count as "lineage-defining"
TB-Profiler + Coll/Napier barcode (Phelan 2019; Coll 2014; Napier 2020)Reference-based SNP call against H37Rv + barcode SNP setMTBC lineage 1-9 + drug resistanceIntegrated lineage + DST; supports Napier 90-SNP barcode covering lineages 7-9Pre-Napier 2020 barcodes miscall lineage 7-9 isolates
Mykrobe (Hunt 2019 Wellcome Open Res 4:191)k-mer presence/absence panelsSpecies + AMR for TB / S. aureus / SalmonellaFast; cross-check for TB-ProfilerPanel may lag WHO catalogue
SISTR (Yoshida 2016 PLoS ONE 11:e0147101)cgMLST + ribosomal MLST + serovar inferenceSalmonella serovar + antigenic formula~94% concordance with traditional sero-typing; monophasic-awareNovel/rare serovars without reference panel; antigen-cluster regulatory mutations are silent
SeqSero2 (Zhang 2019 AEM 85:e01746-19)k-mer + targeted-assembly for SalmonellaSalmonella serovarDesigned for low-coverage / fragmented dataSlightly different output schema than SISTR
SerotypeFinder (Joensen 2015 J Clin Microbiol 53:2410)BLAST against O- and H-antigen biosynthesis genesE. coli O:HStandard for E. coli serotypingMisses novel O / H types; fimH typing is separate
Kleborate + Kaptive (Lam 2021 Nat Commun 12:4188)Integrated MLST + K/O typing + virulence (ICEKp / iuc / ybt / clb / iro) + AMRKlebsiella surveillanceHypervirulence vs classical distinction; K/O loci typed by KaptiveKaptive K-locus DB versioned; KL calls can flip between Kaptive v1 / v2 / v3
spa + SCCmec (Harmsen 2003 J Clin Microbiol 41:5442; Kaya 2018 mSphere 3:e00612-17)spa repeat-region typing + SCCmec cassette typingS. aureus typingHistorical comparability; clinical surveillancespa repeat array fragments at borderline read length; assembler-dependent
SeroBA (Epping 2018 Microb Genom 4:e000186)k-mer-based serotyping from raw readsS. pneumoniae serotype98% concordance; no assembly needed; runs on >=15x coverageVaccine replacement makes serotype the load-bearing surveillance unit

Decision Tree by Scenario

ScenarioRecommendedWhy wrong choices fail
"Is this strain even what we think it is?"Mash / skani ANI triage; ANI >=95% confirms species7-locus MLST cannot answer species; cgMLST schema may not load on wrong species
Routine Salmonella serotypingSISTR (assembled) or SeqSero2 (low coverage / fragmented); flag monophasic 1,4,[5],12:i:- explicitlySingle tool without monophasic awareness; reporting Typhimurium when the antigen cluster is deleted
Klebsiella surveillanceKleborate (integrates MLST + K/O via Kaptive + ICEKp virulence + AMR via AMRFinderPlus)MLST + Kaptive + AMR run separately and not integrated -- missing hypervirulence call
S. pneumoniae surveillanceSeroBA serotype + PopPUNK GPSC; vaccine-replacement is the central post-PCV storyReporting MLST ST without serotype (serotype IS the vaccine-actionable surveillance unit)
S. aureus typingspa (Ridom) + SCCmec (SCCmecFinder) + MLST + clonal complex; flag CC8 USA300 / CC22 EMRSA-15 / CC30 EMRSA-16 / CC398 livestockJust spa without CC; spa repeat assembly artifacts unflagged
M. tuberculosis lineageTB-Profiler primary (Coll/Napier barcode) + Mykrobe cross-check; verify Napier 2020 90-SNP barcode (covers lineages 7-9)Pre-Napier barcodes miscall L7-9; MIRU-VNTR is obsolete for new surveillance
SARS-CoV-2 lineagePangolin with --analysis-mode usher (UShER default since v4; pangoLEARN deprecated mid-2023) + Nextclade cross-check; document pangolin-data versionpangoLEARN alone; reporting lineage without pangolin-data version pin
Define outbreak clustercgMLST allele-distance with pathogen-tuned threshold (Salm <=5; Listeria PulseNet <=4; TB <=12 SNPs core; C. difficile <=2 SNPs core); cite the threshold's source populationUniversal SNP threshold across pathogens (10x variation across taxa); applying Walker 2013 UK thresholds in high-transmission settings
Reproducible nomenclature across yearsPopPUNK GPSC (S. pneumoniae) / Pangolin (SARS-CoV-2) / HierCC (EnteroBase pathogens) / SISTR serovar -- curated stable IDsRaw cgMLST allele profile as the surveillance unit -- not stable across schema updates
Multi-lab cross-comparisonSame schema source AND same schema version AND same tool version; Pathogenwatch / EnteroBase shared platformLocally computed cgMLST profiles compared across labs without schema versioning

Methodology evolves; before any high-stakes typing report, verify Pangolin's current default analysis-mode and the Napier barcode currently bundled in TB-Profiler.

Running 7-Locus MLST and cgMLST

Goal: Produce reproducible ST and cgMLST allele-distance calls for a cohort of bacterial isolates, with schema versioning preserved for cross-lab comparison.

Approach: Run Seemann's mlst for the 7-locus baseline (uses bundled PubMLST schemas); for cgMLST, fetch the schema explicitly with chewBBACA.py DownloadSchema (records the source and date), then AlleleCall, then ExtractCgMLST with the per-organism missing-locus threshold; compute allele distances on the pairwise-complete intersection of called loci.

bash
mlst --threads 8 assemblies/*.fa > cohort.mlst.tsv

chewBBACA.py DownloadSchema \
    -sp "Salmonella enterica" \
    -sc cgMLST \
    -o schema_dir
SCHEMA_DATE=$(date -u +%Y-%m-%d)

chewBBACA.py AlleleCall \
    -i assemblies/ \
    -g schema_dir/cgMLST \
    -o alleles_out/ \
    --cpu 8

chewBBACA.py ExtractCgMLST \
    -i alleles_out/results_alleles.tsv \
    -o cgmlst_profile.tsv \
    --threshold 0.95

echo "schema_source: chewBBACA SalmonellaEnterica cgMLST" > cgmlst_profile.metadata
echo "schema_fetched: ${SCHEMA_DATE}" >> cgmlst_profile.metadata

AlleleCall output uses special codes: LNF (locus not found), PLOT (truncated), NIPH (non-informative paralog), ASM (allele small/short), ALM (allele large), and asterisk (new allele). These are MISSING DATA. Counting "0" or "-" against another "0" or "-" as 1 allelic difference is a quiet correctness bug that affects most cgMLST cluster definitions outside specialist labs.

SNP-Based Outbreak Cluster Definition

Goal: Identify isolates within an outbreak threshold using core-SNP distances on a recombination-aware alignment, with the pathogen-specific threshold cited from its source population.

Approach: Snippy against a high-quality reference; snippy-core to extract core SNPs; for bacteria, Gubbins to mask recombinant tracts before counting (otherwise apparent SNP distance is inflated by recombination); snp-dists for pairwise distance matrix; apply the published pathogen-specific threshold (citing Walker 2013 for TB, Eyre 2013 for C. difficile, EFSA convention for Salmonella).

bash
for r1 in reads/*_R1.fq.gz; do
    sample=$(basename "${r1}" _R1.fq.gz)
    r2="reads/${sample}_R2.fq.gz"
    snippy --outdir snippy_out/${sample} --R1 "${r1}" --R2 "${r2}" --reference reference.fa --cpus 8
done

snippy-core --ref reference.fa --prefix core snippy_out/*

run_gubbins.py --prefix gubbins core.full.aln

snp-dists -c gubbins.filtered_polymorphic_sites.fasta > cohort.snp_dists.csv

run_gubbins.py input MUST be core.full.aln (full-position alignment with reference). Passing core.aln (variable positions only) produces wrong recombination calls because Gubbins cannot estimate background SNP density without invariant positions.

Lineage Calling for Mtb and SARS-CoV-2

Goal: Assign MTBC lineage via the Napier 2020 barcode or SARS-CoV-2 Pango lineage via UShER placement, with explicit version pinning so the call is reproducible.

Approach: TB-Profiler bundles the Coll 2014 + Napier 2020 barcode and reports lineage 1-9 (Napier 2020 added lineages 7-9 that older 62-SNP barcodes miss); Pangolin with --analysis-mode usher performs phylogenetic placement on the daily-updated UShER tree and is preferred over pangoLEARN since v4 (Pongmoragot 2024 Virus Evol 10:vead085); always record pangolin --all-versions output alongside the lineage call.

bash
tb-profiler profile -1 reads_R1.fq.gz -2 reads_R2.fq.gz -p sample --dir tbp_out
mykrobe predict --sample sample --species tb --output sample.mykrobe.json --format json reads_R1.fq.gz reads_R2.fq.gz

pangolin sequences.fasta --analysis-mode usher --outfile lineage_report.csv
pangolin --all-versions > pangolin_versions.txt

nextclade dataset get --name sars-cov-2 --output-dir nc_dataset/sars-cov-2
NC_DATASET_TAG=$(jq -r '.tag' nc_dataset/sars-cov-2/pathogen.json)

nextclade run \
    --input-dataset nc_dataset/sars-cov-2 \
    --output-tsv nextclade.tsv \
    --output-json nextclade.json \
    sequences.fasta

echo "nextclade_dataset_tag: ${NC_DATASET_TAG}" > nextclade.metadata

Per-Method Failure Modes

chewBBACA missing-locus codes counted as allelic differences

Trigger: A naive cgMLST distance computation that treats LNF / PLOT / NIPH / ASM / ALM / "0" / "-" as integer alleles and counts them against other missing codes.

Mechanism: chewBBACA output uses special codes for missing or paralogous loci. The pairwise allele distance must be computed on the intersection of called loci, not the union. Treating two LNFs as "same allele" (=0 difference) or as "different alleles" (=1 difference) are both wrong; the locus must be excluded from the comparison.

Symptom: Outbreak cluster definition flips depending on completeness of the assemblies; samples with more missing data appear artificially close to each other (if missing-vs-missing counts 0) or far from everyone (if missing-vs-allele counts 1).

Fix: Use chewBBACA's ExtractCgMLST with a --threshold (typically 0.95) to drop loci called in fewer than that fraction of samples; for the pairwise distance, restrict to loci called in BOTH samples (pairwise-complete) and report the number of loci compared alongside the distance.

Pangolin lineage call flips between lab A (older pangolin-data) and lab B (current)

Trigger: Two labs submit the same consensus genome to Pangolin with different pangolin-data versions; the lineage call differs (e.g., BA.2 vs BA.2.86 vs JN.1).

Mechanism: Lineage designation happens through pango-designation GitHub issues -- community-driven, often days-to-weeks before pangolin-data releases include the lineage. During this window, the same genome is callable as the parent lineage (older pangolin-data) or the child (current pangolin-data). pangolin-data is updated weekly; lab-to-lab version skew is routine.

Symptom: Cross-lab lineage prevalence comparisons over time show implausible jumps that coincide with pangolin-data release dates rather than biology.

Fix: Pin pangolin-data version explicitly with pangolin --all-versions recorded alongside every call. For published or regulatory output, re-run the WHOLE archive against a single pangolin-data version before reporting. Comparing today's BA.2.86 call to last month's "Unassigned" call is invalid.

MTBC lineage 7-9 miscalled by pre-Napier barcode

Trigger: TB-Profiler / Mykrobe running an older bundled barcode (Coll 2014 62-SNP, pre-2020 builds); isolate is from Ethiopia, Rwanda, or East Africa.

Mechanism: The Coll 2014 62-SNP barcode covers lineages 1-7 but predates the formal designation of lineages 8 (Rwanda) and 9. The Napier 2020 90-SNP barcode (Napier Genome Med 12:114) adds these and refines L4 sublineages. An older bundled barcode silently maps lineage 8 / 9 isolates to "unknown" or to a nearest-barcode-match L4 sub-lineage based on partial SNPs.

Symptom: Cross-lab Mtb surveillance dataset shows Ethiopia / Rwanda isolates as a mix of "unknown" and unexpected L4 sublineages.

Fix: Verify the bundled barcode version (tb-profiler list_db); update to the Napier 2020 barcode or later before any lineage-stratified analysis. Re-run historical Mtb data against the current barcode whenever the barcode is updated.

Beijing-lineage Mtb literature based on spoligotype, not WGS sublineage

Trigger: Citing pre-2014 "Beijing family" findings (hypervirulence claims, vaccine-escape claims, faster-evolution claims) as if they apply to a single modern WGS sublineage.

Mechanism: Pre-WGS, the Mtb "Beijing" family was defined by spoligotype pattern (absence of DR spacers 1-34, presence of 35-43). WGS revealed Beijing is a paraphyletic grouping containing multiple sublineages with distinct phenotypes (modern Beijing = lineage 2.2.1.1; ancestral Beijing = 2.2.1.2; proto-Beijing = 2.1). Much of the pre-2014 Beijing literature was based on the spoligotype-defined paraphyletic group.

Symptom: Beijing-as-a-monolith conclusions inappropriately applied to modern WGS-typed isolates that may fall in different sublineages of L2.

Fix: Always specify the WGS sublineage (e.g., "2.2.1.1 modern Beijing", "2.2.1.2 ancestral Beijing") rather than "Beijing". Treat pre-2014 Beijing claims as hypotheses to be re-validated on WGS-typed cohorts.

Mash distance used to cluster genomes below 80% ANI

Trigger: Mash-distance hierarchical clustering of distantly related genomes (multi-genus or low-identity comparisons).

Mechanism: Mash distance (Ondov 2016 Genome Biol 17:132) is a Jaccard-derived estimator of mutation rate, validated for ANI >=80% (the Mash docs state 95% confidence at ANI >=90%). Below that, the Mash distance becomes non-metric -- A-B + B-C can be < A-C -- and clustering algorithms that assume metric distances produce undefined output.

Symptom: Cross-genus or low-identity comparison clusters do not match phylogenetic expectations; the same cluster definition is unstable to addition of new genomes.

Fix: Use Mash only within the validated ANI range (>=80%, ideally >=90%). For cross-genus or low-identity comparisons use AAI or ANI-from-alignment (skani, pyani). Document the ANI range of the cohort before computing any Mash-based distance.

Show full SKILL.md (1,642 more words)Show less
Cross-schema cgMLST distances pooled as if comparable

Trigger: Combining cgMLST distances from a chewBBACA local schema with distances from a Ridom SeqSphere schema in the same outbreak comparison.

Mechanism: Ridom SeqSphere, chewBBACA-built schemas, and EnteroBase schemas are independently curated and use different locus sets and different allele numbering. A "cgMLST distance of 5" between two isolates does NOT equal a "cgMLST distance of 5" under a different schema.

Symptom: Multi-country outbreak comparison reports irreconcilable cluster definitions; the same isolate pair appears in-cluster in one lab and out-of-cluster in another.

Fix: Document schema source AND schema-fetch date for every cgMLST profile; do not pool distances across schema sources. For multi-country collaboration, agree on a single schema (typically EnteroBase HierCC for Salmonella / E. coli / Listeria) at the outset.

K-locus typing flips between Kaptive versions

Trigger: Comparing K. pneumoniae K-locus prevalence between studies that used different Kaptive database versions.

Mechanism: The Kaptive K-locus DB has been versioned multiple times since release; K-locus calls have flipped (KL1 / KL2 boundary refinements; KL149+ additions) between Kaptive v1 and v2. Kaptive v3 (2024) added O-locus typing and an L-locus scheme.

Symptom: Longitudinal K-locus prevalence trend has implausible jumps coinciding with Kaptive release dates rather than biology.

Fix: Document Kleborate AND Kaptive version with every K/O call. For longitudinal trend analysis, re-run all historical assemblies against a single Kaptive version.

Reconciliation: When Typing Methods Disagree

PatternLikely causeAction
Pangolin "BA.2.86", Nextclade clade "23I"Equivalent at different resolutions -- BA.2.86 is within 23IReport both; Pango lineage for sub-clade resolution
Pangolin "BA.5.2", Nextclade "Unassigned"Nextclade dataset older than pangolin-data; OR Nextclade QC failedUpdate Nextclade dataset; re-run; inspect QC fields
Pangolin UShER and pangoLEARN disagreepangoLEARN is the deprecated decision-tree classifier (deprecated mid-2023)Trust UShER call
SISTR "monophasic Typhimurium 1,4,[5],12:i:-", slide agglutination "Typhimurium"fljB gene deleted in monophasic variant; slide agglutination cannot detectTrust genome call; flag for confirmatory testing if surveillance requires
cgMLST cluster definition flips between two schemasDifferent locus sets; non-comparablePick one schema for the entire analysis; document
TB lineage call differs between TB-Profiler and MykrobeDifferent bundled barcode versions; one may predate Napier 2020Update both; if disagreement persists, manual barcode SNP inspection
Two consecutive pangolin-data releases call the same consensus differentlyLineage definitions revised between releasesPin pangolin-data; re-run whole archive on dataset update

Quantitative Thresholds

QuantityThresholdSource / rationale
Mash / ANI species boundary>=95% ANIANI species-delineation convention
Mash distance validity range>=80% ANI (>=90% for 95% confidence)Ondov 2016 Genome Biol 17:132
cgMLST cluster -- Salmonella (chewBBACA Salm scheme)<=5 allelic differencesEFSA harmonised approach
cgMLST cluster -- Listeria monocytogenes (PulseNet)<=4 allelic differencesPulseNet protocol convention
cgMLST cluster -- E. coli (EnteroBase)<=10 allelic differences (STEC outbreak)EnteroBase convention
Core SNP -- M. tuberculosis<=5 SNPs (recent transmission); <=12 SNPs (likely transmission)Walker 2013 Lancet Infect Dis 13:137 (UK low-transmission setting)
Core SNP -- Staphylococcus aureus<=15 SNPs (within hospital outbreak); <=40 SNPs (broader temporal cluster)Coll 2017 Clin Infect Dis 65:1781
Core SNP -- Klebsiella pneumoniae (KPC outbreak)<=21 SNPsSnitkin 2012 Sci Transl Med 4:148ra116
Core SNP -- Clostridioides difficile (recombination-masked)<=2 SNPs (likely direct transmission); <=10 (plausible within 6 months)Eyre 2013 NEJM 369:1195
Core SNP -- Neisseria gonorrhoeae<=25 SNPs (transmission)UKHSA STI framework
SARS-CoV-2 cluster definitionNOT defined by SNP alone; combine 0-2 SNPs + epi link + sampling windowSARS-CoV-2 within-host diversity literature
chewBBACA cgMLST extraction completeness0.95 (drop loci called in <95% of samples)chewBBACA convention
SeroBA minimum coverage>=15xEpping 2018 Microb Genom 4:e000186

CRITICAL: a number from one pathogen does NOT transfer to another. Substitution rate, recombination, host range, generation interval, and within-host diversity vary by 100x across pathogens. Always cite the threshold's source population, especially for Walker 2013 (UK low-transmission setting) which routinely inflates apparent recent-transmission rates by 2-5x in high-burden settings.

Common Errors

Error / symptomCauseSolution
Outbreak split into spurious clusterscgMLST missing-locus codes counted as allelic differencesRestrict to pairwise-complete loci
Lineage prevalence shows implausible jumppangolin-data version driftPin pangolin-data; re-run archive
Mtb lineage call "unknown" for East African isolatePre-Napier barcodeUpdate to Napier 2020
Salmonella sample called "Typhimurium" when antigen cluster deletedTool did not flag monophasic variantUse SISTR or SeqSero2 with explicit monophasic check
nextclade run --input-dataset rejectedv2 syntax in v3v3 uses --input-dataset for pre-downloaded folder; verify nextclade --version
pangolin --inference usher not recognisedFlag is --analysis-mode usherUse --analysis-mode usher
chewBBACA fails on assembly with broken lociSchema BSR cutoff too strict; assembly too fragmentedDocument quality; consider re-assembly with long reads
GPSC cluster differs between studiesPopPUNK model updated; individual cluster membership can shiftRe-run PopPUNK against current reference DB before comparison
Mash hierarchical clustering unstableCohort spans wide ANI range; below validity thresholdUse ANI-from-alignment (skani / pyANI) for cross-genus comparisons
MOB-suite plasmid context conflicts with cgMLST clustercgMLST is chromosome-focused; plasmid distance is independentReport both with framing

Anticipated Reviewer Pushback

PushbackResponse
"Why cgMLST and not core-SNP?"EFSA harmonised foodborne uses cgMLST; cross-lab comparability via shared schema. For outbreak-internal who-infected-whom, core-SNP supplements
"What threshold was used, and on what population?"Cite Walker 2013 / Eyre 2013 / EFSA per pathogen; caveat the population for non-universal thresholds
"How were missing cgMLST loci handled?"Pairwise-complete distance; locus must be called in BOTH samples to count
"Pangolin version?"pangolin --all-versions recorded; re-run on dataset update
"Why TB-Profiler over Mykrobe?"TB-Profiler is primary (WHO catalogue integration + Napier barcode); Mykrobe is cross-check on R/XDR calls
"Was the Napier 2020 barcode checked?"Verified tb-profiler list_db; lineage 7-9 callable
"Was the Beijing-as-monolith literature cited?"Disaggregated to WGS sublineage (2.2.1.1 modern / 2.2.1.2 ancestral); pre-2014 Beijing claims treated as hypotheses
"Multi-country outbreak: harmonised schema?"EnteroBase HierCC for Salmonella / E. coli / Listeria; documented schema source + date

References

  • Maiden MCJ, van Rensburg MJJ, Bray JE et al (2013) MLST revisited: the gene-by-gene approach to bacterial genomics. Nat Rev Microbiol 11(10):728-736. doi:10.1038/nrmicro3093
  • Silva M, Machado MP, Silva DN et al (2018) chewBBACA: A complete suite for gene-by-gene schema creation and strain identification. Microb Genom 4(3):e000166. doi:10.1099/mgen.0.000166
  • Zhou Z, Charlesworth J, Achtman M (2021) HierCC: a multi-level clustering scheme for population assignments based on core genome MLST. Bioinformatics 37(20):3645-3646. doi:10.1093/bioinformatics/btab234
  • Yoshida CE, Kruczkiewicz P, Laing CR et al (2016) The Salmonella In Silico Typing Resource (SISTR). PLoS ONE 11(1):e0147101. doi:10.1371/journal.pone.0147101
  • Zhang S, den Bakker HC, Li S et al (2019) SeqSero2: rapid and improved Salmonella serotype determination using whole-genome sequencing data. Appl Environ Microbiol 85(23):e01746-19. doi:10.1128/AEM.01746-19
  • Joensen KG, Tetzschner AMM, Iguchi A et al (2015) Rapid and easy in silico serotyping of Escherichia coli isolates by use of whole-genome sequencing data. J Clin Microbiol 53(8):2410-2426. doi:10.1128/JCM.00008-15
  • Lam MMC, Wick RR, Watts SC et al (2021) A genomic surveillance framework and genotyping tool for Klebsiella pneumoniae and its related species complex. Nat Commun 12:4188. doi:10.1038/s41467-021-24448-3
  • Lees JA, Harris SR, Tonkin-Hill G et al (2019) Fast and flexible bacterial genomic epidemiology with PopPUNK. Genome Res 29(2):304-316. doi:10.1101/gr.241455.118
  • Epping L, van Tonder AJ, Gladstone RA et al (2018) SeroBA: rapid high-throughput serotyping of Streptococcus pneumoniae from whole genome sequence data. Microb Genom 4(7):e000186. doi:10.1099/mgen.0.000186
  • Harmsen D, Claus H, Witte W et al (2003) Typing of methicillin-resistant Staphylococcus aureus in a university hospital setting by using novel software for spa repeat determination and database management. J Clin Microbiol 41(12):5442-5448. doi:10.1128/JCM.41.12.5442-5448.2003
  • Kaya H, Hasman H, Larsen J et al (2018) SCCmecFinder, a web-based tool for typing of staphylococcal cassette chromosome mec in Staphylococcus aureus using whole-genome sequence data. mSphere 3(1):e00612-17. doi:10.1128/mSphere.00612-17
  • Coll F, McNerney R, Guerra-Assunção JA et al (2014) A robust SNP barcode for typing Mycobacterium tuberculosis complex strains. Nat Commun 5:4812. doi:10.1038/ncomms5812
  • Coll F, Harrison EM, Toleman MS et al (2017) Longitudinal genomic surveillance of MRSA in the UK reveals transmission patterns in hospitals and the community. Clin Infect Dis 65(11):1781-1789. doi:10.1093/cid/cix645
  • Snitkin ES, Zelazny AM, Thomas PJ et al (2012) Tracking a hospital outbreak of carbapenem-resistant Klebsiella pneumoniae with whole-genome sequencing. Sci Transl Med 4(148):148ra116. doi:10.1126/scitranslmed.3004129
  • Napier G, Campino S, Merid Y et al (2020) Robust barcoding and identification of Mycobacterium tuberculosis lineages for epidemiological and clinical studies. Genome Med 12(1):114. doi:10.1186/s13073-020-00817-3
  • Phelan JE, O'Sullivan DM, Machado D et al (2019) Integrating informatics tools and portable sequencing technology for rapid detection of resistance to anti-tuberculous drugs. Genome Med 11:41. doi:10.1186/s13073-019-0650-x
  • Hunt M, Bradley P, Lapierre SG et al (2019) Antibiotic resistance prediction for Mycobacterium tuberculosis from genome sequence data with Mykrobe. Wellcome Open Res 4:191. doi:10.12688/wellcomeopenres.15603.1
  • Walker TM, Ip CLC, Harrell RH et al (2013) Whole-genome sequencing to delineate Mycobacterium tuberculosis outbreaks: a retrospective observational study. Lancet Infect Dis 13(2):137-146. doi:10.1016/S1473-3099(12)70277-3
  • Eyre DW, Cule ML, Wilson DJ et al (2013) Diverse sources of C. difficile infection identified on whole-genome sequencing. N Engl J Med 369(13):1195-1205. doi:10.1056/NEJMoa1216064
  • Ondov BD, Treangen TJ, Melsted P et al (2016) Mash: fast genome and metagenome distance estimation using MinHash. Genome Biol 17:132. doi:10.1186/s13059-016-0997-x
  • O'Toole Á, Scher E, Underwood A et al (2021) Assignment of epidemiological lineages in an emerging pandemic using the pangolin tool. Virus Evol 7(2):veab064. doi:10.1093/ve/veab064
  • Aksamentov I, Roemer C, Hodcroft EB, Neher RA (2021) Nextclade: clade assignment, mutation calling and quality control for viral genomes. J Open Source Softw 6(67):3773. doi:10.21105/joss.03773
  • Pongmoragot J, Pearson C, Borg ML et al (2024) Comparison of UShER-based and pangoLEARN-based Pangolin lineage assignments for SARS-CoV-2 sequences. Virus Evol 10(1):vead085. doi:10.1093/ve/vead085
  • amr-surveillance - Strain context complements AMR; Kleborate integrates typing + AMR for Klebsiella
  • transmission-inference - SNP-cluster definition from cgMLST or core-SNP feeds outbreak transmission inference
  • phylodynamics - Time-scaled tree from typed isolates for R_e estimation
  • variant-surveillance - SARS-CoV-2 Pango / Nextclade lineage assignment overlaps; this skill owns the typing call, variant-surveillance owns longitudinal frequency tracking
  • comparative-genomics/pangenome-analysis - Core / accessory genome partitioning underlies cgMLST schema design
  • comparative-genomics/whole-genome-alignment - Core-genome alignment for SNP-typing
  • variant-calling/vcf-basics - Per-isolate VCF for SNP-typing
  • variant-calling/variant-calling - Per-isolate variant calling that feeds cgMLST and SNP-typing
  • read-alignment/bwa-alignment - Read mapping upstream of variant calling and snippy
  • alignment/multiple-alignment - Multiple sequence alignment for core SNP extraction
  • database-access/entrez-fetch - Reference genome retrieval for snippy / Snippy-core
  • metagenomics/strain-tracking - Community strain tracking via Kraken2 / StrainPhlAn (NOT isolate-focused)

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in epidemiological-genomics/pathogen-typing of GPTomics/bioSkills.

  • SKILL.md
  • examples/mlst_typing.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Epidemiological Genomics Pathogen Typing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Epidemiological Genomics Pathogen Typing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Epidemiological Genomics Pathogen Typing this skillGPTomics/bioSkills1.2k1 repos~8.7kAutomated safety check: PassMIT
Pathway EnrichmentK-Dense-AI/scientific-agent-skills48k1 repos~4.2kAutomated safety check: PassMIT
Celltype Specificity ProfilerClawBio/ClawBio1.2k—~4.3kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0

Similar skills

  • Pathway Enrichment

    K-Dense-AI/scientific-agent-skills

    Performs pathway and gene-set enrichment analysis on gene lists or ranked gene data and interprets the results.

    48k GitHub starsUsed in 1 repo~4.2k tokens
    Research & ScienceAuto-check passed
  • Given a gene and a single-cell atlas, compute how cell-type-specific its expression is — the tau specificity index, Sarle's expression bimodality coefficient, and the cell types that drive the…

    1.2k GitHub stars~4.3k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Epidemiological Genomics Pathogen Typing

What does Bio Epidemiological Genomics Pathogen Typing do?

Assigns isolate identity at the right resolution for the question -- ANI/Mash species triage, 7-locus MLST historical comparability, cgMLST/wgMLST outbreak resolution (chewBBACA, BIGSdb, Ridom…. Bio Epidemiological Genomics Pathogen Typing is an agent skill from GPTomics/bioSkills. Assigns isolate identity at the right resolution for the question -- ANI/Mash species triage, 7-locus MLST historical comparability, cgMLST/wgMLST outbreak resolution (chewBBACA, BIGSdb, Ridom SeqSphere, EnteroBase HierCC), in-silico serotyping (SISTR/SeqSero2 Salmonella, SerotypeFinder E.

When should I use Bio Epidemiological Genomics Pathogen Typing?

Bio Epidemiological Genomics Pathogen Typing fits situations like: typing bacterial isolates for surveillance; outbreak investigation; choosing between cgMLST allele distance and core-SNP distance for cluster definition; harmonising calls across schemas/database versions.

How do I install Bio Epidemiological Genomics Pathogen Typing in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-epidemiological-genomics-pathogen-typing -a claude-code`. Or copy the skill folder (epidemiological-genomics/pathogen-typing in GPTomics/bioSkills) into .claude/skills/bio-epidemiological-genomics-pathogen-typing in your project. Claude Code loads it when a task matches its description.

How do I install Bio Epidemiological Genomics Pathogen Typing in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-epidemiological-genomics-pathogen-typing -a codex`. Or copy the skill folder (epidemiological-genomics/pathogen-typing in GPTomics/bioSkills) into .agents/skills/bio-epidemiological-genomics-pathogen-typing in your project. Codex loads it when a task matches its description.

Can I use Bio Epidemiological Genomics Pathogen Typing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-epidemiological-genomics-pathogen-typing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-epidemiological-genomics-pathogen-typing, .gemini/skills/bio-epidemiological-genomics-pathogen-typing, .github/skills/bio-epidemiological-genomics-pathogen-typing and .opencode/skills/bio-epidemiological-genomics-pathogen-typing in your project.

What does Bio Epidemiological Genomics Pathogen Typing need to run?

Going by SKILL.md and its folder, Bio Epidemiological Genomics Pathogen Typing needs Python for the scripts in its folder and the command-line tools its instructions call (jq and pip). Our summary lists: Python 3.

Does Bio Epidemiological Genomics Pathogen Typing access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Epidemiological Genomics Pathogen Typing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Epidemiological Genomics Pathogen Typing use?

Bio Epidemiological Genomics Pathogen Typing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Epidemiological Genomics Pathogen Typing use?

About 8.7k tokens (SKILL.md is roughly 35k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Epidemiological Genomics Pathogen Typing?

Skills that share tags, products or a category with Bio Epidemiological Genomics Pathogen Typing: Pathway Enrichment (K-Dense-AI/scientific-agent-skills, 48k stars), Celltype Specificity Profiler (ClawBio/ClawBio, 1.2k stars), Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars) and 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Epidemiological Genomics Pathogen Typing?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.