Agent skill

Bio Metagenomics Kraken

by GPTomics in GPTomics/bioSkills

Classifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation.

MITAuto-check passedResearch & Science

Install Bio Metagenomics Kraken

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-metagenomics-kraken --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/metagenomics/kraken-classification .claude/skills/bio-metagenomics-kraken && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-metagenomics-kraken
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.1k tokens
SKILL.md length
1,796 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Classifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation.

  • Works in 3 steps: Classification is not presence. At… → Read count is not abundance. Counts… → A taxon at the bottom of the report is a…
  • Profiling who-is-there from shotgun reads
  • SKILL.md covers Version Compatibility, The Single Most Important…, Why "Exact K-mer" Is Wrong for… and Tool Taxonomy, plus 10 more sections
  • Runs Shell scripts from its folder; calls pip

What it does

Bio Metagenomics Kraken is an agent skill from GPTomics/bioSkills. Classifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation. Covers why the database (not the algorithm) decides what can be detected, the --confidence and --minimum-hit-groups precision levers, unique-minimizer false-positive control, host-read removal, and why raw Kraken2 read counts are not abundances. Use when profiling who-is-there from shotgun reads, choosing a Kraken2 database, setting a…

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/kraken2_classify.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Profiling who-is-there from shotgun reads
  • Choosing a Kraken2 database
  • Setting a confidence threshold
  • Controlling false positives

Example prompts

  • “Use the bio-metagenomics-kraken skill to classify shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference…”
  • “/bio-metagenomics-kraken”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Classification is not presence. At --confidence 0 with the LCA rule, one shared k-mer labels a read with its nearest database relative…
  2. Read count is not abundance. Counts scale with genome length and copy number, so the report percentage is a fragment fraction, not a cell…
  3. A taxon at the bottom of the report is a hypothesis, not a finding. Single-region hits, hash collisions, and contaminated references…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Metagenomics Kraken loads about 4.1k tokens when it runs. Until then it costs about 196 tokens; SKILL.md has 1,796 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~196
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,796 words, ~4,147 tokens.

Download SKILL.mdSave it as .claude/skills/bio-metagenomics-kraken/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-metagenomics-kraken
description
Classifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation. Covers why the database (not the algorithm) decides what can be detected, the --confidence and --minimum-hit-groups precision levers, unique-minimizer false-positive control, host-read removal, and why raw Kraken2 read counts are not abundances. Use when profiling who-is-there from shotgun reads, choosing a Kraken2 database, setting a confidence threshold, controlling false positives, or feeding reports to Bracken. For marker-gene profiling see metaphlan-profiling; for abundance mechanics see abundance-estimation; for assembly/MAG recovery see genome-assembly/metagenome-assembly.
tool_type
cli
primary_tool
Kraken2

Version Compatibility

Reference examples tested with: Kraken2 2.1.3+, Bracken 2.9+, KrakenTools 1.2+, pandas 2.2+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: kraken2 --version, bracken -h, kraken2 --help to confirm flags and defaults
  • Python: pip show <package> then help(module.function) to check signatures

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

The DATABASE is the version that matters most. A taxon absent from the database is invisible no matter how the binary is configured. Record the Kraken2 database build (Standard / PlusPF / PlusPFP / nt / GTDB, and whether capped to 8/16 GB) and the Bracken databaseRLENmers.kmer_distrib read length, which must match both the Kraken2 database and the actual read length. --minimum-hit-groups enforcement varied across older 2.x point releases; confirm defaults with kraken2 --help on the installed build.

Kraken Classification

"What's in my metagenome?" -> Match each read's k-mers to a reference database by lowest-common-ancestor, then re-estimate abundance with Bracken - because the database, not the algorithm, decides what can be found.

  • CLI: kraken2 --db DB --paired R1.fq.gz R2.fq.gz --report out.kreport --confidence 0.1 --output out.kraken

Scope: read-based, assembly-free taxonomic classification of shotgun reads, plus the Bracken handoff. Marker-gene profiling -> metaphlan-profiling. Bracken command mechanics -> abundance-estimation. Genome/MAG recovery -> genome-assembly/metagenome-assembly. Host removal and read QC -> read-qc/contamination-screening, contamination-controls. Amplicon/16S -> the microbiome category.

The Single Most Important Modern Insight -- A Kraken Report Is a Database-Conditioned Similarity Ledger, Not a Sample Inventory

Kraken reports what each read most resembles in THIS database - never what is truly present, and never how much. Every number is hostage to three choices made before the run: the database, the confidence threshold, and the assumption that read count means abundance. Three corollaries each common misuse violates:

  1. Classification is not presence. At --confidence 0 with the LCA rule, one shared k-mer labels a read with its nearest database relative even when the true organism is absent. Absence-from-database becomes a confident wrong species.
  2. Read count is not abundance. Counts scale with genome length and copy number, so the report percentage is a fragment fraction, not a cell fraction. Bracken fixes the wrong-rank problem; it does not fix this.
  3. A taxon at the bottom of the report is a hypothesis, not a finding. Single-region hits, hash collisions, and contaminated references populate the long tail. Unique-minimizer coverage separates a real low-abundance organism from a phantom.

Organize the analysis around defending against these three, not around listing flags. Kraken2 at defaults over-classifies; Kraken2 tuned (right database + confidence + hit-groups + a unique-k-mer floor + host removal) is competitive with any classifier.

Why "Exact K-mer" Is Wrong for Kraken2

Kraken1 (Wood & Salzberg 2014 Genome Biol 15:R46) stored every exact k-mer. Kraken2 (Wood 2019 Genome Biol 20:257) replaced that with three ideas that make it fast and lean but PROBABILISTIC: (a) minimizers collapse each k-mer (default k=35) to the smallest hashed l=31-mer in its window; (b) a spaced seed (s=7 masked positions) tolerates errors at "don't care" positions; (c) a compact hash table stores only high bits of each key. The compact hash can return a wrong or spurious LCA on collision - which is exactly why the precision levers below exist. Calling Kraken2 "exact k-mer" hides where false positives come from.

Tool Taxonomy

ToolCitationMechanism / roleWhen
Kraken2Wood 2019 Genome Biol 20:257minimizer + spaced-seed + compact-hash LCAfast read classification; database-bound; the default choice
BrackenLu 2017 PeerJ Comput Sci 3:e104Bayesian redistribution of reads stranded at higher ranksalways run after Kraken2 for species/genus estimates
KrakenUniqBreitwieser 2018 Genome Biol 19:198HyperLogLog count of unique k-mers per taxonfalse-positive control on hits of interest
KMCPShen 2023 Bioinformatics 39:btac845genome-coverage pseudo-mappinglow-depth/clinical/viral where conserved-region FPs hurt
sourmash gatherPierce 2019 F1000Res 8:1006FracMinHash containment, min-set-cover"which genomes are present" with calibrated containment
MetaPhlAn 4Blanco-Miguez 2023 Nat Biotechnol 41:1633clade-specific marker genes-> metaphlan-profiling; FP-conservative, abundance directly

Decision Tree by Scenario

ScenarioRecommendedWhy
Who-is-there, fast, custom database possibleKraken2 + Brackenk-mer LCA; Bracken fixes count->rank, not count->cells
Need species relative abundance with no database build-> metaphlan-profilingmarker-based; abundance-conservative; no FP tail
Low-biomass / clinical pathogen IDKraken2 (high confidence + hit-groups) + KrakenUniq unique-k-mer floorthe FP tail is the enemy; one unique-k-mer filter cuts phantoms
Reads not host-depleted / unQC'd-> contamination-controls, read-qc/contamination-screening firsthost reads swamp the profile; references carry human fragments
Want abundance comparable across studiesstate the database + confidence; do not merge with MetaPhlAn percentagesread fraction != cell fraction; different tools = different quantities
Recover genomes / novel taxa / MAGs-> genome-assembly/metagenome-assemblyclassification is assembly-free and database-bound
16S amplicon reads-> microbiome categoryKraken-on-16S works but amplicon analysis lives there

Basic Classification

bash
# Paired-end, with the precision levers that defaults omit
kraken2 --db "$KRAKEN_DB" \
    --paired --gzip-compressed --threads 8 \
    --confidence 0.1 \            # raise from default 0 to suppress single-k-mer false positives
    --minimum-hit-groups 2 \      # require >=2 distinct hit regions (default 2; raise to 3 for clinical)
    --report out.kreport \
    --output out.kraken \
    R1.fq.gz R2.fq.gz

--paired joins mates with a k-mer-breaking N and classifies the pair as one fragment, raising specificity. The per-read --output (large) can be dropped to /dev/null once the .kreport is what feeds Bracken. --memory-mapping runs without loading the database into RAM (slower; for low-memory hosts).

False-Positive Control: Unique Minimizers

Goal: Separate a real low-abundance organism from a single-region phantom before believing any tail taxon.

Approach: Enable --report-minimizer-data so the report carries distinct-minimizer counts; a taxon with many reads but few distinct minimizers is hitting one conserved region and is a red flag. KrakenUniq's HyperLogLog unique-k-mer count is the heavier-weight version of the same signal.

bash
kraken2 --db "$KRAKEN_DB" --paired --confidence 0.1 \
    --report-minimizer-data \    # inserts 2 columns: total + DISTINCT minimizers (shifts later columns)
    --report out.kreport --output /dev/null \
    R1.fq.gz R2.fq.gz
# A species with high reads but low distinct-minimizers = false positive (one region lit up repeatedly).

Calibrate a unique-k-mer floor against negative controls rather than hard-coding one; the clinical ">=1024 unique k-mers" cutoff is dataset/database-specific folklore, not a constant.

Build a Custom Database

Goal: Build a database whose contents define exactly the detectable universe (and include human for host capture).

Approach: Download taxonomy, add the libraries the question needs (including human), build the minimizer index, then build the matching Bracken distributions at the actual read length.

bash
kraken2-build --download-taxonomy --db custom_db
for lib in bacteria archaea viral human UniVec_Core; do
    kraken2-build --download-library "$lib" --db custom_db
done
kraken2-build --build --db custom_db --threads 16    # writes hash.k2d, opts.k2d, taxo.k2d
kraken2-build --clean --db custom_db                 # drop library/ + taxonomy/ to shrink
bracken-build -d custom_db -t 16 -k 35 -l 150        # -k MUST equal the Kraken2 k (35); -l = read length

kraken2-build --special gtdb builds a GTDB-taxonomy database (curated; the greengenes/silva/rdp special downloads have rotted). --max-db-size randomly downsamples k-mers to fit a cap - this is how the prebuilt 8gb/16gb databases are made, and the reason confidence collapses classification on them.

Hand Off to Bracken

Kraken strands reads at the shared genus when species share k-mers; Bracken redistributes them down using genome-derived priors. It fixes the wrong-rank problem only - never genome-size bias, and never false positives (it can amplify or even invent a species by reassigning an absent organism's reads to its nearest congener). Run FP control first. Command mechanics live in abundance-estimation:

bash
bracken -d "$KRAKEN_DB" -i out.kreport -o out.bracken -w out.bracken.kreport \
    -r 150 \   # MUST match a built databaseRLENmers.kmer_distrib AND the actual read length
    -l S -t 10 # species level; -t is a redistribution floor (drops taxa with fewer than 10 clade-level reads, strict <), not a confidence

Per-Method Failure Modes

Over-classification at default confidence

Trigger: running --confidence 0 and reporting the species list. Mechanism: one shared k-mer can classify a read; the LCA labels it with its nearest database relative. Symptom: hundreds of low-abundance species, many biologically implausible. Fix: --confidence 0.1-0.4 (database-dependent) plus --minimum-hit-groups >=2; verify the tail with unique minimizers.

Show full SKILL.md (704 more words)Show less
Counts read as abundance

Trigger: using the .kreport percentage column as relative abundance. Mechanism: read count is proportional to abundance x genome length x copy number. Symptom: large-genome taxa overstated; downstream diversity/ordination on a non-cell-fraction. Fix: run Bracken for the rank problem; treat even Bracken fraction_total_reads as a read fraction and hand off genome-size/copy-number caveats to abundance-estimation.

Bracken read-length mismatch

Trigger: -r 100 on 150 bp reads, or a database built only for a different length. Mechanism: the redistribution model is fragment-length specific. Symptom: biased abundances with no error (silent) if the .kmer_distrib exists, hard crash if it does not. Fix: -r = actual read length AND a matching databaseRLENmers.kmer_distrib must exist (build it or pick a prebuilt database shipping that length).

Capped database plus high confidence

Trigger: a Standard-8/16 database with confidence cranked up. Mechanism: capped databases are random k-mer subsamples; few reads can clear a high threshold. Symptom: classification collapses toward zero; "my sample is mostly novel." Fix: use a full/large database for high confidence, or lower confidence on a capped database and accept lower precision.

Host reads and contaminated references

Trigger: classifying without host depletion, then trusting human and low-level hits. Mechanism: host reads dominate low-biomass samples, and >2 million GenBank entries carry mislabeled human/vector sequence (Steinegger & Salzberg 2020 Genome Biol 21:115). Symptom: confident Homo sapiens plus a long artifactual tail. Fix: include human in the database and/or host-deplete upstream; scrutinize any taxon co-varying with host load; consider Recentrifuge negative-control subtraction.

Quantitative Thresholds

ThresholdSourceRationale
--confidence 0.0 default; use 0.2-0.4 on a comprehensive DBLiu 2024 aBIOTECH 5:465; Lu 2022 Nat Protoc 17:2815species precision rose from ~0.16 to ~0.76 at CS 0.2; default over-classifies
--minimum-hit-groups 2 (raise to 3 for clinical)Kraken2 manuala single lucky minimizer/collision cannot make a call
Bracken -k = 35Lu 2017 PeerJ Comput Sci 3:e104must equal the Kraken2 database k-mer length
Bracken -r = actual read lengthBracken docsredistribution priors are fragment-length specific
Bracken -t 10 defaultBracken docsredistribution floor; too high deletes real rare taxa, not a confidence
Classification rate 30-70% (environmental)communitylow rate = novel taxa OR host contamination OR wrong database - diagnose which
Build k/l/s = 35/31/7 (nuc); 15/12/0 (prot)Kraken2 manualbuild-time only; cannot change k at classify time

Common Errors

Error / symptomCauseSolution
Hundreds of implausible species--confidence 0, no hit-group floorraise confidence; --minimum-hit-groups >=2; unique-k-mer filter
Bracken: "kmer_distrib file not found"-r has no matching built distributionrun bracken-build -l <readlen> or use a prebuilt DB shipping that length
Report parsing columns misaligned--report-minimizer-data inserted 2 columnsparse 8-column layout when the flag is on
Near-zero classification on a small DBcapped (downsampled) DB + high confidencelarger DB, or lower confidence on the capped DB
Confident human + odd tailhost reads + contaminated referenceshost-deplete first; treat human/tail as suspect
Bash example silently truncates flagsinline # comment after a \ line continuationput comments on their own line

References

  • Wood DE, Lu J, Langmead B. 2019. Improved metagenomic analysis with Kraken 2. Genome Biol 20:257.
  • Wood DE, Salzberg SL. 2014. Kraken: ultrafast metagenomic sequence classification using exact alignments. Genome Biol 15:R46.
  • Lu J, Breitwieser FP, Thielen P, Salzberg SL. 2017. Bracken: estimating species abundance in metagenomics data. PeerJ Comput Sci 3:e104.
  • Breitwieser FP, Baker DN, Salzberg SL. 2018. KrakenUniq: confident and fast metagenomics classification using unique k-mer counts. Genome Biol 19:198.
  • Lu J, Rincon N, Wood DE, Breitwieser FP, Pockrandt C, Langmead B, Salzberg SL, Steinegger M. 2022. Metagenome analysis using the Kraken software suite. Nat Protoc 17:2815-2839.
  • Liu Y, Ghaffari MH, Ma T, Tu Y. 2024. Impact of database choice and confidence score on the performance of taxonomic classification using Kraken2. aBIOTECH 5:465-475.
  • Steinegger M, Salzberg SL. 2020. Terminating contamination: large-scale search identifies more than 2,000,000 contaminated entries in GenBank. Genome Biol 21:115.
  • Shen W, Xiang H, Huang T, et al. 2023. KMCP: accurate metagenomic profiling of both prokaryotic and viral populations by pseudo-mapping. Bioinformatics 39:btac845.
  • abundance-estimation - Bracken command mechanics and read-count-to-abundance conversion
  • metaphlan-profiling - Marker-gene alternative; FP-conservative, abundance reported directly
  • metagenome-visualization - Plot and run community stats on the resulting profiles
  • contamination-controls - Host depletion, blanks, and decontam before classification
  • genome-assembly/metagenome-assembly - Assembly/MAG recovery; this category is read-based
  • read-qc/contamination-screening - Host/vector read screening before classification
  • workflows/metagenomics-pipeline - End-to-end shotgun profiling

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in metagenomics/kraken-classification of GPTomics/bioSkills.

  • SKILL.md
  • examples/kraken2_classify.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Metagenomics Kraken next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Metagenomics Kraken compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Metagenomics Kraken this skillGPTomics/bioSkills1.2k1 repos~4.1kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Metagenomics Kraken

What does Bio Metagenomics Kraken do?

Classifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation. Bio Metagenomics Kraken is an agent skill from GPTomics/bioSkills. Classifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation.

When should I use Bio Metagenomics Kraken?

Bio Metagenomics Kraken fits situations like: profiling who-is-there from shotgun reads; choosing a Kraken2 database; setting a confidence threshold; controlling false positives.

How do I install Bio Metagenomics Kraken in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a claude-code`. Or copy the skill folder (metagenomics/kraken-classification in GPTomics/bioSkills) into .claude/skills/bio-metagenomics-kraken in your project. Claude Code loads it when a task matches its description.

How do I install Bio Metagenomics Kraken in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a codex`. Or copy the skill folder (metagenomics/kraken-classification in GPTomics/bioSkills) into .agents/skills/bio-metagenomics-kraken in your project. Codex loads it when a task matches its description.

Can I use Bio Metagenomics Kraken in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-metagenomics-kraken, .gemini/skills/bio-metagenomics-kraken, .github/skills/bio-metagenomics-kraken and .opencode/skills/bio-metagenomics-kraken in your project.

What does Bio Metagenomics Kraken need to run?

Going by SKILL.md and its folder, Bio Metagenomics Kraken needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: A Bash shell.

Does Bio Metagenomics Kraken access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Metagenomics Kraken safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Metagenomics Kraken use?

Bio Metagenomics Kraken is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Metagenomics Kraken use?

About 4.1k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Metagenomics Kraken?

Skills that share tags, products or a category with Bio Metagenomics Kraken: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Metagenomics Kraken?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.