Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Classifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation.
$ npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-metagenomics-kraken --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/metagenomics/kraken-classification .claude/skills/bio-metagenomics-kraken && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-metagenomics-kraken" agent skill from https://github.com/GPTomics/bioSkills/tree/main/metagenomics/kraken-classification into .claude/skills/bio-metagenomics-kraken/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-metagenomics-kraken", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/metagenomics/kraken-classificationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-metagenomics-kraken --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/metagenomics/kraken-classification .agents/skills/bio-metagenomics-kraken && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-metagenomics-kraken" agent skill from https://github.com/GPTomics/bioSkills/tree/main/metagenomics/kraken-classification into .agents/skills/bio-metagenomics-kraken/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-metagenomics-kraken", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-metagenomics-kraken --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/metagenomics/kraken-classification .cursor/skills/bio-metagenomics-kraken && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-metagenomics-kraken" agent skill from https://github.com/GPTomics/bioSkills/tree/main/metagenomics/kraken-classification into .cursor/skills/bio-metagenomics-kraken/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-metagenomics-kraken", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path metagenomics/kraken-classification--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-metagenomics-kraken --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/metagenomics/kraken-classification .gemini/skills/bio-metagenomics-kraken && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-metagenomics-kraken" agent skill from https://github.com/GPTomics/bioSkills/tree/main/metagenomics/kraken-classification into .gemini/skills/bio-metagenomics-kraken/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-metagenomics-kraken", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-metagenomics-krakenInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/metagenomics/kraken-classification .github/skills/bio-metagenomics-kraken && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-metagenomics-kraken" agent skill from https://github.com/GPTomics/bioSkills/tree/main/metagenomics/kraken-classification into .github/skills/bio-metagenomics-kraken/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-metagenomics-kraken", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-metagenomics-kraken --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/metagenomics/kraken-classification .opencode/skills/bio-metagenomics-kraken && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-metagenomics-kraken" agent skill from https://github.com/GPTomics/bioSkills/tree/main/metagenomics/kraken-classification into .opencode/skills/bio-metagenomics-kraken/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-metagenomics-kraken", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-metagenomics-krakenClassifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation.
Bio Metagenomics Kraken is an agent skill from GPTomics/bioSkills. Classifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation. Covers why the database (not the algorithm) decides what can be detected, the --confidence and --minimum-hit-groups precision levers, unique-minimizer false-positive control, host-read removal, and why raw Kraken2 read counts are not abundances. Use when profiling who-is-there from shotgun reads, choosing a Kraken2 database, setting a…
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/kraken2_classify.sh` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Metagenomics Kraken loads about 4.1k tokens when it runs. Until then it costs about 196 tokens; SKILL.md has 1,796 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,796 words, ~4,147 tokens.
.claude/skills/bio-metagenomics-kraken/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: Kraken2 2.1.3+, Bracken 2.9+, KrakenTools 1.2+, pandas 2.2+.
Before using code patterns, verify installed versions match. If versions differ:
kraken2 --version, bracken -h, kraken2 --help to confirm flags and defaultspip show <package> then help(module.function) to check signaturesIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
The DATABASE is the version that matters most. A taxon absent from the database is invisible no matter how the binary is configured. Record the Kraken2 database build (Standard / PlusPF / PlusPFP / nt / GTDB, and whether capped to 8/16 GB) and the Bracken databaseRLENmers.kmer_distrib read length, which must match both the Kraken2 database and the actual read length. --minimum-hit-groups enforcement varied across older 2.x point releases; confirm defaults with kraken2 --help on the installed build.
"What's in my metagenome?" -> Match each read's k-mers to a reference database by lowest-common-ancestor, then re-estimate abundance with Bracken - because the database, not the algorithm, decides what can be found.
kraken2 --db DB --paired R1.fq.gz R2.fq.gz --report out.kreport --confidence 0.1 --output out.krakenScope: read-based, assembly-free taxonomic classification of shotgun reads, plus the Bracken handoff. Marker-gene profiling -> metaphlan-profiling. Bracken command mechanics -> abundance-estimation. Genome/MAG recovery -> genome-assembly/metagenome-assembly. Host removal and read QC -> read-qc/contamination-screening, contamination-controls. Amplicon/16S -> the microbiome category.
Kraken reports what each read most resembles in THIS database - never what is truly present, and never how much. Every number is hostage to three choices made before the run: the database, the confidence threshold, and the assumption that read count means abundance. Three corollaries each common misuse violates:
--confidence 0 with the LCA rule, one shared k-mer labels a read with its nearest database relative even when the true organism is absent. Absence-from-database becomes a confident wrong species.Organize the analysis around defending against these three, not around listing flags. Kraken2 at defaults over-classifies; Kraken2 tuned (right database + confidence + hit-groups + a unique-k-mer floor + host removal) is competitive with any classifier.
Kraken1 (Wood & Salzberg 2014 Genome Biol 15:R46) stored every exact k-mer. Kraken2 (Wood 2019 Genome Biol 20:257) replaced that with three ideas that make it fast and lean but PROBABILISTIC: (a) minimizers collapse each k-mer (default k=35) to the smallest hashed l=31-mer in its window; (b) a spaced seed (s=7 masked positions) tolerates errors at "don't care" positions; (c) a compact hash table stores only high bits of each key. The compact hash can return a wrong or spurious LCA on collision - which is exactly why the precision levers below exist. Calling Kraken2 "exact k-mer" hides where false positives come from.
| Tool | Citation | Mechanism / role | When |
|---|---|---|---|
| Kraken2 | Wood 2019 Genome Biol 20:257 | minimizer + spaced-seed + compact-hash LCA | fast read classification; database-bound; the default choice |
| Bracken | Lu 2017 PeerJ Comput Sci 3:e104 | Bayesian redistribution of reads stranded at higher ranks | always run after Kraken2 for species/genus estimates |
| KrakenUniq | Breitwieser 2018 Genome Biol 19:198 | HyperLogLog count of unique k-mers per taxon | false-positive control on hits of interest |
| KMCP | Shen 2023 Bioinformatics 39:btac845 | genome-coverage pseudo-mapping | low-depth/clinical/viral where conserved-region FPs hurt |
| sourmash gather | Pierce 2019 F1000Res 8:1006 | FracMinHash containment, min-set-cover | "which genomes are present" with calibrated containment |
| MetaPhlAn 4 | Blanco-Miguez 2023 Nat Biotechnol 41:1633 | clade-specific marker genes | -> metaphlan-profiling; FP-conservative, abundance directly |
| Scenario | Recommended | Why |
|---|---|---|
| Who-is-there, fast, custom database possible | Kraken2 + Bracken | k-mer LCA; Bracken fixes count->rank, not count->cells |
| Need species relative abundance with no database build | -> metaphlan-profiling | marker-based; abundance-conservative; no FP tail |
| Low-biomass / clinical pathogen ID | Kraken2 (high confidence + hit-groups) + KrakenUniq unique-k-mer floor | the FP tail is the enemy; one unique-k-mer filter cuts phantoms |
| Reads not host-depleted / unQC'd | -> contamination-controls, read-qc/contamination-screening first | host reads swamp the profile; references carry human fragments |
| Want abundance comparable across studies | state the database + confidence; do not merge with MetaPhlAn percentages | read fraction != cell fraction; different tools = different quantities |
| Recover genomes / novel taxa / MAGs | -> genome-assembly/metagenome-assembly | classification is assembly-free and database-bound |
| 16S amplicon reads | -> microbiome category | Kraken-on-16S works but amplicon analysis lives there |
# Paired-end, with the precision levers that defaults omit
kraken2 --db "$KRAKEN_DB" \
--paired --gzip-compressed --threads 8 \
--confidence 0.1 \ # raise from default 0 to suppress single-k-mer false positives
--minimum-hit-groups 2 \ # require >=2 distinct hit regions (default 2; raise to 3 for clinical)
--report out.kreport \
--output out.kraken \
R1.fq.gz R2.fq.gz--paired joins mates with a k-mer-breaking N and classifies the pair as one fragment, raising specificity. The per-read --output (large) can be dropped to /dev/null once the .kreport is what feeds Bracken. --memory-mapping runs without loading the database into RAM (slower; for low-memory hosts).
Goal: Separate a real low-abundance organism from a single-region phantom before believing any tail taxon.
Approach: Enable --report-minimizer-data so the report carries distinct-minimizer counts; a taxon with many reads but few distinct minimizers is hitting one conserved region and is a red flag. KrakenUniq's HyperLogLog unique-k-mer count is the heavier-weight version of the same signal.
kraken2 --db "$KRAKEN_DB" --paired --confidence 0.1 \
--report-minimizer-data \ # inserts 2 columns: total + DISTINCT minimizers (shifts later columns)
--report out.kreport --output /dev/null \
R1.fq.gz R2.fq.gz
# A species with high reads but low distinct-minimizers = false positive (one region lit up repeatedly).Calibrate a unique-k-mer floor against negative controls rather than hard-coding one; the clinical ">=1024 unique k-mers" cutoff is dataset/database-specific folklore, not a constant.
Goal: Build a database whose contents define exactly the detectable universe (and include human for host capture).
Approach: Download taxonomy, add the libraries the question needs (including human), build the minimizer index, then build the matching Bracken distributions at the actual read length.
kraken2-build --download-taxonomy --db custom_db
for lib in bacteria archaea viral human UniVec_Core; do
kraken2-build --download-library "$lib" --db custom_db
done
kraken2-build --build --db custom_db --threads 16 # writes hash.k2d, opts.k2d, taxo.k2d
kraken2-build --clean --db custom_db # drop library/ + taxonomy/ to shrink
bracken-build -d custom_db -t 16 -k 35 -l 150 # -k MUST equal the Kraken2 k (35); -l = read lengthkraken2-build --special gtdb builds a GTDB-taxonomy database (curated; the greengenes/silva/rdp special downloads have rotted). --max-db-size randomly downsamples k-mers to fit a cap - this is how the prebuilt 8gb/16gb databases are made, and the reason confidence collapses classification on them.
Kraken strands reads at the shared genus when species share k-mers; Bracken redistributes them down using genome-derived priors. It fixes the wrong-rank problem only - never genome-size bias, and never false positives (it can amplify or even invent a species by reassigning an absent organism's reads to its nearest congener). Run FP control first. Command mechanics live in abundance-estimation:
bracken -d "$KRAKEN_DB" -i out.kreport -o out.bracken -w out.bracken.kreport \
-r 150 \ # MUST match a built databaseRLENmers.kmer_distrib AND the actual read length
-l S -t 10 # species level; -t is a redistribution floor (drops taxa with fewer than 10 clade-level reads, strict <), not a confidenceTrigger: running --confidence 0 and reporting the species list. Mechanism: one shared k-mer can classify a read; the LCA labels it with its nearest database relative. Symptom: hundreds of low-abundance species, many biologically implausible. Fix: --confidence 0.1-0.4 (database-dependent) plus --minimum-hit-groups >=2; verify the tail with unique minimizers.
Trigger: using the .kreport percentage column as relative abundance. Mechanism: read count is proportional to abundance x genome length x copy number. Symptom: large-genome taxa overstated; downstream diversity/ordination on a non-cell-fraction. Fix: run Bracken for the rank problem; treat even Bracken fraction_total_reads as a read fraction and hand off genome-size/copy-number caveats to abundance-estimation.
Trigger: -r 100 on 150 bp reads, or a database built only for a different length. Mechanism: the redistribution model is fragment-length specific. Symptom: biased abundances with no error (silent) if the .kmer_distrib exists, hard crash if it does not. Fix: -r = actual read length AND a matching databaseRLENmers.kmer_distrib must exist (build it or pick a prebuilt database shipping that length).
Trigger: a Standard-8/16 database with confidence cranked up. Mechanism: capped databases are random k-mer subsamples; few reads can clear a high threshold. Symptom: classification collapses toward zero; "my sample is mostly novel." Fix: use a full/large database for high confidence, or lower confidence on a capped database and accept lower precision.
Trigger: classifying without host depletion, then trusting human and low-level hits. Mechanism: host reads dominate low-biomass samples, and >2 million GenBank entries carry mislabeled human/vector sequence (Steinegger & Salzberg 2020 Genome Biol 21:115). Symptom: confident Homo sapiens plus a long artifactual tail. Fix: include human in the database and/or host-deplete upstream; scrutinize any taxon co-varying with host load; consider Recentrifuge negative-control subtraction.
| Threshold | Source | Rationale |
|---|---|---|
--confidence 0.0 default; use 0.2-0.4 on a comprehensive DB | Liu 2024 aBIOTECH 5:465; Lu 2022 Nat Protoc 17:2815 | species precision rose from ~0.16 to ~0.76 at CS 0.2; default over-classifies |
--minimum-hit-groups 2 (raise to 3 for clinical) | Kraken2 manual | a single lucky minimizer/collision cannot make a call |
Bracken -k = 35 | Lu 2017 PeerJ Comput Sci 3:e104 | must equal the Kraken2 database k-mer length |
Bracken -r = actual read length | Bracken docs | redistribution priors are fragment-length specific |
Bracken -t 10 default | Bracken docs | redistribution floor; too high deletes real rare taxa, not a confidence |
| Classification rate 30-70% (environmental) | community | low rate = novel taxa OR host contamination OR wrong database - diagnose which |
| Build k/l/s = 35/31/7 (nuc); 15/12/0 (prot) | Kraken2 manual | build-time only; cannot change k at classify time |
| Error / symptom | Cause | Solution |
|---|---|---|
| Hundreds of implausible species | --confidence 0, no hit-group floor | raise confidence; --minimum-hit-groups >=2; unique-k-mer filter |
| Bracken: "kmer_distrib file not found" | -r has no matching built distribution | run bracken-build -l <readlen> or use a prebuilt DB shipping that length |
| Report parsing columns misaligned | --report-minimizer-data inserted 2 columns | parse 8-column layout when the flag is on |
| Near-zero classification on a small DB | capped (downsampled) DB + high confidence | larger DB, or lower confidence on the capped DB |
| Confident human + odd tail | host reads + contaminated references | host-deplete first; treat human/tail as suspect |
| Bash example silently truncates flags | inline # comment after a \ line continuation | put comments on their own line |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in metagenomics/kraken-classification of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Metagenomics Kraken next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Metagenomics Kraken this skillGPTomics/bioSkills | 1.2k | 1 repos | ~4.1k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Classifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation. Bio Metagenomics Kraken is an agent skill from GPTomics/bioSkills. Classifies shotgun metagenomic reads to taxa with Kraken2's minimizer/LCA matching against a chosen reference database, then hands off to Bracken for abundance re-estimation.
Bio Metagenomics Kraken fits situations like: profiling who-is-there from shotgun reads; choosing a Kraken2 database; setting a confidence threshold; controlling false positives.
Run `npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a claude-code`. Or copy the skill folder (metagenomics/kraken-classification in GPTomics/bioSkills) into .claude/skills/bio-metagenomics-kraken in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a codex`. Or copy the skill folder (metagenomics/kraken-classification in GPTomics/bioSkills) into .agents/skills/bio-metagenomics-kraken in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-metagenomics-kraken -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-metagenomics-kraken, .gemini/skills/bio-metagenomics-kraken, .github/skills/bio-metagenomics-kraken and .opencode/skills/bio-metagenomics-kraken in your project.
Going by SKILL.md and its folder, Bio Metagenomics Kraken needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: A Bash shell.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Metagenomics Kraken is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.1k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Metagenomics Kraken: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.