Dbsnp Database
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
Infers exact amplicon sequence variants (ASVs) from demultiplexed 16S rRNA or ITS amplicon FASTQ with DADA2 - removing primers with cutadapt (--discard-untrimmed), learning a per-run error model…
$ npx skills add GPTomics/bioSkills --skill bio-microbiome-amplicon-processing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-microbiome-amplicon-processing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/microbiome/amplicon-processing .claude/skills/bio-microbiome-amplicon-processing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-microbiome-amplicon-processing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/microbiome/amplicon-processing into .claude/skills/bio-microbiome-amplicon-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-microbiome-amplicon-processing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/microbiome/amplicon-processingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-microbiome-amplicon-processing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-microbiome-amplicon-processing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/microbiome/amplicon-processing .agents/skills/bio-microbiome-amplicon-processing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-microbiome-amplicon-processing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/microbiome/amplicon-processing into .agents/skills/bio-microbiome-amplicon-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-microbiome-amplicon-processing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-microbiome-amplicon-processing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-microbiome-amplicon-processing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/microbiome/amplicon-processing .cursor/skills/bio-microbiome-amplicon-processing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-microbiome-amplicon-processing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/microbiome/amplicon-processing into .cursor/skills/bio-microbiome-amplicon-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-microbiome-amplicon-processing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path microbiome/amplicon-processing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-microbiome-amplicon-processing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-microbiome-amplicon-processing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/microbiome/amplicon-processing .gemini/skills/bio-microbiome-amplicon-processing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-microbiome-amplicon-processing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/microbiome/amplicon-processing into .gemini/skills/bio-microbiome-amplicon-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-microbiome-amplicon-processing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-microbiome-amplicon-processingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-microbiome-amplicon-processing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/microbiome/amplicon-processing .github/skills/bio-microbiome-amplicon-processing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-microbiome-amplicon-processing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/microbiome/amplicon-processing into .github/skills/bio-microbiome-amplicon-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-microbiome-amplicon-processing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-microbiome-amplicon-processing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-microbiome-amplicon-processing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/microbiome/amplicon-processing .opencode/skills/bio-microbiome-amplicon-processing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-microbiome-amplicon-processing" agent skill from https://github.com/GPTomics/bioSkills/tree/main/microbiome/amplicon-processing into .opencode/skills/bio-microbiome-amplicon-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-microbiome-amplicon-processing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-microbiome-amplicon-processingInfers exact amplicon sequence variants (ASVs) from demultiplexed 16S rRNA or ITS amplicon FASTQ with DADA2 - removing primers with cutadapt (--discard-untrimmed), learning a per-run error model…
Bio Microbiome Amplicon Processing is an agent skill from GPTomics/bioSkills. Infers exact amplicon sequence variants (ASVs) from demultiplexed 16S rRNA or ITS amplicon FASTQ with DADA2 - removing primers with cutadapt (--discard-untrimmed), learning a per-run error model (filterAndTrim - learnErrors - dada - mergePairs), merging run-level tables with mergeSequenceTables, then one removeBimeraDenovo. Covers why primers come OFF before truncation, why the error model is per-run, truncLen as a merge-overlap detection budget (V4 vs V3-V4), DADA2 vs Deblur and q2-dada2…
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/remove_primers.sh` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (R and Shell), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Microbiome Amplicon Processing loads about 5.7k tokens when it runs. Until then it costs about 251 tokens; SKILL.md has 2,421 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,421 words, ~5,658 tokens.
.claude/skills/bio-microbiome-amplicon-processing/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: DADA2 1.30+, cutadapt 4.6+, ITSxpress 2.0+, QIIME2 2024.2+.
Before using code patterns, verify installed versions match. If versions differ:
packageVersion('<pkg>') then ?function_name to verify parameters<tool> --version then <tool> --help to confirm flagspip show <package> then help(module.function) to check signaturesIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
The error model is a PER-RUN artifact, not a version: learnErrors is fit to one sequencing run (flowcell/chemistry/instrument). Multi-run studies run the per-run inference separately, then mergeSequenceTables, then a single chimera removal - never pool FASTQs across runs before learnErrors. DADA2 dada() defaults (OMEGA_A 1e-40, mergePairs minOverlap 12) and QIIME2 plugin flag spellings drift between releases; confirm with ?dada and qiime dada2 --help.
"Process my 16S amplicon data to get ASVs" -> Strip primers, learn a per-run error model, denoise into exact amplicon sequence variants, merge pairs, and remove chimeras - because an ASV is a model-inferred sequence conditioned on one run, not a clustered consensus or an organism.
dada(filtFs, err=learnErrors(filtFs, multithread=TRUE), multithread=TRUE)cutadapt -g FWD -G REV --discard-untrimmed ... then DADA2, or qiime dada2 denoise-pairedScope: demultiplexed amplicon reads -> chimera-free ASV/feature table + representative sequences. Shotgun reads -> metagenomics/kraken-classification. Taxonomy of the ASVs -> taxonomy-assignment. Diversity/DA of the table -> diversity-analysis, differential-abundance. Compositional/normalization theory (shared) -> metagenomics/abundance-estimation. QIIME2 artifact/provenance/demux mechanics -> qiime2-workflow. Primer-trimming theory -> read-qc/adapter-trimming.
The feature table is not an observation of the community; it is the residue of modeling decisions made BEFORE any result exists - which primers were stripped, where reads were truncated, what error model the run's quality scores supported, what was called a chimera. Turn the knobs differently and the table changes. Three corollaries each common misuse violates:
learnErrors fits one model to a mixture of error structures and denoises wrong. Infer each run separately, then mergeSequenceTables (the exact-sequence string is the join key), then one chimera removal.Organize the work around declaring and defending these knobs - not around running dada() and calling the columns "species."
An ASV (DADA2/Deblur) is an exact inferred sequence at single-nucleotide resolution; a 97% OTU is a centroid of a 3%-identity cluster. Both sides are live (present both, do not declare a winner):
| Tool | Citation | Mechanism / role | When |
|---|---|---|---|
| DADA2 | Callahan 2016 Nat Methods 13:581 | per-run parametric error model, abundance-partition denoising, merge, chimera | the default; variable length, ITS, singleton sensitivity via pseudo-pooling |
| q2-dada2 | (DADA2 engine; Bolyen 2019 Nat Biotechnol 37:852) | QIIME2 wrapper: denoise-paired/single/pyro/ccs | DADA2 inside a QIIME2 artifact/provenance workflow -> qiime2-workflow |
| Deblur | Amir 2017 mSystems 2:e00191-16 | static upper-bound Illumina error profile (positive filter), one fixed length | fast, per-sample-independent, trivially combinable runs; 16S only |
| cutadapt | Martin 2011 EMBnet J 17:10 | primer/adapter trimming (-g/-G, linked adapters) | MUST run before filterAndTrim; primer removal -> read-qc/adapter-trimming |
| ITSxpress | Rivers 2018 F1000Research 7:1418 | HMM-trims the variable-length ITS spacer, keeping quality scores | ITS only; ITS has no valid fixed truncLen |
| VSEARCH | Rognes 2016 PeerJ 4:e2584 | open-source 97% OTU clustering, dereplication, chimera | the OTU path, if a 97% clustering is required (legacy) |
| Scenario | Recommended | Why |
|---|---|---|
| 16S V4 (~253 bp), 2x250 | DADA2 paired, truncLen with comfortable overlap | huge merge slack; truncate to quality freely |
| 16S V3-V4 (~460 bp), 2x250 | DADA2 paired, protect overlap; loosen maxEE R | only ~28 bp slack - the merge budget dominates quality |
| ITS (variable length) | cutadapt + ITSxpress + DADA2 truncLen=0 | fixed truncation slices real biology and breaks merging |
| Full-length 16S (PacBio HiFi/CCS) | DADA2 / qiime dada2 denoise-ccs | resolves to species/strain; single-end CCS, not paired |
| Multiple sequencing runs | per-run inference -> mergeSequenceTables -> one chimera removal | error model is per-run; never pool FASTQs first |
| Want speed, fixed length, many runs, 16S only | Deblur (denoise-16S) | static positive filter; per-sample independent |
| Need singleton/rare-ASV sensitivity | DADA2 dada(..., pool='pseudo') | pseudo-pooling approximates full pooling in linear time |
| NovaSeq/NextSeq/iSeq (binned Q) | inspect plotErrors; enforce monotonic error fit | ~4 quality bins starve the loess fit -> wrong denoising |
| Shotgun (random WGS) reads, not amplicon | -> metagenomics/kraken-classification | no primers/per-run denoising; different category |
Goal: Strip synthetic, often-degenerate primer sequence before any quality/error step.
Approach: Match the forward primer as a 5' adapter on R1 and the reverse primer on R2, discarding pairs where the primer is absent. The order primers -> filter -> learn-errors is non-negotiable: leftover primers corrupt the error model, shift the truncLen frame, and inflate chimeras.
# -g = 515F forward primer (5' adapter on R1); -G = 806R reverse primer (5' adapter on R2);
# --discard-untrimmed drops pairs lacking the primer (a primerless read is suspect).
cutadapt \
-g GTGYCAGCMGCCGCGGTAA \
-G GGACTACNVGGGTWTCTAAT \
--discard-untrimmed \
-o trimmed_R1.fastq.gz -p trimmed_R2.fastq.gz \
sample_R1.fastq.gz sample_R2.fastq.gzThe QIIME2 equivalent is qiime cutadapt trim-paired --p-front-f FWD --p-front-r REV --p-discard-untrimmed.
Goal: Turn one run's primer-trimmed FASTQs into a denoised, merged sequence table.
Approach: Filter on expected errors and truncate within the merge budget, learn the run's error model, denoise each read set against it, merge pairs, then tabulate. Run this block once PER sequencing run.
library(dada2)
out <- filterAndTrim(fnFs, filtFs, fnRs, filtRs,
truncLen=c(240, 160), # region/read-length specific; subject to the merge budget below
maxEE=c(2, 2), truncQ=2, maxN=0, rm.phix=TRUE,
compress=TRUE, multithread=TRUE)
errF <- learnErrors(filtFs, multithread=TRUE) # fit THIS run only
errR <- learnErrors(filtRs, multithread=TRUE)
plotErrors(errF, nominalQ=TRUE) # observed points must track the fitted line and fall with Q
dadaFs <- dada(filtFs, err=errF, multithread=TRUE) # pool='pseudo' for rare-ASV sensitivity
dadaRs <- dada(filtRs, err=errR, multithread=TRUE)
mergers <- mergePairs(dadaFs, filtFs, dadaRs, filtRs, verbose=TRUE)
seqtab_run <- makeSequenceTable(mergers)Paired-end merging needs truncLen_F + truncLen_R >= amplicon_length + ~12 (DADA2 minOverlap default is 12). truncLen is jointly constrained by quality (cut where median Q drops below ~Q30 on plotQualityProfile) AND this overlap budget; the two fight, and for long amplicons the budget wins.
c(240, 200)).maxEE to c(2, 5) to keep low-Q reverse reads.A merge cliff in the read-tracking table is a budget problem, not bad data - the taxa were erased by arithmetic.
Goal: Merge per-run sequence tables into one study table and remove PCR chimeras once.
Approach: Join run-level tables by exact sequence string, then detect bimeras (an ASV reconstructable from two more-abundant parents) across the combined table.
st_all <- mergeSequenceTables(seqtab_run1, seqtab_run2) # exact-sequence string is the join key
seqtab_nochim <- removeBimeraDenovo(st_all, method='consensus', multithread=TRUE, verbose=TRUE)
sum(seqtab_nochim) / sum(st_all) # chimeras = many ASVs but few READS (~0.8-0.99 retained)Carry "run" forward as a batch covariate into differential abundance. A large READ fraction removed as chimeric is a leftover-primer smell (degenerate bases look chimeric), not a real chimera storm.
Goal: Identify and remove reagent/kit ("kitome") contaminant ASVs before any downstream analysis - decisive for low-biomass samples, where contaminants can outnumber real signal.
Approach: Sequence negative controls (extraction blanks, no-template PCR) and a positive mock community alongside the samples, then classify contaminant ASVs with decontam (Davis 2018): the prevalence method when only controls are available, the frequency method when per-sample DNA concentration was measured, combined when both.
library(decontam)
# seqtab_nochim is samples (rows) x ASVs (cols) - decontam's expected orientation.
# is_control: logical, TRUE for negative-control samples; dna_conc: per-sample DNA concentration (qPCR/Qubit).
# prevalence-only threshold 0.1 default; 0.5 = aggressive (ASV more prevalent in controls than samples = contaminant).
contam <- isContaminant(seqtab_nochim, neg = meta$is_control, conc = meta$dna_conc, method = 'combined', threshold = 0.1)
seqtab_clean <- seqtab_nochim[, !contam$contaminant]Low-biomass samples (skin, biopsy, BAL, sterile-site swabs) can be dominated by the kitome, so a "community" there may be mostly contamination - never interpret a low-biomass result without controls. The shotgun analogue is metagenomics/contamination-controls.
Goal: Isolate the biologically variable-length ITS spacer without slicing real sequence.
Approach: Strip primers with cutadapt, then HMM-trim the conserved SSU/5.8S/LSU flanks with ITSxpress (preserving quality scores), then denoise with truncLen=0, filtering on maxEE/minLen only.
itsxpress --fastq r1.fastq.gz --fastq2 r2.fastq.gz \
--region ITS2 --taxa Fungi \ # ITS1/ITS2/ALL; --taxa selects the HMM model
--outfile trimmed.fastq.gz --threads 4out_its <- filterAndTrim(trimmed, filtered, truncLen=0, # NEVER fix-truncate ITS (variable length)
maxEE=2, minLen=50, maxN=0, rm.phix=TRUE, multithread=TRUE)DADA2 inside QIIME2: qiime dada2 denoise-paired --p-trunc-len-f --p-trunc-len-r (also denoise-single, denoise-pyro for 454/Ion Torrent, denoise-ccs with --p-front/--p-adapter/--p-min-len/--p-max-len for PacBio CCS). Deblur (static positive filter, one fixed length, 16S only):
qiime deblur denoise-16S --i-demultiplexed-seqs qc.qza \
--p-trim-length 250 --p-sample-stats \ # ONE fixed length; Deblur cannot handle variable length
--o-representative-sequences rep-seqs.qza --o-table table.qza --o-stats stats.qzaDo not merge a DADA2 ASV table with a Deblur sOTU table - different feature definitions.
Trigger: running filterAndTrim/learnErrors on reads that still carry primers. Mechanism: synthetic, often-degenerate primer bases are read as sequencing error and create spurious split points. Symptom: wrong error fit, a huge READ fraction removed as chimeric, inflated ASV count. Fix: cutadapt --discard-untrimmed first; order is primers -> filter -> learnErrors.
Trigger: concatenating multiple runs' FASTQs into one pipeline. Mechanism: one error model is fit to a mixture of run-specific error structures. Symptom: distorted denoising; ASVs that vanish or appear when runs are split. Fix: per-run inference, then mergeSequenceTables, then one chimera removal; carry run as a batch covariate.
Trigger: truncLen_F + truncLen_R below amplicon length + 12. Mechanism: denoised pairs no longer overlap enough to merge. Symptom: near-zero merged column in read tracking; misread as "low diversity"/"bad data". Fix: compute the budget from amplicon and read length first; for long amplicons keep length and loosen maxEE R.
Trigger: any truncLen on ITS. Mechanism: ITS length is biological (ITS1 ~200-600 bp), so a fixed cut slices real sequence off long variants and merge-fails short ones. Symptom: lost long fungal taxa, poor merging. Fix: cutadapt + ITSxpress, then truncLen=0, filter on maxEE/minLen.
Trigger: default learnErrors on ~4-bin quality data. Mechanism: the loess error-vs-Q fit is starved and can become non-monotonic (error rising at high Q). Symptom: in plotErrors the fitted line diverges from observed points. Fix: enforce monotonicity in the error matrix (nf-core/ampliseq --illumina_novaseq, or set sub-max-Q entries to the max-Q error); never trust the default fit on binned Q.
Trigger: reporting ASV count as richness or each ASV as one organism. Mechanism: intragenomic 16S copy divergence splits one genome into several ASVs (Schloss 2021); reads are not cells (copy number 1-15+). Symptom: inflated richness, "species" that are copies of one organism. Fix: collapse to genus/species (taxonomy-assignment) before richness claims; treat ASV count as an upper bound.
Trigger: analysing low-biomass samples (skin, biopsy, BAL, sterile site) without sequencing controls or running decontam. Mechanism: reagent/kit DNA (the kitome) is amplified alongside scarce template and can dominate the reads. Symptom: a plausible "community" in a near-sterile sample; reagent-associated genera prominent; results track DNA yield. Fix: sequence extraction-blank + no-template-PCR negatives (and a positive mock), run decontam (prevalence or combined), report what was removed (Davis 2018; metagenomics/contamination-controls).
| Threshold | Source | Rationale |
|---|---|---|
maxEE c(2,2) (loosen R to 5 for long amplicons) | Callahan 2016 Nat Methods 13:581 | expected-errors filter beats a hard Q cutoff; computed on the TRUNCATED read, so it interacts with truncLen |
| truncLen budget: truncLen_F + truncLen_R >= amplicon_len + 12 | DADA2 mergePairs minOverlap default | below this, denoised pairs cannot merge; the merge cliff is arithmetic, not data |
| truncLen cut where median Q < ~25-30 | DADA2 docs | quality target, secondary to the merge budget for long amplicons |
maxN = 0 | DADA2 docs | DADA2 cannot model ambiguous bases; mandatory |
| chimera retained-read fraction ~0.8-0.99 | DADA2 docs | chimeras are many ASVs but few reads; a large read loss flags leftover primers |
pool='pseudo' for rare ASVs | DADA2 docs | approximates full pooling (quadratic) in linear time; default FALSE misses cross-sample singletons |
Deblur --p-trim-length one fixed value | Amir 2017 mSystems 2:e00191-16 | the positive filter requires a single read length |
| 16S copy-number correction: report, do not assume | Louca 2018 Microbiome 6:41 | predictable only near reference genomes; correction can ADD error ("unsolved problem") |
| Error / symptom | Cause | Solution |
|---|---|---|
| Near-zero merge rate | truncLen below the overlap budget | recompute budget; keep length, loosen maxEE R |
| Large read fraction "chimeric" | primers not trimmed (degenerate bases) | cutadapt --discard-untrimmed before filtering |
plotErrors fitted line diverges from points | binned quality (NovaSeq/NextSeq) | enforce monotonic error matrix; nf-core/ampliseq --illumina_novaseq |
| ASVs vanish/appear when runs split | one error model fit across runs | per-run learnErrors, then mergeSequenceTables |
| Few reads pass filter | maxEE too strict or truncLen too long (low-Q tail) | loosen maxEE, shorten truncLen within the budget |
| ITS taxa lost / poor merging | fixed truncLen on ITS | cutadapt + ITSxpress, then truncLen=0 |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in microbiome/amplicon-processing of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Microbiome Amplicon Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Microbiome Amplicon Processing this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.7k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 3 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| MFA Pipeline Orchestratoraiming-lab/AutoResearchClaw | 15k | — | ~923 | Automated safety check: Pass | MIT |
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
xuzhougeng/wisp-science
A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Categories
Infers exact amplicon sequence variants (ASVs) from demultiplexed 16S rRNA or ITS amplicon FASTQ with DADA2 - removing primers with cutadapt (--discard-untrimmed), learning a per-run error model…. Bio Microbiome Amplicon Processing is an agent skill from GPTomics/bioSkills. Infers exact amplicon sequence variants (ASVs) from demultiplexed 16S rRNA or ITS amplicon FASTQ with DADA2 - removing primers with cutadapt (--discard-untrimmed), learning a per-run error model (filterAndTrim - learnErrors - dada - mergePairs), merging run-level tables with mergeSequenceTables, then one removeBimeraDenovo.
Bio Microbiome Amplicon Processing fits situations like: turning demultiplexed amplicon reads into an ASV/feature table; choosing truncation lengths; handling multi-run studies.
Run `npx skills add GPTomics/bioSkills --skill bio-microbiome-amplicon-processing -a claude-code`. Or copy the skill folder (microbiome/amplicon-processing in GPTomics/bioSkills) into .claude/skills/bio-microbiome-amplicon-processing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-microbiome-amplicon-processing -a codex`. Or copy the skill folder (microbiome/amplicon-processing in GPTomics/bioSkills) into .agents/skills/bio-microbiome-amplicon-processing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-microbiome-amplicon-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-microbiome-amplicon-processing, .gemini/skills/bio-microbiome-amplicon-processing, .github/skills/bio-microbiome-amplicon-processing and .opencode/skills/bio-microbiome-amplicon-processing in your project.
Going by SKILL.md and its folder, Bio Microbiome Amplicon Processing needs R and a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: A Bash shell.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Microbiome Amplicon Processing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Microbiome Amplicon Processing: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars) and Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.