Agent skill

Bio Ecological Genomics Edna Metabarcoding

by GPTomics in GPTomics/bioSkills

Processes eDNA metabarcoding from raw paired-end reads to species tables, navigating ASV (DADA2, UNOISE3) vs OTU (swarm v2) decision (Callahan 2017 vs Schloss multi-copy-16S critique), marker/primer…

MITAuto-check passedResearch & Science

Install Bio Ecological Genomics Edna Metabarcoding

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-ecological-genomics-edna-metabarcoding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-ecological-genomics-edna-metabarcoding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ecological-genomics/edna-metabarcoding .claude/skills/bio-ecological-genomics-edna-metabarcoding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-ecological-genomics-edna-metabarcoding
GitHub stars
1.2k
Used in
1 other repo
Token cost
~6.6k tokens
SKILL.md length
2,234 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Processes eDNA metabarcoding from raw paired-end reads to species tables, navigating ASV (DADA2, UNOISE3) vs OTU (swarm v2) decision (Callahan 2017 vs Schloss multi-copy-16S critique), marker/primer…

  • Going from raw eDNA FASTQ to species tables
  • SKILL.md covers Version Compatibility, The Single Most Important…, Algorithmic Taxonomy and Decision Tree by Scenario, plus 11 more sections
  • Runs R and Shell scripts from its folder; calls pip
  • Picking marker + denoising pipeline

What it does

Bio Ecological Genomics Edna Metabarcoding is an agent skill from GPTomics/bioSkills. Processes eDNA metabarcoding from raw paired-end reads to species tables, navigating ASV (DADA2, UNOISE3) vs OTU (swarm v2) decision (Callahan 2017 vs Schloss multi-copy-16S critique), marker/primer choice (Leray COI, MiFish 12S, 515F/806R 16S, ITS2) with primer-specific bias, OBITools3 v3 command-name break (obi stats plural; .tar.gz taxonomy), tag-jumping with dual-indexing (Schnell 2015; NovaSeq 10x MiSeq), decontam as screening-not-classifier (Davis 2018), read-counts-not-abundance critique (Lamb 2019)…

Its SKILL.md is about 6.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/obitools3_edna_pipeline.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Performance reviews. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Going from raw eDNA FASTQ to species tables
  • Picking marker + denoising pipeline
  • Deciding whether read counts represent abundance
  • Applying occupancy modeling

Example prompts

  • “Use the bio-ecological-genomics-edna-metabarcoding skill to process eDNA metabarcoding from raw paired-end reads to species tables, navigating ASV…”
  • “/bio-ecological-genomics-edna-metabarcoding”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Ecological Genomics Edna Metabarcoding loads about 6.6k tokens when it runs. Until then it costs about 244 tokens; SKILL.md has 2,234 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~244
When it runs · the whole SKILL.md, loaded when a task matches
~6.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,234 words, ~6,646 tokens.

Download SKILL.mdSave it as .claude/skills/bio-ecological-genomics-edna-metabarcoding/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-ecological-genomics-edna-metabarcoding
description
Processes eDNA metabarcoding from raw paired-end reads to species tables, navigating ASV (DADA2, UNOISE3) vs OTU (swarm v2) decision (Callahan 2017 vs Schloss multi-copy-16S critique), marker/primer choice (Leray COI, MiFish 12S, 515F/806R 16S, ITS2) with primer-specific bias, OBITools3 v3 command-name break (obi stats plural; .tar.gz taxonomy), tag-jumping with dual-indexing (Schnell 2015; NovaSeq 10x MiSeq), decontam as screening-not-classifier (Davis 2018), read-counts-not-abundance critique (Lamb 2019), site-occupancy modeling (Ficetola 2015), Naive-Bayes calibration limits (Bokulich 2018), and eDNA decay (Strickler 2015). Use when going from raw eDNA FASTQ to species tables, picking marker + denoising pipeline, deciding whether read counts represent abundance, applying occupancy modeling, configuring OBITools3 v3, or interpreting decontam output. Not for clinical 16S microbiome (see microbiome/amplicon-processing).
tool_type
mixed
primary_tool
dada2

Version Compatibility

Reference examples tested with: DADA2 1.30+, cutadapt 4.7+, OBITools3 (Python 3), decontam 1.20+, microDecon 1.0+, occumb 1.0+, vsearch 2.27+, swarm 3.1+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

eDNA Metabarcoding

"Process eDNA samples to identify species present" -> Trim primers, denoise to ASVs (or cluster to OTUs), detect chimeras, assign taxonomy, filter contamination with negative controls AND DNA concentration, decompose tag-jumping artifacts, and quantify detection uncertainty via site-occupancy modeling. For the foundational eDNA-for-wildlife review, see Bohmann et al. 2014 Trends Ecol Evol 29:358-367.

  • CLI: cutadapt for primer removal (linked-adapter mode)
  • R: dada2::filterAndTrim() -> dada() -> assignTaxonomy() for ASV pipeline
  • CLI: obi stats / obi clean / obi ecotag for OBITools3 (NOTE: v3 plural commands)
  • R: decontam::isContaminant() for contamination screening
  • R: occumb::occumb() for detection-corrected occurrence

The Single Most Important Modern Insight -- Read Counts Are NOT Abundance

Elbrecht & Leese 2015 PLoS One 10:e0130324 and Lamb et al. 2019 Mol Ecol 28:420-430 (meta-analysis) established that metabarcoding read counts have weak-to-moderate, taxon-specific, NONLINEAR correlation with biomass or DNA input. Primer-binding bias dominates; PCR replicates introduce stochasticity. Reporting read counts as abundance without mock-community calibration is malpractice. Modern practice: report PRESENCE/ABSENCE or relative abundance with explicit calibration; use multiple PCR replicates; apply site-occupancy models for detection correction.

A second cornerstone: the ASV-vs-OTU debate is taxon-specific, not universal. Callahan, McMurdie, Holmes 2017 ISME J 11:2639-2643 argued ASVs replace OTUs because modern denoising resolves single-nucleotide differences. Schloss 2021 mSphere 6:e00191-21 showed that for bacterial 16S with 1-15 intra-genomic rRNA copies, a single E. coli strain produces ~7 distinct ASVs, splitting bacterial genomes across artificial clusters. For COI metazoan metabarcoding, ASVs (DADA2/UNOISE3) are recommended; for bacterial 16S, ASVs inflate alpha-diversity and OTUs may be appropriate.

A third: decontam (Davis 2018) is a SCREENING tool, not a deterministic classifier. It flags candidates; biological plausibility check is required before deletion. The default threshold=0.1 over-flags in low-biomass data.

Algorithmic Taxonomy

MethodOutputStrengthFails when
DADA2Single-nucleotide ASVsHigh resolution; learned error model; standard for COI/12S/18S/fungal-ITSSmall datasets (< 100 samples) for error learning; multi-copy bacterial rRNA
UNOISE3 (USEARCH/VSEARCH; Edgar 2016)zOTUs (essentially ASVs)Fast; algorithmic simplicityLimited Linux/Mac binary distribution under license
Swarm v2 -d 1 --fastidious (Mahé 2015)Abundance-weighted single-linkage OTUsModern OTU pipeline; better than legacy 97% UCLUSTOTUs by design (not single-nt resolution)
97% UCLUSTClassical OTUsLegacy familiarityBiologically arbitrary threshold; supersedes by DADA2/swarm
VSEARCH global pairwiseTaxonomic assignment via best-hitFast, transparent, no trainingConservative; mis-assigns sister species when ref incomplete
Naive Bayes (q2-feature-classifier, RDP)Probabilistic taxonomic assignmentProbabilistic confidence; standard for 16SConfidence values are scikit-learn calibrated, not true probabilities (Bokulich 2018)
SINTAX (Edgar)Bootstrap-supported taxonomyFast; no trainingLess accurate than Naive Bayes for divergent sequences
LCA (BASTA, MEGAN-LCA)Lowest common ancestor of multiple hitsConservative; never over-confidentCan over-merge to high taxonomic ranks
Phylogenetic placement (EPA-ng + gappa)Position on reference treeMost rigorous; phylogenetically explicit10-100x slower; emerging not yet standard
decontamFlagged contaminant candidatesStatistical screening of negative controls and DNA concentration patternsOutput is screening, not classification; needs biological-plausibility check
UCHIME3 (in DADA2/VSEARCH)Chimera detectionStandard for de novo chimera removalSome divergent chimeras escape

Decision Tree by Scenario

ScenarioRecommended approachWhy
Metazoan COI metabarcoding (water, gut content)mlCOIintF/jgHCO2198 (Leray 2013) primers; DADA2 ASVsStandard primer set; ASVs preserve single-nt resolution
Fish eDNA from waterMiFish-U/E (Miya 2015) 12S primers; DADA2 ASVsDominant eDNA fish marker globally
Freshwater macroinvertebrate bioassessmentBF1/BR1 freshwater-optimized COI primersHigher primer-binding inclusivity for aquatic insects
Bacterial community 16S515F/806R (V4) Parada modified; ASVs OR Swarm v2Schloss 2021 caveat applies; ASVs may oversplit multi-copy rRNA
Fungal community ITSITS2 primers; DADA2 or UNITE pipelineUNITE is curated for fungal ITS
Plant community DNAtrnL P6 loop (Taberlet 2007) for degraded DNARobust to degradation
Deciding ASV vs OTUASVs for COI/12S/18S/fungi; OTU consideration for 16S with multi-copy concernTaxon-specific
NovaSeq library (patterned flow cell)Heavier tag-jumping correction; expect 10x higher rates than MiSeqPatterned-cell index hopping
Low-biomass eDNA (deep ocean, ancient)decontam frequency + prevalence methods; explicit reagent-contamination checkReagent contamination dominates
Quantitative comparison across samplesMock-community calibration BEFORE reporting read countsWithout mock, read counts are biased estimators of biomass
Detection probability with replicationSite-occupancy models (occumb, eDNAoccupancy; Ficetola 2015)Read counts alone underestimate occurrence; replicates correct
Taxonomic assignment for marker > 80% coveredNaive Bayes (q2-feature-classifier)Probabilistic; well-supported
Taxonomic assignment for sparse referencePhylogenetic placement (EPA-ng)Robust to incomplete references
OBITools3 pipelineobi stats (NOTE: plural), DMS-based, .tar.gz taxonomyv3 syntax differs from v1

Primer Trimming with cutadapt

Goal: Remove primer sequences while discarding reads that lack primers, before quality filtering.

Approach: Use cutadapt linked-adapter mode with marker-specific 5' and 3' primer pairs. --discard-untrimmed removes reads lacking expected primers; min_overlap prevents false primer detection in random sequence regions.

bash
# COI metazoan (Leray mlCOIintF / jgHCO2198 -> 313 bp)
cutadapt -g 'GGWACWGGWTGAACWGTWTAYCCYCC;min_overlap=20' \
         -G 'TAIACYTCIGGRTGICCRAARAAYCA;min_overlap=20' \
         --discard-untrimmed --pair-filter=any \
         -o trimmed_R1.fastq.gz -p trimmed_R2.fastq.gz \
         raw_R1.fastq.gz raw_R2.fastq.gz

# Fish 12S (MiFish-U -> 163-185 bp)
cutadapt -g 'GTCGGTAAAACTCGTGCCAGC;min_overlap=18' \
         -G 'CATAGTGGGGTATCTAATCCCAGTTTG;min_overlap=18' \
         --discard-untrimmed --pair-filter=any \
         -o trimmed_R1.fastq.gz -p trimmed_R2.fastq.gz \
         raw_R1.fastq.gz raw_R2.fastq.gz

# Fungal ITS2
cutadapt -g 'GTGAATCATCGAATCTTTGAAC;min_overlap=18' \
         -G 'TCCTCCGCTTATTGATATGC;min_overlap=18' \
         --discard-untrimmed --pair-filter=any \
         -o trimmed_R1.fastq.gz -p trimmed_R2.fastq.gz \
         raw_R1.fastq.gz raw_R2.fastq.gz

DADA2 ASV Pipeline

Goal: Denoise paired-end amplicon reads into exact amplicon sequence variants (ASVs) with chimera removal and reference-based taxonomy assignment, per Callahan et al. 2016 Nat Methods 13:581-583.

Approach: Filter to length/quality thresholds, learn error rates per dataset, run dada() to denoise, merge pairs, build sequence table, remove chimeras with UCHIME3-equivalent in DADA2, then assign taxonomy against the marker-appropriate reference DB. CRITICAL: primers must be removed (cutadapt) BEFORE filterAndTrim, OR the error model is corrupted.

r
library(dada2)

# CRITICAL: primer removal MUST precede filterAndTrim
# DADA2's error model assumes primer-free reads
fwd_reads <- sort(list.files('primer_trimmed/', pattern = '_R1', full.names = TRUE))
rev_reads <- sort(list.files('primer_trimmed/', pattern = '_R2', full.names = TRUE))
filt_fwd <- file.path('filtered', basename(fwd_reads))
filt_rev <- file.path('filtered', basename(rev_reads))

# Filter and trim
# maxEE=c(2,2): expected errors per read; tradeoff sensitivity/specificity
# truncLen: set from quality profile inspection; do not guess
out <- filterAndTrim(fwd_reads, filt_fwd, rev_reads, filt_rev,
                     maxN = 0, maxEE = c(2, 2), truncQ = 2,
                     truncLen = c(220, 180),     # data-dependent; inspect plotQualityProfile()
                     minLen = 100, rm.phix = TRUE, multithread = TRUE)

# Learn error rates
# For small datasets (< 100 samples), pool aggressively or use pre-learned model
err_fwd <- learnErrors(filt_fwd, multithread = TRUE)
err_rev <- learnErrors(filt_rev, multithread = TRUE)

# Denoise
dada_fwd <- dada(filt_fwd, err = err_fwd, multithread = TRUE)
dada_rev <- dada(filt_rev, err = err_rev, multithread = TRUE)

# Merge pairs with minimum overlap
merged <- mergePairs(dada_fwd, filt_fwd, dada_rev, filt_rev, minOverlap = 12)

# Build sequence table
seqtab <- makeSequenceTable(merged)

# Remove chimeras
# method='consensus': per-sample then consensus; conservative (default)
# method='pooled': pooled across samples; aggressive; can over-merge real diversity
# Chimera rate >30% typically indicates library prep problems
seqtab_nochim <- removeBimeraDenovo(seqtab, method = 'consensus',
                                     multithread = TRUE)
cat('Chimera rate:', round(1 - sum(seqtab_nochim) / sum(seqtab), 3), '\n')

# Taxonomy assignment
# minBoot=80: standard genus-level confidence; 50 for family-level
# IMPORTANT: pair the marker with the appropriate reference DB
# COI -> MIDORI2 LONGEST_NUC_GB259_CO1 (or BOLD with curation)
# 12S -> MitoFish (Miya lab)
# 16S V4 -> SILVA 138.1+
# 18S V4/V9 -> SILVA 138.1+ or PR2
# Fungal ITS -> UNITE 9.0+
taxa <- assignTaxonomy(seqtab_nochim,
                       'MIDORI2_LONGEST_NUC_GB259_CO1_DADA2.fasta.gz',
                       minBoot = 80, multithread = TRUE)

OBITools3 Pipeline — The v1 -> v3 Command Break

Goal: Process eDNA reads through the Unix-style OBITools v3 pipeline (Boyer et al. 2016 Mol Ecol Resour 16:176-182 introduced OBITools v1; v3 is the post-2018 Python 3 rewrite) with DMS-based sequence management.

Approach: v3 introduces a Database Management System (DMS) abstraction; sequences are imported into a DMS rather than read directly from FASTQ. Commands use spaces (e.g., obi stats plural, not obistat). Taxonomy import expects .tar.gz archive, not a directory.

bash
# v1 -> v3 command-name changes (critical):
# v1: obistat       -> v3: obi stats
# v1: obigrep       -> v3: obi grep
# v1: obiuniq       -> v3: obi uniq
# v1: obitab        -> v3: obi annotate / obi export --tab-output (different semantics)
# v1: ngsfilter     -> v3: obi ngsfilter
# v1: taxdump dir   -> v3: .tar.gz archive

# Import paired FASTQ into DMS
obi import --fastq-input raw_R1.fastq.gz EDNA/reads1
obi import --fastq-input raw_R2.fastq.gz EDNA/reads2

# Paired-end alignment
obi alignpairedend -R EDNA/reads2 EDNA/reads1 EDNA/aligned

# Filter by alignment score and length
obi grep -p 'sequence["score"] >= 50' EDNA/aligned EDNA/filtered
obi grep -p 'len(sequence) >= 100 and len(sequence) <= 500' \
    EDNA/filtered EDNA/length_filtered

# Demultiplex (NGS filter file maps barcodes -> samples)
obi ngsfilter -t ngsfilter.txt -u EDNA/unassigned \
    EDNA/length_filtered EDNA/demux

# Dereplicate (obi uniq creates merged_sample attribute automatically)
obi uniq EDNA/demux EDNA/derep

# Remove suspected error singletons
obi grep -p 'sequence["count"] >= 2' EDNA/derep EDNA/no_singletons

# Denoise via obi clean
obi clean -s merged_sample -r 0.05 -H EDNA/no_singletons EDNA/denoised

# Taxonomy assignment against reference database
obi ecotag -R EDNA/refdb --taxonomy EDNA/taxonomy EDNA/denoised EDNA/assigned

# Export tab-separated species table
obi export --tab-output EDNA/assigned > species_table.tsv

Across most metabarcoding studies, 50-85% of ASVs cannot be assigned to species level due to incomplete references (Wangensteen et al. 2018 PeerJ 6:e4705 documented this for marine COI + 18S). Report this gap honestly; do not infer ecology from "unassigned" reads.

Tag-Jumping Mitigation — Schnell 2015 + NovaSeq Caveat

Goal: Detect and remove sequence-to-sample misassignments arising from chimeric library molecules with mismatched indices.

Approach: Use dual-indexing (different indices at both ends; cross-jumped pairs are discarded). Quantify residual tag-jumping rate from per-ASV cross-sample appearance and apply per-ASV abundance threshold filtering with metabaR::tagjumpslayer. For NovaSeq libraries, expect ~10x higher tag-jumping than MiSeq due to patterned flow cells.

r
library(metabaR)

# metabaR expects an metabarlist object (asv table + sample info + ngsfilter)
# tagjumpslayer applies per-ASV abundance-threshold filter
# threshold: 0.01 (1% of ASV total) is conservative; 0.001 for aggressive removal
# Adjust threshold higher for NovaSeq (~0.005-0.01) than MiSeq (~0.001-0.005)

# Quantify residual tag-jumping rate before filtering:
# Count reads in sample x ASV combinations that should be 0 by experimental design
# (e.g., samples explicitly excluded from a particular condition)
# That rate / total reads = empirical tag-jumping rate
# Report this rate in methods section

Contamination Screening with decontam

Goal: Identify candidate contaminant ASVs from negative controls and DNA-concentration patterns.

Approach: Use decontam::isContaminant with method='combined' when both DNA concentration AND negative controls are available. Treat flagged ASVs as SCREENING CANDIDATES; verify biological plausibility before deletion. The default threshold=0.1 is over-aggressive in low-biomass data.

r
library(decontam)

# Frequency method: contaminants more frequent at LOW DNA concentration
# Prevalence method: contaminants more frequent in negative controls
# Combined: uses both signals (most robust)

contam <- isContaminant(seqtab_nochim,
                        conc = dna_concentration,        # qPCR or Qubit per sample
                        neg = is_negative_control,       # logical: which samples are controls
                        method = 'combined',
                        threshold = 0.1)                 # default; lower for high-confidence calls

# CRITICAL: decontam output is SCREENING, not classification
# Manually inspect each flagged ASV: is the taxonomic assignment plausibly a reagent contaminant?
# Common reagent contaminants: Delftia, Sphingomonas, Burkholderia, Propionibacterium
flagged <- which(contam$contaminant)
cat('Decontam flagged', length(flagged), 'ASVs as candidates\n')

# After manual review, remove confirmed contaminants
confirmed_contam <- intersect(flagged, biological_plausibility_check_result)
seqtab_clean <- seqtab_nochim[, !(colnames(seqtab_nochim) %in% confirmed_contam)]

Site-Occupancy Modeling — Correcting for Imperfect Detection

Goal: Estimate true species occurrence probabilities from replicated eDNA samples, accounting for false negatives in any single PCR replicate.

Approach: Fit a multi-species occupancy model via MCMC (Ficetola 2015 Mol Ecol Resour 15:543-556) on a 3D array of replicated read counts. Output: per-site, per-species occupancy probabilities corrected for detection.

r
library(occumb)

# y: 3D array [species, sites, replicates] of read counts
# spec_cov: species covariates (traits)
# site_cov: site covariates (env)
data_obj <- occumbData(y = count_array, spec_cov = species_covariates,
                       site_cov = site_covariates)

# Fit hierarchical occupancy model
# Requires JAGS installation
# n.iter >= 10000, n.burn >= 2500 for publication-quality posteriors
fit <- occumb(data = data_obj, n.chains = 4, n.iter = 10000,
              n.thin = 5, n.burn = 2500)

# Extract detection-corrected occupancy
summary(fit)

Per-Method Failure Modes

Reporting read counts as biomass without mock-community calibration

Trigger: Comparing read counts of two ASVs and reporting the ratio as a biomass / abundance estimate.

Mechanism: Primer-template binding affinity varies systematically across taxa; PCR amplification is non-linear (saturates); read counts have weak-to-moderate, NONLINEAR correlation with biomass (Elbrecht 2015; Lamb 2019).

Symptom: Reviewer asks "how is it known that reads = biomass?"; cross-study quantitative comparisons fail to replicate.

Fix: Either (a) restrict reporting to presence/absence; (b) report read counts as relative abundances with explicit caveat; or (c) include mock-community of known composition for primer-specific calibration. Do not silently equate reads with biomass.

NovaSeq tag-jumping with MiSeq-tuned filtering

Trigger: Applying tag-jumping filters calibrated on MiSeq libraries to NovaSeq data.

Mechanism: NovaSeq patterned flow cells have ~10x higher index hopping than MiSeq. MiSeq-calibrated thresholds (often ~0.001 fraction) are too permissive on NovaSeq data.

Symptom: Apparent rare-species detections in NovaSeq libraries do not replicate; per-ASV cross-sample appearance is unusually broad.

Fix: Use NovaSeq-appropriate tag-jumping thresholds (~0.005-0.01) and report the empirical tag-jumping rate from explicit-zero combinations.

Show full SKILL.md (872 more words)Show less
decontam threshold over-aggressive in low-biomass data

Trigger: Applying default threshold = 0.1 to ASVs from open-ocean water, ancient sediments, or other dilute samples.

Mechanism: In low-biomass samples, the contaminant signal/background ratio approaches 1; decontam over-flags real but dilute biology as "contaminant" because the statistical pattern looks similar.

Symptom: Many ASVs flagged from low-biomass samples; taxonomic profile of "flagged contaminants" looks biologically realistic.

Fix: Lower threshold (0.05 or 0.01); always manually review flagged ASVs for biological plausibility; cite Salter 2014 BMC Biol 12:87 for the low-biomass reagent-contamination caveat.

Skipping primer removal before DADA2 filterAndTrim

Trigger: Running DADA2's filterAndTrim() on FASTQ files that still contain primer sequences.

Mechanism: DADA2 learns sequencing error from the empirical data; if primer sequences are present, they look like "perfect agreement" and corrupt the error model. ASVs are inferred with primer artifacts attached.

Symptom: DADA2 reports "phix-like contamination" (false; it's primers); ASVs start with the primer sequence; chimera rate elevated.

Fix: Always run cutadapt (or similar) BEFORE filterAndTrim. Verify with head of trimmed FASTQ that primer sequences are gone.

OBITools v3 commands with v1 syntax

Trigger: Running obistat or obigrep on a v3 install.

Mechanism: v1 used concatenated command names (obistat); v3 uses subcommand syntax with a space (obi stats — note plural).

Symptom: Bash error obistat: command not found; tutorial documentation does not match installed version.

Fix: Use obi <subcommand> syntax; consult obi --help for current command list. Taxonomy import requires .tar.gz archive, not unpacked directory.

Quantitative Thresholds

ThresholdValueSource / rationale
DADA2 maxEE per read2Standard sensitivity/specificity balance
DADA2 chimera rate alarm> 30% suggests library issuesEmpirical convention
DADA2 minBoot for taxonomy80 for genus; 50 for familyStandard confidence cutoffs
Tag-jumping filter MiSeq0.001-0.005 fraction of ASV totalSchnell 2015
Tag-jumping filter NovaSeq0.005-0.01 fraction of ASV totalPatterned-cell index hopping ~10x higher
decontam threshold0.1 default; 0.05 for low-biomassDavis 2018; reduce for dilute samples
Per-sample minimum reads1000 (after filtering)Below this rare-species detection unreliable
Singleton removalcount >= 2Singletons often error-driven
Bootstrap nperm for tests999Standard permutation count
Occupancy model iterationsn.iter >= 10000, n.burn >= 2500occumb default for stable posteriors
eDNA decay (20 deg C surface water)half-life ~4-15 hoursStrickler 2015 Biol Conserv 183:85-92

Common errors

ErrorCauseSolution
obistat: command not foundOBITools v3 uses obi stats (plural)Use v3 syntax
DADA2 error rate plot looks pathologicalPrimer sequences still in readsRe-run cutadapt before filterAndTrim
Chimera rate > 30%Library-prep issue or primer dimersInspect raw FASTQ; check PCR conditions
decontam flags many real speciesDefault threshold too aggressive for low-biomassLower threshold; manual review
Naive Bayes confidence 0.95 but species is wrongscikit-learn-calibrated "confidence" not true probabilityUse phylogenetic placement for borderline assignments
occumb JAGS not found errorJAGS not installed system-wideInstall JAGS (CRAN page has platform instructions)
eDNA detections do not replicateRead counts treated as abundanceSwitch to presence/absence; use mock-community calibration
MIDORI2 download path expiredDatabase updated; old URL goneCheck current MIDORI2 / MitoFish download page

References

  • Callahan BJ, McMurdie PJ, Rosen MJ, Han AW, Johnson AJA, Holmes SP (2016) DADA2. Nat Methods 13(7):581-583. doi:10.1038/nmeth.3869
  • Callahan BJ, McMurdie PJ, Holmes SP (2017) Exact sequence variants should replace OTUs. ISME J 11(12):2639-2643. doi:10.1038/ismej.2017.119
  • Edgar RC (2016) UNOISE2 / UNOISE3. bioRxiv preprint. doi:10.1101/081257
  • Mahe F, Rognes T, Quince C, de Vargas C, Dunthorn M (2015) Swarm v2. PeerJ 3:e1420. doi:10.7717/peerj.1420
  • Leray M, Yang JY, Meyer CP et al. (2013) mlCOIintF/jgHCO2198 metazoan COI primer. Front Zool 10:34. doi:10.1186/1742-9994-10-34
  • Miya M, Sato Y, Fukunaga T et al. (2015) MiFish 12S fish eDNA primer. R Soc Open Sci 2(7):150088. doi:10.1098/rsos.150088
  • Schnell IB, Bohmann K, Gilbert MTP (2015) Tag jumps illuminated. Mol Ecol Resour 15(6):1289-1303. doi:10.1111/1755-0998.12402
  • Davis NM, Proctor DM, Holmes SP, Relman DA, Callahan BJ (2018) decontam. Microbiome 6:226. doi:10.1186/s40168-018-0605-2
  • Boyer F, Mercier C, Bonin A, Le Bras Y, Taberlet P, Coissac E (2016) OBITools. Mol Ecol Resour 16(1):176-182. doi:10.1111/1755-0998.12428
  • Elbrecht V, Leese F (2015) DNA-based ecosystem quantification critique. PLoS One 10(7):e0130324. doi:10.1371/journal.pone.0130324
  • Lamb PD, Hunter E, Pinnegar JK, Creer S, Davies RG, Taylor MI (2019) How quantitative is metabarcoding: meta-analysis. Mol Ecol 28(2):420-430. doi:10.1111/mec.14920
  • Ficetola GF, Pansu J, Bonin A et al. (2015) Replication levels and false presences in eDNA. Mol Ecol Resour 15(3):543-556. doi:10.1111/1755-0998.12338
  • Wangensteen OS, Palacin C, Guardiola M, Turon X (2018) COI + 18S marine metabarcoding. PeerJ 6:e4705. doi:10.7717/peerj.4705
  • Strickler KM, Fremier AK, Goldberg CS (2015) eDNA degradation kinetics. Biol Conserv 183:85-92. doi:10.1016/j.biocon.2014.11.038
  • Bokulich NA, Kaehler BD, Rideout JR et al. (2018) Optimizing taxonomic classification with q2-feature-classifier. Microbiome 6:90. doi:10.1186/s40168-018-0470-z
  • Bohmann K, Evans A, Gilbert MTP et al. (2014) eDNA for wildlife and biodiversity. Trends Ecol Evol 29(6):358-367. doi:10.1016/j.tree.2014.04.003
  • Schloss PD (2021) Amplicon sequence variants artificially split bacterial genomes into separate clusters. mSphere 6(4):e00191-21. doi:10.1128/mSphere.00191-21
  • Salter SJ, Cox MJ, Turek EM et al. (2014) Reagent and laboratory contamination can critically impact sequence-based microbiome analyses. BMC Biol 12:87. doi:10.1186/s12915-014-0087-z
  • Taberlet P, Coissac E, Pompanon F et al. (2007) Power and limitations of the chloroplast trnL (UAA) intron for plant DNA barcoding. Nucleic Acids Res 35(3):e14. doi:10.1093/nar/gkl938
  • ecological-genomics/biodiversity-metrics - Diversity analysis from species occurrence tables (Hill numbers, beta partition)
  • ecological-genomics/community-ecology - Environmental gradient analysis of community composition (PERMANOVA + PERMDISP, ordination)
  • microbiome/amplicon-processing - 16S clinical microbiome alternative pipeline
  • read-qc/quality-reports - Upstream read-quality assessment before primer trimming
  • database-access/entrez-fetch - Retrieve reference sequences for custom taxonomy databases

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in ecological-genomics/edna-metabarcoding of GPTomics/bioSkills.

  • SKILL.md
  • examples/dada2_edna_coi.R
  • examples/obitools3_edna_pipeline.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Ecological Genomics Edna Metabarcoding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Ecological Genomics Edna Metabarcoding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Ecological Genomics Edna Metabarcoding this skillGPTomics/bioSkills1.2k1 repos~6.6kAutomated safety check: PassMIT
External Model Validationaipoch/medical-research-skills1.9k—~3.2kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT

Similar skills

  • External Model Validation

    aipoch/medical-research-skills

    A skill your agent uses when validating an existing prognostic risk signature on an external bulk expression cohort with survival outcomes, producing risk scores, Kaplan-Meier curves, risk…

    1.9k GitHub stars~3.2k tokensUpdated 23 days ago
    Research & ScienceAuto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Ecological Genomics Edna Metabarcoding

What does Bio Ecological Genomics Edna Metabarcoding do?

Processes eDNA metabarcoding from raw paired-end reads to species tables, navigating ASV (DADA2, UNOISE3) vs OTU (swarm v2) decision (Callahan 2017 vs Schloss multi-copy-16S critique), marker/primer…. Bio Ecological Genomics Edna Metabarcoding is an agent skill from GPTomics/bioSkills.

When should I use Bio Ecological Genomics Edna Metabarcoding?

Bio Ecological Genomics Edna Metabarcoding fits situations like: going from raw eDNA FASTQ to species tables; picking marker + denoising pipeline; deciding whether read counts represent abundance; applying occupancy modeling.

How do I install Bio Ecological Genomics Edna Metabarcoding in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-ecological-genomics-edna-metabarcoding -a claude-code`. Or copy the skill folder (ecological-genomics/edna-metabarcoding in GPTomics/bioSkills) into .claude/skills/bio-ecological-genomics-edna-metabarcoding in your project. Claude Code loads it when a task matches its description.

How do I install Bio Ecological Genomics Edna Metabarcoding in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-ecological-genomics-edna-metabarcoding -a codex`. Or copy the skill folder (ecological-genomics/edna-metabarcoding in GPTomics/bioSkills) into .agents/skills/bio-ecological-genomics-edna-metabarcoding in your project. Codex loads it when a task matches its description.

Can I use Bio Ecological Genomics Edna Metabarcoding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-ecological-genomics-edna-metabarcoding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-ecological-genomics-edna-metabarcoding, .gemini/skills/bio-ecological-genomics-edna-metabarcoding, .github/skills/bio-ecological-genomics-edna-metabarcoding and .opencode/skills/bio-ecological-genomics-edna-metabarcoding in your project.

What does Bio Ecological Genomics Edna Metabarcoding need to run?

Going by SKILL.md and its folder, Bio Ecological Genomics Edna Metabarcoding needs R and a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Ecological Genomics Edna Metabarcoding access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Ecological Genomics Edna Metabarcoding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Ecological Genomics Edna Metabarcoding use?

Bio Ecological Genomics Edna Metabarcoding is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Ecological Genomics Edna Metabarcoding use?

About 6.6k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Ecological Genomics Edna Metabarcoding?

Skills that share tags, products or a category with Bio Ecological Genomics Edna Metabarcoding: External Model Validation (aipoch/medical-research-skills, 1.9k stars), Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars) and Clinvar Database (google-deepmind/science-skills, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Ecological Genomics Edna Metabarcoding?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.