Agent skill

Bio Workflows Edna Pipeline

by GPTomics in GPTomics/bioSkills

End-to-end eDNA metabarcoding from raw amplicons to community ecology.

MITAuto-check passedProductivity & Automation

Install Bio Workflows Edna Pipeline

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-workflows-edna-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workflows/edna-pipeline .claude/skills/bio-workflows-edna-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-workflows-edna-pipeline
GitHub stars
1.2k
Used in
1 other repo
Token cost
~6.2k tokens
SKILL.md length
1,100 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

End-to-end eDNA metabarcoding from raw amplicons to community ecology.

  • Works in 7 steps: Quality Assessment → Primer Removal → Paired-End Merging and Denoising → …
  • Processing eDNA samples for biodiversity assessment
  • SKILL.md covers Version Compatibility, The governing principle, Made-once commitments and Pipeline Overview, plus 11 more sections
  • Runs R and Shell scripts from its folder

What it does

Bio Workflows Edna Pipeline is an agent skill from GPTomics/bioSkills. End-to-end eDNA metabarcoding from raw amplicons to community ecology. Covers QC, primer removal (mandatory before DADA2 filterAndTrim), denoising with OBITools3 v3 (obi stats plural; DMS-based) or DADA2 ASVs (Callahan 2017), decontam combined method as screening-not-classifier (Davis 2018), tag-jumping (Schnell 2015) with a platform-dependent baseline (NovaSeq patterned flow cells ~10x MiSeq), Hill-number effective species counts with coverage-based rarefaction (Jost 2006; Chao & Jost 2012; doubling rule)…

Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/edna_pipeline.sh` and `usage-guide.md`).

It sits in Productivity & Automation, covering Messaging and chat bots. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Processing eDNA samples for biodiversity assessment
  • Deciding ASV vs OTU
  • Configuring OBITools3 v3
  • Interpreting decontam screening

Example prompts

  • “/bio-workflows-edna-pipeline”

Requirements

  • A Bash shell

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Quality Assessment
  2. Primer Removal
  3. Paired-End Merging and Denoising
  4. Contamination Filtering
  5. Taxonomy Assignment
  6. Diversity Analysis
  7. Community Comparison

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R and Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Workflows Edna Pipeline loads about 6.2k tokens when it runs. Until then it costs about 232 tokens; SKILL.md has 1,100 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~232
When it runs · the whole SKILL.md, loaded when a task matches
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,100 words, ~6,220 tokens.

Download SKILL.mdSave it as .claude/skills/bio-workflows-edna-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-workflows-edna-pipeline
description
End-to-end eDNA metabarcoding from raw amplicons to community ecology. Covers QC, primer removal (mandatory before DADA2 filterAndTrim), denoising with OBITools3 v3 (obi stats plural; DMS-based) or DADA2 ASVs (Callahan 2017), decontam combined method as screening-not-classifier (Davis 2018), tag-jumping (Schnell 2015) with a platform-dependent baseline (NovaSeq patterned flow cells ~10x MiSeq), Hill-number effective species counts with coverage-based rarefaction (Jost 2006; Chao & Jost 2012; doubling rule), beta-diversity decomposition with MANDATORY PERMANOVA + PERMDISP pair (Anderson & Walsh 2013), constrained ordination, and the read-counts-not-abundance critique (Lamb 2019). Use when processing eDNA samples for biodiversity assessment, deciding ASV vs OTU, configuring OBITools3 v3, interpreting decontam screening, or reporting community comparisons with the dispersion confound check.
tool_type
mixed
primary_tool
obitools3
goal_approach_exempt
true
workflow
true
depends_on
ecological-genomics/edna-metabarcoding, ecological-genomics/biodiversity-metrics, ecological-genomics/community-ecology, read-qc/quality-reports

Version Compatibility

Reference examples tested with: DADA2 1.30+, FastQC 0.12+, MultiQC 1.21+, cutadapt 4.4+, phyloseq 1.46+, vegan 2.6+

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

eDNA Metabarcoding Pipeline

"Process my eDNA samples from raw reads to community ecology" -> Orchestrate primer removal, denoising (OBITools3 or DADA2), contamination filtering, taxonomy assignment, Hill number diversity estimation, and constrained ordination for species-environment analysis.

This is a workflow skill: it owns the chaining decisions and hand-offs, not the internals of any one step.

The governing principle

An eDNA community table is a position in a choice-chain (marker -> primer bias -> denoise -> decontam -> tag-jump -> taxonomy DB), never a census; the trustworthy result is decided at these seams.

  1. The marker + primer set + reference DB are committed together (COI/12S/ITS/rbcL/18S). The marker fixes both what taxa amplify and the achievable assignment rate; 50-85% unassigned AT SPECIES LEVEL is TYPICAL (higher ranks assign far better) and must be reported honestly (a high unassigned fraction is a database gap, not a pipeline failure). Report the marker's known primer bias with every result.
  2. Read counts are NOT abundance (Lamb 2019). eDNA read counts reflect biomass AND primer affinity AND copy number AND degradation — commit to presence/occupancy or effective-species-count framing, never raw-read "abundance". This is the eDNA analogue of the metagenomics read-fraction != cell-fraction seam.
  3. Coverage-based (not size-based) rarefaction; Hill numbers as effective species counts; Chao1 is a LOWER bound. Commit to iNEXT coverage-standardization at C~0.95 and the doubling-rule extrapolation bound (Chao & Jost 2012); report Chao1 as ">=" and NEVER compute it after aggressive denoising (singletons stripped -> Chao1 degenerates to observed richness).
  4. The platform sets the tag-jump baseline, committed at sequencing. The tag-jump phenomenon (Schnell 2015) has a platform-dependent rate: MiSeq index-hopping ~0.001-0.005 vs NovaSeq patterned flow cells ~10x higher; quantify and platform-filter the residual rate. And decontam is SCREENING, not a classifier (Davis 2018) — every flagged ASV gets a biological-plausibility review; retain only where statistics AND plausibility agree.

Made-once commitments

CommitmentConsequence inherited downstream
Marker + primer set + reference DBWhat amplifies + the assignment rate; 50-85% unassigned at species level is typical; report the primer bias + DB release
Read-counts-not-abundance framingPresence/occupancy or effective-species-count only; raw-read "abundance" is invalid
Rarefaction baseline (coverage-based iNEXT, C~0.95)Hill effective counts; Chao1 as a lower bound, never post-denoising
Platform (MiSeq vs NovaSeq)The tag-jump baseline; NovaSeq ~10x higher index hopping; quantify + filter

Pipeline Overview

Raw amplicon FASTQ (demultiplexed)
    |
    v
[1. QC] ------------------> FastQC / MultiQC quality assessment
    |
    v
[2. Primer Removal] ------> Cutadapt (remove forward + reverse primers)
    |                            |
    |                            +---> QC: reads per sample >1000
    |
    +--- Path A: OBITools3     +--- Path B: DADA2
    |       |                  |       |
    |       v                  |       v
    |   [3a. obi alignpairedend]   [3b. filterAndTrim]
    |       |                  |       |
    |       v                  |       v
    |   [4a. obi uniq]        |   [4b. learnErrors + dada]
    |       |                  |       |
    |       v                  |       v
    |   [5a. obi ecotag]      |   [5b. assignTaxonomy]
    |                          |
    +-------- Merge -----------+
                |
                v
[6. Contamination Filter] -> decontam / microDecon (negative control removal)
    |
    v
[7. Taxonomy Table] -------> Species x sample matrix
    |
    v
[8. Diversity Analysis] ---> iNEXT Hill numbers (q=0,1,2)
    |
    v
[9. Community Comparison] -> vegan CCA/RDA + indicspecies
    |
    v
Species table + diversity metrics + ordination plots

Step 1: Quality Assessment

bash
fastqc -t 8 -o fastqc_output/ raw_reads/*.fastq.gz
multiqc fastqc_output/ -o multiqc_report/

Step 2: Primer Removal

bash
# Adapter sequences are marker-specific; examples below for common eDNA markers
# --discard-untrimmed: remove reads without primers (likely off-target)
# --minimum-length 50: discard very short fragments after trimming

# COI (Leray primers mlCOIintF / jgHCO2198)
cutadapt -g GGWACWGGWTGAACWGTWTAYCCYCC -G TAIACYTCIGGRTGICCRAARAAYCA \
    --discard-untrimmed --minimum-length 50 -j 8 \
    -o trimmed/{sample}_R1.fastq.gz -p trimmed/{sample}_R2.fastq.gz \
    raw_reads/{sample}_R1.fastq.gz raw_reads/{sample}_R2.fastq.gz

Common primer sets by marker:

MarkerForward PrimerReverse PrimerTarget
COImlCOIintFjgHCO2198Metazoan invertebrates
12S MiFishMiFish-U-FMiFish-U-RFish
ITS2ITS3ITS4Fungi
rbcLrbcLa-FrbcLa-RPlants
18S V91389F1510REukaryotes
QC Checkpoint: Demultiplexing
bash
# Gate: reads per sample >1000; negative controls <100 reads
for f in trimmed/*_R1.fastq.gz; do
    sample=$(basename "$f" _R1.fastq.gz)
    count=$(zcat "$f" | awk 'END{print NR/4}')
    echo "$sample: $count reads"
done

Step 3: Paired-End Merging and Denoising

Path A: OBITools3
bash
# Import, align, and SAMPLE-TAG each demultiplexed sample separately, then concatenate.
# On multiplexed data `obi ngsfilter -t tagfile` assigns the `sample` tag; with pre-demultiplexed
# FASTQ it must be set explicitly (-S takes TAG:PYTHON_EXPRESSION), otherwise `obi uniq -m sample`
# has nothing to merge on and the MERGED_sample tag `obi clean -s` needs is never created.
for s in $(cat samples.txt); do
    obi import --fastq-input trimmed/${s}_R1.fastq.gz reads/${s}_r1
    obi import --fastq-input trimmed/${s}_R2.fastq.gz reads/${s}_r2
    obi alignpairedend -R reads/${s}_r2 reads/${s}_r1 reads/${s}_aligned
    obi annotate -S "sample:'${s}'" reads/${s}_aligned reads/${s}_tagged
done
obi cat $(for s in $(cat samples.txt); do printf -- '-c reads/%s_tagged ' "$s"; done) reads/aligned

# Filter by alignment score
# score >= 50: removes poorly overlapping pairs
obi grep -p 'sequence["score"] >= 50' reads/aligned reads/filtered

# Filter by merged length (marker-dependent range)
# 100-500 bp: typical COI amplicon range
obi grep -p 'len(sequence) >= 100 and len(sequence) <= 500' \
    reads/filtered reads/length_filtered

# Dereplicate
obi uniq -m sample reads/length_filtered reads/dereplicated   # -m sample creates the MERGED_sample tag obi clean -s needs

# Remove singletons
# count >=2: removes sequencing errors; increase to 5-10 for noisy datasets
obi grep -p 'sequence["COUNT"] >= 2' reads/dereplicated reads/denoised   # obi uniq writes the COUNT tag uppercase (case-sensitive)

# Denoise (remove PCR/sequencing errors)
# ratio 0.05: sequences <5% abundance of a 1-mismatch parent are merged
obi clean -s MERGED_sample -r 0.05 -H reads/denoised reads/cleaned
Path B: DADA2 (R)
r
library(dada2)

path <- 'trimmed/'
fnFs <- sort(list.files(path, pattern = '_R1.fastq.gz', full.names = TRUE))
fnRs <- sort(list.files(path, pattern = '_R2.fastq.gz', full.names = TRUE))
sample_names <- gsub('_R1.fastq.gz', '', basename(fnFs))

# Filter and trim
# truncLen: set based on quality profiles; marker-dependent
# maxEE c(2,2): max expected errors; standard for eDNA
# minLen 50: minimum after trimming
filtFs <- file.path('filtered', paste0(sample_names, '_F_filt.fastq.gz'))
filtRs <- file.path('filtered', paste0(sample_names, '_R_filt.fastq.gz'))
out <- filterAndTrim(fnFs, filtFs, fnRs, filtRs,
                     truncLen = c(200, 180), maxEE = c(2, 2),
                     minLen = 50, truncQ = 2, rm.phix = TRUE,
                     multithread = TRUE)

# Learn error rates
errF <- learnErrors(filtFs, multithread = TRUE)
errR <- learnErrors(filtRs, multithread = TRUE)

# Denoise
dadaFs <- dada(filtFs, err = errF, multithread = TRUE)
dadaRs <- dada(filtRs, err = errR, multithread = TRUE)

# Merge paired reads
# minOverlap 20: standard; increase if amplicon has short overlap region
merged <- mergePairs(dadaFs, filtFs, dadaRs, filtRs, minOverlap = 20)

# Build ASV table
seqtab <- makeSequenceTable(merged)

# Remove chimeras
# method 'consensus': standard; 'pooled' for higher sensitivity
seqtab_nochim <- removeBimeraDenovo(seqtab, method = 'consensus', multithread = TRUE)
chimera_rate <- 1 - sum(seqtab_nochim) / sum(seqtab)
message(sprintf('Chimera rate: %.1f%%', chimera_rate * 100))
QC Checkpoint: Denoising
r
# Gate 1: Chimera rate <20%
if (chimera_rate > 0.20) message('WARNING: High chimera rate. Check primer removal and PCR conditions.')

# Gate 2: ASV count reasonable for marker
n_asvs <- ncol(seqtab_nochim)
message(sprintf('ASVs after denoising: %d', n_asvs))
# Typical ranges: COI 500-5000, 12S 50-500, ITS 200-3000

Step 4: Contamination Filtering

R (decontam)
r
library(decontam)
library(phyloseq)

ps <- phyloseq(otu_table(seqtab_nochim, taxa_are_rows = FALSE),
               sample_data(meta))

# Identify negative controls
sample_data(ps)$is_neg <- sample_data(ps)$sample_type == 'negative_control'

# Combined method (Davis 2018): uses BOTH negative controls AND DNA concentration
# threshold=0.1 default; 0.05 for low-biomass samples
# CRITICAL: output is SCREENING, not classification; biological-plausibility check required
if ('dna_concentration' %in% sample_variables(ps)) {
    contam <- isContaminant(ps, method = 'combined', neg = 'is_neg',
                            conc = 'dna_concentration', threshold = 0.1)
} else {
    # Fall back to prevalence-only if no qPCR/Qubit DNA-concentration data
    contam <- isContaminant(ps, method = 'prevalence', neg = 'is_neg',
                            threshold = 0.1)
}
message(sprintf('Flagged candidate contaminant ASVs: %d', sum(contam$contaminant)))
message('Manual review required: verify biological plausibility before deletion')

ps_clean <- prune_taxa(!contam$contaminant, ps)

# Remove negative control samples
ps_clean <- subset_samples(ps_clean, sample_type != 'negative_control')
Tag-jumping removal (Schnell 2015; NovaSeq caveat)
r
# Tag-jumping: cross-contamination from index hopping during library prep / sequencing
# Schnell 2015 Mol Ecol Resour 15:1289-1303 documented 0.1-2% per read pair
# NovaSeq patterned flow cells have ~10x higher rates than MiSeq
# Use platform-appropriate threshold:
#   MiSeq: 0.001-0.005 (0.1-0.5% of ASV total)
#   NovaSeq: 0.005-0.01 (0.5-1% of ASV total)
# Quantify residual rate first from per-ASV cross-sample appearance
otu <- as(otu_table(ps_clean), 'matrix')
max_per_asv <- apply(otu, 2, max)
otu_filtered <- otu
# Set threshold by platform; default below is for MiSeq
tag_jump_frac <- 0.001
for (j in 1:ncol(otu)) {
    threshold <- max_per_asv[j] * tag_jump_frac
    otu_filtered[otu[, j] < threshold, j] <- 0
}
otu_table(ps_clean) <- otu_table(otu_filtered, taxa_are_rows = FALSE)
# Modern alternative: metabaR::tagjumpslayer(metabarlist_obj, threshold = 0.03)
QC Checkpoint: Decontamination
r
# Gate: verify contaminant ASVs were removed from real samples
n_before <- ntaxa(ps)
n_after <- ntaxa(ps_clean)
message(sprintf('ASVs removed as contaminants: %d (%.1f%%)',
                n_before - n_after, (n_before - n_after) / n_before * 100))
if ((n_before - n_after) / n_before > 0.5) {
    message('WARNING: >50% ASVs flagged as contaminants. Review decontam threshold.')
}

Step 5: Taxonomy Assignment

OBITools3 (ecotag)
bash
# ecotag assigns taxonomy using LCA algorithm against reference database
# Reference databases: EMBL, BOLD, MIDORI2, UNITE (marker-dependent)
obi ecotag -R reads/refdb --taxonomy reads/taxonomy reads/cleaned reads/assigned

# Filter by assignment quality (species-level for COI)
obi grep -p 'sequence["BEST_IDENTITY"] >= 0.97' reads/assigned reads/filtered_assigned   # ecotag writes BEST_IDENTITY uppercase (case-sensitive)

obi export --tab-output reads/filtered_assigned > taxonomy_results.tsv
DADA2 (assignTaxonomy)
r
# SILVA for 16S/18S, UNITE for ITS, custom for COI/12S
# Reference databases must be formatted for DADA2
# minBoot 50: minimum bootstrap confidence; 80 for more conservative assignments
taxa <- assignTaxonomy(seqtab_nochim, 'reference_db.fa.gz', multithread = TRUE, minBoot = 50)
taxa <- addSpecies(taxa, 'species_db.fa.gz')
Marker-specific taxonomy databases
MarkerDatabaseTypical Assignment Rate
COIBOLD / Midori2>90% to phylum, 60-80% to species
12SMitoFish / 12S-seqdb>90% to family for fish
ITSUNITE>80% to genus for fungi
rbcLGenBank / NCBI nt>85% to family for plants
18SSILVA / PR2>90% to phylum
QC Checkpoint: Taxonomy
r
# Gate: assignment rate should meet marker expectations
assigned <- !is.na(taxa[, 'Phylum'])
assignment_rate <- sum(assigned) / length(assigned) * 100
message(sprintf('Taxonomy assignment rate (phylum level): %.1f%%', assignment_rate))
if (assignment_rate < 60) message('WARNING: Low assignment rate. Check reference database completeness.')

Step 6: Diversity Analysis

R (iNEXT)
r
library(iNEXT)

otu_matrix <- as(otu_table(ps_clean), 'matrix')

# Hill numbers: q=0 (richness), q=1 (Shannon diversity), q=2 (Simpson diversity).
# Default endpoint = per-sample 2x reference size (the Chao 2014 doubling rule); do NOT hardcode a
# single global 2*max, which over-extrapolates small samples under unequal library sizes.
inext_out <- iNEXT(as.list(as.data.frame(t(otu_matrix))),
                   q = c(0, 1, 2), datatype = 'abundance')

# Coverage-based standardization at C=0.95: the fair cross-sample comparison (equalizes completeness,
# not raw depth). estimateD returns Hill numbers at that coverage.
div_c95 <- estimateD(as.list(as.data.frame(t(otu_matrix))),
                     q = c(0, 1, 2), datatype = 'abundance', base = 'coverage', level = 0.95)

# Sample completeness diagnostic: fraction of estimated diversity observed
completeness <- inext_out$DataInfo$SC
message(sprintf('Sample completeness range: %.1f%% - %.1f%%',
                min(completeness) * 100, max(completeness) * 100))
QC Checkpoint: Diversity
r
# Gate 1: rarefaction approaching asymptote
if (min(completeness) < 0.80) {
    message('WARNING: Some samples have low completeness (<80%). Deeper sequencing recommended.')
}

# Gate 2: richness. Use the COVERAGE-STANDARDIZED q=0, NOT inext_out$AsyEst 'Species richness' (the
# Chao1 asymptotic): denoising stripped singletons (COUNT >= 2), leaving Chao1 no f1, so it degenerates
# to observed richness -- the collapse rule 3 forbids. estimateD returns Order.q / qD.
q0_c95 <- div_c95[div_c95$Order.q == 0, ]
message(sprintf('Coverage-standardized (C=0.95) richness range: %.0f - %.0f effective species',   # qD is fractional; %d errors on doubles
                min(q0_c95$qD), max(q0_c95$qD)))

Step 7: Community Comparison

R (vegan + indicspecies)
r
library(vegan)
library(indicspecies)

otu_matrix <- as(otu_table(ps_clean), 'matrix')
env_data <- as(sample_data(ps_clean), 'data.frame')

# Hellinger transformation: standard for community composition data
# Reduces influence of dominant species
otu_hell <- decostand(otu_matrix, method = 'hellinger')

# DCA on untransformed data to determine gradient length
dca <- decorana(otu_matrix)
gradient_length <- diff(range(scores(dca, display = 'sites', choices = 1)))
message(sprintf('DCA gradient length: %.2f SD', gradient_length))

# RDA: linear response (<=3 SD), uses Hellinger-transformed data
# CCA: unimodal response (>3 SD), uses raw abundances (chi-squared distance)
if (gradient_length <= 3) {
    ord <- rda(otu_hell ~ temperature + depth + season, data = env_data)
    method_name <- 'RDA'
} else {
    ord <- cca(otu_matrix ~ temperature + depth + season, data = env_data)
    method_name <- 'CCA'
}

# Permutation test for significance
# permutations 999: standard; increase to 9999 for publication
anova_result <- anova.cca(ord, permutations = 999)
message(sprintf('%s significance: p = %.4f', method_name, anova_result$`Pr(>F)`[1]))

# MANDATORY companion: PERMANOVA + PERMDISP (Anderson & Walsh 2013 Ecol Monogr 83:557-574)
# adonis2 tests centroid difference; betadisper tests dispersion homogeneity
# If betadisper is also significant, PERMANOVA significance is dispersion-confounded
bray_dist <- vegdist(otu_matrix, method = 'bray')
permanova <- adonis2(bray_dist ~ site, data = env_data,
                     by = 'margin', permutations = 999)
disp <- betadisper(bray_dist, env_data$site)
disp_test <- permutest(disp, permutations = 999)
message(sprintf('PERMANOVA p = %.4f; PERMDISP p = %.4f',
                permanova[['Pr(>F)']][1], disp_test$tab[['Pr(>F)']][1]))
if (permanova[['Pr(>F)']][1] < 0.05 && disp_test$tab[['Pr(>F)']][1] < 0.05) {
    message('WARNING: Both PERMANOVA and PERMDISP significant; location-vs-dispersion confounded')
}

# Indicator species analysis with group-size equalization (NOT basic IndVal)
# func='IndVal.g' corrects for unbalanced group sizes (De Caceres & Legendre 2009)
indval <- multipatt(otu_matrix, env_data$site, func = 'IndVal.g',
                    control = how(nperm = 999))
summary(indval, alpha = 0.05)

Parameter Recommendations

StepParameterRecommendation
Cutadapt--discard-untrimmedAlways use; removes off-target reads. MANDATORY before DADA2 filterAndTrim
Cutadapt--minimum-length50 (general); adjust per expected amplicon size
DADA2truncLenSet from quality profiles; marker-dependent (typical 2x250 COI: c(220,180))
DADA2maxEEc(2,2) standard; c(5,5) for degraded eDNA
DADA2minOverlap20 (standard); increase for short overlaps
DADA2chimera method'consensus' standard (conservative); 'pooled' more aggressive
OBITools3 v3command syntaxobi stats (plural; was obistat in v1); .tar.gz taxonomy
OBITools3--min-count2 (removes singletons); 5-10 for noisy datasets
decontammethod'combined' if concentration AND controls; 'prevalence' fallback
decontamthreshold0.1 default; 0.05 for low-biomass samples
Tag-jumping MiSeqthreshold0.001-0.005 fraction of ASV total
Tag-jumping NovaSeqthreshold0.005-0.01 (~10x MiSeq; patterned flow cells)
TaxonomyminBoot50 (sensitive); 80 (conservative)
ecotag--minimum-identity (-m)0.97 (COI species); 0.95 (genus); marker-dependent
iNEXTendpointper-sample 2x reference size (default doubling rule); do not set a global 2*max
veganpermutations999 (standard); 9999 (publication)
Show full SKILL.md (367 more words)Show less

Common Errors

SymptomCauseFix
Over-called rare-species presence (worst on NovaSeq)Tag-jump rate never quantified ("we used dual indexing" and stop)Report the residual rate (reads in should-be-zero cells / total) and platform-filter
Real low-biomass taxa deleteddecontam output treated as ground truthBiological-plausibility review of every flag; retain only where statistics AND plausibility agree
Chao1 collapses to observed richnessChao1 computed after aggressive denoising (singletons stripped)Use incidence-based Chao2, or estimators only on data with a real singleton/doubleton distribution
A "community shift" that is really dispersionPERMANOVA reported without PERMDISPAlways pair adonis2 with betadisper/permutest; if dispersion is significant, the location conclusion is not supported
Read counts interpreted as abundanceeDNA reads treated as biomassPresence/occupancy or effective-species-count framing only (Lamb 2019)
Few reads after primer removalWrong primer sequences or orientationVerify primer sequences; try --revcomp
High chimera rate (>20%)Excessive PCR cycles or low-quality inputReduce PCR cycles; improve DNA extraction
Many unassigned ASVsIncomplete reference databaseUse marker-specific database; lower minBoot
Contamination in negativesTag-jumping or lab contaminationApply tag-jump filter; review extraction protocol
Low sample completenessInsufficient sequencing depthIncrease sequencing; pool fewer samples
Ordination axes not significantWeak environmental gradientsAdd more environmental variables; check sample size
Unexpected taxa (e.g., human)Sample contaminationFilter known contaminants; review field protocols
Very few ASVsOver-aggressive filteringRelax truncLen, maxEE, or min-count thresholds
  • ecological-genomics/edna-metabarcoding - Detailed eDNA processing
  • ecological-genomics/biodiversity-metrics - Diversity analysis details
  • ecological-genomics/community-ecology - Ordination and indicator species
  • read-qc/quality-reports - Raw read quality assessment
  • reporting/automated-qc-reports - Aggregate FastQC across samples with MultiQC (a triage snapshot, not a pass/fail gate)
  • microbiome/amplicon-processing - 16S clinical alternative

References

  • Lamb PD, Hunter E, Pinnegar JK, et al (2019) How quantitative is metabarcoding? A meta-analytical approach. Molecular Ecology 28:420-430. DOI 10.1111/mec.14920. (read counts are not abundance.)
  • Schnell IB, Bohmann K, Gilbert MTP (2015) Tag jumps illuminated - reducing sequence-to-sample misidentifications in metabarcoding studies. Molecular Ecology Resources 15:1289-1303. DOI 10.1111/1755-0998.12402. (tag jumping.)
  • Davis NM, Proctor DM, Holmes SP, Relman DA, Callahan BJ (2018) Simple statistical identification and removal of contaminant sequences in marker-gene and metagenomics data. Microbiome 6:226. DOI 10.1186/s40168-018-0605-2. (decontam.)
  • Anderson MJ, Walsh DCI (2013) PERMANOVA, ANOSIM, and the Mantel test in the face of heterogeneous dispersions. Ecological Monographs 83:557-574. DOI 10.1890/12-2010.1. (PERMANOVA + PERMDISP.)

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in workflows/edna-pipeline of GPTomics/bioSkills.

  • SKILL.md
  • examples/edna_pipeline.R
  • examples/edna_pipeline.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Workflows Edna Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Workflows Edna Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Workflows Edna Pipeline this skillGPTomics/bioSkills1.2k1 repos~6.2kAutomated safety check: PassMIT
Feishu Docopenclaw/openclaw392k—~516Automated safety check: PassMIT
She Love Me863401402/she-love-me9311 repos~1.3kAutomated safety check: PassMIT
Feishu Docraucvr/Group-Goki1123 repos~592Automated safety check: PassMIT
Wechat Article Extractorfreestylefly/wechat-article-extractor-skill1361 repos~1kAutomated safety check: PassNone
Wechat Miniprogram Builderchenjin-cmd/wechat-miniprogram-builder356—~634Automated safety check: PassMIT

Similar skills

  • Feishu Doc

    openclaw/openclaw

    Feishu document read/write workflows. An agent skill from openclaw/openclaw.

    392k GitHub stars~516 tokensUpdated today
    Productivity & AutomationAuto-check passed
  • She Love Me

    863401402/she-love-me

    Acquire, import, and analyze WeChat or QQ chat histories, including installing supported exporters, guiding required login or contact selection, converting exports, assessing relationship dynamics…

    931 GitHub starsUsed in 1 repo~1.3k tokens
    Productivity & AutomationAuto-check passed
  • Feishu Doc

    raucvr/Group-Goki

    Feishu document read/write operations. An agent skill from raucvr/Group-Goki.

    112 GitHub starsUsed in 3 repos~592 tokens
    Productivity & AutomationAuto-check passed
  • Wechat Article Extractor

    freestylefly/wechat-article-extractor-skill

    Extract metadata and content from WeChat Official Account articles.

    136 GitHub starsUsed in 1 repo~1k tokens
    Productivity & AutomationAuto-check passed
  • Wechat Miniprogram Builder

    chenjin-cmd/wechat-miniprogram-builder

    This skill should be used when the user wants to build, launch, monetize, or promote a WeChat mini-program with AI (vibe coding) — including topic selection, account registration & ICP filing…

    356 GitHub stars~634 tokensUpdated 2 mo ago
    Productivity & AutomationAuto-check passed
  • Slack

    huangruiteng/CS-Notes

    A skill your agent uses when you need to control Slack from Clawdbot via the slack tool, including reacting to messages or pinning/unpinning items in Slack channels or DMs.

    4k GitHub starsUsed in 10 repos~578 tokens
    Productivity & AutomationAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Workflows Edna Pipeline

What does Bio Workflows Edna Pipeline do?

End-to-end eDNA metabarcoding from raw amplicons to community ecology. Bio Workflows Edna Pipeline is an agent skill from GPTomics/bioSkills. End-to-end eDNA metabarcoding from raw amplicons to community ecology.

When should I use Bio Workflows Edna Pipeline?

Bio Workflows Edna Pipeline fits situations like: processing eDNA samples for biodiversity assessment; deciding ASV vs OTU; configuring OBITools3 v3; interpreting decontam screening.

How do I install Bio Workflows Edna Pipeline in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a claude-code`. Or copy the skill folder (workflows/edna-pipeline in GPTomics/bioSkills) into .claude/skills/bio-workflows-edna-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Bio Workflows Edna Pipeline in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a codex`. Or copy the skill folder (workflows/edna-pipeline in GPTomics/bioSkills) into .agents/skills/bio-workflows-edna-pipeline in your project. Codex loads it when a task matches its description.

Can I use Bio Workflows Edna Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-workflows-edna-pipeline, .gemini/skills/bio-workflows-edna-pipeline, .github/skills/bio-workflows-edna-pipeline and .opencode/skills/bio-workflows-edna-pipeline in your project.

What does Bio Workflows Edna Pipeline need to run?

Going by SKILL.md and its folder, Bio Workflows Edna Pipeline needs R and a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Bio Workflows Edna Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Workflows Edna Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Workflows Edna Pipeline use?

Bio Workflows Edna Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Workflows Edna Pipeline use?

About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Workflows Edna Pipeline?

Skills that share tags, products or a category with Bio Workflows Edna Pipeline: Feishu Doc (openclaw/openclaw, 392k stars), She Love Me (863401402/she-love-me, 931 stars), Feishu Doc (raucvr/Group-Goki, 112 stars) and Wechat Article Extractor (freestylefly/wechat-article-extractor-skill, 136 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Workflows Edna Pipeline?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.