Feishu Doc
openclaw/openclaw
Feishu document read/write workflows. An agent skill from openclaw/openclaw.
End-to-end eDNA metabarcoding from raw amplicons to community ecology.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-edna-pipeline --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workflows/edna-pipeline .claude/skills/bio-workflows-edna-pipeline && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-workflows-edna-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/edna-pipeline into .claude/skills/bio-workflows-edna-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-edna-pipeline", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/workflows/edna-pipelineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-edna-pipeline --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/workflows/edna-pipeline .agents/skills/bio-workflows-edna-pipeline && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-workflows-edna-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/edna-pipeline into .agents/skills/bio-workflows-edna-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-edna-pipeline", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-edna-pipeline --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/workflows/edna-pipeline .cursor/skills/bio-workflows-edna-pipeline && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-workflows-edna-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/edna-pipeline into .cursor/skills/bio-workflows-edna-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-edna-pipeline", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path workflows/edna-pipeline--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-edna-pipeline --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/workflows/edna-pipeline .gemini/skills/bio-workflows-edna-pipeline && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-workflows-edna-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/edna-pipeline into .gemini/skills/bio-workflows-edna-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-edna-pipeline", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-workflows-edna-pipelineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/workflows/edna-pipeline .github/skills/bio-workflows-edna-pipeline && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-workflows-edna-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/edna-pipeline into .github/skills/bio-workflows-edna-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-edna-pipeline", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-edna-pipeline --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/workflows/edna-pipeline .opencode/skills/bio-workflows-edna-pipeline && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-workflows-edna-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/edna-pipeline into .opencode/skills/bio-workflows-edna-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-edna-pipeline", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-workflows-edna-pipelineEnd-to-end eDNA metabarcoding from raw amplicons to community ecology.
Bio Workflows Edna Pipeline is an agent skill from GPTomics/bioSkills. End-to-end eDNA metabarcoding from raw amplicons to community ecology. Covers QC, primer removal (mandatory before DADA2 filterAndTrim), denoising with OBITools3 v3 (obi stats plural; DMS-based) or DADA2 ASVs (Callahan 2017), decontam combined method as screening-not-classifier (Davis 2018), tag-jumping (Schnell 2015) with a platform-dependent baseline (NovaSeq patterned flow cells ~10x MiSeq), Hill-number effective species counts with coverage-based rarefaction (Jost 2006; Chao & Jost 2012; doubling rule)…
Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/edna_pipeline.sh` and `usage-guide.md`).
It sits in Productivity & Automation, covering Messaging and chat bots. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (R and Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Workflows Edna Pipeline loads about 6.2k tokens when it runs. Until then it costs about 232 tokens; SKILL.md has 1,100 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,100 words, ~6,220 tokens.
.claude/skills/bio-workflows-edna-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: DADA2 1.30+, FastQC 0.12+, MultiQC 1.21+, cutadapt 4.4+, phyloseq 1.46+, vegan 2.6+
Before using code patterns, verify installed versions match. If versions differ:
packageVersion('<pkg>') then ?function_name to verify parameters<tool> --version then <tool> --help to confirm flagsIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Process my eDNA samples from raw reads to community ecology" -> Orchestrate primer removal, denoising (OBITools3 or DADA2), contamination filtering, taxonomy assignment, Hill number diversity estimation, and constrained ordination for species-environment analysis.
This is a workflow skill: it owns the chaining decisions and hand-offs, not the internals of any one step.
An eDNA community table is a position in a choice-chain (marker -> primer bias -> denoise -> decontam -> tag-jump -> taxonomy DB), never a census; the trustworthy result is decided at these seams.
| Commitment | Consequence inherited downstream |
|---|---|
| Marker + primer set + reference DB | What amplifies + the assignment rate; 50-85% unassigned at species level is typical; report the primer bias + DB release |
| Read-counts-not-abundance framing | Presence/occupancy or effective-species-count only; raw-read "abundance" is invalid |
| Rarefaction baseline (coverage-based iNEXT, C~0.95) | Hill effective counts; Chao1 as a lower bound, never post-denoising |
| Platform (MiSeq vs NovaSeq) | The tag-jump baseline; NovaSeq ~10x higher index hopping; quantify + filter |
Raw amplicon FASTQ (demultiplexed)
|
v
[1. QC] ------------------> FastQC / MultiQC quality assessment
|
v
[2. Primer Removal] ------> Cutadapt (remove forward + reverse primers)
| |
| +---> QC: reads per sample >1000
|
+--- Path A: OBITools3 +--- Path B: DADA2
| | | |
| v | v
| [3a. obi alignpairedend] [3b. filterAndTrim]
| | | |
| v | v
| [4a. obi uniq] | [4b. learnErrors + dada]
| | | |
| v | v
| [5a. obi ecotag] | [5b. assignTaxonomy]
| |
+-------- Merge -----------+
|
v
[6. Contamination Filter] -> decontam / microDecon (negative control removal)
|
v
[7. Taxonomy Table] -------> Species x sample matrix
|
v
[8. Diversity Analysis] ---> iNEXT Hill numbers (q=0,1,2)
|
v
[9. Community Comparison] -> vegan CCA/RDA + indicspecies
|
v
Species table + diversity metrics + ordination plotsfastqc -t 8 -o fastqc_output/ raw_reads/*.fastq.gz
multiqc fastqc_output/ -o multiqc_report/# Adapter sequences are marker-specific; examples below for common eDNA markers
# --discard-untrimmed: remove reads without primers (likely off-target)
# --minimum-length 50: discard very short fragments after trimming
# COI (Leray primers mlCOIintF / jgHCO2198)
cutadapt -g GGWACWGGWTGAACWGTWTAYCCYCC -G TAIACYTCIGGRTGICCRAARAAYCA \
--discard-untrimmed --minimum-length 50 -j 8 \
-o trimmed/{sample}_R1.fastq.gz -p trimmed/{sample}_R2.fastq.gz \
raw_reads/{sample}_R1.fastq.gz raw_reads/{sample}_R2.fastq.gzCommon primer sets by marker:
| Marker | Forward Primer | Reverse Primer | Target |
|---|---|---|---|
| COI | mlCOIintF | jgHCO2198 | Metazoan invertebrates |
| 12S MiFish | MiFish-U-F | MiFish-U-R | Fish |
| ITS2 | ITS3 | ITS4 | Fungi |
| rbcL | rbcLa-F | rbcLa-R | Plants |
| 18S V9 | 1389F | 1510R | Eukaryotes |
# Gate: reads per sample >1000; negative controls <100 reads
for f in trimmed/*_R1.fastq.gz; do
sample=$(basename "$f" _R1.fastq.gz)
count=$(zcat "$f" | awk 'END{print NR/4}')
echo "$sample: $count reads"
done# Import, align, and SAMPLE-TAG each demultiplexed sample separately, then concatenate.
# On multiplexed data `obi ngsfilter -t tagfile` assigns the `sample` tag; with pre-demultiplexed
# FASTQ it must be set explicitly (-S takes TAG:PYTHON_EXPRESSION), otherwise `obi uniq -m sample`
# has nothing to merge on and the MERGED_sample tag `obi clean -s` needs is never created.
for s in $(cat samples.txt); do
obi import --fastq-input trimmed/${s}_R1.fastq.gz reads/${s}_r1
obi import --fastq-input trimmed/${s}_R2.fastq.gz reads/${s}_r2
obi alignpairedend -R reads/${s}_r2 reads/${s}_r1 reads/${s}_aligned
obi annotate -S "sample:'${s}'" reads/${s}_aligned reads/${s}_tagged
done
obi cat $(for s in $(cat samples.txt); do printf -- '-c reads/%s_tagged ' "$s"; done) reads/aligned
# Filter by alignment score
# score >= 50: removes poorly overlapping pairs
obi grep -p 'sequence["score"] >= 50' reads/aligned reads/filtered
# Filter by merged length (marker-dependent range)
# 100-500 bp: typical COI amplicon range
obi grep -p 'len(sequence) >= 100 and len(sequence) <= 500' \
reads/filtered reads/length_filtered
# Dereplicate
obi uniq -m sample reads/length_filtered reads/dereplicated # -m sample creates the MERGED_sample tag obi clean -s needs
# Remove singletons
# count >=2: removes sequencing errors; increase to 5-10 for noisy datasets
obi grep -p 'sequence["COUNT"] >= 2' reads/dereplicated reads/denoised # obi uniq writes the COUNT tag uppercase (case-sensitive)
# Denoise (remove PCR/sequencing errors)
# ratio 0.05: sequences <5% abundance of a 1-mismatch parent are merged
obi clean -s MERGED_sample -r 0.05 -H reads/denoised reads/cleanedlibrary(dada2)
path <- 'trimmed/'
fnFs <- sort(list.files(path, pattern = '_R1.fastq.gz', full.names = TRUE))
fnRs <- sort(list.files(path, pattern = '_R2.fastq.gz', full.names = TRUE))
sample_names <- gsub('_R1.fastq.gz', '', basename(fnFs))
# Filter and trim
# truncLen: set based on quality profiles; marker-dependent
# maxEE c(2,2): max expected errors; standard for eDNA
# minLen 50: minimum after trimming
filtFs <- file.path('filtered', paste0(sample_names, '_F_filt.fastq.gz'))
filtRs <- file.path('filtered', paste0(sample_names, '_R_filt.fastq.gz'))
out <- filterAndTrim(fnFs, filtFs, fnRs, filtRs,
truncLen = c(200, 180), maxEE = c(2, 2),
minLen = 50, truncQ = 2, rm.phix = TRUE,
multithread = TRUE)
# Learn error rates
errF <- learnErrors(filtFs, multithread = TRUE)
errR <- learnErrors(filtRs, multithread = TRUE)
# Denoise
dadaFs <- dada(filtFs, err = errF, multithread = TRUE)
dadaRs <- dada(filtRs, err = errR, multithread = TRUE)
# Merge paired reads
# minOverlap 20: standard; increase if amplicon has short overlap region
merged <- mergePairs(dadaFs, filtFs, dadaRs, filtRs, minOverlap = 20)
# Build ASV table
seqtab <- makeSequenceTable(merged)
# Remove chimeras
# method 'consensus': standard; 'pooled' for higher sensitivity
seqtab_nochim <- removeBimeraDenovo(seqtab, method = 'consensus', multithread = TRUE)
chimera_rate <- 1 - sum(seqtab_nochim) / sum(seqtab)
message(sprintf('Chimera rate: %.1f%%', chimera_rate * 100))# Gate 1: Chimera rate <20%
if (chimera_rate > 0.20) message('WARNING: High chimera rate. Check primer removal and PCR conditions.')
# Gate 2: ASV count reasonable for marker
n_asvs <- ncol(seqtab_nochim)
message(sprintf('ASVs after denoising: %d', n_asvs))
# Typical ranges: COI 500-5000, 12S 50-500, ITS 200-3000library(decontam)
library(phyloseq)
ps <- phyloseq(otu_table(seqtab_nochim, taxa_are_rows = FALSE),
sample_data(meta))
# Identify negative controls
sample_data(ps)$is_neg <- sample_data(ps)$sample_type == 'negative_control'
# Combined method (Davis 2018): uses BOTH negative controls AND DNA concentration
# threshold=0.1 default; 0.05 for low-biomass samples
# CRITICAL: output is SCREENING, not classification; biological-plausibility check required
if ('dna_concentration' %in% sample_variables(ps)) {
contam <- isContaminant(ps, method = 'combined', neg = 'is_neg',
conc = 'dna_concentration', threshold = 0.1)
} else {
# Fall back to prevalence-only if no qPCR/Qubit DNA-concentration data
contam <- isContaminant(ps, method = 'prevalence', neg = 'is_neg',
threshold = 0.1)
}
message(sprintf('Flagged candidate contaminant ASVs: %d', sum(contam$contaminant)))
message('Manual review required: verify biological plausibility before deletion')
ps_clean <- prune_taxa(!contam$contaminant, ps)
# Remove negative control samples
ps_clean <- subset_samples(ps_clean, sample_type != 'negative_control')# Tag-jumping: cross-contamination from index hopping during library prep / sequencing
# Schnell 2015 Mol Ecol Resour 15:1289-1303 documented 0.1-2% per read pair
# NovaSeq patterned flow cells have ~10x higher rates than MiSeq
# Use platform-appropriate threshold:
# MiSeq: 0.001-0.005 (0.1-0.5% of ASV total)
# NovaSeq: 0.005-0.01 (0.5-1% of ASV total)
# Quantify residual rate first from per-ASV cross-sample appearance
otu <- as(otu_table(ps_clean), 'matrix')
max_per_asv <- apply(otu, 2, max)
otu_filtered <- otu
# Set threshold by platform; default below is for MiSeq
tag_jump_frac <- 0.001
for (j in 1:ncol(otu)) {
threshold <- max_per_asv[j] * tag_jump_frac
otu_filtered[otu[, j] < threshold, j] <- 0
}
otu_table(ps_clean) <- otu_table(otu_filtered, taxa_are_rows = FALSE)
# Modern alternative: metabaR::tagjumpslayer(metabarlist_obj, threshold = 0.03)# Gate: verify contaminant ASVs were removed from real samples
n_before <- ntaxa(ps)
n_after <- ntaxa(ps_clean)
message(sprintf('ASVs removed as contaminants: %d (%.1f%%)',
n_before - n_after, (n_before - n_after) / n_before * 100))
if ((n_before - n_after) / n_before > 0.5) {
message('WARNING: >50% ASVs flagged as contaminants. Review decontam threshold.')
}# ecotag assigns taxonomy using LCA algorithm against reference database
# Reference databases: EMBL, BOLD, MIDORI2, UNITE (marker-dependent)
obi ecotag -R reads/refdb --taxonomy reads/taxonomy reads/cleaned reads/assigned
# Filter by assignment quality (species-level for COI)
obi grep -p 'sequence["BEST_IDENTITY"] >= 0.97' reads/assigned reads/filtered_assigned # ecotag writes BEST_IDENTITY uppercase (case-sensitive)
obi export --tab-output reads/filtered_assigned > taxonomy_results.tsv# SILVA for 16S/18S, UNITE for ITS, custom for COI/12S
# Reference databases must be formatted for DADA2
# minBoot 50: minimum bootstrap confidence; 80 for more conservative assignments
taxa <- assignTaxonomy(seqtab_nochim, 'reference_db.fa.gz', multithread = TRUE, minBoot = 50)
taxa <- addSpecies(taxa, 'species_db.fa.gz')| Marker | Database | Typical Assignment Rate |
|---|---|---|
| COI | BOLD / Midori2 | >90% to phylum, 60-80% to species |
| 12S | MitoFish / 12S-seqdb | >90% to family for fish |
| ITS | UNITE | >80% to genus for fungi |
| rbcL | GenBank / NCBI nt | >85% to family for plants |
| 18S | SILVA / PR2 | >90% to phylum |
# Gate: assignment rate should meet marker expectations
assigned <- !is.na(taxa[, 'Phylum'])
assignment_rate <- sum(assigned) / length(assigned) * 100
message(sprintf('Taxonomy assignment rate (phylum level): %.1f%%', assignment_rate))
if (assignment_rate < 60) message('WARNING: Low assignment rate. Check reference database completeness.')library(iNEXT)
otu_matrix <- as(otu_table(ps_clean), 'matrix')
# Hill numbers: q=0 (richness), q=1 (Shannon diversity), q=2 (Simpson diversity).
# Default endpoint = per-sample 2x reference size (the Chao 2014 doubling rule); do NOT hardcode a
# single global 2*max, which over-extrapolates small samples under unequal library sizes.
inext_out <- iNEXT(as.list(as.data.frame(t(otu_matrix))),
q = c(0, 1, 2), datatype = 'abundance')
# Coverage-based standardization at C=0.95: the fair cross-sample comparison (equalizes completeness,
# not raw depth). estimateD returns Hill numbers at that coverage.
div_c95 <- estimateD(as.list(as.data.frame(t(otu_matrix))),
q = c(0, 1, 2), datatype = 'abundance', base = 'coverage', level = 0.95)
# Sample completeness diagnostic: fraction of estimated diversity observed
completeness <- inext_out$DataInfo$SC
message(sprintf('Sample completeness range: %.1f%% - %.1f%%',
min(completeness) * 100, max(completeness) * 100))# Gate 1: rarefaction approaching asymptote
if (min(completeness) < 0.80) {
message('WARNING: Some samples have low completeness (<80%). Deeper sequencing recommended.')
}
# Gate 2: richness. Use the COVERAGE-STANDARDIZED q=0, NOT inext_out$AsyEst 'Species richness' (the
# Chao1 asymptotic): denoising stripped singletons (COUNT >= 2), leaving Chao1 no f1, so it degenerates
# to observed richness -- the collapse rule 3 forbids. estimateD returns Order.q / qD.
q0_c95 <- div_c95[div_c95$Order.q == 0, ]
message(sprintf('Coverage-standardized (C=0.95) richness range: %.0f - %.0f effective species', # qD is fractional; %d errors on doubles
min(q0_c95$qD), max(q0_c95$qD)))library(vegan)
library(indicspecies)
otu_matrix <- as(otu_table(ps_clean), 'matrix')
env_data <- as(sample_data(ps_clean), 'data.frame')
# Hellinger transformation: standard for community composition data
# Reduces influence of dominant species
otu_hell <- decostand(otu_matrix, method = 'hellinger')
# DCA on untransformed data to determine gradient length
dca <- decorana(otu_matrix)
gradient_length <- diff(range(scores(dca, display = 'sites', choices = 1)))
message(sprintf('DCA gradient length: %.2f SD', gradient_length))
# RDA: linear response (<=3 SD), uses Hellinger-transformed data
# CCA: unimodal response (>3 SD), uses raw abundances (chi-squared distance)
if (gradient_length <= 3) {
ord <- rda(otu_hell ~ temperature + depth + season, data = env_data)
method_name <- 'RDA'
} else {
ord <- cca(otu_matrix ~ temperature + depth + season, data = env_data)
method_name <- 'CCA'
}
# Permutation test for significance
# permutations 999: standard; increase to 9999 for publication
anova_result <- anova.cca(ord, permutations = 999)
message(sprintf('%s significance: p = %.4f', method_name, anova_result$`Pr(>F)`[1]))
# MANDATORY companion: PERMANOVA + PERMDISP (Anderson & Walsh 2013 Ecol Monogr 83:557-574)
# adonis2 tests centroid difference; betadisper tests dispersion homogeneity
# If betadisper is also significant, PERMANOVA significance is dispersion-confounded
bray_dist <- vegdist(otu_matrix, method = 'bray')
permanova <- adonis2(bray_dist ~ site, data = env_data,
by = 'margin', permutations = 999)
disp <- betadisper(bray_dist, env_data$site)
disp_test <- permutest(disp, permutations = 999)
message(sprintf('PERMANOVA p = %.4f; PERMDISP p = %.4f',
permanova[['Pr(>F)']][1], disp_test$tab[['Pr(>F)']][1]))
if (permanova[['Pr(>F)']][1] < 0.05 && disp_test$tab[['Pr(>F)']][1] < 0.05) {
message('WARNING: Both PERMANOVA and PERMDISP significant; location-vs-dispersion confounded')
}
# Indicator species analysis with group-size equalization (NOT basic IndVal)
# func='IndVal.g' corrects for unbalanced group sizes (De Caceres & Legendre 2009)
indval <- multipatt(otu_matrix, env_data$site, func = 'IndVal.g',
control = how(nperm = 999))
summary(indval, alpha = 0.05)| Step | Parameter | Recommendation |
|---|---|---|
| Cutadapt | --discard-untrimmed | Always use; removes off-target reads. MANDATORY before DADA2 filterAndTrim |
| Cutadapt | --minimum-length | 50 (general); adjust per expected amplicon size |
| DADA2 | truncLen | Set from quality profiles; marker-dependent (typical 2x250 COI: c(220,180)) |
| DADA2 | maxEE | c(2,2) standard; c(5,5) for degraded eDNA |
| DADA2 | minOverlap | 20 (standard); increase for short overlaps |
| DADA2 | chimera method | 'consensus' standard (conservative); 'pooled' more aggressive |
| OBITools3 v3 | command syntax | obi stats (plural; was obistat in v1); .tar.gz taxonomy |
| OBITools3 | --min-count | 2 (removes singletons); 5-10 for noisy datasets |
| decontam | method | 'combined' if concentration AND controls; 'prevalence' fallback |
| decontam | threshold | 0.1 default; 0.05 for low-biomass samples |
| Tag-jumping MiSeq | threshold | 0.001-0.005 fraction of ASV total |
| Tag-jumping NovaSeq | threshold | 0.005-0.01 (~10x MiSeq; patterned flow cells) |
| Taxonomy | minBoot | 50 (sensitive); 80 (conservative) |
| ecotag | --minimum-identity (-m) | 0.97 (COI species); 0.95 (genus); marker-dependent |
| iNEXT | endpoint | per-sample 2x reference size (default doubling rule); do not set a global 2*max |
| vegan | permutations | 999 (standard); 9999 (publication) |
| Symptom | Cause | Fix |
|---|---|---|
| Over-called rare-species presence (worst on NovaSeq) | Tag-jump rate never quantified ("we used dual indexing" and stop) | Report the residual rate (reads in should-be-zero cells / total) and platform-filter |
| Real low-biomass taxa deleted | decontam output treated as ground truth | Biological-plausibility review of every flag; retain only where statistics AND plausibility agree |
| Chao1 collapses to observed richness | Chao1 computed after aggressive denoising (singletons stripped) | Use incidence-based Chao2, or estimators only on data with a real singleton/doubleton distribution |
| A "community shift" that is really dispersion | PERMANOVA reported without PERMDISP | Always pair adonis2 with betadisper/permutest; if dispersion is significant, the location conclusion is not supported |
| Read counts interpreted as abundance | eDNA reads treated as biomass | Presence/occupancy or effective-species-count framing only (Lamb 2019) |
| Few reads after primer removal | Wrong primer sequences or orientation | Verify primer sequences; try --revcomp |
| High chimera rate (>20%) | Excessive PCR cycles or low-quality input | Reduce PCR cycles; improve DNA extraction |
| Many unassigned ASVs | Incomplete reference database | Use marker-specific database; lower minBoot |
| Contamination in negatives | Tag-jumping or lab contamination | Apply tag-jump filter; review extraction protocol |
| Low sample completeness | Insufficient sequencing depth | Increase sequencing; pool fewer samples |
| Ordination axes not significant | Weak environmental gradients | Add more environmental variables; check sample size |
| Unexpected taxa (e.g., human) | Sample contamination | Filter known contaminants; review field protocols |
| Very few ASVs | Over-aggressive filtering | Relax truncLen, maxEE, or min-count thresholds |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in workflows/edna-pipeline of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Workflows Edna Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Workflows Edna Pipeline this skillGPTomics/bioSkills | 1.2k | 1 repos | ~6.2k | Automated safety check: Pass | MIT | |
| Feishu Docopenclaw/openclaw | 392k | — | ~516 | Automated safety check: Pass | MIT | |
| She Love Me863401402/she-love-me | 931 | 1 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Feishu Docraucvr/Group-Goki | 112 | 3 repos | ~592 | Automated safety check: Pass | MIT | |
| Wechat Article Extractorfreestylefly/wechat-article-extractor-skill | 136 | 1 repos | ~1k | Automated safety check: Pass | None | |
| Wechat Miniprogram Builderchenjin-cmd/wechat-miniprogram-builder | 356 | — | ~634 | Automated safety check: Pass | MIT |
openclaw/openclaw
Feishu document read/write workflows. An agent skill from openclaw/openclaw.
863401402/she-love-me
Acquire, import, and analyze WeChat or QQ chat histories, including installing supported exporters, guiding required login or contact selection, converting exports, assessing relationship dynamics…
raucvr/Group-Goki
Feishu document read/write operations. An agent skill from raucvr/Group-Goki.
freestylefly/wechat-article-extractor-skill
Extract metadata and content from WeChat Official Account articles.
chenjin-cmd/wechat-miniprogram-builder
This skill should be used when the user wants to build, launch, monetize, or promote a WeChat mini-program with AI (vibe coding) — including topic selection, account registration & ICP filing…
huangruiteng/CS-Notes
A skill your agent uses when you need to control Slack from Clawdbot via the slack tool, including reacting to messages or pinning/unpinning items in Slack channels or DMs.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
End-to-end eDNA metabarcoding from raw amplicons to community ecology. Bio Workflows Edna Pipeline is an agent skill from GPTomics/bioSkills. End-to-end eDNA metabarcoding from raw amplicons to community ecology.
Bio Workflows Edna Pipeline fits situations like: processing eDNA samples for biodiversity assessment; deciding ASV vs OTU; configuring OBITools3 v3; interpreting decontam screening.
Run `npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a claude-code`. Or copy the skill folder (workflows/edna-pipeline in GPTomics/bioSkills) into .claude/skills/bio-workflows-edna-pipeline in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a codex`. Or copy the skill folder (workflows/edna-pipeline in GPTomics/bioSkills) into .agents/skills/bio-workflows-edna-pipeline in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-workflows-edna-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-workflows-edna-pipeline, .gemini/skills/bio-workflows-edna-pipeline, .github/skills/bio-workflows-edna-pipeline and .opencode/skills/bio-workflows-edna-pipeline in your project.
Going by SKILL.md and its folder, Bio Workflows Edna Pipeline needs R and a shell for the scripts in its folder. Our summary lists: A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Workflows Edna Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Workflows Edna Pipeline: Feishu Doc (openclaw/openclaw, 392k stars), She Love Me (863401402/she-love-me, 931 stars), Feishu Doc (raucvr/Group-Goki, 112 stars) and Wechat Article Extractor (freestylefly/wechat-article-extractor-skill, 136 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.