Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Orchestrates an end-to-end long-read structural-variant pipeline - basecalling to minimap2 alignment (platform-matched preset) to Sniffles2/cuteSV/pbsv calling to optional assembly-based calling…
$ npx skills add GPTomics/bioSkills --skill bio-workflows-longread-sv-pipeline -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-longread-sv-pipeline --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workflows/longread-sv-pipeline .claude/skills/bio-workflows-longread-sv-pipeline && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-workflows-longread-sv-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/longread-sv-pipeline into .claude/skills/bio-workflows-longread-sv-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-longread-sv-pipeline", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/workflows/longread-sv-pipelineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-longread-sv-pipeline -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-longread-sv-pipeline --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/workflows/longread-sv-pipeline .agents/skills/bio-workflows-longread-sv-pipeline && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-workflows-longread-sv-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/longread-sv-pipeline into .agents/skills/bio-workflows-longread-sv-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-longread-sv-pipeline", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-longread-sv-pipeline -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-longread-sv-pipeline --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/workflows/longread-sv-pipeline .cursor/skills/bio-workflows-longread-sv-pipeline && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-workflows-longread-sv-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/longread-sv-pipeline into .cursor/skills/bio-workflows-longread-sv-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-longread-sv-pipeline", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path workflows/longread-sv-pipeline--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-longread-sv-pipeline -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-longread-sv-pipeline --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/workflows/longread-sv-pipeline .gemini/skills/bio-workflows-longread-sv-pipeline && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-workflows-longread-sv-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/longread-sv-pipeline into .gemini/skills/bio-workflows-longread-sv-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-longread-sv-pipeline", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-workflows-longread-sv-pipelineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-workflows-longread-sv-pipeline -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/workflows/longread-sv-pipeline .github/skills/bio-workflows-longread-sv-pipeline && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-workflows-longread-sv-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/longread-sv-pipeline into .github/skills/bio-workflows-longread-sv-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-longread-sv-pipeline", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-longread-sv-pipeline -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-longread-sv-pipeline --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/workflows/longread-sv-pipeline .opencode/skills/bio-workflows-longread-sv-pipeline && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-workflows-longread-sv-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/longread-sv-pipeline into .opencode/skills/bio-workflows-longread-sv-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-longread-sv-pipeline", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-workflows-longread-sv-pipelineOrchestrates an end-to-end long-read structural-variant pipeline - basecalling to minimap2 alignment (platform-matched preset) to Sniffles2/cuteSV/pbsv calling to optional assembly-based calling…
Bio Workflows Longread Sv Pipeline is an agent skill from GPTomics/bioSkills. Orchestrates an end-to-end long-read structural-variant pipeline - basecalling to minimap2 alignment (platform-matched preset) to Sniffles2/cuteSV/pbsv calling to optional assembly-based calling (dipcall/PAV) to two-step .snf cohort merging to Truvari benchmarking - chaining ONT and PacBio HiFi runs while handing the SV signal mechanism off to the component skills. Use when running a long-read SV workflow from reads to a benchmarked callset, choosing the minimap2 preset and SV caller by platform and goal…
Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/ont_sv_calling.sh` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
Shell commands in SKILL.md call:
makeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Workflows Longread Sv Pipeline loads about 5.4k tokens when it runs. Until then it costs about 225 tokens; SKILL.md has 2,067 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,067 words, ~5,367 tokens.
.claude/skills/bio-workflows-longread-sv-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: minimap2 2.28+, Sniffles 2.2+, cuteSV 2.1+, pbsv 2.9+, dipcall 0.3+, bcftools 1.19+, samtools 1.19+, truvari 4.0+
Before using code patterns, verify installed versions match. If versions differ:
<tool> --version then <tool> --help to confirm flagsUse minimap2 >= 2.28: the lr:hq accurate-read preset was added in 2.27, and 2.28 fixes the 2.27 --MD regression. Supply a reference-matched tandem-repeat BED to the caller - it is the single biggest false-positive lever in repeats. Truvari renamed the alt-sequence-similarity param from --pctsim to --pctseq at v4; confirm against truvari bench --help.
If code throws an error, introspect the installed tool and adapt the example to the actual API rather than retrying.
"Detect structural variants from my long-read sequencing data" -> Chain basecalling, platform-matched minimap2 alignment, an SV caller selected by platform and goal, an optional assembly-based branch, cohort merging, and a parameterized Truvari benchmark - with the SV mechanism delegated to the component skills.
This is a workflow (orchestration) skill: it makes the stage-to-stage decisions and quality gates that connect the component skills. It does NOT re-teach how a caller sees an SV or the VCF representation minefield - that lives in variant-calling/structural-variant-calling and long-read-sequencing/structural-variants.
Run this pipeline instead of a short-read SV workflow for one mechanistic reason: a single long read (PacBio HiFi ~15-25 kb, ONT tens of kb to >Mb ultralong) physically spans the SV and both flanks in one molecule, turning SV detection from an inference-over-fragments problem into near-direct observation. Two consequences decide whether long reads earn their per-sample cost:
The governing pipeline principle: an SV call is a representation artifact of choices made upstream. The aligner preset, the tandem-repeat BED handed to the caller, and the Truvari matching parameters move precision/recall as much as the caller does. Chaining decisions - not the caller name - are what this skill is about.
POD5/FAST5 (ONT only)
| [Step 0] Dorado basecall -> long-read-sequencing/basecalling
v (model + methylation are IRREVERSIBLE choices; sup model for SV)
FASTQ (ONT / PacBio HiFi)
| [QC] NanoPlot / NanoComp -> long-read-sequencing/long-read-qc
v gate: read N50 >10 kb, sane quality, chimera screen
[Step 1] minimap2 alignment -> long-read-sequencing/long-read-alignment
| preset by platform/chemistry; -Y keeps breakpoint seq on split reads
v gate: mapping rate >90%, coverage >=15x
[Step 2] SV calling -> long-read-sequencing/structural-variants
| Sniffles2 / cuteSV / pbsv, caller by platform + goal
| + tandem-repeat BED (biggest FP lever) variant-calling/structural-variant-calling
v
[Step 3, optional] assembly-based SV (dipcall / PAV against a phased diploid assembly)
| highest-quality callset; the way truth sets are built
v
[Step 4] cohort merge -> two-step Sniffles2 .snf (per-sample -> combine)
|
v
[Step 5] benchmark -> Truvari vs GIAB HG002 Tier 1 + CMRG
an F1 is meaningless without refdist/pctsize/pctseqThe platform sets the preset, the caller options, and what bonus channels are available. Decide before basecalling.
| Dimension | ONT (R10.4.1) | PacBio HiFi |
|---|---|---|
| Per-base accuracy | ~Q20+ simplex, higher duplex | ~Q30+ (circular consensus) |
| Read length | tens of kb; ultralong >Mb achievable | ~15-25 kb |
| minimap2 preset | lr:hq (R10/Q20 accurate) or map-ont (older R9) | map-hifi |
| Best for | ultralong spans, repeat/centromere traversal, native methylation | highest base accuracy, small variants + SV in one run |
| SV caller | Sniffles2 or cuteSV | Sniffles2, cuteSV, or pbsv (official, TR-aware) |
| Bonus channel | 5mCG/6mA methylation if requested AT basecall time | 5mCG via kinetics; phasing native from HiFi length |
R10.4 chemistry moved ONT simplex to ~Q20, which is why lr:hq (not the noisy-read map-ont) is the right preset for modern ONT - it rewrites the scoring/chaining model for accurate reads and runs faster at equal accuracy. Older R9 data still needs map-ont.
Do not re-derive the caller mechanism here; pick by goal and hand tuning to the component skill.
| Goal / platform | Caller | Why / cross-reference |
|---|---|---|
| ONT/HiFi germline, cohorts, mosaic | Sniffles2 | field standard; two-step .snf population merge scales linearly in N; --mosaic for low-VAF (Smolka 2024) |
| Highest recall on noisy ONT | cuteSV | signature clustering; MUST pass the per-platform param set and --genotype (Jiang 2020) |
| PacBio HiFi, official, TR-aware | pbsv | expects pbmm2 alignments; single- and joint-sample modes |
| tandem-vs-interspersed DUP detail | SVIM | reports origin AND destination of duplications (Heller 2019) |
| Highest-quality callset / truth set | dipcall or PAV | assembly-vs-reference from a phased diploid assembly (Li 2018; Ebert 2021) |
| Somatic (tumor-normal) | Severus / nanomonsv | matched-normal subtraction; do NOT use Sniffles --mosaic for somatic (Keskus 2025) |
Methods evolve; verify current best practice against each tool's docs before committing. Deeper caller tuning (cuteSV per-platform params, the tandem-repeat BED, aligner effects) lives in long-read-sequencing/structural-variants.
Goal: Convert raw POD5/FAST5 signal into reads suitable for SV calling, capturing methylation if it will ever be needed.
Approach: Basecall with Dorado using a chemistry-matched sup (super-accuracy) model; request modified bases at basecall time because methylation cannot be recovered later. PacBio HiFi arrives as reads already, so this step is skipped.
# sup model maximizes accuracy for SV; 5mCG_5hmCG requested now (irreversible if omitted).
# See long-read-sequencing/basecalling for model selection and duplex.
dorado basecaller sup pod5_dir/ --modified-bases 5mCG_5hmCG > reads.bam
samtools fastq -T MM,ML reads.bam | gzip > reads.fastq.gz # -T carries methylation tags throughGoal: Produce a sorted, indexed BAM whose split (supplementary) alignments retain the breakpoint sequence SV callers reconstruct from.
Approach: Align with the platform-matched minimap2 preset; keep -Y so supplementary alignments are soft-clipped (not hard-clipped), which preserves the junction bases on split reads. SV calling rides on supplementary, not secondary, alignments.
# ONT R10/Q20: lr:hq (accurate reads). Older R9: map-ont. HiFi: map-hifi. PacBio CLR: map-pb.
minimap2 -ax lr:hq -t 16 --MD -Y reference.fa reads.fastq.gz | \
samtools sort -@ 4 -o aligned.bam
samtools index aligned.bamQC checkpoint (gate before spending compute on calling):
samtools flagstat aligned.bam # mapping rate should be >90%
samtools depth -a aligned.bam | awk '{s+=$3} END{print "mean cov:", s/NR}'
# Gate: >=15x for confident SV calling; below ~10x callers drift toward false negatives.Goal: Call SVs (>=50 bp DEL/INS/DUP/INV/BND) from the aligned reads with a caller matched to the platform.
Approach: Run Sniffles2 (the default) with a reference-matched tandem-repeat BED - it clusters the repeat-driven false positives that otherwise dominate the callset. --minsvlen 50 enforces the GIAB >=50 bp SV convention (Sniffles2 defaults to 35).
# Sniffles2: the tandem-repeat BED is the single biggest false-positive lever in repeats.
sniffles --input aligned.bam --reference reference.fa \
--tandem-repeats human_GRCh38_TR.bed \
--vcf svs.vcf.gz --threads 8 --minsvlen 50 --output-rnamescuteSV as an alternative - its defaults are NOT platform-appropriate, and --genotype is off by default:
# ONT param set shown. HiFi: 1000/0.9/1000/0.5. CLR: 100/0.3/200/0.5. See structural-variants.
mkdir -p work_dir # cuteSV requires the work dir to pre-exist; it does not create it
cuteSV aligned.bam reference.fa svs.vcf work_dir/ --threads 8 --genotype \
--max_cluster_bias_INS 100 --diff_ratio_merging_INS 0.3 \
--max_cluster_bias_DEL 100 --diff_ratio_merging_DEL 0.3Goal: Produce the highest-quality SV callset by comparing a phased diploid assembly to the reference, rather than inferring from read alignments.
Approach: Assemble the genome (hifiasm/verkko), then call variants from the two haplotype assemblies aligned to the reference. dipcall (the syndip method) and PAV (the HGSVC method) are the assembly-vs-reference callers; this is how the GIAB and HGSVC truth sets themselves are built. Use it when an assembly already exists or when callset quality outranks turnaround.
# dipcall needs two haplotype assemblies (hap1/hap2) plus minimap2/k8/htsbox on PATH.
run-dipcall prefix reference.fa hap1.fa hap2.fa > prefix.mak
make -j2 -f prefix.mak # emits prefix.dip.vcf.gz + prefix.dip.bedGoal: Build a joint-genotyped multi-sample SV matrix, not a union of per-sample discovery VCFs.
Approach: Sniffles2's population design processes each sample independently into a compact .snf, then combines the .snf files in a second pass - scaling linearly in N. A union of per-sample discovery VCFs is wrong: a sample recorded 0/0 may simply not have had that event discovered in it (a false missing), which corrupts allele frequencies. The two-step .snf combine force-genotypes every sample at every merged site.
# Pass 1: per-sample .snf (each sample processed once, independently).
for s in sample1 sample2 sample3; do
sniffles --input ${s}.bam --reference reference.fa \
--tandem-repeats human_GRCh38_TR.bed --snf ${s}.snf
done
# Pass 2: combine into a jointly genotyped cohort VCF (linear in N).
sniffles --input sample1.snf sample2.snf sample3.snf --vcf cohort.vcf.gzFor sequence-aware AF work across callsets, prefer Truvari collapse over position-only merging (position-only mergers inflate allele frequency by up to 2.2x; English 2022) - see variant-calling/structural-variant-calling for the merger decision table.
Goal: Report a defensible, reproducible accuracy figure - not a number that looks good because of loose matching.
Approach: Truvari bench counts a call as a true positive only if it matches a truth variant under ALL of --refdist, --pctsize, and --pctseq simultaneously. Every one of these moves the score, so a bare F1 is uninterpretable. Report the full parameter set, run truvari refine for a harmonized re-comparison, and stratify by region.
# Report EVERY parameter. --pctseq 0 disables alt-sequence checking and quietly inflates INS scores.
truvari bench -b HG002_SV_Tier1.vcf.gz -c svs.vcf.gz -o bench_tier1/ --passonly \
-f reference.fa --refdist 500 --pctsize 0.70 --pctseq 0.70 --sizemin 50 # -f persists reference to params.json
truvari refine bench_tier1/ # harmonized breakpoint re-comparison (needs the reference bench recorded)
# CMRG is NOT optional: Tier 1 EXCLUDES the medically relevant repetitive genes.
truvari bench -b HG002_CMRG_SV.vcf.gz -c svs.vcf.gz -o bench_cmrg/ --passonly \
--refdist 500 --pctsize 0.70 --pctseq 0.70 --sizemin 50Three escalating bars are routinely conflated, and a pipeline can pass the first while failing the ones that matter:
--refdist, --pctseq on). Matters at exon/splice boundaries.Region stratification is decisive: Tier 1 (Zook 2020 Nat Biotechnol 38:1347) is conservative isolated SVs, while CMRG (Wagner 2022 Nat Biotechnol 40:672) covers the repetitive medically relevant genes Tier 1 leaves out - where GRCh38 false duplications cause reference-specific misses that masking raised from 8% to 100% recall. A good Tier 1 F1 certifies nothing about the genes clinicians care about; run both.
bcftools view -i 'QUAL>=20 && ABS(SVLEN)>=50' svs.vcf.gz -Oz -o svs.filtered.vcf.gz
bcftools index svs.filtered.vcf.gz # ABS() is mandatory: DEL SVLEN is negative by convention
bcftools stats svs.filtered.vcf.gz > sv_stats.txt
AnnotSV -SVinputFile svs.filtered.vcf.gz -genomeBuild GRCh38 -outputFile annotated_svs
# gene overlap, DGV/gnomAD-SV population AF, ClinVar pathogenicityHeterozygous variants on the same long read are physically phased, so SVs can be assigned to haplotypes with no statistical phasing. Haplotag the BAM (whatshap/sniffles --phase) before or during calling to get haplotype-resolved SVs; see long-read-sequencing/haplotype-phasing. If methylation was requested at basecall time (Step 0), the MM/ML tags ride through alignment (via minimap2 -y / samtools fastq -T) and give a per-haplotype methylation channel alongside the SV call at no extra sequencing cost - useful for imprinting and allele-specific silencing, but it must be captured at basecall time or it is gone.
| Type | ALT | Notes for long reads |
|---|---|---|
| Deletion | DEL | excellent recall; breakpoints base-precise when a read spans the junction |
| Insertion | INS | the reason to use long reads; the read carries the inserted sequence |
| Duplication | DUP | tandem vs interspersed distinguishable (SVIM reports origin + destination) |
| Inversion | INV | resolved when unique anchors flank the repeat-embedded breakpoints |
| Translocation | BND | paired breakend records linked by MATEID; complex events are BND graphs |
| Symptom | Cause | Fix |
|---|---|---|
| Few SVs / missing known INS | coverage <10x or missing tandem-repeat BED | raise depth to >=15x; pass --tandem-repeats |
| Many false positives in repeats | no tandem-repeat BED supplied | provide a reference-matched TR BED (biggest FP lever) |
map-ont on R10 data is slow/less accurate | wrong preset for accurate reads | use lr:hq for R10/Q20 ONT; map-ont only for R9 |
| Split reads lost breakpoint sequence | aligned without -Y (hard-clipped supplementaries) | re-align with -Y |
| Methylation channel gone | not requested at basecall time | rebasecall with --modified-bases; it is irreversible |
| cuteSV recall poor / no genotypes | ran defaults; --genotype off | pass the per-platform param set and --genotype |
| Cohort "0/0" wrong, AF too low | took a union of per-sample discovery VCFs | use the two-step .snf combine (force-genotypes all sites) |
ABS(SVLEN)>=50 filter drops all deletions | filtered raw SVLEN (DEL is negative) | always wrap in ABS() |
| Truvari F1 not reproducible / suspiciously high | reported without params, or --pctseq 0 | state refdist/pctsize/pctseq/sizemin; never disable pctseq to look good |
| Passed Tier 1 but clinical genes fail | benchmarked only on Tier 1 | also run CMRG (Tier 1 excludes those genes) |
-Y soft-clipping, MM/ML tag passthrough<DEL>/<INS> alleles are not directly consensus-able© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in workflows/longread-sv-pipeline of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Workflows Longread Sv Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Workflows Longread Sv Pipeline this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.4k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Orchestrates an end-to-end long-read structural-variant pipeline - basecalling to minimap2 alignment (platform-matched preset) to Sniffles2/cuteSV/pbsv calling to optional assembly-based calling…. Bio Workflows Longread Sv Pipeline is an agent skill from GPTomics/bioSkills.snf cohort merging to Truvari benchmarking - chaining ONT and PacBio HiFi runs while handing the SV signal mechanism off to the component skills.
Bio Workflows Longread Sv Pipeline fits situations like: running a long-read SV workflow from reads to a benchmarked callset; choosing the minimap2 preset and SV caller by platform and goal; deciding when long reads are worth it for the insertions and repeat-mediated SVs short reads physically miss; building a joint-genotyped cohort with the two-step .snf design.
Run `npx skills add GPTomics/bioSkills --skill bio-workflows-longread-sv-pipeline -a claude-code`. Or copy the skill folder (workflows/longread-sv-pipeline in GPTomics/bioSkills) into .claude/skills/bio-workflows-longread-sv-pipeline in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-workflows-longread-sv-pipeline -a codex`. Or copy the skill folder (workflows/longread-sv-pipeline in GPTomics/bioSkills) into .agents/skills/bio-workflows-longread-sv-pipeline in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-workflows-longread-sv-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-workflows-longread-sv-pipeline, .gemini/skills/bio-workflows-longread-sv-pipeline, .github/skills/bio-workflows-longread-sv-pipeline and .opencode/skills/bio-workflows-longread-sv-pipeline in your project.
Going by SKILL.md and its folder, Bio Workflows Longread Sv Pipeline needs a shell for the scripts in its folder and the command-line tools its instructions call (make). Our summary lists: A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Workflows Longread Sv Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.4k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Workflows Longread Sv Pipeline: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.