Bio Variant Calling Structural Variant Calling
FreedomIntelligence/OpenClaw-Medical-Skills
Call structural variants (SVs) from short-read sequencing using Manta, Delly, and LUMPY.
Chains a somatic (tumor-normal) SNV/indel and structural-variant pipeline end to end with GATK Mutect2 (or Strelka2), wiring the somatic-specific machinery - panel-of-normals and gnomAD…
$ npx skills add GPTomics/bioSkills --skill bio-workflows-somatic-variant-pipeline -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-somatic-variant-pipeline --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/workflows/somatic-variant-pipeline .claude/skills/bio-workflows-somatic-variant-pipeline && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-workflows-somatic-variant-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/somatic-variant-pipeline into .claude/skills/bio-workflows-somatic-variant-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-somatic-variant-pipeline", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/workflows/somatic-variant-pipelineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-somatic-variant-pipeline -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-somatic-variant-pipeline --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/workflows/somatic-variant-pipeline .agents/skills/bio-workflows-somatic-variant-pipeline && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-workflows-somatic-variant-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/somatic-variant-pipeline into .agents/skills/bio-workflows-somatic-variant-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-somatic-variant-pipeline", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-somatic-variant-pipeline -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-somatic-variant-pipeline --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/workflows/somatic-variant-pipeline .cursor/skills/bio-workflows-somatic-variant-pipeline && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-workflows-somatic-variant-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/somatic-variant-pipeline into .cursor/skills/bio-workflows-somatic-variant-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-somatic-variant-pipeline", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path workflows/somatic-variant-pipeline--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-somatic-variant-pipeline -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-somatic-variant-pipeline --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/workflows/somatic-variant-pipeline .gemini/skills/bio-workflows-somatic-variant-pipeline && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-workflows-somatic-variant-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/somatic-variant-pipeline into .gemini/skills/bio-workflows-somatic-variant-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-somatic-variant-pipeline", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-workflows-somatic-variant-pipelineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-workflows-somatic-variant-pipeline -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/workflows/somatic-variant-pipeline .github/skills/bio-workflows-somatic-variant-pipeline && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-workflows-somatic-variant-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/somatic-variant-pipeline into .github/skills/bio-workflows-somatic-variant-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-somatic-variant-pipeline", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-workflows-somatic-variant-pipeline -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-workflows-somatic-variant-pipeline --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/workflows/somatic-variant-pipeline .opencode/skills/bio-workflows-somatic-variant-pipeline && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-workflows-somatic-variant-pipeline" agent skill from https://github.com/GPTomics/bioSkills/tree/main/workflows/somatic-variant-pipeline into .opencode/skills/bio-workflows-somatic-variant-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-workflows-somatic-variant-pipeline", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-workflows-somatic-variant-pipelineChains a somatic (tumor-normal) SNV/indel and structural-variant pipeline end to end with GATK Mutect2 (or Strelka2), wiring the somatic-specific machinery - panel-of-normals and gnomAD…
Bio Workflows Somatic Variant Pipeline is an agent skill from GPTomics/bioSkills. Chains a somatic (tumor-normal) SNV/indel and structural-variant pipeline end to end with GATK Mutect2 (or Strelka2), wiring the somatic-specific machinery - panel-of-normals and gnomAD germline-resource priors, GetPileupSummaries/CalculateContamination, and LearnReadOrientationModel FFPE/oxoG orientation-bias filtering fed into FilterMutectCalls. Use when calling somatic mutations from a tumor-normal pair (or tumor-only with PoN caveats), deciding which artifact filter removes which class of false positive…
Its SKILL.md is about 6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/run_mutect2.sh` and `usage-guide.md`).
The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Workflows Somatic Variant Pipeline loads about 6k tokens when it runs. Until then it costs about 191 tokens; SKILL.md has 2,162 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,162 words, ~6,031 tokens.
.claude/skills/bio-workflows-somatic-variant-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: GATK 4.5+, Strelka2 2.9+, Manta 1.6+, Ensembl VEP 111+, bcftools 1.19+
Before using code patterns, verify installed versions match. If versions differ:
<tool> --version then <tool> --help to confirm flagsIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
Note: this is a WORKFLOW skill - it wires the somatic-specific chain and the DECISIONS between steps. Component mechanism (caller internals, filter thresholds, annotation, CNV) lives in the cross-referenced component skills; interpretation lives in variant-calling/clinical-interpretation and uses the AMP/ASCO/CAP tier system, NOT germline ACMG.
"Call somatic mutations from my tumor-normal pair" -> Orchestrate somatic SNV/indel calling (Mutect2 or Strelka2) with panel-of-normals + germline-resource priors, contamination and orientation-bias filtering, then somatic SV/CNV, then tier/oncogenicity interpretation.
gatk Mutect2 (+ FilterMutectCalls chain), configureStrelkaSomaticWorkflow.py (Strelka2), configManta.py (Manta SV), VEP/Funcotator for annotation.A somatic callset is not a fact about the tumor; it is a joint property of (tumor, matched normal, caller, filters, reference). Three consequences drive every decision in this pipeline:
Tumor BAM + matched Normal BAM (aligned, dedup'd, BQSR - see read-alignment/*)
|
├── SNV/indel calling
│ Mutect2 (GATK) - somatic likelihood, tumor+normal in one command
│ Strelka2 - faster; pair with Manta candidateSmallIndels
│
├── Somatic-specific filtering (the four artifact/germline removers, below)
│ PoN + germline-resource -> Mutect2 call
│ LearnReadOrientationModel (FFPE/oxoG)
│ GetPileupSummaries -> CalculateContamination
│ FilterMutectCalls (consumes all three) -> PASS somatic VCF
│
├── Structural variants -> variant-calling/structural-variant-calling (Manta somatic mode)
├── Copy number + purity/ploidy -> copy-number/cnvkit-analysis (purity feeds VAF reasoning)
├── Annotation (normalize FIRST) -> variant-calling/variant-annotation (VEP/Funcotator)
│
└── Interpretation -> variant-calling/clinical-interpretation
AMP/ASCO/CAP Tier I-IV + oncogenicity (Horak) + TMB/MSI/signaturesThe machinery that separates somatic calling from germline is four independent artifact/germline removers. Knowing WHICH false-positive class each addresses is the core decision - do not treat them as interchangeable boilerplate.
| Filter | Removes | Built from | When it is critical |
|---|---|---|---|
| Panel of Normals (PoN) | Recurrent technical/site artifacts + common germline that reproduce across normals | 40+ unrelated normals from the SAME assay/platform, CreateSomaticPanelOfNormals | Always; the ONLY artifact defense in tumor-only mode |
| Germline resource (gnomAD AF-only) | Germline variants, via a population-AF prior in the somatic model | af-only-gnomad.vcf.gz (AF field only) | Always; carries the germline burden alone in tumor-only mode |
| Contamination estimate | Cross-individual sample contamination masquerading as low-VAF somatic | GetPileupSummaries on common biallelic SNPs -> CalculateContamination | Any sample with suspected swap/contamination; low-VAF calls |
| Orientation-bias model | FFPE deamination (C>T/G>A) and oxoG (C>A/G>T) library artifacts, detected as F1R2/F2R1 strand imbalance | --f1r2-tar-gz from Mutect2 -> LearnReadOrientationModel | FFPE, archival, or oxidatively damaged input; ALWAYS for FFPE |
PoN and germline-resource act at CALL time (priors passed to Mutect2); contamination and orientation-bias are learned separately and injected at FILTER time (FilterMutectCalls). The tumor-only trap: without a matched normal, the PoN and germline resource are the ONLY things removing germline and artifacts, so both must be assay-matched and current, and the false-positive rate is materially higher.
The end-to-end chained script is in examples/run_mutect2.sh; the steps and their decisions:
# Each normal called in tumor-only mode; --max-mnp-distance 0 is REQUIRED for GenomicsDBImport
for normal in normal1.bam normal2.bam normal3.bam; do
s=$(basename "$normal" .bam)
gatk Mutect2 -R reference.fa -I "$normal" --max-mnp-distance 0 -O "${s}.vcf.gz"
done
gatk GenomicsDBImport -R reference.fa --genomicsdb-workspace-path pon_db \
-V normal1.vcf.gz -V normal2.vcf.gz -V normal3.vcf.gz -L intervals.bed
gatk CreateSomaticPanelOfNormals -R reference.fa -V gendb://pon_db -O pon.vcf.gzA PoN needs 40+ normals from the same platform/chemistry to capture recurrent artifacts; a PoN from a different assay imports the wrong artifact profile and misses real ones. Do NOT build a PoN from tumor-adjacent normals if they may carry tumor-in-normal contamination.
gatk Mutect2 -R reference.fa \
-I tumor.bam -I normal.bam -normal normal_sample_name \
--germline-resource af-only-gnomad.vcf.gz \
--panel-of-normals pon.vcf.gz \
--f1r2-tar-gz f1r2.tar.gz \
-O unfiltered.vcf.gz-normal takes the normal read-group SM name (not the filename). --f1r2-tar-gz collects the read-orientation counts needed in Step 3 - omit it and orientation-bias filtering is impossible. Mutect2 also writes unfiltered.vcf.gz.stats, which FilterMutectCalls reads automatically.
gatk LearnReadOrientationModel -I f1r2.tar.gz -O read-orientation-model.tar.gzModels the strand-orientation artifacts (oxoG C>A/G>T from oxidative shearing damage; FFPE cytosine-deamination C>T/G>A). These masquerade as low-VAF somatic SNVs; the model lets FilterMutectCalls down-weight them by their F1R2/F2R1 imbalance.
gatk GetPileupSummaries -I tumor.bam \
-V small_exac_common_3.vcf.gz -L small_exac_common_3.vcf.gz -O tumor_pileups.table
gatk GetPileupSummaries -I normal.bam \
-V small_exac_common_3.vcf.gz -L small_exac_common_3.vcf.gz -O normal_pileups.table
gatk CalculateContamination -I tumor_pileups.table -matched normal_pileups.table \
-O contamination.table --tumor-segmentation segments.tableGetPileupSummaries uses a COMMON biallelic-SNP sites resource (e.g. small_exac_common_3.vcf.gz), not the af-only-gnomAD used for the germline prior - it needs sites with a known population AF where reference/alt read counts reveal foreign DNA. --tumor-segmentation also captures allelic-copy segments that feed the filter.
gatk FilterMutectCalls -R reference.fa -V unfiltered.vcf.gz \
--contamination-table contamination.table \
--tumor-segmentation segments.table \
--ob-priors read-orientation-model.tar.gz \
-O filtered.vcf.gz
bcftools view -f PASS filtered.vcf.gz -Oz -o somatic_final.vcf.gz
bcftools index -t somatic_final.vcf.gzFilterMutectCalls applies a single joint model (contamination + orientation + segmentation + the built-in weak-evidence/germline/strand filters) and sets a per-variant FILTER. Never hand-tune individual thresholds first - the filter is calibrated to balance them together.
When no matched normal exists (archival FFPE, cell lines, legacy cohorts):
gatk Mutect2 -R reference.fa -I tumor.bam \
--germline-resource af-only-gnomad.vcf.gz \
--panel-of-normals pon.vcf.gz \
-O tumor_only.vcf.gzDecision framing: with no normal, every germline variant is a candidate somatic call, and only the PoN (artifacts) and gnomAD prior (germline) remove them - so both must be assay-matched and the false-positive rate rises sharply. High-VAF (~50% or ~100%) calls are especially suspect for germline. Tumor-only cannot cleanly separate somatic from germline, which creates a disclosure problem (a germline pathogenic finding surfaced as "somatic") - see the interpretation section. Paired tumor-normal is the defensible design; tumor-only requires explicit germline-subtraction logic and reporting caveats.
Somatic VAF is not a genotype - it is (clonal_fraction x mutation_copies) / local_total_copies, scaled by tumor purity. Reasoning consequences:
# Run Manta first; its small-indel candidates sharpen Strelka2 indel calls
configManta.py --normalBam normal.bam --tumorBam tumor.bam \
--referenceFasta reference.fa --runDir manta_run
manta_run/runWorkflow.py -m local -j 16
configureStrelkaSomaticWorkflow.py \
--normalBam normal.bam --tumorBam tumor.bam --referenceFasta reference.fa \
--indelCandidates manta_run/results/variants/candidateSmallIndels.vcf.gz \
--runDir strelka_run
strelka_run/runWorkflow.py -m local -j 16
bcftools concat \
strelka_run/results/variants/somatic.snvs.vcf.gz \
strelka_run/results/variants/somatic.indels.vcf.gz \
-a -Oz -o strelka_somatic.vcf.gzStrelka2 is faster and strong on indels (Kim 2018 Nat Methods 15:591); passing Manta's candidateSmallIndels.vcf.gz via --indelCandidates is the documented coupling that improves Strelka2 indel recall.
Given the Alioto/DREAM discordance, running independent callers and requiring agreement raises precision - the concordant core is the reproducible callset:
# Intersect PASS calls from callers with uncorrelated error modes (2/3 agreement)
bcftools isec -n+2 -p consensus_dir \
mutect2_pass.vcf.gz strelka2_pass.vcf.gz muse_pass.vcf.gzStrict all-agree intersection sacrifices too much recall; union admits too many false positives; majority voting (e.g. 2 of 3) balances the two. Normalize every caller's VCF to the same representation first (bcftools norm -f ref.fa -m-) or the intersection undercounts because an indel left-aligned differently in two callers will not match. Consensus is a precision tool, not a substitute for orthogonal validation on the assay.
A complete somatic profile is more than SNVs/indels - hand each off to its component skill:
somaticSV.vcf.gz. For complex rearrangements/fusions use GRIDSS -> GRIPSS -> LINX. See variant-calling/structural-variant-calling.bcftools norm -f reference.fa -m- somatic_final.vcf.gz -Oz -o somatic_norm.vcf.gz # left-align + split multiallelics
gatk Funcotator -R reference.fa -V somatic_norm.vcf.gz -O annotated.vcf.gz \
--output-file-format VCF --data-sources-path funcotator_dataSources.v1.7 --ref-version hg38
vep -i somatic_norm.vcf.gz -o annotated_vep.vcf --vcf --cache --offline \
--assembly GRCh38 --everything \
--custom cosmic.vcf.gz,COSMIC,vcf,exact,0,CNT --fork 4Normalize BEFORE annotating (annotate-then-normalize is an order error): un-normalized records attach annotations to a non-canonical representation and fail to match COSMIC/gnomAD. Full annotation mechanics (transcript choice, MANE, HGVS 3'-shift) live in variant-calling/variant-annotation.
This is the critical decision of the whole pipeline and the most common category error. Route the annotated somatic VCF to variant-calling/clinical-interpretation, which applies TWO orthogonal cancer frameworks:
AMP/ASCO/CAP four-tier clinical actionability (Li 2017 J Mol Diagn 19:4-23):
| Tier | Meaning | Example |
|---|---|---|
| I | Strong clinical significance - FDA-approved therapy or in professional guidelines for THIS tumor type | BRAF V600E in melanoma |
| II | Potential significance - therapy in a different tumor type, or clinical-trial/multi-study evidence | same variant in a non-approved tumor type |
| III | Unknown clinical significance (the somatic "VUS") | rare novel missense, no actionability |
| IV | Benign / likely benign - high population frequency, no oncogenic role | common polymorphism |
Tier is tumor-type-specific - the SAME variant can be Tier I in one cancer and Tier II/III in another (no germline-ACMG analog).
Oncogenicity (Horak 2022 Genet Med 24:986): a SEPARATE points-based ClinGen/CGC/VICC axis (Oncogenic / Likely Oncogenic / VUS / Likely Benign / Benign) from hotspot recurrence, functional data, and tumor-type frequency. Oncogenicity (is it a driver) and actionability (is there a drug) are complementary: an oncogenic driver may still be Tier III if no therapy exists.
bcftools query -f '%FILTER\n' filtered.vcf.gz | sort | uniq -c # counts by filter status
# Substitution spectrum: excess C>A/G>T flags residual oxoG; excess C>T/G>A flags FFPE deamination
bcftools query -f '%REF>%ALT\n' somatic_final.vcf.gz | sort | uniq -c
# VAF distribution (a spike near 0.5/1.0 in tumor-only suggests residual germline).
# -s takes the read-group SM name, NOT the filename -- read it from the BAM header, as Step 2 does.
TUMOR_SM=$(samtools view -H tumor.bam | awk '/^@RG/{for(i=1;i<=NF;i++) if($i ~ /^SM:/){sub(/^SM:/,"",$i); print $i; exit}}')
# -s restricts to the tumor sample: a bare [%AF] iterates BOTH samples and concatenates
# them with no separator (0.25 + 0.01 -> "0.250.01"), which awk then silently truncates to 0.25.
bcftools query -s "$TUMOR_SM" -f '[%AF]\n' somatic_final.vcf.gz | awk '{print int($1*100)/100}' | sort -n | uniq -cDo NOT judge somatic SNVs by a fixed germline-like Ti/Tv (~2-3); the somatic spectrum is signature-dependent, and a collapse toward transversion excess signals artifact contamination, not a target value.
| Symptom | Cause | Fix |
|---|---|---|
| Flood of germline variants in output | Missing/wrong germline-resource, or tumor-only without a good PoN | Pass --germline-resource af-only-gnomad.vcf.gz; use an assay-matched 40+ normal PoN |
| Excess C>A/G>T (or C>T/G>A) low-VAF calls | oxoG (or FFPE) artifacts, orientation-bias filter not applied | Emit --f1r2-tar-gz, run LearnReadOrientationModel, pass --ob-priors to FilterMutectCalls |
FilterMutectCalls errors on missing stats | unfiltered.vcf.gz.stats not alongside the VCF | Keep the .stats Mutect2 wrote next to the VCF, or pass --stats |
| GetPileupSummaries gives nonsense contamination | Used af-only-gnomAD instead of a common biallelic-SNP resource | Use small_exac_common_3.vcf.gz (common SNP sites with AF) |
| Consensus intersection drops real shared calls | VCFs not normalized before bcftools isec | bcftools norm -f ref.fa -m- every caller's VCF first |
| Low-purity tumor: expected drivers missing | VAF below detection at that purity/depth | Get purity from copy-number; increase depth; correct VAF to cancer-cell fraction |
| Applied ACMG PVS1/PM2 to a tumor variant | Germline framework used for somatic | Use AMP/ASCO/CAP tiers + oncogenicity via variant-calling/clinical-interpretation |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in workflows/somatic-variant-pipeline of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Workflows Somatic Variant Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Workflows Somatic Variant Pipeline this skillGPTomics/bioSkills | 1.2k | 1 repos | ~6k | Automated safety check: Pass | MIT | |
| Bio Variant Calling Structural Variant CallingFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | — | ~1.7k | Automated safety check: Pass | None | |
| Bio Variant NormalizationFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | — | ~2.1k | Automated safety check: Pass | None | |
| Bio Longread Structural VariantsFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | 1 repos | ~1.4k | Automated safety check: Pass | None | |
| Canonical Chainthedaviddias/Front-End-Checklist | 74k | — | ~417 | Automated safety check: Pass | MIT | |
| Structured Datathedaviddias/Front-End-Checklist | 74k | — | ~420 | Automated safety check: Pass | MIT |
FreedomIntelligence/OpenClaw-Medical-Skills
Call structural variants (SVs) from short-read sequencing using Manta, Delly, and LUMPY.
FreedomIntelligence/OpenClaw-Medical-Skills
Normalize indel representation and split multiallelic variants using bcftools norm.
FreedomIntelligence/OpenClaw-Medical-Skills
Detect structural variants from long-read alignments using Sniffles, cuteSV, and SVIM.
thedaviddias/Front-End-Checklist
A skill your agent uses when auditing metadata, crawlability, structured data, or indexability related to Avoid redirect chains on canonical URLs.
thedaviddias/Front-End-Checklist
A skill your agent uses when auditing metadata, crawlability, structured data, or indexability related to Add structured data markup.
wu-yc/LabClaw
Comprehensive structural variant (SV) analysis skill for clinical genomics.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Chains a somatic (tumor-normal) SNV/indel and structural-variant pipeline end to end with GATK Mutect2 (or Strelka2), wiring the somatic-specific machinery - panel-of-normals and gnomAD…. Bio Workflows Somatic Variant Pipeline is an agent skill from GPTomics/bioSkills. Chains a somatic (tumor-normal) SNV/indel and structural-variant pipeline end to end with GATK Mutect2 (or Strelka2), wiring the somatic-specific machinery - panel-of-normals and gnomAD germline-resource priors, GetPileupSummaries/CalculateContamination, and LearnReadOrientationModel FFPE/oxoG orientation-bias filtering fed into FilterMutectCalls.
Bio Workflows Somatic Variant Pipeline fits situations like: calling somatic mutations from a tumor-normal pair (or tumor-only with PoN caveats); deciding which artifact filter removes which class of false positive; reasoning about VAF/purity/ploidy and clonal-vs-subclonal detection; adding somatic SV/CNV.
Run `npx skills add GPTomics/bioSkills --skill bio-workflows-somatic-variant-pipeline -a claude-code`. Or copy the skill folder (workflows/somatic-variant-pipeline in GPTomics/bioSkills) into .claude/skills/bio-workflows-somatic-variant-pipeline in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-workflows-somatic-variant-pipeline -a codex`. Or copy the skill folder (workflows/somatic-variant-pipeline in GPTomics/bioSkills) into .agents/skills/bio-workflows-somatic-variant-pipeline in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-workflows-somatic-variant-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-workflows-somatic-variant-pipeline, .gemini/skills/bio-workflows-somatic-variant-pipeline, .github/skills/bio-workflows-somatic-variant-pipeline and .opencode/skills/bio-workflows-somatic-variant-pipeline in your project.
Going by SKILL.md and its folder, Bio Workflows Somatic Variant Pipeline needs a shell for the scripts in its folder. Our summary lists: A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Workflows Somatic Variant Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Workflows Somatic Variant Pipeline: Bio Variant Calling Structural Variant Calling (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Bio Variant Normalization (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Bio Longread Structural Variants (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars) and Canonical Chain (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.