Agent skill

Bio Copy Number Gatk Cnv

by GPTomics in GPTomics/bioSkills

Call copy number variants with the GATK best-practices workflows — the somatic CNV pipeline (CollectReadCounts, DenoiseReadCounts with tangent normalization, ModelSegments, CallCopyRatioSegments)…

MITAuto-check passedDatabases

Install Bio Copy Number Gatk Cnv

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-copy-number-gatk-cnv -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-copy-number-gatk-cnv --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/copy-number/gatk-cnv .claude/skills/bio-copy-number-gatk-cnv && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-copy-number-gatk-cnv
GitHub stars
1.2k
Used in
2 other repos
Token cost
~3.9k tokens
SKILL.md length
1,413 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Call copy number variants with the GATK best-practices workflows — the somatic CNV pipeline (CollectReadCounts, DenoiseReadCounts with tangent normalization, ModelSegments, CallCopyRatioSegments)…

  • Integrating CNV calling into a GATK variant pipeline
  • SKILL.md covers Version Compatibility, Critical: What GATK Somatic…, Somatic vs Germline — Choosing… and Decision Tree by Scenario, plus 8 more sections
  • Runs Shell scripts from its folder
  • Calling rare germline CNVs from an exome cohort

What it does

Bio Copy Number Gatk Cnv is an agent skill from GPTomics/bioSkills. Call copy number variants with the GATK best-practices workflows — the somatic CNV pipeline (CollectReadCounts, DenoiseReadCounts with tangent normalization, ModelSegments, CallCopyRatioSegments) and the germline GATK-gCNV pipeline (DetermineGermlineContigPloidy, GermlineCNVCaller cohort/case mode, PostprocessGermlineCNVCalls). Covers panel-of-normals construction, AnnotateIntervals/FilterIntervals, allelic-count integration, and QS-based filtering. Use when integrating CNV calling into a GATK variant pipeline…

Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/run_gatk_cnv.sh` and `usage-guide.md`).

It sits in Databases, covering Database schema design. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Integrating CNV calling into a GATK variant pipeline
  • Calling rare germline CNVs from an exome cohort
  • Deciding between the somatic and germline GATK workflows
  • Diagnosing why tangent normalization removed a real event

Example prompts

  • “/bio-copy-number-gatk-cnv”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Copy Number Gatk Cnv loads about 3.9k tokens when it runs. Until then it costs about 187 tokens; SKILL.md has 1,413 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~187
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,413 words, ~3,857 tokens.

Download SKILL.mdSave it as .claude/skills/bio-copy-number-gatk-cnv/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-copy-number-gatk-cnv
description
Call copy number variants with the GATK best-practices workflows — the somatic CNV pipeline (CollectReadCounts, DenoiseReadCounts with tangent normalization, ModelSegments, CallCopyRatioSegments) and the germline GATK-gCNV pipeline (DetermineGermlineContigPloidy, GermlineCNVCaller cohort/case mode, PostprocessGermlineCNVCalls). Covers panel-of-normals construction, AnnotateIntervals/FilterIntervals, allelic-count integration, and QS-based filtering. Use when integrating CNV calling into a GATK variant pipeline, calling rare germline CNVs from an exome cohort, deciding between the somatic and germline GATK workflows, or diagnosing why tangent normalization removed a real event or why gCNV output has low precision.
tool_type
cli
primary_tool
gatk

Version Compatibility

Reference examples tested with: GATK 4.5+ (gatk4), Python 3.10+ (gcnv conda env), R 4.3+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: gatk --version then gatk <ToolName> --help to confirm arguments
  • gCNV requires a working gatkcondaenv (theano/tensorflow stack) — gatk will report if the Python environment is missing

GATK 4.5+ gCNV inference defaults are tuned for whole-exome data; whole-genome runs generally need parameter changes. If a tool reports an unrecognized argument, check the help for that exact GATK version rather than retrying.

GATK CNV Workflows

"Call CNVs the GATK way" -> GATK has two separate CNV workflows that share almost no tools. Picking the wrong one is the most common mistake.

  • Somatic CNV: CollectReadCounts -> DenoiseReadCounts -> ModelSegments -> CallCopyRatioSegments. Tumor copy-ratio segments, optionally allele-aware.
  • Germline gCNV: DetermineGermlineContigPloidy -> GermlineCNVCaller -> PostprocessGermlineCNVCalls. Per-sample germline CN genotypes (VCF).

Critical: What GATK Somatic CNV Does NOT Provide

ModelSegments + CallCopyRatioSegments produce copy-ratio segments and a minor-allele fraction per segment, and the "call" is a simple t-test emitting + / - / 0. This is not integer allele-specific copy number, not tumor purity, and not ploidy. Practitioners routinely assume parity with ASCAT/FACETS and there is none. For integer allele-specific CN, purity, ploidy, LOH state, or whole-genome-doubling status, use allele-specific-copy-number (ASCAT, Sequenza, FACETS, or PureCN — PureCN can even reuse the GATK ModelSegments segmentation as input).

Somatic vs Germline — Choosing the Workflow

QuestionSomatic CNVGermline gCNV
InputOne tumor (+ optional matched normal)A cohort of constitutional samples
OutputCopy-ratio segments, +/-/0 call, minor-allele fractionInteger germline CN genotype VCF per sample
NormalizationTangent (projection onto PoN subspace)PCA batching + Bayesian read-depth model
Cohort neededPoN of normals for denoising>= ~100 technically matched samples (cohort mode)
Use forTumor SCNAs, focal amplifications/deletionsRare/de novo germline CNVs, NDD/Mendelian cohorts

Decision Tree by Scenario

ScenarioWorkflowKey parameters
Tumor-normal WGS/WES, want SCNAsSomatic, with matched-normal allelic countsPreprocessIntervals --bin-length 1000 (WGS) or 0 (WES)
Tumor-only somatic CNVSomatic, no matched-normal allelic countsGenotype hets in the case sample; expect more no-calls
Rare germline CNV, exome cohort >= 100gCNV cohort modeRun DetermineGermlineContigPloidy cohort first
New sample vs an existing gCNV modelgCNV case modeMust reuse identical scatter count and interval list
Need integer ASCN / purity / ploidyNeither — escalateUse allele-specific-copy-number
Targeted panel (< few hundred genes)Prefer CNVkitGATK interval models are unstable on tiny panels

Somatic CNV Pipeline

bash
# 1. Preprocess and annotate intervals (WES: bin-length 0 = use exome targets as-is)
gatk PreprocessIntervals -R ref.fa -L targets.interval_list \
    --bin-length 0 --interval-merging-rule OVERLAPPING_ONLY -O preprocessed.interval_list
gatk AnnotateIntervals -R ref.fa -L preprocessed.interval_list \
    --interval-merging-rule OVERLAPPING_ONLY -O annotated.tsv     # GC content for FilterIntervals

# 2. Collect read counts (each BAM)
gatk CollectReadCounts -R ref.fa -I sample.bam -L preprocessed.interval_list \
    --interval-merging-rule OVERLAPPING_ONLY -O sample.counts.hdf5

# 3. Build the panel of normals (tangent-normalization basis)
# --minimum-interval-median-percentile 5.0 is the GATK CNV tutorial value (tool default 10.0)
gatk CreateReadCountPanelOfNormals \
    -I normal1.counts.hdf5 -I normal2.counts.hdf5 -I normalN.counts.hdf5 \
    --annotated-intervals annotated.tsv \
    --minimum-interval-median-percentile 5.0 -O cnv.pon.hdf5

# 4. Denoise tumor against the PoN (tangent normalization)
gatk DenoiseReadCounts -I tumor.counts.hdf5 --count-panel-of-normals cnv.pon.hdf5 \
    --standardized-copy-ratios tumor.standardizedCR.tsv \
    --denoised-copy-ratios tumor.denoisedCR.tsv

# 5. Allelic counts at common biallelic SNPs (tumor and matched normal)
gatk CollectAllelicCounts -R ref.fa -I tumor.bam -L common_snps.interval_list \
    -O tumor.allelicCounts.tsv
gatk CollectAllelicCounts -R ref.fa -I normal.bam -L common_snps.interval_list \
    -O normal.allelicCounts.tsv

# 6. Joint segmentation of copy ratio and allele fraction
gatk ModelSegments --denoised-copy-ratios tumor.denoisedCR.tsv \
    --allelic-counts tumor.allelicCounts.tsv \
    --normal-allelic-counts normal.allelicCounts.tsv \
    --output-prefix tumor -O segments/

# 7. Call each segment +/-/0 (simple t-test against the copy-ratio baseline)
gatk CallCopyRatioSegments -I segments/tumor.cr.seg -O segments/tumor.called.seg

AnnotateIntervals (step 1) and supplying --annotated-intervals to the PoN are frequently skipped — they enable explicit GC-bias correction and are recommended.

Germline gCNV Pipeline

bash
# 1. Determine contig ploidy across the cohort (karyotype + global depth)
gatk DetermineGermlineContigPloidy -L preprocessed.interval_list \
    --interval-merging-rule OVERLAPPING_ONLY \
    -I sample1.counts.hdf5 -I sampleN.counts.hdf5 \
    --contig-ploidy-priors ploidy_priors.tsv --output-prefix cohort -O ploidy-calls/

# 2. FilterIntervals — remove low-mappability / extreme-GC / low-count intervals
gatk FilterIntervals -L preprocessed.interval_list --annotated-intervals annotated.tsv \
    -I sample1.counts.hdf5 -I sampleN.counts.hdf5 \
    --interval-merging-rule OVERLAPPING_ONLY -O filtered.interval_list

# 3. GermlineCNVCaller, cohort mode (builds the model AND calls the cohort)
gatk GermlineCNVCaller --run-mode COHORT -L filtered.interval_list \
    --interval-merging-rule OVERLAPPING_ONLY \
    --contig-ploidy-calls ploidy-calls/cohort-calls \
    -I sample1.counts.hdf5 -I sampleN.counts.hdf5 \
    --output-prefix cohort -O gcnv-calls/

# 4. Post-process per sample into a genotyped VCF
gatk PostprocessGermlineCNVCalls \
    --calls-shard-path gcnv-calls/cohort-calls \
    --model-shard-path gcnv-calls/cohort-model \
    --contig-ploidy-calls ploidy-calls/cohort-calls \
    --sample-index 0 \
    --output-genotyped-intervals sample0.intervals.vcf.gz \
    --output-genotyped-segments sample0.segments.vcf.gz \
    --output-denoised-copy-ratios sample0.denoisedCR.tsv

Case mode (--run-mode CASE) scores a new sample against the cohort *-model shards; it must use the identical filtered.interval_list and the same scatter count as the cohort run, or it fails or produces incomparable calls.

Failure Modes

Tangent normalization removes a real CNV

Trigger: PoN is small (< ~20 normals) or contains samples that share a recurrent CNV (e.g. a common germline CNV, or a PoN accidentally built from tumors).

Mechanism: DenoiseReadCounts projects the tumor coverage profile onto the subspace spanned by the PoN's principal components. Any copy-number pattern present in that subspace is treated as "systematic noise" and subtracted. A CNV shared by PoN members is therefore normalized out of the tumor.

Symptom: A known event (recurrent amplification/deletion, or a common germline CNV) is absent from denoisedCR.tsv; denoised profile is suspiciously flat at that locus.

Fix: Build the PoN from >= 20-40 unrelated, tumor-free, process-matched normals. Never put tumors in the PoN. Cross-check against the standardizedCR.tsv (pre-tangent) profile — if the event is there but gone after denoising, the PoN ate it.

Mistaking ModelSegments output for allele-specific integer CN

Trigger: Treating tumor.modelFinal.seg minor-allele fraction as integer minor copy number, or expecting a purity/ploidy field.

Mechanism: GATK somatic CNV models copy ratio and allele fraction but never fits the purity/ploidy grid that converts log-ratio to integer absolute CN.

Symptom: No purity/ploidy in any output; "copy number" is continuous log2; LOH is a low minor-allele fraction, not an explicit CN-LOH state.

Fix: Accept GATK somatic CNV as a relative caller. For integer ASCN, feed the data to PureCN (segmentationGATK4 reuses GATK segments), FACETS, ASCAT, or Sequenza — see allele-specific-copy-number.

FilterIntervals silently drops intervals containing real variants

Trigger: Aggressive mappability or segmental-duplication cutoffs in FilterIntervals.

Mechanism: A minimum mappability > 0 or a maximum segmental-duplication content < 1 excludes intervals overlapping segdups and low-mappability regions — exactly where many disease-relevant CNVs (e.g. recurrent genomic-disorder loci flanked by segdups) live.

Symptom: Known recurrent CNVs at segdup-mediated loci are never called; the gene of interest has no intervals in filtered.interval_list.

Fix: Inspect filtered.interval_list for genes of interest before calling. Relax mappability/segdup cutoffs for targeted analyses; for genomic-disorder loci, depth-based callers are inherently limited near segdups — confirm with an orthogonal assay.

Raw gCNV output has ~22% precision

Trigger: Using unfiltered GermlineCNVCaller / PostprocessGermlineCNVCalls output for association or de novo analysis.

Mechanism: gCNV is tuned for high recall (~95% of rare coding CNVs >= 2 exons) at the cost of precision; raw calls are dominated by false positives.

Symptom: Implausibly many rare CNVs per sample; de novo CNV rate far above the expected ~0.01-0.02/genome.

Fix: Apply the QS (quality score) filter. QS > 100 is a common starting threshold; QS > 1000 reaches ~96% precision. Also apply sample-level filters (call rate, number of CNVs per sample) per Babadi 2023.

Show full SKILL.md (552 more words)Show less
gCNV cohort too small or mismatched

Trigger: Cohort mode with < ~100 samples, or a cohort spanning multiple capture kits / library protocols.

Mechanism: The Bayesian model needs enough technically similar samples to learn coverage bias; mixed protocols are not separable and the model misattributes batch effects to copy number.

Symptom: Unstable calls; many CNVs tracking sequencing batch; model fails to converge.

Fix: Use >= 100 process-matched samples per cohort model; split heterogeneous cohorts by capture kit. For < 100 samples, gCNV case mode against an external compatible model, or a cohort caller like ExomeDepth, is more appropriate — see germline-cnv-interpretation.

Reconciliation: GATK vs Other Callers

PatternLikely causeAction
GATK somatic flat where CNVkit calls a focal eventTangent normalization absorbed itCheck standardizedCR.tsv; rebuild PoN without the event
GATK and ASCAT disagree on a "deletion"GATK has no purity model; the event is subclonal or impureTrust ASCAT/FACETS integer ASCN
gCNV calls a CNV ExomeDepth missesDifferent sensitivity profiles; both have poor inter-tool concordanceRequire QS filtering + a second caller for rare-CNV claims
gCNV CNV count tracks batchCohort mixes protocolsRe-batch by capture kit

Operational rule: GATK somatic CNV output is relative copy ratio — report it as gain/loss/neutral, not absolute CN, unless downstream-fit by an allele-specific tool. gCNV calls are reportable only after QS and sample-level filtering, and rare-CNV or de novo claims need orthogonal confirmation.

Quantitative Thresholds

ThresholdValueSource / Rationale
PoN size (somatic)>= 20-40 normalsLarger PoN = stabler tangent subspace; small PoN over-fits
gCNV cohort size>= ~100 technically matchedBabadi 2023 Nat Genet; model needs coverage-bias signal
gCNV QS for high precisionQS > 1000 -> ~96% precisionBabadi 2023; raw output ~22% precision, ~95% recall
minimum-interval-median-percentile10.0 default; 5.0 in the GATK CNV tutorialDrops the lowest-coverage intervals from the PoN
WGS bin length~1000 bpPreprocessIntervals --bin-length 1000; WES uses 0 (targets as-is)
Het sites for somatic ModelSegments>= ~10,000 (WGS)Sparse hets give noisy minor-allele-fraction segmentation

Common Errors

Error / symptomCauseSolution
gatkcondaenv / theano error in gCNVgCNV Python env not installedInstall the GATK conda env; gCNV cannot run without it
"At least one interval must remain" in FilterIntervalsAll intervals filtered outRelax mappability/GC/count cutoffs; check annotated.tsv
Case-mode gCNV fails or gives odd callsDifferent scatter count or interval list vs cohort modelReuse the exact cohort scatter count and filtered.interval_list
Denoised profile flat at a known eventTangent normalization removed itRebuild PoN larger, tumor-free; inspect standardizedCR
ModelSegments has few het sitesSNP interval list misses captured regionsUse a common-SNP list intersected with the capture targets
Expecting purity/ploidy in outputSomatic CNV does not estimate themUse allele-specific-copy-number

References

  • GATK Best Practices: Somatic copy number variant discovery (CNV). Broad Institute documentation.
  • Babadi M et al 2023. GATK-gCNV enables the discovery of rare copy number variants from exome sequencing data. Nat Genet 55:1589
  • Gao GF, Oh C, Saksena G, Tabak B, Beroukhim R, Getz G et al 2022. Tangent normalization for somatic copy-number inference in cancer genome analysis. Bioinformatics 38:4677 (Tabak is a middle, not first, author).
  • copy-number/allele-specific-copy-number - Integer ASCN, purity, ploidy (ASCAT/Sequenza/FACETS/PureCN)
  • copy-number/copy-ratio-segmentation - Segmentation algorithms and depth normalization theory
  • copy-number/cnvkit-analysis - Read-depth CNV calling for panels and exomes
  • copy-number/germline-cnv-interpretation - ACMG/ClinGen classification of germline CNV calls
  • copy-number/cnv-visualization - Plotting GATK denoised ratios and modeled segments
  • copy-number/recurrent-cnv - Cohort-level recurrent and driver CNV
  • variant-calling/gatk-variant-calling - GATK SNV/indel pipeline for allelic-count SNP sites

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in copy-number/gatk-cnv of GPTomics/bioSkills.

  • SKILL.md
  • examples/run_gatk_cnv.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Copy Number Gatk Cnv next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Copy Number Gatk Cnv compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Copy Number Gatk Cnv this skillGPTomics/bioSkills1.2k2 repos~3.9kAutomated safety check: PassMIT
SQL Optimization Patternsynulihao/AgentSkillOS61811 repos~3.3kAutomated safety check: PassNone
Datamodellmnimbalyst/nimbalyst1.9k—~713Automated safety check: PassMIT
Add Mpk Taskmirage-project/mirage2.5k—~4.5kAutomated safety check: PassApache-2.0
B200 Flash Attention4 Plannermirage-project/mirage2.5k—~1.9kAutomated safety check: PassApache-2.0
Experiment Auditwanshuiyin/Auto-claude-code-research-in-sleep17k1 repos~2.7kAutomated safety check: NotesMIT

Similar skills

  • SQL Optimization Patterns

    ynulihao/AgentSkillOS

    Master SQL query optimization, indexing strategies, and EXPLAIN analysis to dramatically improve database performance and eliminate slow queries.

    618 GitHub starsUsed in 11 repos~3.3k tokens
    DatabasesAuto-check passed
  • Datamodellm

    nimbalyst/nimbalyst

    Create visual data models for database schemas using Nimbalyst's DataModelLM editor.

    1.9k GitHub stars~713 tokensUpdated today
    DatabasesAuto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated 3 days ago
    DatabasesAuto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated 3 days ago
    DatabasesAuto-check passed
  • Experiment Audit

    wanshuiyin/Auto-claude-code-research-in-sleep

    Audit experiment integrity before claiming results. An agent skill from wanshuiyin/Auto-claude-code-research-in-sleep.

    17k GitHub starsUsed in 1 repo~2.7k tokens
    DatabasesAuto-check: notes
  • Sqlite Schema Design

    fastrepl/anarlog

    Design or review schemas for crates/cloudsync using SQLite Sync constraints, not generic SQLite advice.

    9.5k GitHub stars~1.9k tokensUpdated today
    DatabasesAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Categories

Questions about Bio Copy Number Gatk Cnv

What does Bio Copy Number Gatk Cnv do?

Call copy number variants with the GATK best-practices workflows — the somatic CNV pipeline (CollectReadCounts, DenoiseReadCounts with tangent normalization, ModelSegments, CallCopyRatioSegments)…. Bio Copy Number Gatk Cnv is an agent skill from GPTomics/bioSkills. Call copy number variants with the GATK best-practices workflows — the somatic CNV pipeline (CollectReadCounts, DenoiseReadCounts with tangent normalization, ModelSegments, CallCopyRatioSegments) and the germline GATK-gCNV pipeline (DetermineGermlineContigPloidy, GermlineCNVCaller cohort/case mode, PostprocessGermlineCNVCalls).

When should I use Bio Copy Number Gatk Cnv?

Bio Copy Number Gatk Cnv fits situations like: integrating CNV calling into a GATK variant pipeline; calling rare germline CNVs from an exome cohort; deciding between the somatic and germline GATK workflows; diagnosing why tangent normalization removed a real event.

How do I install Bio Copy Number Gatk Cnv in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-copy-number-gatk-cnv -a claude-code`. Or copy the skill folder (copy-number/gatk-cnv in GPTomics/bioSkills) into .claude/skills/bio-copy-number-gatk-cnv in your project. Claude Code loads it when a task matches its description.

How do I install Bio Copy Number Gatk Cnv in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-copy-number-gatk-cnv -a codex`. Or copy the skill folder (copy-number/gatk-cnv in GPTomics/bioSkills) into .agents/skills/bio-copy-number-gatk-cnv in your project. Codex loads it when a task matches its description.

Can I use Bio Copy Number Gatk Cnv in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-copy-number-gatk-cnv -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-copy-number-gatk-cnv, .gemini/skills/bio-copy-number-gatk-cnv, .github/skills/bio-copy-number-gatk-cnv and .opencode/skills/bio-copy-number-gatk-cnv in your project.

What does Bio Copy Number Gatk Cnv need to run?

Going by SKILL.md and its folder, Bio Copy Number Gatk Cnv needs a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.

Does Bio Copy Number Gatk Cnv access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Copy Number Gatk Cnv safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Copy Number Gatk Cnv use?

Bio Copy Number Gatk Cnv is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Copy Number Gatk Cnv use?

About 3.9k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Copy Number Gatk Cnv?

Skills that share tags, products or a category with Bio Copy Number Gatk Cnv: SQL Optimization Patterns (ynulihao/AgentSkillOS, 618 stars), Datamodellm (nimbalyst/nimbalyst, 1.9k stars), Add Mpk Task (mirage-project/mirage, 2.5k stars) and B200 Flash Attention4 Planner (mirage-project/mirage, 2.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Copy Number Gatk Cnv?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.