Agent skill

Cnvkit Copy Number

by jaechang-hits in jaechang-hits/SciAgent-Skills

Detect somatic CNVs from WES/WGS/targeted BAMs (CNVkit v0.9.x).

Apache-2.0Auto-check passedResearch & Science

Install Cnvkit Copy Number

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills cnvkit-copy-number --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/cnvkit-copy-number .claude/skills/cnvkit-copy-number && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cnvkit-copy-number
GitHub stars
374
Used in
1 other repo
Token cost
~5.9k tokens
SKILL.md length
1,092 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
Apache-2.0

At a glance

Detect somatic CNVs from WES/WGS/targeted BAMs (CNVkit v0.9.x).

  • Works in 8 steps: Create Copy Number Reference → Calculate Coverage in Target and… → Normalize and Correct Copy Ratios → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Pre-flight Interview, plus 7 more sections
  • Calls conda and pip

What it does

Cnvkit Copy Number is an agent skill from jaechang-hits/SciAgent-Skills. Detect somatic CNVs from WES/WGS/targeted BAMs (CNVkit v0.9.x). Bin coverage in target/antitarget regions, normalize vs reference, segment with CBS/HMM, call amps/dels, scatter/diagram plots, purity/ploidy, VCF/SEG export. CLI plus Python API (cnvlib). Use GATK CNV for deep WGS with population controls; use CNVkit for targeted/exome where antitarget bins matter.

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics. It works with Python. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/cnvkit-copy-number”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Create Copy Number Reference
  2. Calculate Coverage in Target and Antitarget Bins
  3. Normalize and Correct Copy Ratios
  4. Segment Copy Number Ratios
  5. Call CNV States
  6. Visualize CNV Profile
  7. Estimate Tumor Purity and Ploidy
  8. Export to VCF, BED, and SEG Formats

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • conda
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • cnvkit.readthedocs.io
    • doi.org
    • github.com
    • hgdownload.soe.ucsc.edu

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cnvkit Copy Number loads about 5.9k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 1,092 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 1,092 words, ~5,871 tokens.

Download SKILL.mdSave it as .claude/skills/cnvkit-copy-number/SKILL.md (or your agent's skills folder).
name
cnvkit-copy-number
description
Detect somatic CNVs from WES/WGS/targeted BAMs (CNVkit v0.9.x). Bin coverage in target/antitarget regions, normalize vs reference, segment with CBS/HMM, call amps/dels, scatter/diagram plots, purity/ploidy, VCF/SEG export. CLI plus Python API (cnvlib). Use GATK CNV for deep WGS with population controls; use CNVkit for targeted/exome where antitarget bins matter.
license
Apache-2.0

CNVkit Copy Number Analysis

Overview

CNVkit detects somatic copy number variants (CNVs) from whole-exome sequencing (WES), whole-genome sequencing (WGS), or targeted panel BAM files. It calculates read depth in both on-target (capture) bins and off-target (antitarget) bins, corrects for GC bias and library depth, segments the log2 copy ratio profile with circular binary segmentation (CBS) or a hidden Markov model (HMM), and calls amplifications and deletions. CNVkit provides both a CLI (cnvkit.py) and a Python API (cnvlib) for integration into analysis pipelines, and produces scatter plots, chromosome diagrams, heatmaps, and export files in VCF, BED, and SEG formats.

When to Use

  • Calling somatic copy number variants from tumor-normal paired exome (WES) or targeted panel sequencing
  • Detecting copy number alterations in tumor-only samples using a pooled normal reference
  • Running CNV analysis on whole-genome sequencing (WGS) data with the --method wgs mode
  • Estimating tumor purity and ploidy for samples where purity is unknown, to interpret copy ratio calls
  • Generating SEG format copy number files for GISTIC2, cBioPortal, or IGV visualization
  • Identifying focal amplifications (e.g., ERBB2, MYC) or homozygous deletions (e.g., CDKN2A, RB1)
  • Use omics-plotting SKILL for generic coverage/log2-ratio figures from exported tables; genome-wide CNV views use cnvkit.py scatter/diagram
  • Use GATK CNV (gatk DenoiseReadCounts / gatk ModelSegments) instead for deep WGS cohorts with large matched panel-of-normals (PoN); CNVkit is better suited for targeted/exome data
  • Use Control-FREEC instead when you need allele-frequency-based B-allele fraction modeling alongside CNV calling

Prerequisites

  • Software: CNVkit v0.9.x, Python 3.8+, R (for CBS segmentation), samtools
  • Python packages: cnvlib (installed as part of CNVkit), matplotlib, pandas
  • Input files: sorted, indexed BAM files (tumor ± matched normal); BED file of capture targets; reference genome FASTA; access to R with DNAcopy package for CBS
  • Data requirements: minimum ~50× mean target coverage for WES; WGS works at 20-30×

Check before installing: The tool may already be available in the current environment (e.g., inside a pixi / conda env). Run command -v cnvkit.py first and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool via pixi run cnvkit.py rather than bare cnvkit.py.

bash
# Install CNVkit via conda (recommended — handles R/DNAcopy dependency)
conda install -c bioconda cnvkit

# Or via pip (requires R + DNAcopy already installed)
pip install cnvkit

# Verify
cnvkit.py version
# cnvkit 0.9.10

# Install R DNAcopy (for CBS segmentation)
Rscript -e 'if (!requireNamespace("BiocManager")) install.packages("BiocManager"); BiocManager::install("DNAcopy")'

# Index BAM files if not already indexed
samtools index tumor.bam
samtools index normal.bam

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: sequencingMethod
    kind: required
    source: data
    ask: "Was this whole-genome, hybrid-capture exome, or amplicon sequencing?"
    default: null

  - id: D2
    param: referenceNormals
    kind: required
    source: user
    ask: "Which normal samples should define the expected coverage baseline?"
    default: null

  - id: D3
    param: targetRegions
    kind: required
    source: user
    depends_on: [D1]
    ask: "Which BED file describes the captured regions?"
    default: null
    skip_if: "whole-genome sequencing - there are no capture targets"

  - id: D4
    param: tumorPurity
    kind: required
    source: user
    ask: "What fraction of the sample is tumour rather than contaminating normal tissue?"
    default: "1.0 - assumes a pure sample, which understates every copy-number change"

  - id: D5
    param: ploidy
    kind: required
    source: user
    ask: "What baseline ploidy should absolute copy number be called against?"
    default: 2

  - id: D6
    param: segmentMethod
    kind: optional
    source: user
    ask: "Which algorithm should join bins into segments?"
    default: "circular binary segmentation"

  - id: D7
    param: binSizes
    kind: optional_conditional
    source: user
    depends_on: [D1]
    ask: "Do the default target and antitarget bin sizes suit this capture design?"
    default: "200 bp targets, 150 kb antitargets"

  - id: D8
    param: processes
    kind: never_ask
    source: data
    reason: "Affects runtime only, not the copy-number calls"
    default: "min(8, available_cores)"

D2 and D4 are the two that decide whether the calls mean anything. A flat reference instead of matched normals leaves capture bias in the log2 ratios, and a purity of 1.0 on a 40%-tumour sample compresses every real gain and loss toward neutral. Both produce a complete, plausible-looking segment table.

Quick Start

bash
# One-command paired tumor/normal CNV analysis (WES)
cnvkit.py batch tumor.bam \
    --normal normal.bam \
    --targets targets.bed \
    --fasta GRCh38.fa \
    --output-dir cnvkit_results/ \
    --diagram --scatter \
    --method hybrid

# Output files in cnvkit_results/:
#   tumor.targetcoverage.cnn   — target bin coverage
#   tumor.antitargetcoverage.cnn — antitarget coverage
#   tumor.cnr                  — copy number ratios
#   tumor.cns                  — segmented copy numbers
#   tumor-scatter.png          — genome-wide scatter plot
#   tumor-diagram.pdf          — chromosome diagram
echo "CNV analysis complete"

Workflow

Step 1: Create Copy Number Reference

Build a reference from one or more normal BAM files. This corrects for systematic biases (GC content, mappability) and sets the neutral baseline.

bash
# Option A: Paired normal reference (single matched normal)
cnvkit.py reference normal.targetcoverage.cnn normal.antitargetcoverage.cnn \
    --fasta GRCh38.fa \
    -o reference_normal.cnn

# Option B: Flat reference (no normal; uses GC/mappability correction only)
# Use when no matched normal is available
cnvkit.py reference \
    --targets targets.bed \
    --fasta GRCh38.fa \
    --output flat_reference.cnn

# Option C: Pooled normal reference from multiple normals (most robust)
cnvkit.py batch \
    normal1.bam normal2.bam normal3.bam \
    --normal \
    --targets targets.bed \
    --fasta GRCh38.fa \
    --output-reference pooled_reference.cnn \
    --output-dir normals_cov/

echo "Reference created: pooled_reference.cnn"
Step 2: Calculate Coverage in Target and Antitarget Bins

Bin the target BED file and compute per-bin read depth for tumor and normal samples.

bash
# First, create accessible bins from the target BED
cnvkit.py target targets.bed \
    --annotate refFlat.txt \
    --split \
    -o targets.split.bed

cnvkit.py antitarget targets.bed \
    --access data/access-5k-mappable.hg38.bed \
    -o antitargets.bed

# Calculate coverage for tumor sample
cnvkit.py coverage tumor.bam targets.split.bed \
    -o tumor.targetcoverage.cnn

cnvkit.py coverage tumor.bam antitargets.bed \
    -o tumor.antitargetcoverage.cnn

echo "Coverage files:"
echo "  tumor.targetcoverage.cnn"
echo "  tumor.antitargetcoverage.cnn"
python
# Python API equivalent: compute coverage with cnvlib
import cnvlib

# Load and inspect coverage files
target_cov = cnvlib.read("tumor.targetcoverage.cnn")
antitarget_cov = cnvlib.read("tumor.antitargetcoverage.cnn")

print(f"Target bins: {len(target_cov)}")
print(f"Antitarget bins: {len(antitarget_cov)}")
print(f"Mean target depth: {target_cov['depth'].mean():.1f}×")
print(f"Median target depth: {target_cov['depth'].median():.1f}×")

# Export target-bin depths; render the coverage histogram with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`)
# "Box / Violin / Bar" recipe (or a histogram) -> figures/coverage_distribution.png
target_cov["depth"].clip(upper=500).to_csv("target_depths.csv", index=False)
print("Saved: target_depths.csv")
Step 3: Normalize and Correct Copy Ratios

Normalize tumor coverage against the reference (correcting for GC bias, library depth, and target efficiency).

bash
# Fix sample-level biases relative to the reference
cnvkit.py fix \
    tumor.targetcoverage.cnn \
    tumor.antitargetcoverage.cnn \
    pooled_reference.cnn \
    -o tumor.cnr

echo "Copy number ratios: tumor.cnr"
# tumor.cnr columns: chromosome, start, end, gene, log2, depth, weight
python
# Inspect copy ratio file with cnvlib Python API
import cnvlib
import pandas as pd

cnr = cnvlib.read("tumor.cnr")
print(f"Total bins: {len(cnr)}")
print(f"Chromosomes: {sorted(cnr.chromosome.unique())}")

# Convert to DataFrame and inspect
df = cnr.data
print(f"\nLog2 copy ratio summary:")
print(df["log2"].describe().round(3))

# Flag high-amplitude events
high_amp = df[df["log2"] >= 2.0]
hom_del  = df[df["log2"] <= -3.0]
print(f"\nHigh amplitude bins (log2 >= 2.0): {len(high_amp)}")
print(f"Homozygous deletion bins (log2 <= -3.0): {len(hom_del)}")
if not high_amp.empty:
    print(high_amp[["chromosome", "start", "end", "gene", "log2"]].head())
Step 4: Segment Copy Number Ratios

Identify contiguous regions of similar copy ratio using CBS (Circular Binary Segmentation) or HMM segmentation.

bash
# CBS segmentation (default; requires R DNAcopy)
cnvkit.py segment tumor.cnr \
    -o tumor.cns \
    --method cbs

# HMM segmentation (no R required; faster)
cnvkit.py segment tumor.cnr \
    -o tumor.hmm.cns \
    --method hmm

echo "Segments: tumor.cns"
# tumor.cns columns: chromosome, start, end, gene, log2, cn, depth, p_ttest, weight, probes
python
# Python API: segment and inspect
import cnvlib
import subprocess

# Run segmentation via Python subprocess (mirrors CLI)
subprocess.run(
    ["cnvkit.py", "segment", "tumor.cnr", "-o", "tumor.cns", "--method", "cbs"],
    check=True
)

# Load and analyze segments
cns = cnvlib.read("tumor.cns")
df_seg = cns.data
print(f"Total segments: {len(df_seg)}")
print(f"\nSegment log2 summary:")
print(df_seg["log2"].describe().round(3))

# Large segments (>5 Mb) with copy gain or loss
large_events = df_seg[
    ((df_seg["end"] - df_seg["start"]) > 5_000_000) &
    (df_seg["log2"].abs() > 0.3)
].copy()
large_events["size_mb"] = (large_events["end"] - large_events["start"]) / 1e6
print(f"\nLarge CNV segments (>5 Mb, |log2|>0.3): {len(large_events)}")
print(large_events[["chromosome", "start", "end", "log2", "size_mb"]].head(8).to_string(index=False))
Step 5: Call CNV States

Assign integer copy number states and classify amplifications and deletions.

bash
# Call with default thresholds (diploid normal)
cnvkit.py call tumor.cns \
    -o tumor.call.cns \
    --ploidy 2

# Call with tumor purity estimate (if known)
cnvkit.py call tumor.cns \
    --purity 0.7 \
    --ploidy 2 \
    -o tumor.call.purity.cns

echo "Called CNVs: tumor.call.cns"
python
# Parse called CNV file and classify events
import cnvlib
import pandas as pd

cns_called = cnvlib.read("tumor.call.cns")
df = cns_called.data

# Classify by log2 thresholds (diploid assumed)
# log2 >= 1.0 = high-level amplification (CN >= 4)
# log2 0.2–1.0 = copy gain (CN = 3)
# log2 -1.0–-0.2 = heterozygous deletion (CN = 1)
# log2 <= -3.5 = homozygous deletion (CN = 0)

def classify_cnv(log2):
    if log2 >= 1.0:    return "AMP"
    if log2 >= 0.2:    return "GAIN"
    if log2 <= -3.5:   return "HOMDEL"
    if log2 <= -1.0:   return "LOSS"
    return "NEUTRAL"

df["cnv_class"] = df["log2"].apply(classify_cnv)
print("CNV class counts:")
print(df["cnv_class"].value_counts().to_string())

# Focal amplifications in known oncogenes
oncogenes = ["ERBB2", "MYC", "EGFR", "CCND1", "CDK6", "MDM2", "KRAS"]
focal_amps = df[(df["cnv_class"] == "AMP") &
                (df["gene"].str.split(",").apply(
                    lambda genes: any(g in oncogenes for g in genes)))]
if not focal_amps.empty:
    print(f"\nFocal amplifications in oncogenes:")
    print(focal_amps[["chromosome", "start", "end", "gene", "log2", "cn"]].to_string(index=False))
Step 6: Visualize CNV Profile

Generate scatter plots and chromosome diagrams to review the copy number landscape.

bash
# Genome-wide scatter plot (CNR bins + segments)
cnvkit.py scatter tumor.cnr \
    -s tumor.cns \
    -o tumor-scatter.png

# Chromosome diagram (color-coded by CN state)
cnvkit.py diagram tumor.cnr \
    -s tumor.cns \
    -o tumor-diagram.pdf

# Heatmap across multiple samples
cnvkit.py heatmap sample1.cns sample2.cns sample3.cns \
    -o cohort_heatmap.pdf

echo "Plots saved: tumor-scatter.png, tumor-diagram.pdf, cohort_heatmap.pdf"
python
# Extract a single chromosome's bins + segments for a custom log2-ratio scatter
import cnvlib

cnr = cnvlib.read("tumor.cnr")
cns = cnvlib.read("tumor.cns")

chrom = "chr7"   # EGFR locus
cnr_chr = cnr.data[cnr.data["chromosome"] == chrom][["start", "log2"]]
cns_chr = cns.data[cns.data["chromosome"] == chrom][["start", "end", "log2"]]
cnr_chr.to_csv(f"{chrom}_bins.csv", index=False)
cns_chr.to_csv(f"{chrom}_segments.csv", index=False)
print(f"{chrom}: {len(cnr_chr)} bins, {len(cns_chr)} segments")
# Render a per-chromosome log2 scatter with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) scatter recipe
# (overlay segment hlines and a gene-locus axvspan for EGFR) -> figures/chr7_cnv_scatter.png
Step 7: Estimate Tumor Purity and Ploidy

Use CNVkit's purity/ploidy estimation to interpret absolute copy numbers.

bash
# Estimate purity and ploidy from the segmented CNV profile
cnvkit.py call tumor.cns \
    --purity auto \
    --ploidy 2 \
    --method clonal \
    -o tumor.call.auto.cns \
    --center median

# Print purity/ploidy estimate embedded in header
head -5 tumor.call.auto.cns
python
# Python API: purity/ploidy estimation with cnvlib
import cnvlib
from cnvlib import segmetrics

cns = cnvlib.read("tumor.cns")
cnr = cnvlib.read("tumor.cnr")

# Compute segment-level statistics
cns_with_stats = segmetrics.do_segmetrics(cnr, cns,
    location_stats=["mean", "median"],
    spread_stats=["stdev"])

# Export stats to CSV for review
df = cns_with_stats.data
df.to_csv("tumor_segment_stats.csv", index=False)
print(f"Segment stats saved: tumor_segment_stats.csv")
print(f"Segments: {len(df)}")
print(df[["chromosome", "start", "end", "log2", "cn", "mean", "stdev"]].head(8).to_string(index=False))
Step 8: Export to VCF, BED, and SEG Formats

Export CNV calls for downstream tools (GISTIC2, cBioPortal, IGV, clinical reporting).

bash
# Export to VCF (for clinical variant databases)
cnvkit.py export vcf tumor.call.cns \
    -o tumor.cnv.vcf

# Export to SEG format (for GISTIC2 and cBioPortal)
cnvkit.py export seg tumor.call.cns \
    -o tumor.seg

# Export to BED format (for bedtools/IGV)
cnvkit.py export bed tumor.call.cns \
    -o tumor.cnv.bed

echo "Exported:"
echo "  tumor.cnv.vcf   — VCF format"
echo "  tumor.seg       — SEG format (GISTIC2 / cBioPortal)"
echo "  tumor.cnv.bed   — BED format (IGV / bedtools)"
python
# Parse and summarize SEG file
import pandas as pd

seg = pd.read_csv("tumor.seg", sep="\t",
                  names=["sample", "chrom", "start", "end", "n_probes", "log2"],
                  comment="#")
print(f"SEG file: {len(seg)} segments")
print(seg.head())

# Count amplifications and deletions by chromosome arm
amp_count = (seg["log2"] >= 0.585).sum()
del_count  = (seg["log2"] <= -1.0).sum()
print(f"\nAmplifications (log2 >= 0.585): {amp_count}")
print(f"Deletions (log2 <= -1.0):       {del_count}")
seg.to_csv("tumor_seg_annotated.csv", index=False)

Key Parameters

ParameterDefaultRange / OptionsEffect
--method (batch/coverage)"hybrid""hybrid", "wgs", "amplicon"Sequencing method; selects target binning strategy
--segment-method (segment)"cbs""cbs", "hmm", "haar", "none"Segmentation algorithm; CBS requires R DNAcopy
--ploidy (call)21–6Assumed baseline ploidy for absolute CN calling
--purity (call)1.00.1–1.0 or "auto"Tumor cell fraction; corrects log2 ratios for admixed normal
--target-avg-size (target)200 (WES)50–500 bpDesired mean target bin size after splitting
--antitarget-avg-size15000010000–500000 bpAntitarget bin size (larger = fewer bins, less noise)
--drop-low-coverage (segment)offflagDrop bins with depth < 5× before segmentation
-p / --processes (coverage)11–CPU countParallel processes for coverage calculation
--scatter (batch)offflagAutomatically generate scatter plot
--diagram (batch)offflagAutomatically generate chromosome diagram
Show full SKILL.md (397 more words)Show less

Common Recipes

Recipe: Tumor-Only Analysis with Flat Reference

When to use: No matched normal is available; use a flat GC/mappability-corrected reference.

bash
# Step 1: Create flat reference
cnvkit.py reference \
    --targets targets.bed \
    --fasta GRCh38.fa \
    --output flat_reference.cnn

# Step 2: Run batch on tumor-only
cnvkit.py batch tumor.bam \
    --reference flat_reference.cnn \
    --output-dir tumor_only_results/

echo "Tumor-only results: tumor_only_results/"
Recipe: WGS Mode for Whole-Genome Data

When to use: Input is whole-genome sequencing (no capture BED needed).

bash
# WGS mode uses a uniform genome-wide bin grid
cnvkit.py batch tumor_wgs.bam \
    --normal normal_wgs.bam \
    --method wgs \
    --fasta GRCh38.fa \
    --output-dir wgs_results/ \
    --scatter

echo "WGS CNV analysis complete: wgs_results/"
Recipe: Snakemake Integration for Multi-Sample Cohort

When to use: Automate CNVkit for a cohort of paired tumor/normal samples.

python
# Snakefile — CNVkit batch pipeline
configfile: "config.yaml"
SAMPLES = config["tumor_samples"]   # list of sample names
GENOME  = config["genome_fasta"]
TARGETS = config["targets_bed"]
POOLED  = config["pooled_reference"]  # pre-built pooled normal reference

rule all:
    input:
        expand("cnvkit/{sample}.call.cns", sample=SAMPLES),
        expand("cnvkit/{sample}-scatter.png", sample=SAMPLES),

rule cnvkit_batch:
    input:
        tumor  = "bam/{sample}.tumor.bam",
        ref    = POOLED,
    output:
        cnr    = "cnvkit/{sample}.cnr",
        cns    = "cnvkit/{sample}.cns",
        call   = "cnvkit/{sample}.call.cns",
        scatter = "cnvkit/{sample}-scatter.png",
    threads: 4
    shell:
        """
        cnvkit.py batch {input.tumor} \
            --reference {input.ref} \
            --targets {TARGETS} \
            --fasta {GENOME} \
            --output-dir cnvkit/ \
            --scatter --processes {threads}
        cnvkit.py call cnvkit/{wildcards.sample}.cns \
            --ploidy 2 -o {output.call}
        """
Recipe: Export Gene-Level CNV Table

When to use: Summarize copy number status for a panel of cancer genes from a called CNS file.

python
import cnvlib
import pandas as pd

cns = cnvlib.read("tumor.call.cns")
df = cns.data

# Cancer genes of interest
cancer_genes = ["ERBB2", "MYC", "EGFR", "CDKN2A", "RB1", "TP53",
                "KRAS", "PTEN", "BRCA1", "BRCA2", "MDM2", "CDK4"]

rows = []
for gene in cancer_genes:
    hits = df[df["gene"].str.contains(gene, na=False)]
    if hits.empty:
        rows.append({"gene": gene, "log2_mean": float("nan"), "cn": "n/a", "status": "not_detected"})
    else:
        best = hits.loc[hits["log2"].abs().idxmax()]
        status = ("AMP" if best["log2"] >= 1.0 else
                  "GAIN" if best["log2"] >= 0.2 else
                  "HOMDEL" if best["log2"] <= -3.5 else
                  "LOSS" if best["log2"] <= -1.0 else "NEUTRAL")
        rows.append({"gene": gene, "log2_mean": round(best["log2"], 3),
                     "cn": best.get("cn", "?"), "status": status})

result_df = pd.DataFrame(rows)
print(result_df.to_string(index=False))
result_df.to_csv("cancer_gene_cnv_summary.csv", index=False)
print("\nSaved: cancer_gene_cnv_summary.csv")

Expected Outputs

Output FileFormatDescription
tumor.targetcoverage.cnnCNNPer-bin target region read depth and log2 coverage
tumor.antitargetcoverage.cnnCNNPer-bin antitarget region coverage (off-target reads)
tumor.cnrCNRNormalized log2 copy ratio per bin (GC/depth corrected)
tumor.cnsCNSSegmented copy ratios (CBS/HMM output); one row per segment
tumor.call.cnsCNSCalled copy number states with integer CN and purity adjustment
tumor-scatter.pngPNGGenome-wide scatter plot of bins + segment overlays
tumor-diagram.pdfPDFChromosome arm diagram color-coded by CN state
tumor.segSEGGISTIC2/cBioPortal-format segment file
tumor.cnv.vcfVCFVCF-format CNV calls for clinical databases

Troubleshooting

ProblemCauseSolution
High noise in antitarget binsLow off-target coverage (<0.1× mean)Increase --antitarget-avg-size to 500kb; use --drop-low-coverage to exclude low bins before segmentation
CBS segmentation fails with Error in DNAcopyR DNAcopy not installed or incompatibleInstall via BiocManager::install("DNAcopy"); alternatively use --method hmm (no R required)
All segments near 0 (no CNVs detected)Low tumor purity (<20%) or shallow coverageVerify purity with cnvkit.py call --purity auto; check target coverage depth with cnvkit.py coverage
Wavy GC bias across chromosomesGC normalization failed or flat reference used with biased tumorRebuild reference with matched normal BAMs; check cnvkit.py reference includes --fasta for GC correction
Many very short segments (over-segmentation)CBS threshold too sensitiveAdd --threshold 0.2 or --smooth-cbs to reduce false segment boundaries; increase minimum segment size
ImportError: No module named cnvlibCNVkit not installed in active environmentActivate the correct conda env: conda activate cnvkit_env; verify cnvkit.py version
Chromosome naming mismatchBAM uses 1 but reference uses chr1 or vice versaEnsure BAM, BED, and FASTA all use the same chromosome naming convention

References

© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/genomics-bioinformatics/variant/cnvkit-copy-number of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Cnvkit Copy Number next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cnvkit Copy Number compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cnvkit Copy Number this skilljaechang-hits/SciAgent-Skills3741 repos~5.9kAutomated safety check: PassApache-2.0
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Singlecell Qcxuzhougeng/wisp-science1k—~1.6kAutomated safety check: PassAGPL-3.0
Trackplotygidtu/trackplot109—~1.9kAutomated safety check: PassBSD-3-Clause
UniProt Database Accessdavila7/claude-code-templates33k14 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Singlecell Qc

    xuzhougeng/wisp-science

    A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.

    1k GitHub stars~1.6k tokensUpdated today
    Research & ScienceAuto-check passed
  • Trackplot

    ygidtu/trackplot

    Generate sashimi-style genome visualization plots (coverage, line, heatmap, IGV read-by-read, HiC, circRNA, motif) from BAM/bigWig/depth/HiC inputs.

    109 GitHub stars~1.9k tokensUpdated 15 days ago
    Research & ScienceAuto-check passed
  • UniProt Database Access

    davila7/claude-code-templates

    Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.

    33k GitHub starsUsed in 14 repos~1.7k tokens
    Research & ScienceAuto-check passed
  • End-to-end 10x Visium spatial transcriptomics analysis workflow with staged execution and human review gates.

    101 GitHub stars~1.4k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Works with

Questions about Cnvkit Copy Number

What does Cnvkit Copy Number do?

Detect somatic CNVs from WES/WGS/targeted BAMs (CNVkit v0.9.x). Cnvkit Copy Number is an agent skill from jaechang-hits/SciAgent-Skills.x).

When should I use Cnvkit Copy Number?

Cnvkit Copy Number fits situations like: tasks that involve Bioinformatics.

How do I install Cnvkit Copy Number in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/variant/cnvkit-copy-number in jaechang-hits/SciAgent-Skills) into .claude/skills/cnvkit-copy-number in your project. Claude Code loads it when a task matches its description.

How do I install Cnvkit Copy Number in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/variant/cnvkit-copy-number in jaechang-hits/SciAgent-Skills) into .agents/skills/cnvkit-copy-number in your project. Codex loads it when a task matches its description.

Can I use Cnvkit Copy Number in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cnvkit-copy-number, .gemini/skills/cnvkit-copy-number, .github/skills/cnvkit-copy-number and .opencode/skills/cnvkit-copy-number in your project.

What does Cnvkit Copy Number need to run?

Going by SKILL.md and its folder, Cnvkit Copy Number needs the command-line tools its instructions call (conda and pip). Our summary lists: Python 3.

Does Cnvkit Copy Number access the network?

SKILL.md names 4 domains. As links in the text: cnvkit.readthedocs.io, doi.org, github.com and hgdownload.soe.ucsc.edu. This is read from the text; nothing was executed.

Is Cnvkit Copy Number safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cnvkit Copy Number use?

Cnvkit Copy Number is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cnvkit Copy Number use?

About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cnvkit Copy Number?

Skills that share tags, products or a category with Cnvkit Copy Number: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Singlecell Qc (xuzhougeng/wisp-science, 1k stars) and Trackplot (ygidtu/trackplot, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cnvkit Copy Number?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.