Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Detect somatic CNVs from WES/WGS/targeted BAMs (CNVkit v0.9.x).
$ npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills cnvkit-copy-number --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/cnvkit-copy-number .claude/skills/cnvkit-copy-number && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cnvkit-copy-number" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/cnvkit-copy-number into .claude/skills/cnvkit-copy-number/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cnvkit-copy-number", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/cnvkit-copy-numberType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills cnvkit-copy-number --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/cnvkit-copy-number .agents/skills/cnvkit-copy-number && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cnvkit-copy-number" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/cnvkit-copy-number into .agents/skills/cnvkit-copy-number/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cnvkit-copy-number", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills cnvkit-copy-number --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/cnvkit-copy-number .cursor/skills/cnvkit-copy-number && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cnvkit-copy-number" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/cnvkit-copy-number into .cursor/skills/cnvkit-copy-number/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cnvkit-copy-number", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/variant/cnvkit-copy-number--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills cnvkit-copy-number --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/cnvkit-copy-number .gemini/skills/cnvkit-copy-number && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cnvkit-copy-number" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/cnvkit-copy-number into .gemini/skills/cnvkit-copy-number/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cnvkit-copy-number", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills cnvkit-copy-numberInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/cnvkit-copy-number .github/skills/cnvkit-copy-number && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cnvkit-copy-number" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/cnvkit-copy-number into .github/skills/cnvkit-copy-number/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cnvkit-copy-number", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills cnvkit-copy-number --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/cnvkit-copy-number .opencode/skills/cnvkit-copy-number && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cnvkit-copy-number" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/variant/cnvkit-copy-number into .opencode/skills/cnvkit-copy-number/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cnvkit-copy-number", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cnvkit-copy-numberDetect somatic CNVs from WES/WGS/targeted BAMs (CNVkit v0.9.x).
Cnvkit Copy Number is an agent skill from jaechang-hits/SciAgent-Skills. Detect somatic CNVs from WES/WGS/targeted BAMs (CNVkit v0.9.x). Bin coverage in target/antitarget regions, normalize vs reference, segment with CBS/HMM, call amps/dels, scatter/diagram plots, purity/ploidy, VCF/SEG export. CLI plus Python API (cnvlib). Use GATK CNV for deep WGS with population controls; use CNVkit for targeted/exome where antitarget bins matter.
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. It works with Python. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
condapipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
cnvkit.readthedocs.iodoi.orggithub.comhgdownload.soe.ucsc.eduFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cnvkit Copy Number loads about 5.9k tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 1,092 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 1,092 words, ~5,871 tokens.
.claude/skills/cnvkit-copy-number/SKILL.md (or your agent's skills folder).CNVkit detects somatic copy number variants (CNVs) from whole-exome sequencing (WES), whole-genome sequencing (WGS), or targeted panel BAM files. It calculates read depth in both on-target (capture) bins and off-target (antitarget) bins, corrects for GC bias and library depth, segments the log2 copy ratio profile with circular binary segmentation (CBS) or a hidden Markov model (HMM), and calls amplifications and deletions. CNVkit provides both a CLI (cnvkit.py) and a Python API (cnvlib) for integration into analysis pipelines, and produces scatter plots, chromosome diagrams, heatmaps, and export files in VCF, BED, and SEG formats.
--method wgs modecnvkit.py scatter/diagramgatk DenoiseReadCounts / gatk ModelSegments) instead for deep WGS cohorts with large matched panel-of-normals (PoN); CNVkit is better suited for targeted/exome datacnvlib (installed as part of CNVkit), matplotlib, pandasCheck before installing: The tool may already be available in the current environment (e.g., inside a
pixi/condaenv). Runcommand -v cnvkit.pyfirst and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool viapixi run cnvkit.pyrather than barecnvkit.py.
# Install CNVkit via conda (recommended — handles R/DNAcopy dependency)
conda install -c bioconda cnvkit
# Or via pip (requires R + DNAcopy already installed)
pip install cnvkit
# Verify
cnvkit.py version
# cnvkit 0.9.10
# Install R DNAcopy (for CBS segmentation)
Rscript -e 'if (!requireNamespace("BiocManager")) install.packages("BiocManager"); BiocManager::install("DNAcopy")'
# Index BAM files if not already indexed
samtools index tumor.bam
samtools index normal.bamSettle these with the user before writing any analysis code.
decisions:
- id: D1
param: sequencingMethod
kind: required
source: data
ask: "Was this whole-genome, hybrid-capture exome, or amplicon sequencing?"
default: null
- id: D2
param: referenceNormals
kind: required
source: user
ask: "Which normal samples should define the expected coverage baseline?"
default: null
- id: D3
param: targetRegions
kind: required
source: user
depends_on: [D1]
ask: "Which BED file describes the captured regions?"
default: null
skip_if: "whole-genome sequencing - there are no capture targets"
- id: D4
param: tumorPurity
kind: required
source: user
ask: "What fraction of the sample is tumour rather than contaminating normal tissue?"
default: "1.0 - assumes a pure sample, which understates every copy-number change"
- id: D5
param: ploidy
kind: required
source: user
ask: "What baseline ploidy should absolute copy number be called against?"
default: 2
- id: D6
param: segmentMethod
kind: optional
source: user
ask: "Which algorithm should join bins into segments?"
default: "circular binary segmentation"
- id: D7
param: binSizes
kind: optional_conditional
source: user
depends_on: [D1]
ask: "Do the default target and antitarget bin sizes suit this capture design?"
default: "200 bp targets, 150 kb antitargets"
- id: D8
param: processes
kind: never_ask
source: data
reason: "Affects runtime only, not the copy-number calls"
default: "min(8, available_cores)"D2 and D4 are the two that decide whether the calls mean anything. A flat reference instead of matched normals leaves capture bias in the log2 ratios, and a purity of 1.0 on a 40%-tumour sample compresses every real gain and loss toward neutral. Both produce a complete, plausible-looking segment table.
# One-command paired tumor/normal CNV analysis (WES)
cnvkit.py batch tumor.bam \
--normal normal.bam \
--targets targets.bed \
--fasta GRCh38.fa \
--output-dir cnvkit_results/ \
--diagram --scatter \
--method hybrid
# Output files in cnvkit_results/:
# tumor.targetcoverage.cnn — target bin coverage
# tumor.antitargetcoverage.cnn — antitarget coverage
# tumor.cnr — copy number ratios
# tumor.cns — segmented copy numbers
# tumor-scatter.png — genome-wide scatter plot
# tumor-diagram.pdf — chromosome diagram
echo "CNV analysis complete"Build a reference from one or more normal BAM files. This corrects for systematic biases (GC content, mappability) and sets the neutral baseline.
# Option A: Paired normal reference (single matched normal)
cnvkit.py reference normal.targetcoverage.cnn normal.antitargetcoverage.cnn \
--fasta GRCh38.fa \
-o reference_normal.cnn
# Option B: Flat reference (no normal; uses GC/mappability correction only)
# Use when no matched normal is available
cnvkit.py reference \
--targets targets.bed \
--fasta GRCh38.fa \
--output flat_reference.cnn
# Option C: Pooled normal reference from multiple normals (most robust)
cnvkit.py batch \
normal1.bam normal2.bam normal3.bam \
--normal \
--targets targets.bed \
--fasta GRCh38.fa \
--output-reference pooled_reference.cnn \
--output-dir normals_cov/
echo "Reference created: pooled_reference.cnn"Bin the target BED file and compute per-bin read depth for tumor and normal samples.
# First, create accessible bins from the target BED
cnvkit.py target targets.bed \
--annotate refFlat.txt \
--split \
-o targets.split.bed
cnvkit.py antitarget targets.bed \
--access data/access-5k-mappable.hg38.bed \
-o antitargets.bed
# Calculate coverage for tumor sample
cnvkit.py coverage tumor.bam targets.split.bed \
-o tumor.targetcoverage.cnn
cnvkit.py coverage tumor.bam antitargets.bed \
-o tumor.antitargetcoverage.cnn
echo "Coverage files:"
echo " tumor.targetcoverage.cnn"
echo " tumor.antitargetcoverage.cnn"# Python API equivalent: compute coverage with cnvlib
import cnvlib
# Load and inspect coverage files
target_cov = cnvlib.read("tumor.targetcoverage.cnn")
antitarget_cov = cnvlib.read("tumor.antitargetcoverage.cnn")
print(f"Target bins: {len(target_cov)}")
print(f"Antitarget bins: {len(antitarget_cov)}")
print(f"Mean target depth: {target_cov['depth'].mean():.1f}×")
print(f"Median target depth: {target_cov['depth'].median():.1f}×")
# Export target-bin depths; render the coverage histogram with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`)
# "Box / Violin / Bar" recipe (or a histogram) -> figures/coverage_distribution.png
target_cov["depth"].clip(upper=500).to_csv("target_depths.csv", index=False)
print("Saved: target_depths.csv")Normalize tumor coverage against the reference (correcting for GC bias, library depth, and target efficiency).
# Fix sample-level biases relative to the reference
cnvkit.py fix \
tumor.targetcoverage.cnn \
tumor.antitargetcoverage.cnn \
pooled_reference.cnn \
-o tumor.cnr
echo "Copy number ratios: tumor.cnr"
# tumor.cnr columns: chromosome, start, end, gene, log2, depth, weight# Inspect copy ratio file with cnvlib Python API
import cnvlib
import pandas as pd
cnr = cnvlib.read("tumor.cnr")
print(f"Total bins: {len(cnr)}")
print(f"Chromosomes: {sorted(cnr.chromosome.unique())}")
# Convert to DataFrame and inspect
df = cnr.data
print(f"\nLog2 copy ratio summary:")
print(df["log2"].describe().round(3))
# Flag high-amplitude events
high_amp = df[df["log2"] >= 2.0]
hom_del = df[df["log2"] <= -3.0]
print(f"\nHigh amplitude bins (log2 >= 2.0): {len(high_amp)}")
print(f"Homozygous deletion bins (log2 <= -3.0): {len(hom_del)}")
if not high_amp.empty:
print(high_amp[["chromosome", "start", "end", "gene", "log2"]].head())Identify contiguous regions of similar copy ratio using CBS (Circular Binary Segmentation) or HMM segmentation.
# CBS segmentation (default; requires R DNAcopy)
cnvkit.py segment tumor.cnr \
-o tumor.cns \
--method cbs
# HMM segmentation (no R required; faster)
cnvkit.py segment tumor.cnr \
-o tumor.hmm.cns \
--method hmm
echo "Segments: tumor.cns"
# tumor.cns columns: chromosome, start, end, gene, log2, cn, depth, p_ttest, weight, probes# Python API: segment and inspect
import cnvlib
import subprocess
# Run segmentation via Python subprocess (mirrors CLI)
subprocess.run(
["cnvkit.py", "segment", "tumor.cnr", "-o", "tumor.cns", "--method", "cbs"],
check=True
)
# Load and analyze segments
cns = cnvlib.read("tumor.cns")
df_seg = cns.data
print(f"Total segments: {len(df_seg)}")
print(f"\nSegment log2 summary:")
print(df_seg["log2"].describe().round(3))
# Large segments (>5 Mb) with copy gain or loss
large_events = df_seg[
((df_seg["end"] - df_seg["start"]) > 5_000_000) &
(df_seg["log2"].abs() > 0.3)
].copy()
large_events["size_mb"] = (large_events["end"] - large_events["start"]) / 1e6
print(f"\nLarge CNV segments (>5 Mb, |log2|>0.3): {len(large_events)}")
print(large_events[["chromosome", "start", "end", "log2", "size_mb"]].head(8).to_string(index=False))Assign integer copy number states and classify amplifications and deletions.
# Call with default thresholds (diploid normal)
cnvkit.py call tumor.cns \
-o tumor.call.cns \
--ploidy 2
# Call with tumor purity estimate (if known)
cnvkit.py call tumor.cns \
--purity 0.7 \
--ploidy 2 \
-o tumor.call.purity.cns
echo "Called CNVs: tumor.call.cns"# Parse called CNV file and classify events
import cnvlib
import pandas as pd
cns_called = cnvlib.read("tumor.call.cns")
df = cns_called.data
# Classify by log2 thresholds (diploid assumed)
# log2 >= 1.0 = high-level amplification (CN >= 4)
# log2 0.2–1.0 = copy gain (CN = 3)
# log2 -1.0–-0.2 = heterozygous deletion (CN = 1)
# log2 <= -3.5 = homozygous deletion (CN = 0)
def classify_cnv(log2):
if log2 >= 1.0: return "AMP"
if log2 >= 0.2: return "GAIN"
if log2 <= -3.5: return "HOMDEL"
if log2 <= -1.0: return "LOSS"
return "NEUTRAL"
df["cnv_class"] = df["log2"].apply(classify_cnv)
print("CNV class counts:")
print(df["cnv_class"].value_counts().to_string())
# Focal amplifications in known oncogenes
oncogenes = ["ERBB2", "MYC", "EGFR", "CCND1", "CDK6", "MDM2", "KRAS"]
focal_amps = df[(df["cnv_class"] == "AMP") &
(df["gene"].str.split(",").apply(
lambda genes: any(g in oncogenes for g in genes)))]
if not focal_amps.empty:
print(f"\nFocal amplifications in oncogenes:")
print(focal_amps[["chromosome", "start", "end", "gene", "log2", "cn"]].to_string(index=False))Generate scatter plots and chromosome diagrams to review the copy number landscape.
# Genome-wide scatter plot (CNR bins + segments)
cnvkit.py scatter tumor.cnr \
-s tumor.cns \
-o tumor-scatter.png
# Chromosome diagram (color-coded by CN state)
cnvkit.py diagram tumor.cnr \
-s tumor.cns \
-o tumor-diagram.pdf
# Heatmap across multiple samples
cnvkit.py heatmap sample1.cns sample2.cns sample3.cns \
-o cohort_heatmap.pdf
echo "Plots saved: tumor-scatter.png, tumor-diagram.pdf, cohort_heatmap.pdf"# Extract a single chromosome's bins + segments for a custom log2-ratio scatter
import cnvlib
cnr = cnvlib.read("tumor.cnr")
cns = cnvlib.read("tumor.cns")
chrom = "chr7" # EGFR locus
cnr_chr = cnr.data[cnr.data["chromosome"] == chrom][["start", "log2"]]
cns_chr = cns.data[cns.data["chromosome"] == chrom][["start", "end", "log2"]]
cnr_chr.to_csv(f"{chrom}_bins.csv", index=False)
cns_chr.to_csv(f"{chrom}_segments.csv", index=False)
print(f"{chrom}: {len(cnr_chr)} bins, {len(cns_chr)} segments")
# Render a per-chromosome log2 scatter with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) scatter recipe
# (overlay segment hlines and a gene-locus axvspan for EGFR) -> figures/chr7_cnv_scatter.pngUse CNVkit's purity/ploidy estimation to interpret absolute copy numbers.
# Estimate purity and ploidy from the segmented CNV profile
cnvkit.py call tumor.cns \
--purity auto \
--ploidy 2 \
--method clonal \
-o tumor.call.auto.cns \
--center median
# Print purity/ploidy estimate embedded in header
head -5 tumor.call.auto.cns# Python API: purity/ploidy estimation with cnvlib
import cnvlib
from cnvlib import segmetrics
cns = cnvlib.read("tumor.cns")
cnr = cnvlib.read("tumor.cnr")
# Compute segment-level statistics
cns_with_stats = segmetrics.do_segmetrics(cnr, cns,
location_stats=["mean", "median"],
spread_stats=["stdev"])
# Export stats to CSV for review
df = cns_with_stats.data
df.to_csv("tumor_segment_stats.csv", index=False)
print(f"Segment stats saved: tumor_segment_stats.csv")
print(f"Segments: {len(df)}")
print(df[["chromosome", "start", "end", "log2", "cn", "mean", "stdev"]].head(8).to_string(index=False))Export CNV calls for downstream tools (GISTIC2, cBioPortal, IGV, clinical reporting).
# Export to VCF (for clinical variant databases)
cnvkit.py export vcf tumor.call.cns \
-o tumor.cnv.vcf
# Export to SEG format (for GISTIC2 and cBioPortal)
cnvkit.py export seg tumor.call.cns \
-o tumor.seg
# Export to BED format (for bedtools/IGV)
cnvkit.py export bed tumor.call.cns \
-o tumor.cnv.bed
echo "Exported:"
echo " tumor.cnv.vcf — VCF format"
echo " tumor.seg — SEG format (GISTIC2 / cBioPortal)"
echo " tumor.cnv.bed — BED format (IGV / bedtools)"# Parse and summarize SEG file
import pandas as pd
seg = pd.read_csv("tumor.seg", sep="\t",
names=["sample", "chrom", "start", "end", "n_probes", "log2"],
comment="#")
print(f"SEG file: {len(seg)} segments")
print(seg.head())
# Count amplifications and deletions by chromosome arm
amp_count = (seg["log2"] >= 0.585).sum()
del_count = (seg["log2"] <= -1.0).sum()
print(f"\nAmplifications (log2 >= 0.585): {amp_count}")
print(f"Deletions (log2 <= -1.0): {del_count}")
seg.to_csv("tumor_seg_annotated.csv", index=False)| Parameter | Default | Range / Options | Effect |
|---|---|---|---|
--method (batch/coverage) | "hybrid" | "hybrid", "wgs", "amplicon" | Sequencing method; selects target binning strategy |
--segment-method (segment) | "cbs" | "cbs", "hmm", "haar", "none" | Segmentation algorithm; CBS requires R DNAcopy |
--ploidy (call) | 2 | 1–6 | Assumed baseline ploidy for absolute CN calling |
--purity (call) | 1.0 | 0.1–1.0 or "auto" | Tumor cell fraction; corrects log2 ratios for admixed normal |
--target-avg-size (target) | 200 (WES) | 50–500 bp | Desired mean target bin size after splitting |
--antitarget-avg-size | 150000 | 10000–500000 bp | Antitarget bin size (larger = fewer bins, less noise) |
--drop-low-coverage (segment) | off | flag | Drop bins with depth < 5× before segmentation |
-p / --processes (coverage) | 1 | 1–CPU count | Parallel processes for coverage calculation |
--scatter (batch) | off | flag | Automatically generate scatter plot |
--diagram (batch) | off | flag | Automatically generate chromosome diagram |
When to use: No matched normal is available; use a flat GC/mappability-corrected reference.
# Step 1: Create flat reference
cnvkit.py reference \
--targets targets.bed \
--fasta GRCh38.fa \
--output flat_reference.cnn
# Step 2: Run batch on tumor-only
cnvkit.py batch tumor.bam \
--reference flat_reference.cnn \
--output-dir tumor_only_results/
echo "Tumor-only results: tumor_only_results/"When to use: Input is whole-genome sequencing (no capture BED needed).
# WGS mode uses a uniform genome-wide bin grid
cnvkit.py batch tumor_wgs.bam \
--normal normal_wgs.bam \
--method wgs \
--fasta GRCh38.fa \
--output-dir wgs_results/ \
--scatter
echo "WGS CNV analysis complete: wgs_results/"When to use: Automate CNVkit for a cohort of paired tumor/normal samples.
# Snakefile — CNVkit batch pipeline
configfile: "config.yaml"
SAMPLES = config["tumor_samples"] # list of sample names
GENOME = config["genome_fasta"]
TARGETS = config["targets_bed"]
POOLED = config["pooled_reference"] # pre-built pooled normal reference
rule all:
input:
expand("cnvkit/{sample}.call.cns", sample=SAMPLES),
expand("cnvkit/{sample}-scatter.png", sample=SAMPLES),
rule cnvkit_batch:
input:
tumor = "bam/{sample}.tumor.bam",
ref = POOLED,
output:
cnr = "cnvkit/{sample}.cnr",
cns = "cnvkit/{sample}.cns",
call = "cnvkit/{sample}.call.cns",
scatter = "cnvkit/{sample}-scatter.png",
threads: 4
shell:
"""
cnvkit.py batch {input.tumor} \
--reference {input.ref} \
--targets {TARGETS} \
--fasta {GENOME} \
--output-dir cnvkit/ \
--scatter --processes {threads}
cnvkit.py call cnvkit/{wildcards.sample}.cns \
--ploidy 2 -o {output.call}
"""When to use: Summarize copy number status for a panel of cancer genes from a called CNS file.
import cnvlib
import pandas as pd
cns = cnvlib.read("tumor.call.cns")
df = cns.data
# Cancer genes of interest
cancer_genes = ["ERBB2", "MYC", "EGFR", "CDKN2A", "RB1", "TP53",
"KRAS", "PTEN", "BRCA1", "BRCA2", "MDM2", "CDK4"]
rows = []
for gene in cancer_genes:
hits = df[df["gene"].str.contains(gene, na=False)]
if hits.empty:
rows.append({"gene": gene, "log2_mean": float("nan"), "cn": "n/a", "status": "not_detected"})
else:
best = hits.loc[hits["log2"].abs().idxmax()]
status = ("AMP" if best["log2"] >= 1.0 else
"GAIN" if best["log2"] >= 0.2 else
"HOMDEL" if best["log2"] <= -3.5 else
"LOSS" if best["log2"] <= -1.0 else "NEUTRAL")
rows.append({"gene": gene, "log2_mean": round(best["log2"], 3),
"cn": best.get("cn", "?"), "status": status})
result_df = pd.DataFrame(rows)
print(result_df.to_string(index=False))
result_df.to_csv("cancer_gene_cnv_summary.csv", index=False)
print("\nSaved: cancer_gene_cnv_summary.csv")| Output File | Format | Description |
|---|---|---|
tumor.targetcoverage.cnn | CNN | Per-bin target region read depth and log2 coverage |
tumor.antitargetcoverage.cnn | CNN | Per-bin antitarget region coverage (off-target reads) |
tumor.cnr | CNR | Normalized log2 copy ratio per bin (GC/depth corrected) |
tumor.cns | CNS | Segmented copy ratios (CBS/HMM output); one row per segment |
tumor.call.cns | CNS | Called copy number states with integer CN and purity adjustment |
tumor-scatter.png | PNG | Genome-wide scatter plot of bins + segment overlays |
tumor-diagram.pdf | Chromosome arm diagram color-coded by CN state | |
tumor.seg | SEG | GISTIC2/cBioPortal-format segment file |
tumor.cnv.vcf | VCF | VCF-format CNV calls for clinical databases |
| Problem | Cause | Solution |
|---|---|---|
| High noise in antitarget bins | Low off-target coverage (<0.1× mean) | Increase --antitarget-avg-size to 500kb; use --drop-low-coverage to exclude low bins before segmentation |
CBS segmentation fails with Error in DNAcopy | R DNAcopy not installed or incompatible | Install via BiocManager::install("DNAcopy"); alternatively use --method hmm (no R required) |
| All segments near 0 (no CNVs detected) | Low tumor purity (<20%) or shallow coverage | Verify purity with cnvkit.py call --purity auto; check target coverage depth with cnvkit.py coverage |
| Wavy GC bias across chromosomes | GC normalization failed or flat reference used with biased tumor | Rebuild reference with matched normal BAMs; check cnvkit.py reference includes --fasta for GC correction |
| Many very short segments (over-segmentation) | CBS threshold too sensitive | Add --threshold 0.2 or --smooth-cbs to reduce false segment boundaries; increase minimum segment size |
ImportError: No module named cnvlib | CNVkit not installed in active environment | Activate the correct conda env: conda activate cnvkit_env; verify cnvkit.py version |
| Chromosome naming mismatch | BAM uses 1 but reference uses chr1 or vice versa | Ensure BAM, BED, and FASTA all use the same chromosome naming convention |
© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/variant/cnvkit-copy-number of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Cnvkit Copy Number next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cnvkit Copy Number this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~5.9k | Automated safety check: Pass | Apache-2.0 | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Singlecell Qcxuzhougeng/wisp-science | 1k | — | ~1.6k | Automated safety check: Pass | AGPL-3.0 | |
| Trackplotygidtu/trackplot | 109 | — | ~1.9k | Automated safety check: Pass | BSD-3-Clause | |
| UniProt Database Accessdavila7/claude-code-templates | 33k | 14 repos | ~1.7k | Automated safety check: Pass | MIT |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
xuzhougeng/wisp-science
A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.
ygidtu/trackplot
Generate sashimi-style genome visualization plots (coverage, line, heatmap, IGV read-by-read, HiC, circRNA, motif) from BAM/bigWig/depth/HiC inputs.
davila7/claude-code-templates
Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.
QING1105/ezST
End-to-end 10x Visium spatial transcriptomics analysis workflow with staged execution and human review gates.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
Detect somatic CNVs from WES/WGS/targeted BAMs (CNVkit v0.9.x). Cnvkit Copy Number is an agent skill from jaechang-hits/SciAgent-Skills.x).
Cnvkit Copy Number fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/variant/cnvkit-copy-number in jaechang-hits/SciAgent-Skills) into .claude/skills/cnvkit-copy-number in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/variant/cnvkit-copy-number in jaechang-hits/SciAgent-Skills) into .agents/skills/cnvkit-copy-number in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill cnvkit-copy-number -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cnvkit-copy-number, .gemini/skills/cnvkit-copy-number, .github/skills/cnvkit-copy-number and .opencode/skills/cnvkit-copy-number in your project.
Going by SKILL.md and its folder, Cnvkit Copy Number needs the command-line tools its instructions call (conda and pip). Our summary lists: Python 3.
SKILL.md names 4 domains. As links in the text: cnvkit.readthedocs.io, doi.org, github.com and hgdownload.soe.ucsc.edu. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cnvkit Copy Number is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cnvkit Copy Number: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Singlecell Qc (xuzhougeng/wisp-science, 1k stars) and Trackplot (ygidtu/trackplot, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.