Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
Infer large-scale copy-number alterations from tumor single-cell or single-nucleus RNA-seq to separate malignant from normal cells and call subclones, using inferCNV, copyKAT, Numbat, and SCEVAN.
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-cnv-inference -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-single-cell-cnv-inference --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/single-cell/cnv-inference .claude/skills/bio-single-cell-cnv-inference && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-single-cell-cnv-inference" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/cnv-inference into .claude/skills/bio-single-cell-cnv-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-cnv-inference", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/single-cell/cnv-inferenceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-cnv-inference -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-single-cell-cnv-inference --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/single-cell/cnv-inference .agents/skills/bio-single-cell-cnv-inference && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-single-cell-cnv-inference" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/cnv-inference into .agents/skills/bio-single-cell-cnv-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-cnv-inference", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-cnv-inference -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-single-cell-cnv-inference --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/single-cell/cnv-inference .cursor/skills/bio-single-cell-cnv-inference && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-single-cell-cnv-inference" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/cnv-inference into .cursor/skills/bio-single-cell-cnv-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-cnv-inference", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path single-cell/cnv-inference--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-cnv-inference -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-single-cell-cnv-inference --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/single-cell/cnv-inference .gemini/skills/bio-single-cell-cnv-inference && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-single-cell-cnv-inference" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/cnv-inference into .gemini/skills/bio-single-cell-cnv-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-cnv-inference", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-single-cell-cnv-inferenceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-cnv-inference -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/single-cell/cnv-inference .github/skills/bio-single-cell-cnv-inference && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-single-cell-cnv-inference" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/cnv-inference into .github/skills/bio-single-cell-cnv-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-cnv-inference", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-single-cell-cnv-inference -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-single-cell-cnv-inference --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/single-cell/cnv-inference .opencode/skills/bio-single-cell-cnv-inference && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-single-cell-cnv-inference" agent skill from https://github.com/GPTomics/bioSkills/tree/main/single-cell/cnv-inference into .opencode/skills/bio-single-cell-cnv-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-single-cell-cnv-inference", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-single-cell-cnv-inferenceInfer large-scale copy-number alterations from tumor single-cell or single-nucleus RNA-seq to separate malignant from normal cells and call subclones, using inferCNV, copyKAT, Numbat, and SCEVAN.
Bio Single Cell Cnv Inference is an agent skill from GPTomics/bioSkills. Infer large-scale copy-number alterations from tumor single-cell or single-nucleus RNA-seq to separate malignant from normal cells and call subclones, using inferCNV, copyKAT, Numbat, and SCEVAN. Use when separating malignant from normal cells in a tumor scRNA-seq dataset, inferring chromosome-arm CNVs or aneuploidy from expression, calling tumor subclones from single cells, choosing a CNV-inference method (reference-based vs reference-free, expression-only vs allele-aware), or deciding which cells are tumor…
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (R), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Single Cell Cnv Inference loads about 4.7k tokens when it runs. Until then it costs about 143 tokens; SKILL.md has 2,127 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,127 words, ~4,654 tokens.
.claude/skills/bio-single-cell-cnv-inference/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: inferCNV 1.18+, copyKAT 1.1+, numbat 1.4+, SCEVAN 1.0+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturespackageVersion('<pkg>') then ?function_name to verify parametersIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Which cells are tumor, and what CNVs and subclones do they carry?" -> Estimate large-scale copy-number from smoothed expression across genomic windows, compare against a normal reference, and cluster cells into malignant vs normal and into subclones.
inferCNV - smooth expression along chromosomes against a defined normal reference, optional HMM for discrete CNV statescopyKAT, SCEVAN - estimate the diploid baseline internally and segment, then classify aneuploid vs diploidNumbat - add phased B-allele frequency for the best subclone and copy-neutral-LOH resolutionAveraged expression over a genomic window is a PROXY for DNA copy number, not a measurement of it. A chromosome-arm gain raises the average expression of the many genes sitting on that arm, and a loss lowers it, so smoothing expression across long runs of contiguous genes reconstructs a coarse copy-number landscape. Three consequences follow and drive every decision. First, the signal is noisy and indirect: it must be smoothed over windows of dozens to hundreds of neighboring genes, and its resolution is chromosome-arm or large-segment (around 5 Mb), never focal genes or exons - this is the hard line separating it from DNA-based copy-number (copy-number/cnvkit-analysis, copy-number/gatk-cnv), which measures DNA read depth and resolves focal amplifications and deletions. Second, the value is RELATIVE: copy-number is only defined against a copy-neutral baseline, so a NORMAL reference is required, and the choice of that reference is the single most consequential decision - a wrong or mismatched reference fabricates CNVs out of ordinary cell-type expression differences. Third, calling a cell malignant is a CLUSTERING decision on this noisy proxy, so it is a hypothesis that needs orthogonal support (lineage markers, mutations, allele evidence), and subclone calls are even softer hypotheses.
Tumor CNV profiles are patient-PRIVATE, so CNV inference runs PER SAMPLE / per patient. Integrating or batch-correcting cells across patients before inference mixes distinct private karyotypes and erases the per-patient signal the analysis depends on (single-cell/batch-integration notes this). Integrate across patients only for shared transcriptional-state analysis, never as input to CNV calling.
CNV-quiet tumors are invisible to this approach. Many hematologic and low-grade tumors carry little large-scale CNV, so expression-based inference returns a near-flat profile - absence of inferred CNV is NOT evidence that cells are normal, and a confident malignant call then requires allele evidence or orthogonal markers. Balanced whole-genome doubling is invisible for the same reason and to all expression methods, not just copyKAT: per-cell library-size normalization removes a uniform ploidy multiple, so a 4N WGD reads copy-neutral against a 2N reference for inferCNV and Numbat's expression channel alike, and only allele or ploidy evidence reveals it.
| Method | Model / assumption | Reference | Allele-aware | Use when | Fails when |
|---|---|---|---|---|---|
| inferCNV | Smooth expression along chromosomes vs a normal reference; optional HMM for discrete states | Reference-based (needs normal cells) | No | A clean in-sample normal reference exists; want interpretable arm-level heatmap + HMM states | No trustworthy reference; CNV-quiet tumor; needs focal resolution |
| copyKAT | Bayesian segmentation + hierarchical clustering + GMM; estimates diploid baseline internally | Reference-free (optional known normals) | No | No reference cells annotated; want a quick aneuploid-vs-diploid call at ~5 Mb | Mostly-aneuploid sample with no diploid baseline; whole-genome doubling confuses the root |
| SCEVAN | Variational multichannel segmentation sharing breakpoints across a clone | Reference-free | No | Want automatic malignant/non-malignant + subclones in one call | Same baseline ambiguity as copyKAT; very sparse data |
| Numbat | Joint expression + phased B-allele frequency + population haplotypes; iterative phylogeny | Reference (expression) + population phasing | Yes | Best subclone resolution needed; copy-neutral LOH matters; allele counts obtainable | No BAM/phasing available; very low SNP coverage (shallow or snRNA) |
Reference-based (inferCNV, Numbat expression side) is the most reliable when a clean normal reference is available; reference-free (copyKAT, SCEVAN) trades that for not needing one but is vulnerable to baseline ambiguity. Expression-only methods (inferCNV, copyKAT, SCEVAN) are simpler and need only counts; the allele-aware method (Numbat) is the most powerful for subclones and copy-neutral events but requires per-cell allele counts and phasing. Run an expression method to get the malignant/normal split, then Numbat when subclone structure is the question, and reconcile. When methods compete, verify current best practice against the installed package documentation before committing to one.
Goal: Pick a copy-neutral baseline that defines what "no CNV" looks like, since the reference choice determines whether the inferred CNVs are real or artifacts.
Approach: Prefer non-malignant cells from the SAME sample (T cells, B cells, myeloid, endothelial, fibroblasts identified by lineage markers) because they share the patient, protocol, and ambient-RNA background; fall back to an external normal only when no in-sample normals exist, and expect batch artifacts.
The reference cells must be confidently non-malignant and abundant enough for a stable baseline mean: aim for tens to hundreds of reference cells, not a handful, because a few cells give a noisy baseline that fabricates CNVs in every observation cell. A reference that is itself a malignant or stressed population, or a single mismatched cell type, will make every other cell look aneuploid relative to it. The reference must also MATCH the tumor's sex: because inferCNV works on expression, X-inactivation dosage-compensates most chrX expression so there is no uniform two-fold chrX shift, but an external or cross-individual normal (including shipped or pooled-donor references) of the opposite sex still fabricates a convincing uniform sex-chromosome CNV through chrY genes present in XY versus near-absent in XX, XIST high in XX versus silent in XY, and the attenuated X-inactivation escape genes - match sexes or drop chrX and chrY before inference. With no annotated normals, use a reference-free method (copyKAT, SCEVAN) rather than guessing a reference; do not pass tumor-contaminated cells as the reference.
Goal: Build a chromosome-ordered expression heatmap against a defined normal reference and call discrete CNV states per region.
Approach: Assemble a raw counts matrix, a cell-to-group annotation file, and a gene-ordering file with genomic coordinates, name the normal groups as the reference, then run with droplet-appropriate cutoff, denoising, and the HMM.
library(infercnv)
infercnv_obj <- CreateInfercnvObject(
raw_counts_matrix = 'counts.matrix',
annotations_file = 'cell_annotations.txt',
delim = '\t',
gene_order_file = 'gene_ordering.txt',
ref_group_names = c('Tcell', 'Myeloid'))
infercnv_obj <- infercnv::run(
infercnv_obj,
cutoff = 0.1,
out_dir = 'infercnv_out',
cluster_by_groups = TRUE,
denoise = TRUE,
HMM = TRUE,
num_threads = 4)ref_group_names lists the normal groups from the annotation file; set it to NULL only when no reference exists (less reliable). cutoff = 0.1 suits 10x and other droplet data; use cutoff = 1 for full-length Smart-seq. cluster_by_groups = TRUE clusters within annotated groups rather than forcing one global tree. The HMM (HMM_type = 'i6' default, six copy states; 'i3' for a simpler deletion/neutral/amplification model) yields per-region discrete states under out_dir.
Goal: Classify cells as aneuploid (tumor) or diploid (normal) without an annotated reference, and obtain a per-cell copy-number matrix.
Approach: Pass the raw gene-by-cell matrix; copyKAT estimates the diploid baseline by segmentation and clustering, then labels each cell, optionally anchored by any known-normal barcodes.
library(copykat)
res <- copykat(
rawmat = exp_rawdata,
id.type = 'S',
ngene.chr = 5,
win.size = 25,
KS.cut = 0.1,
sam.name = 'tumor1',
distance = 'euclidean',
norm.cell.names = '',
genome = 'hg20',
n.cores = 4)
pred <- res$prediction
cna <- res$CNAmatres$prediction$copykat.pred is aneuploid, diploid, or not.defined per cell; res$CNAmat holds smoothed copy-number values in ~220 kb bins (the output bin size; effective detection resolution is still ~5 Mb). KS.cut controls segmentation stringency (raise it for fewer, larger segments). Supplying confident normal barcodes via norm.cell.names anchors the diploid baseline and improves accuracy when the sample is mostly aneuploid.
Goal: Resolve subclones and copy-neutral LOH by combining smoothed expression with phased B-allele frequencies.
Approach: Generate per-cell allele counts and population phasing with the pileup_and_phase.R preprocessing script, build an expression reference from matched normals (or the shipped ref_hca), then run the joint model.
Rscript pileup_and_phase.R --label tumor1 --samples tumor1 \
--bams tumor1.bam --barcodes barcodes.tsv \
--gmap genetic_map_hg38_withX.txt.gz --snpvcf genome1K.phase3.SNP.vcf \
--paneldir 1000G_panel/ --outdir numbat_out/ --ncores 8library(numbat)
ref <- aggregate_counts(count_mat_normal, cell_annot)
out <- run_numbat(
count_mat,
lambdas_ref = ref,
df_allele = df_allele,
genome = 'hg38',
t = 1e-5,
ncores = 4,
plot = TRUE,
out_dir = 'numbat_out')
nb <- Numbat$new(out_dir = 'numbat_out')df_allele is the allele dataframe written by pileup_and_phase.R (columns include cell, snp_id, CHROM, POS, AD, DP, GT). lambdas_ref is a gene-by-cell-type expression reference from aggregate_counts(count_mat, cell_annot) where cell_annot has cell and group columns, or the package-shipped ref_hca. t is the HMM transition probability. The loaded Numbat object exposes clone_post (clone assignments) and per-cell copy-number posteriors. Numbat needs no paired-normal DNA but does need a BAM and phasing reference.
Goal: Get a per-cell malignant-vs-normal label from inferCNV, which (unlike copyKAT and SCEVAN) returns a heatmap and HMM region states but no automatic per-cell class.
Approach: Subcluster the observation cells on their CNV signal and/or compute a per-cell CNV score, then threshold; carry the result onto the cells with add_to_seurat and validate the cut against lineage markers and allele/mutation evidence.
infercnv_obj <- infercnv::run(
infercnv_obj, cutoff = 0.1, out_dir = 'infercnv_out',
cluster_by_groups = FALSE, analysis_mode = 'subclusters',
denoise = TRUE, HMM = TRUE, num_threads = 4)
seurat_obj <- infercnv::add_to_seurat(
seurat_obj = seurat_obj, infercnv_output_path = 'infercnv_out', top_n = 10)
obs <- read.table('infercnv_out/infercnv.observations.txt', header = TRUE, row.names = 1)
cnv_score <- colSums((obs - 1)^2)
malignant <- cnv_score > quantile(cnv_score, 0.5)analysis_mode = 'subclusters' (with cluster_by_groups = FALSE) partitions the observation cells by CNV signal so a malignant subcluster separates from a copy-neutral one. The per-cell CNV score (sum of squared deviation of the denoised infercnv.observations.txt profile from the copy-neutral value, or each cell's correlation to the mean putative-tumor profile) gives a continuous malignancy axis to threshold; add_to_seurat writes per-cell and per-chromosome CNA metadata back onto the object for plotting. The malignant/normal cut is a hypothesis: confirm it against lineage markers (single-cell/cell-annotation) and allele or mutation evidence, never treat the threshold as ground truth.
| Parameter | Tool | Default / typical | Rationale |
|---|---|---|---|
| cutoff | inferCNV | 0.1 (droplet), 1 (Smart-seq) | Drops genes below mean expression; droplet data are sparse so the threshold is lower |
| HMM_type | inferCNV | i6 | Six copy states (0 to 3+) give finer state calls; i3 (del/neutral/amp) is simpler and more robust |
| ref_group_names | inferCNV | named normals | The copy-neutral baseline; NULL only when no reference exists, at a reliability cost |
| ngene.chr | copyKAT | 5 | Minimum genes per chromosome per cell so a cell has signal on each chromosome |
| win.size | copyKAT | 25 | Genes per segment for smoothing; larger windows denoise but blur small events |
| KS.cut | copyKAT | 0.1 | Segmentation stringency; raise for fewer, larger segments in noisy data |
| genome | copyKAT | hg20 | Must match the assembly the gene coordinates came from (hg20 or mm10) |
| t | Numbat | 1e-5 | HMM transition probability; lower favors longer segments |
| resolution (~5 Mb) | all expression methods | ~5 Mb | The floor of expression-based CNV; focal events below this are invisible |
| Symptom | Cause | Fix |
|---|---|---|
| Every cell looks aneuploid, including the immune cells | Reference cells are tumor-contaminated or a mismatched cell type | Re-pick confident non-malignant in-sample normals by lineage markers; do not use a stressed/malignant population as reference |
| Flat profile, no CNVs called, but cells are clearly tumor | CNV-quiet / low-grade tumor, or snRNA sparsity | Absence of CNV is not normality; switch to allele-aware Numbat or confirm malignancy with markers/mutations |
| copyKAT cannot find a diploid baseline / labels almost all aneuploid | Mostly-aneuploid sample with no internal diploid cells, or balanced whole-genome doubling that library-size normalization hides | Supply known-normal barcodes via norm.cell.names, or use inferCNV with an external reference; confirm ploidy with allele/DNA evidence since balanced WGD looks copy-neutral to all expression methods |
| Uniform chrX gain/loss and a chrY call across all tumor cells | Reference and tumor are opposite sex (external/shipped/pooled-donor reference) | Match reference and tumor sex, or drop chrX and chrY before inference |
| Recurrent localized segment over 6p, 14q, 22q, or 2p that tracks lineage | High, variable HLA (MHC 6p) and immunoglobulin/TCR expression clusters genomically and mimics a segment, especially with immune references | Mask or distrust segments over HLA and Ig/TCR loci; do not call them CNV |
| Stripe artifacts that track proliferating cells | Cell-cycle and high-expression gene programs mimic CNV | Account for / regress cell cycle, enable denoising, and treat cycle-correlated bands skeptically |
| Subclones from inferCNV/copyKAT do not replicate | Expression-only subclone calls are weak hypotheses | Confirm with Numbat allele evidence or DNA; report subclones as hypotheses |
| CNV signal vanishes after integrating samples | Cross-patient integration erased patient-private karyotypes | Run CNV inference per sample BEFORE any cross-patient integration |
| Genes silently dropped / wrong chromosome bands | Gene-order file or genome build mismatched to the counts | Match gene_order_file and genome to the same assembly and gene IDs as the matrix |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in single-cell/cnv-inference of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Single Cell Cnv Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Single Cell Cnv Inference this skillGPTomics/bioSkills | 1.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Infer large-scale copy-number alterations from tumor single-cell or single-nucleus RNA-seq to separate malignant from normal cells and call subclones, using inferCNV, copyKAT, Numbat, and SCEVAN. Bio Single Cell Cnv Inference is an agent skill from GPTomics/bioSkills. Infer large-scale copy-number alterations from tumor single-cell or single-nucleus RNA-seq to separate malignant from normal cells and call subclones, using inferCNV, copyKAT, Numbat, and SCEVAN.
Bio Single Cell Cnv Inference fits situations like: separating malignant from normal cells in a tumor scRNA-seq dataset; inferring chromosome-arm CNVs; aneuploidy from expression; calling tumor subclones from single cells.
Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-cnv-inference -a claude-code`. Or copy the skill folder (single-cell/cnv-inference in GPTomics/bioSkills) into .claude/skills/bio-single-cell-cnv-inference in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-single-cell-cnv-inference -a codex`. Or copy the skill folder (single-cell/cnv-inference in GPTomics/bioSkills) into .agents/skills/bio-single-cell-cnv-inference in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-single-cell-cnv-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-single-cell-cnv-inference, .gemini/skills/bio-single-cell-cnv-inference, .github/skills/bio-single-cell-cnv-inference and .opencode/skills/bio-single-cell-cnv-inference in your project.
Going by SKILL.md and its folder, Bio Single Cell Cnv Inference needs R for the scripts in its folder and the command-line tools its instructions call (pip).
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Single Cell Cnv Inference is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Single Cell Cnv Inference: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.