Agent skill

Bio Atac Seq Single Cell Atac

by GPTomics in GPTomics/bioSkills

Process and analyze single-cell ATAC-seq data with Signac, ArchR, SnapATAC2, or Cell Ranger ATAC.

MITAuto-check passedResearch & Science

Install Bio Atac Seq Single Cell Atac

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-atac-seq-single-cell-atac -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-atac-seq-single-cell-atac --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/atac-seq/single-cell-atac .claude/skills/bio-atac-seq-single-cell-atac && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-atac-seq-single-cell-atac
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6k tokens
SKILL.md length
2,246 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Process and analyze single-cell ATAC-seq data with Signac, ArchR, SnapATAC2, or Cell Ranger ATAC.

  • Handling 10X scATAC
  • SKILL.md covers Version Compatibility, Ecosystem Choice (The Most…, 10X Multiome Caveat (Paired… and Per-Cell QC Thresholds, plus 18 more sections
  • Runs R scripts from its folder; calls pip
  • 10X Multiome (paired RNA+ATAC) data

What it does

Bio Atac Seq Single Cell Atac is an agent skill from GPTomics/bioSkills. Process and analyze single-cell ATAC-seq data with Signac, ArchR, SnapATAC2, or Cell Ranger ATAC. Use when handling 10X scATAC or 10X Multiome (paired RNA+ATAC) data, performing per-cell QC, choosing between ArchR/Signac/SnapATAC2 ecosystems, building per-cluster consensus peaksets, integrating with paired scRNA-seq, doublet detection (AMULET vs ArchR vs scDblFinder), or running pseudobulk differential accessibility per cluster.

Its SKILL.md is about 6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Handling 10X scATAC
  • 10X Multiome (paired RNA+ATAC) data
  • Performing per-cell QC
  • Choosing between ArchR/Signac/SnapATAC2 ecosystems

Example prompts

  • “/bio-atac-seq-single-cell-atac”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Atac Seq Single Cell Atac loads about 6k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 2,246 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,246 words, ~5,974 tokens.

Download SKILL.mdSave it as .claude/skills/bio-atac-seq-single-cell-atac/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-atac-seq-single-cell-atac
description
Process and analyze single-cell ATAC-seq data with Signac, ArchR, SnapATAC2, or Cell Ranger ATAC. Use when handling 10X scATAC or 10X Multiome (paired RNA+ATAC) data, performing per-cell QC, choosing between ArchR/Signac/SnapATAC2 ecosystems, building per-cluster consensus peaksets, integrating with paired scRNA-seq, doublet detection (AMULET vs ArchR vs scDblFinder), or running pseudobulk differential accessibility per cluster.
tool_type
mixed
primary_tool
Signac

Version Compatibility

Reference examples tested with: Cell Ranger ATAC 2.1+, Signac 1.13+, Seurat 5.0+, ArchR 1.0.2+, SnapATAC2 2.8+, AMULET 1.1+, scDblFinder 1.16+, scater 1.30+, scvi-tools 1.1+, GenomicRanges 1.54+, JASPAR2024 0.99+, BSgenome.Hsapiens.UCSC.hg38 1.4+, EnsDb.Hsapiens.v86 2.99+, MACS3 3.0+. SnapATAC2 2.8+ uses pp.import_fragments; older 2.5-2.7 used pp.import_data (renamed/removed in 2.9).

Verify before use:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws unexpected errors, introspect the installed package and adapt rather than retrying.

Single-Cell ATAC-seq

"Process my 10X scATAC data from cellranger output" -> Build a per-cell fragment matrix, compute per-cell QC, dimensionality reduction (TF-IDF + LSI / spectral / autoencoder), cluster, call cluster-level pseudobulk peaks, annotate cell types via gene-activity scores, and integrate with paired scRNA-seq if Multiome.

  • R: Signac::CreateChromatinAssay() -> Seurat workflow (TF-IDF + SVD + UMAP + Leiden)
  • R: ArchR::createArrowFiles() -> ArchR project (TileMatrix + LSI + UMAP)
  • Python: snapatac2.pp.import_fragments() -> SnapATAC2 (spectral / diffusion-map clustering)
  • CLI (preprocessing): cellranger-atac count (10X) or chromap (alignment-only fragment files)

Ecosystem Choice (The Most Important Decision)

EcosystemLanguageStrengthFails whenBest for
Signac (Stuart 2021)R, Seurat-basedTightest scRNA-seq integration; Seurat ecosystem matureMemory hungry on >100K cells; slower than ArchRMultiome RNA+ATAC; small-to-medium datasets; Seurat user
ArchR (Granja 2021)R, Arrow/HDF5Memory-efficient (Arrow files); fast on 100K-1M cells; built-in trajectory + doubletLess RNA-seq integration; ArchR-specific formatLarge bulk-cohort scATAC; trajectory analysis; ATAC-only
SnapATAC2 (Zhang 2024)Python, AnnDataMemory-efficient; modern Python ecosystem; spectral clustering performantNewer; benchmarks evolving; ecosystem smaller than RPython-first labs; very large datasets (>1M cells)
Cell Ranger ATACCLI (10X-specific)10X official preprocessingClosed; fixed pipelineOnly as preprocessing step; analysis happens elsewhere
scATAC-proCLI-based pipelineAlternative preprocessingLess maintainedLegacy; not recommended for new projects

Methodology evolves; verify against Granja 2021 (ArchR), Stuart 2021 (Signac), Zhang 2024 (SnapATAC2), Heumos 2023 (best practices) before locking pipelines.

10X Multiome Caveat (Paired RNA + ATAC)

10X Multiome chemistry profiles RNA AND ATAC from the same cell. Outputs are joined by a shared barcode. Multiome workflows use Signac for ATAC and Seurat for RNA, integrated through the same Seurat object via WNN (Weighted Nearest Neighbor; Hao 2021).

Single-modality 10X scATAC (chemistry v1, v2) does NOT produce paired RNA. Verify the chemistry on the cellranger summary before assuming Multiome.

Per-Cell QC Thresholds

MetricDefinitionPassCautionRejectSource
Fragment count per celln_fragments after dedup3000-500001000-3000< 1000 or > 8000010X recommendation; high = doublet
TSS enrichment per cellSignal at TSS / flanks (per cell)>= 42-4< 2ArchR / Signac default; lower than bulk because per-cell
Nucleosome signalMono / NFR fragment ratio<= 44-10> 10Signac default; high = poor library
% reads in peaks (per cell)FRiP per cell at consensus>= 0.150.10-0.15< 0.05ArchR / Signac defaults
Mitochondrial fraction (per cell)chrM / total per cell< 0.050.05-0.15> 0.20Lower than bulk; per-cell more sensitive
Doublet scoreAMULET / ArchR doublet< 0.50.5-0.7> 0.7Tool-dependent threshold
Blacklist ratioReads in ENCODE blacklist / total< 0.050.05-0.10> 0.10Standard

Per-cell thresholds are looser than bulk because individual cells have orders of magnitude less signal; the population aggregate is what matters.

Doublet Detection: Three Approaches

ToolMethodStrengthFails when
AMULET (Thibodeau 2021)Collision-based: detects cells with too many fragments at same position (mathematically impossible from single cell because of 2-allele limit)Specific to ATAC biology; orthogonal to clusteringRequires high depth; recall drops sharply below ~15-16K valid read pairs/cell (peaks ~90% near ~25K)
ArchR addDoubletScoresSynthetic doublet simulation + projection into LSIBuilt into ArchR; auto-thresholdsTied to ArchR's LSI; not portable
scDblFinder (Germain 2021)Synthetic doublets + classifierWorks on Signac and SCE objects; well-benchmarkedRNA-developed; ATAC adaptation requires careful settings

Operational rule: Run AMULET as the primary ATAC-specific check; verify with ArchR or scDblFinder as orthogonal evidence. Double-flagged cells are high-confidence doublets.

Per-Tool Failure Modes

Signac TF-IDF + SVD -- First component is depth

Trigger: Running RunTFIDF -> RunSVD and using all components for clustering.

Mechanism: Component 1 of LSI is highly correlated with sequencing depth per cell, not biology. Including it pulls clusters along depth axis.

Symptom: UMAP shows linear cell-density gradient that tracks fragment count.

Fix: Skip component 1 in downstream UMAP and clustering: RunUMAP(object, dims=2:30), FindNeighbors(object, dims=2:30). ArchR and SnapATAC2 do this automatically.

ArchR -- TileMatrix vs PeakMatrix confusion

Trigger: Using TileMatrix for differential testing or motif analysis.

Mechanism: TileMatrix is regular fixed bins (default 500 bp) covering the whole genome; useful for clustering but not for biology because ATAC signal is at peaks, not arbitrary bins.

Fix: Use TileMatrix for embedding/clustering (it's faster); use PeakMatrix (after addReproduciblePeakSet) for differential, motif, and gene-activity analysis.

SnapATAC2 -- Memory layout assumes integer counts

Trigger: Loading non-integer or negative-valued matrices.

Fix: SnapATAC2 expects raw fragment counts (Int32). Convert from float matrices before loading.

Cell Ranger ATAC -- Empty droplet detection

Trigger: Default cellranger-atac calls cells based on UMI count but ATAC has no UMIs; uses fragment-based heuristic. Sometimes calls < 1000-fragment cells as real.

Fix: Re-filter cellranger output to require fragment count >= 1000 AND TSS enrichment >= 4 BEFORE downstream analysis.

Multiome WNN -- ATAC weighting

Trigger: Default WNN equally weights RNA and ATAC modalities.

Mechanism: ATAC's per-cell signal is much sparser than RNA; equal weighting can swamp the joint embedding with ATAC noise.

Fix: Inspect per-modality weights with IntegrateLayers; if ATAC weights dominate noise, manually adjust or use FindMultiModalNeighbors carefully.

Per-cluster pseudobulk peak calling -- Empty clusters

Trigger: Calling MACS3 per cluster when one cluster has < 200 cells.

Mechanism: MACS3 needs >= 1M reads in pseudobulk to call peaks reliably; small clusters do not produce enough reads.

Fix: Aggregate small clusters into a "rare" group OR drop them from per-cluster calling. Use the union of larger-cluster peaks as a fallback for rare-cell-type analysis.

Decision Tree by Goal

GoalRecommended pipeline
Standard 10X scATAC analysis (R user)Signac: CreateChromatinAssay -> RunTFIDF -> RunSVD -> RunUMAP (dims 2:30) -> FindClusters
Standard 10X scATAC analysis (Python user)SnapATAC2: pp.import_fragments -> add_tile_matrix -> spectral -> UMAP -> leiden
Large dataset (> 100K cells)ArchR (memory-efficient Arrow files)
10X Multiome (paired RNA + ATAC)Signac + Seurat: per-modality embedding then WNN integration
Trajectory / pseudotimeArchR getTrajectory; or Signac + Cicero
Differential accessibility per clusterPseudobulk per cluster -> consensus peakset -> DESeq2
Cell-type annotationGene-activity scores via ArchR or Signac; then run single-cell/markers-annotation
Multimodal trajectories (RNA + ATAC)MOFA+, ArchR + scRNA integration, or SCENIC+
Plant / non-modelSignac with custom EnsDb / TxDb; ArchR with custom annotations

Standard Signac Workflow

Goal: Process 10X scATAC output into a clustered, annotated Seurat object with per-cell QC, embedding, and gene-activity scores.

Approach: Load fragments and peak counts into a ChromatinAssay, compute per-cell QC (TSS enrichment, nucleosome signal, FRiP, blacklist), subset to passing cells, run TF-IDF + SVD + UMAP skipping LSI component 1, then derive a gene-activity assay for marker-based annotation.

r
library(Signac); library(Seurat); library(EnsDb.Hsapiens.v86)
library(BSgenome.Hsapiens.UCSC.hg38)

# 1. Load 10X output
counts <- Read10X_h5('outs/filtered_peak_bc_matrix.h5')
metadata <- read.csv('outs/singlecell.csv', row.names=1)

chrom_assay <- CreateChromatinAssay(
    counts=counts,
    sep=c(':', '-'),
    fragments='outs/fragments.tsv.gz',
    annotation=GetGRangesFromEnsDb(EnsDb.Hsapiens.v86),
    genome='hg38')

obj <- CreateSeuratObject(counts=chrom_assay, assay='ATAC', meta.data=metadata)

# 2. Per-cell QC
obj <- NucleosomeSignal(obj)              # Mono/NFR ratio
obj <- TSSEnrichment(obj, fast=FALSE)
obj$pct_reads_in_peaks <- obj$peak_region_fragments / obj$passed_filters * 100
obj$blacklist_ratio <- obj$blacklist_region_fragments / obj$peak_region_fragments

# 3. QC filter (looser per-cell thresholds)
obj <- subset(obj,
    subset = peak_region_fragments > 1000 & peak_region_fragments < 20000 &
             pct_reads_in_peaks > 15 & blacklist_ratio < 0.05 &
             nucleosome_signal < 4 & TSS.enrichment > 4)

# 4. Dimensionality reduction (skip component 1 - it's depth)
obj <- RunTFIDF(obj)
obj <- FindTopFeatures(obj, min.cutoff='q0')
obj <- RunSVD(obj)
obj <- RunUMAP(obj, reduction='lsi', dims=2:30)
obj <- FindNeighbors(obj, reduction='lsi', dims=2:30)
obj <- FindClusters(obj, algorithm=4, resolution=0.5)    # Leiden

# 5. Gene-activity for annotation
gene.activities <- GeneActivity(obj)
obj[['ACT']] <- CreateAssayObject(counts=gene.activities)
DefaultAssay(obj) <- 'ACT'
obj <- NormalizeData(obj, normalization.method='LogNormalize',
                     scale.factor=median(obj$nCount_ACT))

Per-Cluster Pseudobulk Peak Calling

r
# Signac wrapper around MACS3
peaks <- CallPeaks(obj, group.by='seurat_clusters',
                   macs2.path='/path/to/macs3', cleanup=FALSE,
                   format='BED', shift=-75, extsize=150,   # Tn5 cut-site recipe; shift only applies to format='BED'
                   additional.args='-p 0.01')
# Then iterative-overlap consensus across clusters (atac-seq/consensus-peakset)

ArchR Workflow

Goal: Build an ArchR project from fragment files, filter doublets, cluster, and call per-cluster reproducible peaks.

Approach: Create Arrow files with minimum TSS and fragment thresholds, score doublets and filter, run iterative LSI + UMAP + clustering on the TileMatrix, then build per-cluster group coverages and a reproducible peakset via MACS3.

r
library(ArchR)
addArchRGenome('hg38')

# 1. Create Arrow files from fragment files
ArrowFiles <- createArrowFiles(
    inputFiles=c('fragments_rep1.tsv.gz', 'fragments_rep2.tsv.gz'),
    sampleNames=c('rep1', 'rep2'),
    minTSS=4, minFrags=1000,
    addTileMat=TRUE, addGeneScoreMat=TRUE)

# 2. Doublet detection (built-in)
doubletScores <- addDoubletScores(input=ArrowFiles, k=10, knnMethod='UMAP')

# 3. Project + filter
proj <- ArchRProject(ArrowFiles=ArrowFiles, outputDirectory='ArchR_out')
proj <- filterDoublets(proj)

# 4. LSI + UMAP + clustering
proj <- addIterativeLSI(proj, useMatrix='TileMatrix', name='IterativeLSI')
proj <- addClusters(proj, reducedDims='IterativeLSI', method='Seurat', resolution=0.5)
proj <- addUMAP(proj, reducedDims='IterativeLSI')

# 5. Reproducible peakset (per cluster)
proj <- addGroupCoverages(proj, groupBy='Clusters')
proj <- addReproduciblePeakSet(proj, groupBy='Clusters', pathToMacs2='/path/macs3')
proj <- addPeakMatrix(proj)

SnapATAC2 Workflow (Python)

Goal: Run a Python-native scATAC pipeline from fragments through clusters, per-cluster peaks, and gene activity.

Approach: Import fragments to a backed AnnData, compute per-cell TSS enrichment and fragment-size metrics, filter cells, build a tile matrix, run spectral embedding + UMAP + Leiden, call per-cluster peaks via MACS3, and derive a gene-activity matrix.

python
import snapatac2 as snap

# 1. Load 10X fragments. SnapATAC2 uses snap.pp.import_fragments (NOT snap.read_10x, which doesn't exist).
data = snap.pp.import_fragments(
    fragment_file='outs/fragments.tsv.gz',
    chrom_sizes=snap.genome.hg38,
    file='out.h5ad',                        # backed AnnData; backend handled by file path
    sorted_by_barcode=False)
data.obs['sample_id'] = 'rep1'

# 2. Per-cell QC
snap.metrics.tsse(data, gene_anno=snap.genome.hg38)
snap.metrics.frag_size_distr(data)
snap.pp.filter_cells(data, min_counts=1000, min_tsse=4)

# 3. Tile matrix + spectral
snap.pp.add_tile_matrix(data, bin_size=500)
snap.pp.select_features(data, n_features=250000)
snap.tl.spectral(data)
snap.tl.umap(data)
snap.tl.leiden(data)

# 4. Per-cluster peak calling (uses MACS3)
snap.tl.macs3(data, groupby='leiden')

# 5. Gene activity (gene score) for annotation
gene_mat = snap.pp.make_gene_matrix(data, gene_anno=snap.genome.hg38)

Reconciliation Across Tools

PatternLikely causeAction
Signac UMAP shows tight clusters; ArchR UMAP shows looseDifferent LSI implementation; ArchR uses iterative LSI by defaultBoth valid; biology should match in cluster labels
Different tools call different doublet ratesDifferent algorithms (collision vs simulation)Use intersection (cells flagged by 2+ tools) as high-confidence doublets
Cluster boundaries differDifferent clustering algorithm or resolutionStandardize on Leiden algorithm 4 with same resolution
Per-cluster peak count differsDifferent peak callers or pseudobulk depthEnsure same MACS3 parameters; pool small clusters

Operational rule: For high-confidence cell-type annotation, agree across two ecosystems (e.g., Signac + ArchR clusters) and report agreement metrics.

Show full SKILL.md (943 more words)Show less

cellranger-atac vs cellranger-arc

PipelineUse forOutput
cellranger-atacSingle-modality 10X scATAC (chemistry v1, v2)per-cell ATAC barcodes; fragments.tsv.gz; per-barcode metadata
cellranger-arc10X Multiome (paired RNA + ATAC same cell, "Multiome chemistry")joint barcodes for RNA + ATAC; separate fragment / count files; one barcode whitelist

Trigger: Loading 10X output without checking which pipeline produced it.

Mechanism: cellranger-atac and cellranger-arc produce different output structures. cellranger-arc fragments.tsv has barcodes paired with the RNA matrix; cellranger-atac fragments are ATAC-only with their own barcode universe.

Fix: Verify the chemistry on the cellranger summary (look for "Multiome" in the run config). Use Read10X_h5 for cellranger-arc Multiome RNA output; use CreateChromatinAssay with the matched fragments for the ATAC. Mixing barcodes across pipelines fails silently.

Cell-Cycle Correction for scATAC

Trigger: Proliferating cell types in the dataset; cells distributed across cell cycle phases.

Mechanism: Replication-associated chromatin opening adds 5-10% global accessibility per cell as it moves G1 -> S -> G2/M; this confounds clustering and DA.

Detection: Score cells with a chromatin-adapted S-phase signature (regulated origin loci, replication-stress-response gene accessibility) analogous to Seurat's CellCycleScoring on RNA. Or compute a Repli-seq peak overlap score.

Fix: Regress S-phase score on the TF-IDF residuals before downstream LSI: ScaleData(obj, vars.to.regress='S.score') analogous to RNA workflow. For DA between cell-cycle-mismatched conditions, add S-phase as covariate in pseudobulk DESeq2.

Sex-Chromosome QC for scATAC

Trigger: Mixed-sex donors in the dataset; sample-swap detection.

Mechanism: XIST locus accessibility is high in female cells (X-inactivation); chrY peak count is essentially zero in females. Per-cell or per-sample, the XIST/chrY ratio identifies sex unambiguously.

Detection: In a per-cell counts matrix, compute fraction of fragments at XIST locus (chrX:73820651-73852753 hg38) and chrY peak count; classify cells; flag cells/samples where assignment disagrees with metadata.

XCI escapees: Genes that escape X-inactivation (KDM6A, DDX3X, EIF1AX) are biallelically accessible in female cells but not male; can be used as additional sex confirmation if XIST is ambiguous.

scArches Reference Mapping

Trigger: Projecting query scATAC onto a reference atlas; cross-study integration without batch effects.

Mechanism: scArches (Lotfollahi 2022) provides transfer-learning to project a new dataset onto an existing reference's latent space without retraining the reference. For ATAC, the relevant model is PEAKVI (Ashuach 2022). PEAKVI lives in scvi-tools (scvi.model.PEAKVI); the load_query_data classmethod implements the scArches algorithm directly, so an explicit import scarches is not required.

python
import scvi

# Pre-trained reference model (e.g. PBMC scATAC atlas)
query_model = scvi.model.PEAKVI.load_query_data(adata_query, reference_path)
query_model.train(max_epochs=200)
adata_query.obsm['X_emb'] = query_model.get_latent_representation()

Reference atlases for ATAC are still emerging; the most-developed are PBMC (Granja 2021) and brain (BRAIN Initiative).

chromBPNet Per-Cluster Pseudobulk

Trigger: Cell-type-specific variant effect prediction; per-cluster bias-corrected calling.

Mechanism: Train chromBPNet (atac-seq/deep-learning-atac) per pseudobulk cluster; outputs are bias-corrected per-base profiles + variant effect predictions specific to that cell type.

Workflow: Aggregate fragments per cluster into pseudobulk BAMs; run chromBPNet pipeline per cluster (~24h GPU per cluster); use the resulting model for in silico variant scoring at GWAS / rare-variant SNPs in that cell type.

For most studies, this is reserved for top 5-10 priority clusters; running chromBPNet per cluster on >20 cell types is computationally heavy.

AMULET Depth Threshold

AMULET reaches maximum recall (~90%) near ~25,000 valid read pairs (~fragments) per cell and degrades sharply below ~15,000-16,000, where the collision-based detection has insufficient power: the expected number of collisions per cell is too small for the binomial test to distinguish doublet from singleton. For lower-depth libraries, use synthetic-doublet methods (ArchR addDoubletScores, scDblFinder) instead.

Multiome WNN Integration (Signac)

Goal: Build a joint RNA + ATAC embedding from a 10X Multiome dataset using Weighted Nearest Neighbors.

Approach: Run per-modality embeddings (PCA on RNA, TF-IDF + SVD on ATAC skipping LSI-1), then use FindMultiModalNeighbors to learn per-cell modality weights and project a joint UMAP.

r
# Assume `obj` has both 'RNA' and 'ATAC' assays from same Multiome experiment
DefaultAssay(obj) <- 'RNA'
obj <- NormalizeData(obj) %>% FindVariableFeatures() %>% ScaleData() %>% RunPCA()

DefaultAssay(obj) <- 'ATAC'
obj <- RunTFIDF(obj) %>% FindTopFeatures(min.cutoff='q0') %>% RunSVD()

# Joint embedding
obj <- FindMultiModalNeighbors(obj, reduction.list=list('pca', 'lsi'),
                               dims.list=list(1:30, 2:30))
obj <- RunUMAP(obj, nn.name='weighted.nn', reduction.name='wnn.umap')

Common Errors

Error / symptomCauseSolution
UMAP shows depth gradientLSI component 1 included in clusteringdims=2:30 instead of 1:30
Cell Ranger output has many low-quality cellscellranger ATAC uses lenient cell callingRe-filter at fragment count >= 1000 + TSS enrichment >= 4
ArchR "TileMatrix not found"Forgot addTileMat=TRUE in createArrowFilesRe-create Arrow files with the flag
Signac CreateChromatinAssay fails on missing fragments filePath to fragments.tsv.gz incorrect; or missing tabix indexProvide full path; run tabix -p bed fragments.tsv.gz
MACS3 fails on tiny pseudobulkCluster too smallUse cluster aggregation; require >= 200 cells per cluster
AMULET reports 100% doubletsThreshold mis-set or input is technical replicatesCheck fragment-count distribution; AMULET needs high depth (recall peaks near ~25K valid read pairs/cell)
Multiome WNN clusters dominated by ATAC noiseEqual modality weightingInspect modality weights; manually adjust if needed
chromVAR / motif assay all NARun before peakset finalizedRe-run AddMotifs / RunChromVAR after peaks stable
EnsDb / BSgenome version mismatchhg38 BSgenome with wrong-build EnsDbMatch builds; EnsDb.Hsapiens.v86 is GRCh38 (Ensembl v86); use EnsDb.Hsapiens.v75 for hg19. Newer hg38 EnsDb releases (v98+) reflect newer GENCODE annotations

References

  • Stuart T et al 2021 Nat Methods 18:1333 (Signac)
  • Granja JM et al 2021 Nat Genet 53:403 (ArchR)
  • Zhang K et al 2024 Nat Methods 21:217 (SnapATAC2)
  • Hao Y et al 2021 Cell 184:3573 (Seurat WNN)
  • Thibodeau A et al 2021 Genome Biol 22:252 (AMULET doublet detection)
  • Germain PL et al 2021 F1000Res 10:979 (scDblFinder)
  • Cusanovich DA et al 2015 Science 348:910 (sciATAC; LSI for sc data)
  • Chen H et al 2019 Genome Biol 20:241 (scATAC analysis benchmark)
  • Heumos L et al 2023 Nat Rev Genet 24:550 (Best practices for single-cell)
  • 10X Genomics Cell Ranger ATAC documentation
  • atac-seq/atac-qc - Bulk QC patterns adapted for per-cell
  • atac-seq/atac-peak-calling - Pseudobulk peak calling per cluster
  • atac-seq/consensus-peakset - Across-cluster consensus
  • atac-seq/differential-accessibility - Pseudobulk DA per cluster
  • atac-seq/motif-deviation - chromVAR for per-cell TF activity
  • atac-seq/footprinting - scprinter for sc footprinting
  • atac-seq/co-accessibility - Cicero for cis-regulatory connections
  • atac-seq/deep-learning-atac - chromBPNet / scBasset for per-cluster bias correction and variant effects
  • atac-seq/enhancer-gene-linking - Per-cell-type enhancer-gene maps
  • atac-seq/allele-specific-accessibility - sc allelic imbalance for cis-effects
  • single-cell/preprocessing - General sc QC patterns
  • single-cell/clustering - Cluster definition
  • single-cell/cell-annotation - Automated reference-based label transfer
  • single-cell/multimodal-integration - Multiome RNA+ATAC integration
  • single-cell/scatac-analysis - Cross-reference single-cell ATAC-specific patterns
  • single-cell/batch-integration - scArches reference mapping; Harmony

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in atac-seq/single-cell-atac of GPTomics/bioSkills.

  • SKILL.md
  • examples/signac_workflow.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Atac Seq Single Cell Atac next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Atac Seq Single Cell Atac compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Atac Seq Single Cell Atac this skillGPTomics/bioSkills1.2k2 repos~6kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Atac Seq Single Cell Atac

What does Bio Atac Seq Single Cell Atac do?

Process and analyze single-cell ATAC-seq data with Signac, ArchR, SnapATAC2, or Cell Ranger ATAC. Bio Atac Seq Single Cell Atac is an agent skill from GPTomics/bioSkills. Process and analyze single-cell ATAC-seq data with Signac, ArchR, SnapATAC2, or Cell Ranger ATAC.

When should I use Bio Atac Seq Single Cell Atac?

Bio Atac Seq Single Cell Atac fits situations like: handling 10X scATAC; 10X Multiome (paired RNA+ATAC) data; performing per-cell QC; choosing between ArchR/Signac/SnapATAC2 ecosystems.

How do I install Bio Atac Seq Single Cell Atac in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-atac-seq-single-cell-atac -a claude-code`. Or copy the skill folder (atac-seq/single-cell-atac in GPTomics/bioSkills) into .claude/skills/bio-atac-seq-single-cell-atac in your project. Claude Code loads it when a task matches its description.

How do I install Bio Atac Seq Single Cell Atac in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-atac-seq-single-cell-atac -a codex`. Or copy the skill folder (atac-seq/single-cell-atac in GPTomics/bioSkills) into .agents/skills/bio-atac-seq-single-cell-atac in your project. Codex loads it when a task matches its description.

Can I use Bio Atac Seq Single Cell Atac in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-atac-seq-single-cell-atac -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-atac-seq-single-cell-atac, .gemini/skills/bio-atac-seq-single-cell-atac, .github/skills/bio-atac-seq-single-cell-atac and .opencode/skills/bio-atac-seq-single-cell-atac in your project.

What does Bio Atac Seq Single Cell Atac need to run?

Going by SKILL.md and its folder, Bio Atac Seq Single Cell Atac needs R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Atac Seq Single Cell Atac access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Atac Seq Single Cell Atac safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Atac Seq Single Cell Atac use?

Bio Atac Seq Single Cell Atac is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Atac Seq Single Cell Atac use?

About 6k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Atac Seq Single Cell Atac?

Skills that share tags, products or a category with Bio Atac Seq Single Cell Atac: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Atac Seq Single Cell Atac?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.