Agent skill

Bio Expression Matrix Sparse Handling

by GPTomics in GPTomics/bioSkills

Stores and operates on sparse expression matrices for single-cell and large bulk RNA-seq, covering dgCMatrix/dgRMatrix/dgTMatrix when-each-is-fast, the dgCMatrix (CSC, R) <- CSR (Python) implicit…

MITAuto-check passedResearch & Science

Install Bio Expression Matrix Sparse Handling

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-expression-matrix-sparse-handling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-expression-matrix-sparse-handling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/expression-matrix/sparse-handling .claude/skills/bio-expression-matrix-sparse-handling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-expression-matrix-sparse-handling
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.6k tokens
SKILL.md length
2,178 words
Files
3
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Stores and operates on sparse expression matrices for single-cell and large bulk RNA-seq, covering dgCMatrix/dgRMatrix/dgTMatrix when-each-is-fast, the dgCMatrix (CSC, R) <- CSR (Python) implicit…

  • Works in 2 steps: Sparse becomes inefficient above ~10-15%… → AnnData backed mode quietly differs from…
  • Choosing sparse format
  • SKILL.md covers Version Compatibility, The Single Most Important…, Algorithmic Taxonomy and Decision Tree by Scenario, plus 16 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Bio Expression Matrix Sparse Handling is an agent skill from GPTomics/bioSkills. Stores and operates on sparse expression matrices for single-cell and large bulk RNA-seq, covering dgCMatrix/dgRMatrix/dgTMatrix when-each-is-fast, the dgCMatrix (CSC, R) <- CSR (Python) implicit transpose, AnnData (cells-rows) <- SingleCellExperiment (cells-cols) orientation flip, HDF5/h5ad vs Zarr cloud-native shift, HDF5SummarizedExperiment + DelayedArray for out-of-memory bulk, scanpy backed mode for large h5ad, the ~10-15% density crossover where dense beats sparse, 10X format proliferation (MTX vs…

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/sparse_operations.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and DataFrames. It works with Python, Zarr, AnnData and Scanpy. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Choosing sparse format
  • Working with single-cell-sized matrices
  • Importing/exporting 10X
  • Debugging R/Python interop transposes

Example prompts

  • “Use the bio-expression-matrix-sparse-handling skill to store and operates on sparse expression matrices for single-cell and large bulk RNA-seq…”
  • “/bio-expression-matrix-sparse-handling”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Sparse becomes inefficient above ~10-15% density. Sparse iteration has cache-unfriendly indirection; dense iteration is sequential. Above…
  2. AnnData backed mode quietly differs from in-memory in important ways. sc.read_h5ad('file.h5ad', backed='r') returns an AnnData where .X is…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Expression Matrix Sparse Handling loads about 5.6k tokens when it runs. Until then it costs about 217 tokens; SKILL.md has 2,178 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~217
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,178 words, ~5,606 tokens.

Download SKILL.mdSave it as .claude/skills/bio-expression-matrix-sparse-handling/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-expression-matrix-sparse-handling
description
Stores and operates on sparse expression matrices for single-cell and large bulk RNA-seq, covering dgCMatrix/dgRMatrix/dgTMatrix when-each-is-fast, the dgCMatrix (CSC, R) <-> CSR (Python) implicit transpose, AnnData (cells-rows) <-> SingleCellExperiment (cells-cols) orientation flip, HDF5/h5ad vs Zarr cloud-native shift, HDF5SummarizedExperiment + DelayedArray for out-of-memory bulk, scanpy backed mode for large h5ad, the ~10-15% density crossover where dense beats sparse, 10X format proliferation (MTX vs CellRanger H5 vs h5ad), the dense-conversion memory blow-up, and Dask + Zarr for consortium-scale matrices. Use when choosing sparse format, working with single-cell-sized matrices, importing/exporting 10X, debugging R/Python interop transposes, processing matrices too large for RAM, or building cloud-native pipelines.
tool_type
python
primary_tool
scipy.sparse

Version Compatibility

Reference examples tested with: numpy 1.26+, scipy 1.12+, pandas 2.2+, anndata 0.10+, scanpy 1.10+, Matrix R package 1.6+, HDF5Array 1.30+ (Bioconductor), DelayedArray 0.28+, zellkonverter 1.12+, zarr-python 2.18+, dask 2024.1+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Sparse Matrix Handling

"Store / compute on a single-cell matrix without blowing up memory" -> Pick the sparse format that matches the access pattern (CSC for column ops, CSR for row ops), respect the R/Python convention difference (Bioconductor stores cells in columns; AnnData stores cells in rows), and use HDF5/Zarr backed mode for matrices too large for RAM.

The Single Most Important Modern Insight -- The R/Python interop transpose is silent and catastrophic

R Bioconductor (Seurat, SingleCellExperiment) stores cells in COLUMNS and uses dgCMatrix (Column-Compressed Sparse). Python scverse (AnnData, scanpy) stores cells in ROWS and defaults to CSR (Compressed Sparse Row). Round-tripping with anndata2ri, zellkonverter, or rpy2 triggers a TRANSPOSE under the hood -- once per direction. Two flips silently cancel. A debugging session that converts back and forth multiple times can end up with mysteriously transposed data and no error.

ConversionImplicit transpose
Python CSR -> R dgCMatrixYES (rows <-> cols)
AnnData .X (cells x genes) -> SingleCellExperiment counts (genes x cells)YES
Seurat @assays$RNA@counts (genes x cells) -> AnnData .X (cells x genes)YES
scipy.sparse.csr_matrix(dense)None (just format conversion)
csr.tocsc() / csc.tocsr()None (semantically same matrix, different layout)

The safe pattern for one-shot conversions: file-based intermediate. adata.write('file.h5ad') then zellkonverter::readH5AD('file.h5ad') -- avoids the in-memory rpy2/reticulate gymnastics, and the file roundtrip makes orientation explicit.

For a 1M-cell single-cell matrix, the implicit transpose is non-trivial -- minutes of wall time and a temporary memory peak roughly equal to nnz x 12 bytes (CSC) or x 16 bytes (CSR with int64). Avoid unnecessary transposes by aligning the format to the consumer.

Two adjacent insights:

  1. Sparse becomes inefficient above ~10-15% density. Sparse iteration has cache-unfriendly indirection; dense iteration is sequential. Above ~10-15% nonzero, dense is often faster for most operations even though it uses more memory.
  2. AnnData backed mode quietly differs from in-memory in important ways. sc.read_h5ad('file.h5ad', backed='r') returns an AnnData where .X is a wrapped HDF5 dataset, read-only. Many scanpy functions silently load to memory; some functions error or hang on backed mode.

Algorithmic Taxonomy

FormatLayoutFast forSlow for
dgCMatrix (CSC, R)i (row indices), p (col pointers), x (values)Column slicing; per-cell ops in single-cell (cells in cols); matrix-vector with col vectorRow slicing
dgRMatrix (CSR, R)j (col indices), p (row pointers), x (values)Row slicing; per-gene ops when genes in rowsColumn slicing
dgTMatrix (COO, R triplet)i, j, xRandom insertion when building; reading MTXMost operations -- convert to dgCMatrix after build
scipy csc_matrixSame as dgCMatrixColumn ops in PythonRow ops
scipy csr_matrixSame as dgRMatrixRow ops (NumPy convention); matmul; sklearn defaultsColumn ops
scipy coo_matrixTriplet (i, j, data)Construction; MTX I/OMost ops -- convert after build
HDF5 (h5ad, h5)Single-file binary chunkedRandom access via chunks; compression; widely supportedCloud / parallel writes
ZarrChunked array, per-chunk file (or S3 object)Cloud-native; parallel writes; Dask integrationSingle-file simplicity
DelayedArray + HDF5ArrayBioc lazy evaluation over HDF5Out-of-memory bulk ops in RSpeed of in-memory

Decision Tree by Scenario

ScenarioRecommended approach
Single-cell (>10k cells), per-cell opsdgCMatrix in R (Bioconductor); CSR adata.X in Python (scanpy)
Bulk RNA-seq (60-80% density)Dense -- sparse overhead exceeds benefit
Single-cell pseudobulk (after donor aggregation)Dense -- now 60-80% density typically
1M+ cells, can't fit in RAMscanpy backed mode OR HDF5SummarizedExperiment + DelayedArray
TCGA + GTEx + recount3 scale (100k+ samples)HDF5Array / Zarr + Dask
10X CellRanger 3.0+ outputsc.read_10x_h5() or Read10X_h5() -- the .h5 is faster than the .mtx triplet
Cloud-native (anndata on S3, dask compute)Zarr
Local workstation, single-machineHDF5 -- faster, more widely supported
Building sparse matrix incrementallyCOO (dgTMatrix / coo_matrix); convert to CSC/CSR after
R <-> Python conversionFile-based intermediate (adata.write then zellkonverter::readH5AD); aware of the transpose

Check Sparsity

Goal: Decide whether sparse is the right format given the actual data density.

Approach: Compute nonzero fraction; rule of thumb is sparse > ~85% (single-cell). Below that, dense often wins.

python
import numpy as np
import scipy.sparse as sp

def sparsity(m):
    if sp.issparse(m):
        return 1 - m.nnz / (m.shape[0] * m.shape[1])
    return (m == 0).mean()

s = sparsity(adata.X)
print(f'{s:.1%} sparse')

Memory math:

DataFormatBytes
30k x 100k single-cell matrix, 5% densitydgCMatrix(5% * 3e9) * 12 bytes ~= 1.8 GB
SameDense double24 GB
60k x 100k bulk, 70% densitydgCMatrix50 GB (worse than dense!)
SameDense double48 GB

For single-cell (typically 90-95% sparse), sparse is essential. For bulk RNA-seq (typically 60-80% density), dense is faster and not appreciably larger.

dgCMatrix / scipy CSC / CSR

Goal: Construct, query, and convert sparse matrices in the format matching the consumer's expected layout.

Approach: Matrix::sparseMatrix(i, j, x, dims=...) (R) or scipy.sparse.csr_matrix((data, (i, j))) (Python); preserve row/column names; convert layout (CSC <-> CSR) without changing semantics.

python
import scipy.sparse as sp
import pandas as pd

dense_df = pd.read_csv('counts.csv', index_col=0)
sparse_csr = sp.csr_matrix(dense_df.values)
sparse_csc = sp.csc_matrix(dense_df.values)

gene_names = dense_df.index.tolist()
sample_names = dense_df.columns.tolist()

sparse_csr.tocsc()
sparse_csc.tocsr()
r
library(Matrix)

dense_mat <- as.matrix(read.csv('counts.csv', row.names = 1))
sparse_dgc <- as(dense_mat, 'CsparseMatrix')

class(sparse_dgc)
rownames(sparse_dgc) <- rownames(dense_mat)
colnames(sparse_dgc) <- colnames(dense_mat)

dgTMatrix is best for building matrices incrementally (reading MTX, parsing per-row); convert to dgCMatrix for downstream ops:

r
mat_t <- as(triplet_data, 'TsparseMatrix')
mat_c <- as(mat_t, 'CsparseMatrix')

HDF5 vs Zarr -- The Cloud-Native Shift

HDF5 (Hierarchical Data Format 5): hierarchical, single-file binary. Random access via chunks; supports compression (gzip, blosc, lz4). On-disk format for AnnData .h5ad, MuData .h5mu, 10x Genomics .h5, HDF5SummarizedExperiment.

Zarr: cloud-native, chunked array storage. Each chunk is a separate file (or S3 object). Parallel-write friendly; splittable by Dask. Format used by recent AnnData (anndata.write_zarr), SpatialData, and large-cohort consortia.

CriterionHDF5Zarr
File structureSingle binary fileDirectory of chunk files
Parallel writesLimited (process-level locks)Native
S3 / cloud object storageWorkarounds (h5cloud); often slowNative; first-class
Compression optionsgzip, blosc, lz4, szipgzip, blosc, lz4, zstd, custom
Local workstation speedFasterSlightly slower (many small files)
Wide ecosystem supportYes (mature)Growing; modern scverse

For local workstation work, HDF5 is faster and more widely supported. For cloud-mounted analysis (anndata on S3 with dask-distributed compute), Zarr wins because of object-storage friendliness.

HDF5SummarizedExperiment + DelayedArray (Bioconductor)

Goal: Work with bulk SummarizedExperiment objects too large to fit in RAM by keeping the matrix on disk.

Approach: HDF5Array wraps an HDF5 dataset as a DelayedArray. Subsetting builds a delayed operation tree -- no I/O until realization. DelayedMatrixStats provides delayed-friendly stat functions.

r
library(HDF5Array)
library(SummarizedExperiment)
library(DelayedMatrixStats)

se <- loadHDF5SummarizedExperiment('saved_se_dir')

s <- se[1:1000, 1:50]
row_means <- rowMeans2(assay(se))

saveHDF5SummarizedExperiment(se, 'saved_se_dir', replace = TRUE)

The assay(se) returns a DelayedMatrix backed by HDF5. DESeq2, edgeR, limma have varying levels of DelayedArray support; consult package docs before assuming all ops work in delayed mode.

For TCGA + GTEx + recount3 scale (100k+ samples, 60k genes), a dense matrix is ~48 GB (double); dgCMatrix at 70% density is ~50 GB (sparse loses). HDF5Array + chunk-aware ops keeps memory at whatever-fits-in-RAM.

scanpy Backed Mode

Goal: Work with h5ad files too large for memory by loading only accessed slices on demand.

Approach: sc.read_h5ad(..., backed='r') returns an AnnData with .X as a wrapped HDF5 dataset. Subset operations are lazy; .to_memory() realizes.

python
import scanpy as sc

adata = sc.read_h5ad('large_dataset.h5ad', backed='r')
print(f'Shape: {adata.shape}, X type: {type(adata.X)}')

t_cells = adata[adata.obs['cell_type'] == 'T_cell', :].to_memory()

Limitations:

  • .X is read-only in backed='r'. Use backed='r+' for in-place updates, but only .X updates supported.
  • .obs and .var are fully loaded -- only .X supports backed access.
  • Very large sparse h5ad (>35 GB) can still cause memory issues even in backed mode (anndata library overhead).
  • Many scanpy functions internally load to memory; check ?function docs for backed compatibility.
  • Functions like sc.tl.pca, sc.pp.neighbors typically require in-memory; subset first with .to_memory().

For datasets too large for backed mode, process in chunks:

python
import anndata as ad

def process_in_chunks(h5ad_path, chunk_size=10000, func=None):
    adata = sc.read_h5ad(h5ad_path, backed='r')
    n_cells = adata.shape[0]
    results = []
    for start in range(0, n_cells, chunk_size):
        end = min(start + chunk_size, n_cells)
        chunk = adata[start:end].to_memory()
        if func:
            chunk = func(chunk)
        results.append(chunk)
    return ad.concat(results)

10X Genomics Format Proliferation

FormatFilesNotes
MTX (pre-CellRanger 3.0)matrix.mtx + barcodes.tsv + features.tsv (or genes.tsv)Triplet format; slow to read for large matrices
H5 (CellRanger 3.0+)filtered_feature_bc_matrix.h5HDF5 with /matrix/data, /matrix/indices, /matrix/indptr, /matrix/shape; single file, fast
H5ADdata.h5adAnnData; convert on import
kallistobustools outputoutput.bus + barcode and gene mappings
python
import scanpy as sc

adata = sc.read_10x_h5('filtered_feature_bc_matrix.h5')
adata = sc.read_10x_mtx('filtered_feature_bc_matrix/')
r
library(Seurat)
mat <- Read10X_h5('filtered_feature_bc_matrix.h5')
mat <- Read10X(data.dir = 'filtered_feature_bc_matrix/')

library(DropletUtils)
sce <- read10xCounts('filtered_feature_bc_matrix/')

For 10X output, prefer the .h5 over the MTX triplet -- typically 5-10x faster for large matrices.

Show full SKILL.md (916 more words)Show less

Dense Conversion -- The Memory Blow-Up

adata.X.toarray() (Python) or as.matrix(seurat_obj@assays$RNA@counts) (R) on a 30k x 100k single-cell matrix instantiates a ~24 GB dense double array. Common triggers:

  • Passing sparse to a function that internally calls as.matrix() (older R cor() implementations).
  • Heatmap functions (pheatmap, ComplexHeatmap) that require dense.
  • ML libraries with no sparse support (some sklearn models; XGBoost requires specific sparse API).
  • Plotting functions (plot(), ggplot2) called on the full matrix.

Defensive pattern: subset to a manageable gene/cell set BEFORE dense conversion.

For per-cell PCA-style ops, use sparse-aware solvers:

python
from scipy.sparse.linalg import svds
U, s, Vt = svds(adata.X, k=50)
r
library(irlba)
svd_res <- irlba(sparse_mat, nv = 50)

irlba (R) and scipy.sparse.linalg.svds (Python) compute truncated SVD without densifying.

SCE vs AnnData vs MuData -- Where Bulk Fits

ContainerLibraryCells/samplesMulti-modalBulk fit
SummarizedExperiment / RangedSummarizedExperimentBioconductorn/a; bulkNoYES -- standard for DESeq2/edgeR/limma bulk
SingleCellExperiment (Amezquita 2020)Bioconductorcells in colsVia altExpsscRNA-seq with spike-ins, ADT
AnnDatascverse/Pythoncells in rowsVia layersscRNA-seq; bulk is unusual
MuDatascverse/Pythoncells in rowsYes, multiple AnnDataMulti-modal scRNA + ATAC + protein
MultiAssayExperimentBioconductorsamplesYesR-side multi-modal analog

Bulk RNA-seq rarely uses AnnData -- it shines on the single-cell dimensionality reduction / neighbors / clustering machinery. For bulk in R, use SummarizedExperiment. For bulk in Python, a tidy DataFrame + numpy array is usually sufficient.

Dask + Zarr for Consortium-Scale Matrices

For TCGA + GTEx + recount3 (Wilks 2021 Genome Biol 22:323) or pancancer assemblies:

python
import zarr
import dask.array as da

z = zarr.open('counts.zarr', mode='r')
da_arr = da.from_zarr(z)

col_sums = da_arr.sum(axis=0).compute()
filtered = da_arr[da_arr.sum(axis=1) > 100, :]
python
import anndata as ad
adata_disk = ad.read_zarr('large_data.zarr')

For R: HDF5Array + DelayedArray is the equivalent, but R doesn't have a true Dask analog. BiocParallel can parallelize chunks, but lazy planning is more manual.

Sparse Operations

python
import numpy as np
import scipy.sparse as sp

row_sums = np.array(sparse_matrix.sum(axis=1)).flatten()
col_sums = np.array(sparse_matrix.sum(axis=0)).flatten()

keep_rows = row_sums > 10
sparse_filt = sparse_matrix[keep_rows, :]

sparse_log = sparse_matrix.copy()
sparse_log.data = np.log1p(sparse_log.data)

Subsetting: select genes (rows) or samples (cols) by index:

python
gene_idx = [gene_names.index(g) for g in ['TP53', 'BRCA1', 'MYC'] if g in gene_names]
subset = sparse_matrix[gene_idx, :]

CPM Normalization on Sparse

Goal: Apply CPM normalization without densifying.

Approach: Compute library sizes from column sums; broadcast scaling factors with sparse multiply for CPM; transform only the nonzero data array in-place with log1p.

python
import numpy as np
import scipy.sparse as sp

def normalize_sparse_cpm(sparse_matrix):
    lib_sizes = np.array(sparse_matrix.sum(axis=0)).flatten()
    scaling = 1e6 / lib_sizes
    return sparse_matrix.multiply(scaling)

def log1p_inplace(sparse_matrix):
    out = sparse_matrix.copy()
    out.data = np.log1p(out.data)
    return out

cpm = normalize_sparse_cpm(adata.X)
log_cpm = log1p_inplace(cpm)

After log-transformation, sparsity is PRESERVED (log1p(0) = 0). After CPM with pseudocount, zeros become nonzero -- check sparsity and convert to dense if density drops below ~15%.

Save / Load Sparse Matrices

python
import scipy.sparse as sp
import numpy as np

sp.save_npz('counts_sparse.npz', sparse_matrix)
loaded = sp.load_npz('counts_sparse.npz')

np.savez('counts_with_meta.npz',
    data    = sparse_matrix.data,
    indices = sparse_matrix.indices,
    indptr  = sparse_matrix.indptr,
    shape   = sparse_matrix.shape,
    genes   = np.array(gene_names),
    samples = np.array(sample_names))

For interop and durable storage, prefer h5ad or zarr:

python
adata.write_h5ad('counts.h5ad')
adata.write_zarr('counts.zarr')

Per-Method Failure Modes

Implicit transpose in R/Python conversion

Trigger: AnnData with cells in rows passed to a SingleCellExperiment workflow that expects cells in cols; downstream colSums returns gene-level totals.

Mechanism: AnnData stores cells in rows; SCE in cols. The conversion auto-transposes ONCE per direction; two roundtrips silently restore.

Symptom: Per-cell stats look like per-gene stats; QC plots have wrong axes.

Fix: Use file-based intermediate (adata.write('file.h5ad'); zellkonverter::readH5AD('file.h5ad')). Always verify dimensions and orientation after conversion.

Dense conversion blew up memory

Trigger: as.matrix(seurat_obj@assays$RNA@counts) on a 100k-cell dataset; R session crashes with OOM.

Mechanism: 100k cells x 30k genes = 3e9 entries; double precision = 24 GB.

Symptom: R session killed; "cannot allocate vector of size N GB".

Fix: Don't densify the full matrix. Subset to genes/cells of interest first. For dimensionality reduction, use irlba::irlba() (sparse SVD).

scanpy backed mode silently loaded to memory

Trigger: adata = sc.read_h5ad(path, backed='r') then sc.tl.pca(adata); memory spikes to dense-equivalent.

Mechanism: Many scanpy functions internally call .to_memory() because they cannot operate on backed mode. sc.tl.pca, sc.pp.neighbors, sc.tl.umap all materialize.

Symptom: OOM despite backed mode.

Fix: Subset first (adata[mask].to_memory()), then operate. Or use a streaming-aware alternative (Dask + Zarr).

Sparse stored where dense would be faster

Trigger: Bulk RNA-seq with 70% density stored as dgCMatrix; per-gene rowVars is slow.

Mechanism: Sparse iteration has cache-unfriendly indirection; above ~10-15% density, dense wins.

Symptom: Operations notably slower than expected; profiler shows time in sparse indexing.

Fix: Convert to dense for the hot path: as.matrix(sparse_mat) (R) or sparse_matrix.toarray() (Python). Memory may go up but speed improves substantially.

10X MTX read is slow

Trigger: Reading a 100k-cell 10X dataset via the MTX three-file format; takes 10+ minutes.

Mechanism: MTX is a text format; parsing is slow for large matrices.

Symptom: Long load times; user kills the process before completion.

Fix: Use the CellRanger H5 (.h5) instead -- typically 5-10x faster.

Common errors

Error / symptomCauseFix
cannot allocate vector of size N GBImplicit dense conversionSubset first; use sparse-aware solver (irlba)
Sparse-dense arithmetic returns numpy.matrixDeprecated NumPy type from sparse+densenp.asarray(sparse + dense) to force ndarray
KeyError: '_index' reading h5adanndata version mismatchUpdate anndata; or sc.read_h5ad(..., backed=None)
Empty rows/cols after sparse subsetSubset removed all dataVerify the index list; cross-check sample/gene names
Backed mode AnnData crash on sc.tl.umapFunction not backed-compatible.to_memory() on the subset first
Per-cell totals look wrong after R<->Python conversionImplicit transposeVerify dimensions; use file-based intermediate
CSR (Python) <-> dgCMatrix (R) treated as sameConvention differenceThey're transposes of each other; verify shape and a known cell-gene pair

References

  • Amezquita RA, Lun ATL, Becht E et al. 2020. Orchestrating single-cell analysis with Bioconductor. Nat Methods 17:137-145. doi:10.1038/s41592-019-0654-x
  • Wolf FA, Angerer P, Theis FJ. 2018. SCANPY: large-scale single-cell gene expression data analysis. Genome Biol 19:15. doi:10.1186/s13059-017-1382-0
  • Wilks C et al. 2021. recount3: summaries and queries for large-scale RNA-seq expression and splicing. Genome Biol 22:323. doi:10.1186/s13059-021-02533-6
  • Bates D, Maechler M. 2023. Matrix: Sparse and Dense Matrix Classes and Methods. R package version 1.6-x.
  • Pages H et al. 2020. HDF5Array: HDF5 backend for DelayedArray objects. Bioconductor package.
  • Miles A et al. 2020. zarr-python. Python package documentation.
  • Rocklin M. 2015. Dask: Parallel Computation with Blocked algorithms and Task Scheduling. Proc Python Sci Conf. (canonical Dask reference)
  • Lachmann A et al. 2018. Massive mining of publicly available RNA-seq data from human and mouse. Nat Commun 9:1366. doi:10.1038/s41467-018-03751-6
  • counts-ingest - Reading 10X formats; building sparse matrices from quantification output
  • gene-id-mapping - Var (gene) metadata in AnnData
  • metadata-joins - Obs (sample) metadata in AnnData
  • normalization - log1p and CPM patterns on sparse
  • differential-expression/deseq2-basics - Pseudobulk aggregation makes dense
  • single-cell/data-io - Single-cell file format ecosystem
  • single-cell/preprocessing - Standard single-cell sparse pipeline

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in expression-matrix/sparse-handling of GPTomics/bioSkills.

  • SKILL.md
  • examples/sparse_operations.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Expression Matrix Sparse Handling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Expression Matrix Sparse Handling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Expression Matrix Sparse Handling this skillGPTomics/bioSkills1.2k1 repos~5.6kAutomated safety check: PassMIT
AnndataK-Dense-AI/scientific-agent-skills48k1 repos~3.9kAutomated safety check: NotesBSD-3-Clause
Lamindb Data Managementjaechang-hits/SciAgent-Skills3702 repos~4kAutomated safety check: PassApache-2.0
Anndatadavila7/claude-code-templates32k12 repos~2.5kAutomated safety check: PassMIT
Anndata Data Structurejaechang-hits/SciAgent-Skills3702 repos~5.8kAutomated safety check: PassBSD-3-Clause
ScanpyK-Dense-AI/scientific-agent-skills48k1 repos~5.1kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Anndata

    K-Dense-AI/scientific-agent-skills

    Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem.

    48k GitHub starsUsed in 1 repo~3.9k tokens
    Research & ScienceAuto-check: notes
  • Lamindb Data Management

    jaechang-hits/SciAgent-Skills

    Open-source FAIR biology data framework. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 2 repos~4k tokens
    Research & ScienceAuto-check passed
  • Anndata

    davila7/claude-code-templates

    This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…

    32k GitHub starsUsed in 12 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Anndata Data Structure

    jaechang-hits/SciAgent-Skills

    Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.

    370 GitHub starsUsed in 2 repos~5.8k tokens
    Research & ScienceAuto-check passed
  • Scanpy

    K-Dense-AI/scientific-agent-skills

    Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…

    48k GitHub starsUsed in 1 repo~5.1k tokens
    Research & ScienceAuto-check passed
  • Bio Single Cell Data Io

    FreedomIntelligence/OpenClaw-Medical-Skills

    Read, write, and create single-cell data objects using Seurat (R) and Scanpy (Python).

    3.1k GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Expression Matrix Sparse Handling

What does Bio Expression Matrix Sparse Handling do?

Stores and operates on sparse expression matrices for single-cell and large bulk RNA-seq, covering dgCMatrix/dgRMatrix/dgTMatrix when-each-is-fast, the dgCMatrix (CSC, R) <- CSR (Python) implicit…. Bio Expression Matrix Sparse Handling is an agent skill from GPTomics/bioSkills.

When should I use Bio Expression Matrix Sparse Handling?

Bio Expression Matrix Sparse Handling fits situations like: choosing sparse format; working with single-cell-sized matrices; importing/exporting 10X; debugging R/Python interop transposes.

How do I install Bio Expression Matrix Sparse Handling in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-expression-matrix-sparse-handling -a claude-code`. Or copy the skill folder (expression-matrix/sparse-handling in GPTomics/bioSkills) into .claude/skills/bio-expression-matrix-sparse-handling in your project. Claude Code loads it when a task matches its description.

How do I install Bio Expression Matrix Sparse Handling in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-expression-matrix-sparse-handling -a codex`. Or copy the skill folder (expression-matrix/sparse-handling in GPTomics/bioSkills) into .agents/skills/bio-expression-matrix-sparse-handling in your project. Codex loads it when a task matches its description.

Can I use Bio Expression Matrix Sparse Handling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-expression-matrix-sparse-handling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-expression-matrix-sparse-handling, .gemini/skills/bio-expression-matrix-sparse-handling, .github/skills/bio-expression-matrix-sparse-handling and .opencode/skills/bio-expression-matrix-sparse-handling in your project.

What does Bio Expression Matrix Sparse Handling need to run?

Going by SKILL.md and its folder, Bio Expression Matrix Sparse Handling needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Expression Matrix Sparse Handling access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Expression Matrix Sparse Handling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Expression Matrix Sparse Handling use?

Bio Expression Matrix Sparse Handling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Expression Matrix Sparse Handling use?

About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Expression Matrix Sparse Handling?

Skills that share tags, products or a category with Bio Expression Matrix Sparse Handling: Anndata (K-Dense-AI/scientific-agent-skills, 48k stars), Lamindb Data Management (jaechang-hits/SciAgent-Skills, 370 stars), Anndata (davila7/claude-code-templates, 32k stars) and Anndata Data Structure (jaechang-hits/SciAgent-Skills, 370 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Expression Matrix Sparse Handling?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.