Agent skill

Single-Cell Initial Analysis

by LigphiDonk in LigphiDonk/Oh-my--paper

Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

MITAuto-check passedResearch & Science

Install Single-Cell Initial Analysis

skills CLI
$ npx skills add LigphiDonk/Oh-my--paper --skill init-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LigphiDonk/Oh-my--paper init-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LigphiDonk/Oh-my--paper.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bioinformatics-init-analysis/skills/init-analysis .claude/skills/init-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
init-analysis
GitHub stars
738
Used in
1 other repo
Token cost
~1.4k tokens
SKILL.md length
473 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
MIT

At a glance

Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found.

  • Works in 7 steps: Data Loading → Quality Control → Normalization → …
  • Running first-pass quality control on a new single-cell dataset
  • SKILL.md covers Supported Data Types, Approach 1: Full Pipeline…, Approach 2: Modular Steps and Pipeline Steps, plus 2 more sections
  • Calls python3

What it does

The pipeline works out whether your input is scRNA-seq, CyTOF or flow cytometry data from the file format and contents, accepting `.h5ad`, 10X `.h5`, `.mtx` with barcodes, `.csv` and `.fcs` files or a folder of CSVs. You run it with `scripts/run_pipeline.py`, choosing an optional data type override, a per-group cell cap (500 by default), an output directory and a report style, either clinical or technical, with clinical as the default.

Quality control adapts to the data type: marker-level outlier detection and batch checks for CyTOF, and mitochondrial percentage, genes and counts per cell and doublet detection for scRNA-seq. For CyTOF, normalization verifies the existing transform before scaling. The description also lists dimensionality reduction, clustering and marker analysis. Results land in an `analysis_output` folder with PNG figures and processed data, and each step can also be imported as a module that takes and returns an AnnData object.

When your agent uses it

  • Running first-pass quality control on a new single-cell dataset
  • Exploring CyTOF, flow cytometry or scRNA-seq data before deeper analysis
  • Producing a plain-language summary of a dataset for clinicians or collaborators

Example prompts

  • “QC my scRNA-seq data in ./data/pbmc.h5ad and generate an analysis report.”
  • “Run the bioinformatics pipeline on the CyTOF CSV files in ./cytof_runs with the technical report style.”
  • “Do exploratory data analysis on ./panel.fcs and subsample to 500 cells per group.”

Requirements

  • Python 3 to run scripts/run_pipeline.py
  • Input data as .h5ad, .h5, .mtx, .csv or .fcs files

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Data Loading
  2. Quality Control
  3. Normalization
  4. Dimensionality Reduction
  5. Clustering
  6. Marker Analysis
  7. Report Generation

What it can do on your machine

Read from SKILL.md and the folder at commit 6baece9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Single-Cell Initial Analysis loads about 1.4k tokens when it runs. Until then it costs about 109 tokens; SKILL.md has 473 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LigphiDonk/Oh-my--paper at commit 6baece9, republished under its MIT licence (© LigphiDonk). 473 words, ~1,356 tokens.

Download SKILL.mdSave it as .claude/skills/init-analysis/SKILL.md (or your agent's skills folder).
name
init-analysis
description
This skill should be used when the user asks to "run initial analysis", "analyze single-cell data", "QC my data", "run bioinformatics pipeline", "generate analysis report", "explore my dataset", "do exploratory data analysis", "initial data analysis", or needs to perform quality control, dimensionality reduction, clustering, or marker analysis on single-cell biology data (CyTOF, scRNA-seq, flow cytometry, proteomics).
version
0.1.0

Bioinformatics Initial Data Analysis

Automated 7-step analysis pipeline for high-dimensional single-cell biology data with plain-language report generation.

Supported Data Types

The pipeline auto-detects input data type:

Data TypeFile FormatsDetection Pattern
scRNA-seq.h5ad, .h5 (10X), .mtx + barcodesGene names, count matrix
CyTOF.csv, .h5adPhospho-markers (p.ERK, p.AKT, etc.)
Flow cytometry.fcs, .csvSurface markers, scatter channels

Run the complete 7-step analysis:

bash
python3 scripts/run_pipeline.py <input_path> \
    [--data-type auto|cytof|scrnaseq|flow] \
    [--subsample 500] \
    [--output-dir ./analysis_output] \
    [--report-style clinical|technical]

Arguments:

  • input_path: Path to data file (.h5ad, .csv, .h5) or directory of CSV files
  • --data-type: Data type override (default: auto for auto-detection)
  • --subsample: Max cells per group for tractable analysis (default: 500)
  • --output-dir: Output directory (default: ./analysis_output)
  • --report-style: clinical for plain-language medical summaries, technical for bioinformatics detail (default: clinical)

Output Files:

analysis_output/
├── figures/                    # All generated plots (PNG)
├── processed/
│   └── adata_processed.h5ad   # Processed AnnData object
├── report.html                 # Complete analysis report
└── analysis_summary.json       # Machine-readable summary statistics

Approach 2: Modular Steps

For custom workflows, import individual step modules:

python
from step1_load_data import load_data
from step2_qc import run_qc
from step3_normalize import normalize_data
from step4_dim_reduction import run_dim_reduction
from step5_clustering import run_clustering
from step6_marker_analysis import run_marker_analysis
from step7_report import generate_report

Each step function accepts an AnnData object and returns the modified AnnData plus a dictionary of results/figures.

Pipeline Steps

Step 1: Data Loading

Load data from various formats into AnnData. For directories of CSVs (e.g., CyTOF per-cell-line files), automatically concatenate with metadata. Apply subsampling if dataset is large.

Step 2: Quality Control

Data-type-aware QC:

  • CyTOF: MAD-based outlier detection per marker, signal distributions, batch effects across cell lines
  • scRNA-seq: Mitochondrial %, genes per cell, counts per cell, doublet detection
  • Universal: Missing value assessment, distribution violin plots, outlier flagging
Step 3: Normalization
  • CyTOF: Data is typically pre-transformed (arcsinh). Verify transformation, apply z-score for dim. reduction.
  • scRNA-seq: Library size normalization (CPM) -> log1p -> HVG selection -> z-score scaling
  • Store raw values in adata.raw for downstream differential analysis.
Step 4: Dimensionality Reduction

PCA with scree plot and loadings analysis, followed by UMAP visualization colored by all available metadata and key markers.

Show full SKILL.md (200 more words)Show less
Step 5: Clustering

Leiden graph-based clustering at multiple resolutions. Evaluate with ARI, NMI, Silhouette scores if reference labels exist. Visualize cluster composition across metadata categories.

Step 6: Marker Analysis

Wilcoxon rank-sum differential expression per cluster. Marker correlation heatmap. If treatment/condition metadata exists: treatment response analysis with boxplots and effect heatmaps. If time metadata exists: time-course dynamics.

Step 7: Report Generation

Generate HTML report with embedded figures and interpretations. Two styles:

  • Clinical: Written for medical doctors and non-bioinformaticians. Each plot includes "What this shows", "Key findings", and "Clinical relevance" sections.
  • Technical: Standard bioinformatics report with methods, parameters, and statistical details.

Reference Files

For detailed guidance, consult:

  • references/plot_interpretation_guide.md - How to explain each plot type to non-experts
  • references/cytof_specifics.md - CyTOF-specific QC, normalization, and markers
  • references/scrnaseq_specifics.md - scRNA-seq-specific processing details
  • references/statistical_methods.md - Plain-language glossary of statistical methods

Important Notes

  • CyTOF data is often pre-transformed: Check value ranges before applying arcsinh. Values in range [-1, 15] indicate arcsinh-transformed data.
  • Large datasets: Always subsample for initial exploration. Use 500-2000 cells per group.
  • NaN handling: Always run nan_to_num after z-score scaling — constant-value features produce NaN.
  • Leiden timeout: For >50K cells, Leiden clustering may take >10 minutes. Reduce subsample size.
  • Report audience: Default to clinical style unless user requests technical detail.

© LigphiDonk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/bioinformatics-init-analysis/skills/init-analysis of LigphiDonk/Oh-my--paper.

Open the folder on GitHubat commit 6baece9

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in LigphiDonk/Oh-my--paper, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Single-Cell Initial Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Single-Cell Initial Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Single-Cell Initial Analysis this skillLigphiDonk/Oh-my--paper7381 repos~1.4kAutomated safety check: PassMIT
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k16 repos~2.8kAutomated safety check: PassMIT
PyDESeq2 Differential Expressiondavila7/claude-code-templates32k12 repos~4kAutomated safety check: PassMIT
Ukb Ppp Region FetchClawBio/ClawBio1.2k—~4.6kAutomated safety check: PassMIT
Gwas PipelineClawBio/ClawBio1.2k1 repos~1.4kAutomated safety check: PassMIT
Tooluniverse Polygenic Risk Scorewu-yc/LabClaw1.1k2 repos~3.7kAutomated safety check: PassNone

Similar skills

  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 16 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • PyDESeq2 Differential Expression

    davila7/claude-code-templates

    Runs differential gene expression analysis on bulk RNA-seq counts with PyDESeq2: design formulas, Wald tests, FDR correction and volcano or MA plots.

    32k GitHub starsUsed in 12 repos~4k tokens
    Research & ScienceAuto-check passed
  • Ukb Ppp Region Fetch

    ClawBio/ClawBio

    Fetch a regional slice of plasma pQTL summary statistics from the UK Biobank Pharma Proteomics Project (UKB-PPP; Sun 2023 Nature) for a specific (protein, ancestry) measurement.

    1.2k GitHub stars~4.6k tokensUpdated today
    Research & ScienceAuto-check passed
  • Gwas Pipeline

    ClawBio/ClawBio

    End-to-end GWAS automation wrapping PLINK2 for genotype QC and REGENIE for two-step whole-genome regression association testing.

    1.2k GitHub starsUsed in 1 repo~1.4k tokens
    Research & ScienceAuto-check passed
  • Build and interpret polygenic risk scores (PRS) for complex diseases using GWAS summary statistics.

    1.1k GitHub starsUsed in 2 repos~3.7k tokens
    Research & ScienceAuto-check passed
  • Use this bioinformatics data analysis skill to construct a database-driven lncRNA-mRNA regulatory network from target lncRNA and/or gene lists by projecting shared miRNA evidence from local ceRNA…

    2k GitHub stars~2.7k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed

More from LigphiDonk/Oh-my--paper

All 27 skills in this repo
  • Preprint Search on bioRxiv

    LigphiDonk/Oh-my--paper

    Searches bioRxiv life sciences preprints by keyword, author, date range or category with a Python script, returning JSON metadata and optional PDF downloads.

    738 GitHub starsUsed in 13 repos~3.7k tokens
    Auto-check passed
  • Literature PDF OCR Library Builder

    LigphiDonk/Oh-my--paper

    Searches and downloads legally accessible academic PDFs, OCRs them to Markdown, and organizes the results into a traceable, AI-readable literature library.

    738 GitHub stars~1.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Inno Code Survey

    LigphiDonk/Oh-my--paper

    Finds and clones missing code repositories for a chosen research idea, then writes a survey that maps academic concepts to their implementations.

    738 GitHub stars~3.6k tokensUpdated 5 mo ago
    Auto-check passed
  • Turns experimental data such as CSV, JSON or TensorBoard logs into statistical significance tests, visualizations and a drafted Results section.

    738 GitHub stars~3k tokensUpdated 5 mo ago
    Auto-check passed
  • Citation Verification Guide

    LigphiDonk/Oh-my--paper

    Lays out principles for catching fake, mismatched, or inconsistently formatted citations in academic writing, checked through live web search.

    738 GitHub stars~2.2k tokensUpdated 5 mo ago
    Auto-check passed
  • Making Academic Presentations

    LigphiDonk/Oh-my--paper

    Create academic presentation slide decks and optionally demo videos from research papers.

    738 GitHub stars~2k tokensUpdated 5 mo ago
    Auto-check passed

Works with

Questions about Single-Cell Initial Analysis

What does Single-Cell Initial Analysis do?

Runs a seven-step quality-control and exploration pipeline on scRNA-seq, CyTOF or flow cytometry data and writes a plain-language report of what it found. fcs` files or a folder of CSVs.py`, choosing an optional data type override, a per-group cell cap (500 by default), an output directory and a report style, either clinical or technical, with clinical as the default.

When should I use Single-Cell Initial Analysis?

Single-Cell Initial Analysis fits situations like: running first-pass quality control on a new single-cell dataset; exploring CyTOF, flow cytometry or scRNA-seq data before deeper analysis; producing a plain-language summary of a dataset for clinicians or collaborators.

How do I install Single-Cell Initial Analysis in Claude Code?

Run `npx skills add LigphiDonk/Oh-my--paper --skill init-analysis -a claude-code`. Or copy the skill folder (skills/bioinformatics-init-analysis/skills/init-analysis in LigphiDonk/Oh-my--paper) into .claude/skills/init-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Single-Cell Initial Analysis in Codex?

Run `npx skills add LigphiDonk/Oh-my--paper --skill init-analysis -a codex`. Or copy the skill folder (skills/bioinformatics-init-analysis/skills/init-analysis in LigphiDonk/Oh-my--paper) into .agents/skills/init-analysis in your project. Codex loads it when a task matches its description.

Can I use Single-Cell Initial Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LigphiDonk/Oh-my--paper --skill init-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/init-analysis, .gemini/skills/init-analysis, .github/skills/init-analysis and .opencode/skills/init-analysis in your project.

What does Single-Cell Initial Analysis need to run?

Going by SKILL.md and its folder, Single-Cell Initial Analysis needs the command-line tools its instructions call (python3). Our summary lists: Python 3 to run scripts/run_pipeline.py; Input data as .h5ad, .h5, .mtx, .csv or .fcs files.

Does Single-Cell Initial Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Single-Cell Initial Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Single-Cell Initial Analysis use?

Single-Cell Initial Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Single-Cell Initial Analysis use?

About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Single-Cell Initial Analysis?

Skills that share tags, products or a category with Single-Cell Initial Analysis: Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars), PyDESeq2 Differential Expression (davila7/claude-code-templates, 32k stars), Ukb Ppp Region Fetch (ClawBio/ClawBio, 1.2k stars) and Gwas Pipeline (ClawBio/ClawBio, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Single-Cell Initial Analysis?

LigphiDonk (a GitHub user) maintains it in LigphiDonk/Oh-my--paper, which has 738 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on April 15, 2026.

Source: LigphiDonk/Oh-my--paper on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.