Agent skill

Gene Protein Expression Matrix Normalization

by aipoch in aipoch/medical-research-skills

A skill your agent uses when normalizing bulk gene or protein expression matrices with log2 transform, z-score standardization, or min-max scaling before downstream visualization or exploratory…

MITAuto-check passedResearch & Science

Install Gene Protein Expression Matrix Normalization

skills CLI
$ npx skills add aipoch/medical-research-skills --skill gene-protein-expression-matrix-normalization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills gene-protein-expression-matrix-normalization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'awesome-med-research-skills/Data Analysis/gene-protein-expression-matrix-normalization' .claude/skills/gene-protein-expression-matrix-normalization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gene-protein-expression-matrix-normalization
GitHub stars
2k
Token cost
~1.5k tokens
SKILL.md length
596 words
Files
15 (incl. scripts, references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when normalizing bulk gene or protein expression matrices with log2 transform, z-score standardization, or min-max scaling before downstream visualization or exploratory…

  • Normalizing bulk gene
  • SKILL.md covers When to Use, When Not to Use, When to Read External Files and Usage, plus 6 more sections
  • Runs R scripts from its folder
  • Protein expression matrices with log2 transform

What it does

Gene Protein Expression Matrix Normalization is an agent skill from aipoch/medical-research-skills. Use when normalizing bulk gene or protein expression matrices with log2 transform, z-score standardization, or min-max scaling before downstream visualization or exploratory analysis. NOT for count-model normalization such as TPM/DESeq2 size factors, batch correction, or single-cell preprocessing.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including scripts and reference files (for example `eval_report_gene-protein-expression-matrix-normalization_result.json`, `references/algorithm.md` and `references/cli-guide.md`).

It sits in Research & Science, covering Database schema design and Bioinformatics. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Normalizing bulk gene
  • Protein expression matrices with log2 transform
  • Z-score standardization
  • Min-max scaling before downstream visualization

Example prompts

  • “/gene-protein-expression-matrix-normalization”

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gene Protein Expression Matrix Normalization loads about 1.5k tokens when it runs, and up to ~2.9k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 596 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 596 words, ~1,530 tokens.

Download SKILL.mdSave it as .claude/skills/gene-protein-expression-matrix-normalization/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
gene-protein-expression-matrix-normalization
description
Use when normalizing bulk gene or protein expression matrices with log2 transform, z-score standardization, or min-max scaling before downstream visualization or exploratory analysis. NOT for count-model normalization such as TPM/DESeq2 size factors, batch correction, or single-cell preprocessing.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

Gene Protein Expression Matrix Normalization

When to Use

Use this skill when the user wants to normalize a numeric expression matrix before plotting, clustering, or exploratory comparison.

Typical requests:

  • "Normalize this gene expression matrix with log2"
  • "Do z-score scaling across samples"
  • "Map protein abundance values into 0 to 1"

When Not to Use

Do not use this skill for:

  • Count-model normalization such as CPM, TPM, TMM, or DESeq2 size factors
  • Batch correction or covariate adjustment
  • Single-cell preprocessing workflows
  • Matrices that contain missing, Inf, or NaN values unless they are cleaned first

When to Read External Files

When executing the analysis, run:

bash
Rscript scripts/main.R --input_file <matrix.csv> --output_dir <output_dir> --method <log2|zscore|minmax>
SituationFile to ReadPurpose
Need to execute the workflowscripts/main.RCLI entry point
Need algorithm detailsreferences/algorithm.mdMethod definitions and assumptions
Encounter an errorreferences/troubleshooting.mdStandard error codes and fixes
Need examples or baseline run detailsreferences/cli-guide.mdReady-to-run commands and test record
Need dependency declarationsDESCRIPTIONRuntime package list

Usage

bash
Rscript scripts/main.R \
  --input_file tests/data/expression_matrix.csv \
  --output_dir ./output \
  --method log2 \
  --pseudo_count 1 \
  --seed 42

Arguments

ShortLongTypeDefaultDescription
-i--input_filefilerequiredExpression matrix in CSV or TSV format
-o--output_dirdir./outputOutput directory
-m--methodstringlog2Normalization method: log2, zscore, minmax
-r--marginstringcolumnApply normalization by row or column
-p--pseudo_countnumeric1Added before log2 transformation
-c--centerbooleantrueCenter values for z-score
-s--scale_valuesbooleantrueScale values for z-score
-t--timeout_secondsinteger0Optional timeout; 0 disables it
-d--delimiterstringautoInput delimiter: auto, csv, or tsv
--seedinteger42Random seed
--verbosebooleantruePrint progress logs

Input Format

The first column must contain feature identifiers. Remaining columns must be finite numeric sample values.

Missing values and non-finite values such as NA, NaN, Inf, and -Inf are rejected.

csv
feature,S1,S2,S3
TP53,10,20,30
EGFR,3,5,9

This skill accepts gene or protein expression matrices. It does not infer count-model normalization such as CPM, TPM, TMM, or DESeq2 size factors.

Output Files

If --output_dir already exists, result files with the same names are overwritten. When --verbose=true, the workflow prints a warning before writing into a non-empty output directory.

For single-sample inputs, feature_summary.csv reports per-feature standard deviations as 0 by design because each feature contributes one observed value.

FileDescription
table/normalized_matrix.csvNormalized matrix with the original feature column preserved
table/feature_summary.csvPer-feature min, max, mean, and SD before and after normalization
table/sample_summary.csvPer-sample min, max, mean, and SD before and after normalization
data/normalized_matrix.rdsSerialized normalized matrix and run metadata
run_record.txtStructured execution record
output_manifest.txtOutput file manifest
session_info.txtR session information
Show full SKILL.md (193 more words)Show less

Methods

log2

Computes log2(x + pseudo_count) for each numeric value.

zscore

Centers and scales along the selected margin. margin=column standardizes each sample; margin=row standardizes each feature.

When center=false and scale_values=true, the workflow divides by standard deviation without subtracting the mean first.

minmax

Rescales values to [0, 1] along the selected margin. Constant vectors are returned as zeros to avoid division-by-zero errors.

Error Handling

ErrorCauseSolution
SKILL_FILE_NOT_FOUNDInput file path is invalidCheck the input path
SKILL_MISSING_COLUMNSMatrix has fewer than two columnsProvide one feature column and at least one sample column
SKILL_INVALID_PARAMETERCLI value is unsupported or malformed, or the matrix contains non-finite valuesReview the argument table and inspect the matrix values
SKILL_TIMEOUTThe run exceeded --timeout_secondsIncrease the timeout or simplify the input size
SKILL_EMPTY_DATANo usable rows or columns remainCheck the input matrix

Testing

bash
Rscript scripts/main.R --help

Rscript tests/run_tests.R

Rscript tests/run_tests.R audit_output_check

Rscript tests/test_skill.R

Rscript tests/test_skill.R audit_output_check --skip-prepare

tests/run_tests.R executes bundled log2, zscore, and minmax runs and writes their outputs under tests/output/.

When you pass a relative directory name such as audit_output_check, the test runner writes outputs under tests/output/audit_output_check/.

Run tests/run_tests.R before tests/test_skill.R when you want to validate pre-generated outputs explicitly. The validation script can also prepare missing outputs on its own.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files (scripts, references) in awesome-med-research-skills/Data Analysis/gene-protein-expression-matrix-normalization of aipoch/medical-research-skills.

  • SKILL.md
  • DESCRIPTION
  • eval_report_gene-protein-expression-matrix-normalization_result.json
  • references/algorithm.md
  • references/cli-guide.md
  • references/troubleshooting.md
  • scripts/cli_options.R
  • scripts/functions.R
  • scripts/io.R
  • scripts/main.R
  • scripts/recording.R
  • scripts/run_analysis.R
  • scripts/utils.R
  • tests/data/expression_matrix.csv
  • tests/run_tests.R

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Gene Protein Expression Matrix Normalization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gene Protein Expression Matrix Normalization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gene Protein Expression Matrix Normalization this skillaipoch/medical-research-skills2k—~1.5kAutomated safety check: PassMIT
Bio Chipseq Differential BindingGPTomics/bioSkills1.2k2 repos~5.1kAutomated safety check: PassMIT
Bio Geo DataGPTomics/bioSkills1.2k2 repos~4.4kAutomated safety check: PassMIT
Tooluniverse Metabolomics Analysiswu-yc/LabClaw1.1k2 repos~5.9kAutomated safety check: PassNone
Bio Spatial Transcriptomics Spatial PreprocessingFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~2kAutomated safety check: PassNone
Bio Expression Matrix NormalizationGPTomics/bioSkills1.2k1 repos~6.2kAutomated safety check: PassMIT

Similar skills

  • Identifies differentially bound ChIP-seq regions between conditions using DiffBind, csaw (sliding windows), DESeq2/edgeR/PyDESeq2 on count matrices, NormR (control-aware), or MAnorm2.

    1.2k GitHub starsUsed in 2 repos~5.1k tokens
    Research & ScienceAuto-check passed
  • Bio Geo Data

    GPTomics/bioSkills

    Query and download from NCBI Gene Expression Omnibus (GEO) and EMBL-EBI's BioStudies/ArrayExpress mirror.

    1.2k GitHub starsUsed in 2 repos~4.4k tokens
    Research & ScienceAuto-check passed
  • Analyze metabolomics data including metabolite identification, quantification, pathway analysis, and metabolic flux.

    1.1k GitHub starsUsed in 2 repos~5.9k tokens
    Research & ScienceAuto-check passed
  • Bio Spatial Transcriptomics Spatial Preprocessing

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quality control, filtering, normalization, and feature selection for spatial transcriptomics data.

    3.1k GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Normalizes and transforms RNA-seq count matrices for DE, visualization, clustering, and ML.

    1.2k GitHub starsUsed in 1 repo~6.2k tokens
    Research & ScienceAuto-check passed
  • Harmonizes already-normalized per-omic matrices onto a common footing before joint integration - assembling a MultiAssayExperiment, choosing the per-omic variance-stabilizing transform, deciding…

    1.2k GitHub starsUsed in 1 repo~5.4k tokens
    Research & ScienceAuto-check passed

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 22 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 22 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 22 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 22 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 22 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 22 days ago
    Auto-check passed

Questions about Gene Protein Expression Matrix Normalization

What does Gene Protein Expression Matrix Normalization do?

A skill your agent uses when normalizing bulk gene or protein expression matrices with log2 transform, z-score standardization, or min-max scaling before downstream visualization or exploratory…. Gene Protein Expression Matrix Normalization is an agent skill from aipoch/medical-research-skills. Use when normalizing bulk gene or protein expression matrices with log2 transform, z-score standardization, or min-max scaling before downstream visualization or exploratory analysis.

When should I use Gene Protein Expression Matrix Normalization?

Gene Protein Expression Matrix Normalization fits situations like: normalizing bulk gene; protein expression matrices with log2 transform; Z-score standardization; min-max scaling before downstream visualization.

How do I install Gene Protein Expression Matrix Normalization in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill gene-protein-expression-matrix-normalization -a claude-code`. Or copy the skill folder (awesome-med-research-skills/Data Analysis/gene-protein-expression-matrix-normalization in aipoch/medical-research-skills) into .claude/skills/gene-protein-expression-matrix-normalization in your project. Claude Code loads it when a task matches its description.

How do I install Gene Protein Expression Matrix Normalization in Codex?

Run `npx skills add aipoch/medical-research-skills --skill gene-protein-expression-matrix-normalization -a codex`. Or copy the skill folder (awesome-med-research-skills/Data Analysis/gene-protein-expression-matrix-normalization in aipoch/medical-research-skills) into .agents/skills/gene-protein-expression-matrix-normalization in your project. Codex loads it when a task matches its description.

Can I use Gene Protein Expression Matrix Normalization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill gene-protein-expression-matrix-normalization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gene-protein-expression-matrix-normalization, .gemini/skills/gene-protein-expression-matrix-normalization, .github/skills/gene-protein-expression-matrix-normalization and .opencode/skills/gene-protein-expression-matrix-normalization in your project.

What does Gene Protein Expression Matrix Normalization need to run?

Going by SKILL.md and its folder, Gene Protein Expression Matrix Normalization needs R for the scripts in its folder.

Does Gene Protein Expression Matrix Normalization access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gene Protein Expression Matrix Normalization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gene Protein Expression Matrix Normalization use?

Gene Protein Expression Matrix Normalization is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gene Protein Expression Matrix Normalization use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Gene Protein Expression Matrix Normalization?

Skills that share tags, products or a category with Gene Protein Expression Matrix Normalization: Bio Chipseq Differential Binding (GPTomics/bioSkills, 1.2k stars), Bio Geo Data (GPTomics/bioSkills, 1.2k stars), Tooluniverse Metabolomics Analysis (wu-yc/LabClaw, 1.1k stars) and Bio Spatial Transcriptomics Spatial Preprocessing (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gene Protein Expression Matrix Normalization?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,978 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.