Agent skill

Bio Proteomics Proteomics Qc

by GPTomics in GPTomics/bioSkills

Quality control for bottom-up proteomics across three levels -- instrument/raw-signal (mass accuracy, RT/iRT fit, FWHM, TIC vs injection time, % MS2 identified), identification/run (missed…

MITAuto-check passedData & Analytics

Install Bio Proteomics Proteomics Qc

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-proteomics-qc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-proteomics-proteomics-qc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/proteomics/proteomics-qc .claude/skills/bio-proteomics-proteomics-qc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-proteomics-proteomics-qc
GitHub stars
1.2k
Used in
1 other repo
Token cost
~6.4k tokens
SKILL.md length
2,651 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Quality control for bottom-up proteomics across three levels -- instrument/raw-signal (mass accuracy, RT/iRT fit, FWHM, TIC vs injection time, % MS2 identified), identification/run (missed…

  • Works in 3 steps: QC is a three-level funnel of silent… → Almost no metric has a universal… → Median normalization HIDES loading…
  • Assessing proteomics data quality
  • SKILL.md covers Version Compatibility, The Single Most Important…, The Three Levels in One Table and Tool Taxonomy, plus 11 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Bio Proteomics Proteomics Qc is an agent skill from GPTomics/bioSkills. Quality control for bottom-up proteomics across three levels -- instrument/raw-signal (mass accuracy, RT/iRT fit, FWHM, TIC vs injection time, % MS2 identified), identification/run (missed cleavages, charge states, PTM handling artifacts, contaminants), and experiment/quantitative (replicate correlation on log2, CV on the linear scale, completeness, MNAR-vs-MCAR missingness, PCA/batch, TMT channel balance, DIA q-values). Frames QC as a control chart against a per-instrument rolling baseline, not fixed cutoffs…

Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/qc_analysis.py` and `usage-guide.md`).

It sits in Data & Analytics, covering Bioinformatics and Database schema design. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Assessing proteomics data quality
  • Diagnosing outlier samples
  • Deciding which samples to exclude before differential testing

Example prompts

  • “/bio-proteomics-proteomics-qc”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. QC is a three-level funnel of silent failures, and the deliverable matrix is the LAST place a problem becomes visible. Faults originate at…
  2. Almost no metric has a universal pass/fail cutoff; the defensible practice is a per-instrument, per-method control chart. Deviation from a…
  3. Median normalization HIDES loading problems that must be SEEN first. Median/quantile normalization works by forcing a chosen summary…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Proteomics Proteomics Qc loads about 6.4k tokens when it runs. Until then it costs about 249 tokens; SKILL.md has 2,651 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~249
When it runs · the whole SKILL.md, loaded when a task matches
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,651 words, ~6,364 tokens.

Download SKILL.mdSave it as .claude/skills/bio-proteomics-proteomics-qc/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-proteomics-proteomics-qc
description
Quality control for bottom-up proteomics across three levels -- instrument/raw-signal (mass accuracy, RT/iRT fit, FWHM, TIC vs injection time, % MS2 identified), identification/run (missed cleavages, charge states, PTM handling artifacts, contaminants), and experiment/quantitative (replicate correlation on log2, CV on the linear scale, completeness, MNAR-vs-MCAR missingness, PCA/batch, TMT channel balance, DIA q-values). Frames QC as a control chart against a per-instrument rolling baseline, not fixed cutoffs, and mandates inspecting raw boxplots, per-sample ID counts, total signal, and contaminant removal BEFORE normalizing -- because median normalization erases loading failures. Use when assessing proteomics data quality, diagnosing outlier samples, or deciding which samples to exclude before differential testing. The statistical test itself is differential-abundance; normalization mechanics are quantification; DIA q-value internals are dia-analysis.
tool_type
mixed
primary_tool
pandas

Version Compatibility

Reference examples tested with: pandas 2.2+, numpy 1.26+, scipy 1.12+, matplotlib 3.8+, scikit-learn 1.4+, limma 3.58+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Proteomics Quality Control -- A Three-Level Funnel Where the Matrix Sees the Failure Last

"Check the quality of my proteomics data" -> Read instrument, identification, and quantitative metrics as a descending funnel of silent failures, and inspect raw signal BEFORE normalizing -- because by the time a fault reaches the deliverable matrix, normalization has usually erased the evidence.

  • Python: pandas for matrix QC; matplotlib/seaborn for raw boxplots, correlation heatmaps, PCA
  • R: PTXQC createReport() for MaxQuant search-table QC; limma::plotMDS()/plotDensities(); MSstatsTMT dataProcessPlotsTMT() for TMT channel balance

Scope: This skill OWNS QC diagnosis across all three levels -- which metric localizes which fault, what threshold means trouble, and the mandatory inspect-before-normalize ordering. Normalization mechanics route to quantification; the differential test routes to differential-abundance; DIA q-value computation routes to dia-analysis. OUT OF SCOPE: running the statistical test, the normalization algorithms themselves, and DIA q-value/FDR internals.

The Single Most Important Modern Insight -- QC Is a Three-Level Funnel and Normalization Hides the Evidence

  1. QC is a three-level funnel of silent failures, and the deliverable matrix is the LAST place a problem becomes visible. Faults originate at the instrument (spray, calibration, column) or in identification (digestion, contamination, PTM artifacts), but a protein matrix only shows the downstream symptom -- a low correlation or an outlier sample. A matrix-only QC pass is one-third of the job and blind to where faults actually start. Localize by descending: read every metric together with its co-readouts, never alone.

  2. Almost no metric has a universal pass/fail cutoff; the defensible practice is a per-instrument, per-method control chart. Deviation from a lab's own rolling baseline (Levey-Jennings, +/-2 SD warn, +/-3 SD action) detects faults that a constant threshold misses or false-flags (Neely and Palmblad 2024). The numbers below seed a control chart, they are not standards-body limits.

  3. Median normalization HIDES loading problems that must be SEEN first. Median/quantile normalization works by forcing a chosen summary statistic of every sample equal. A sample that genuinely loaded 3x low sits visibly shifted down in a RAW boxplot -- an obvious, diagnosable defect. The instant the matrix is median-normalized, the algorithm shifts that sample up by a constant to match everyone's median; boxplots line up perfectly; the evidence is mathematically erased. Worse, the low-loaded sample's noisy low-abundance signal gets stretched up to mid-range and injected into the differential test while QC plots look pristine. MANDATE: inspect raw/un-normalized boxplots plus per-sample ID counts, total signal, and missing fraction BEFORE normalizing; remove loading/injection failures and contaminants; THEN normalize and re-plot on the survivors. The identical principle governs TMT channel-loading balance.

The Three Levels in One Table

LevelQuestionInputsFaults localized
1. Instrument / raw-signalIs the LC-MS hardware performing?Vendor .raw/.d; RawTools/RawBeans/rawrr/rawDiag/QuaMeter; Panorama AutoQCColumn, spray/emitter, mass analyzer/calibration
2. Identification / runDid this run identify peptides correctly?Search tables (MaxQuant txt/, FragPipe *.tsv, DIA-NN report); PTXQCDigestion, sample-handling PTM artifacts, contamination, FDR efficiency
3. Experiment / quantitativeAre the numbers reproducible and comparable?Protein/peptide intensity matrix; MSstatsTMT, pandas/limmaLoading/pipetting, batch, outliers, missingness, sample swaps

The most integrative metric (% MS2 identified / ID count) is the first alarm but the LEAST specific -- it moves whenever anything upstream degrades. The same protein-count drop means spray (erratic TIC + maxed injection time), column (lost RT + broad peaks + rising backpressure), or sample (high contaminant fraction) depending on what co-moves.

Tool Taxonomy

Tool / methodLevelCitationMechanism / roleWhen
PTXQC (R, CRAN)2+3Bielow 2016createReport() over MaxQuant txt/ or mzTab; per-metric scores in [0,1], QC heatmap PDFMaxQuant output, fast multi-metric report
RawTools / RawBeans1Kovalchik 2019; Morgenstern 2021Parse Thermo .raw for IT, TIC, FWHM, scan timingDiagnose instrument faults from raw files
rawrr / rawDiag1Kockmann 2021; Trachsel 2018R access to Orbitrap scan metadataCustom Level-1 plots / method optimization
QuaMeter1+2Ma 2012Vendor-independent ID-free and ID-based metricsCross-vendor Level-1 QC
Skyline + Panorama AutoQC1 (longitudinal)Bereman 2016Levey-Jennings + CUSUM/Moving-Range, SD-band flaggingSystem-suitability trending over time
MSstatsTMT3 (TMT)Huang 2020proteinSummarization(), dataProcessPlotsTMT(); filters isolation interference on importTMT channel-balance and QC plots
pandas / limma matrix QC3(this skill)Correlation, CV, completeness, PCA on the matrixExperiment-level QC (the code below)
differential-abundance(route OUT)--The moderated test itselfHit calling after QC passes
quantification(route OUT)--Normalization and imputation mechanicsThe how of normalizing
dia-analysis(route OUT)--DIA q-value/FDR internalsDIA-NN/Spectronaut report computation

Decision Tree by Scenario

ScenarioRecommendedWhy
MaxQuant txt/ folder, want fast multi-metric reportPTXQC createReport(txt_folder=...)Scores Level-2/3 metrics vs a representative file; one PDF
Protein-count drop, cause unknownDescend to Level 1: read TIC + injection time + RT/FWHM togetherCo-readouts localize spray vs column vs sample
Replicate correlation low for one sampleCheck if it correlates better with a DIFFERENT groupDistinguishes sample swap from prep failure
Boxplots flat but a sample feels wrongRe-plot the RAW (un-normalized) matrixNormalization erased the loading evidence
Deciding how to imputeDiagnose MNAR (left tail) vs MCAR (all-abundance) from the histogram FIRSTWrong imputer corrupts present/absent calls
TMT data, channel looks offMSstatsTMT QC plots on RAW reporter intensitiesSee the imbalance before global median rescales it
DIA matrix, how many proteins are realFilter Global.Q.Value and Global.PG.Q.Value, route q internals to dia-analysisPrecursor q != protein q; both needed
Long sample queue, drift suspectedInterspersed QC every 4th-5th injection + Levey-JenningsTurns one check into a time series

Default when uncertain: plot the RAW per-sample boxplots, ID counts, total signal, and missing fraction first; remove loading/injection failures and contaminants; only then normalize, re-plot, and proceed to correlation/CV/PCA on the survivors.

Inspect Raw Signal and Remove Contaminants Before Normalizing

Goal: Catch loading/injection failures and strip contaminant/decoy rows while they are still visible -- before normalization erases them.

Approach: Load the un-normalized matrix, plot per-sample boxplots plus ID counts and total signal, filter MaxQuant Potential contaminant/Reverse/Only identified by site rows, THEN log-transform and normalize on the survivors.

python
import pandas as pd
import numpy as np

contaminant_flags = ['Potential contaminant', 'Reverse', 'Only identified by site']

def strip_contaminant_rows(protein_groups):
    keep = pd.Series(True, index=protein_groups.index)
    for col in contaminant_flags:
        match = next((c for c in protein_groups.columns if c.lower() == col.lower()), None)  # MaxQuant casing varies by version -- match case-insensitively
        if match is not None:
            keep &= protein_groups[match].fillna('') != '+'  # MaxQuant marks flagged rows with a literal '+'
    return protein_groups[keep]

def raw_sample_qc(raw_intensities):
    return pd.DataFrame({
        'n_quantified': raw_intensities.notna().sum(),
        'total_signal': raw_intensities.sum(),
        'median_intensity': raw_intensities.median(),
        'missing_pct': 100 * raw_intensities.isna().sum() / len(raw_intensities)})

Read the boxplots before normalizing: a sample shifted >=2-3x below its group median is a loading/injection failure to exclude, not to rescale. The contaminant fraction of summed intensity should be small (PTXQC default flags >1%); keratin and trypsin autolysis dominate LOW-INPUT samples (single-cell, IPs, gel bands) because they are a roughly fixed absolute amount whose fractional share explodes as load shrinks.

Replicate Correlation on log2

Goal: Quantify reproducibility without letting a few abundant proteins fake agreement.

Approach: Correlate on log2 intensities (variance-stabilized, high-abundance tail compressed), report within-group pairs, and flag a sample correlating better with another group as a possible swap.

python
from itertools import combinations

def replicate_correlation(log2_intensities, sample_groups):
    corr = log2_intensities.corr(method='pearson')  # log2 first: Pearson on raw is a high-abundance artifact
    rows = []
    for group in sample_groups.unique():
        members = sample_groups[sample_groups == group].index
        for s1, s2 in combinations(members, 2):
            rows.append({'group': group, 's1': s1, 's2': s2, 'r': corr.loc[s1, s2]})
    return pd.DataFrame(rows)

Technical replicates r > 0.98 (instrument noise only); biological r ~ 0.90-0.98 (genuine variance, lower is expected and correct); soft floor r > 0.8 to retain a biological replicate. A Spearman check is a robustness aid only -- ranks discard the magnitude that quant QC cares about.

Coefficient of Variation on the Linear Scale

Goal: Summarize per-condition precision with a number that means what it says.

Approach: Compute CV = SD/mean on LINEAR (non-log) intensities; if only logged values exist use the geometric-CV formula. Report the median CV per condition (the per-protein distribution is right-skewed).

python
def median_cv_linear(linear_intensities, sample_groups):
    rows = []
    for group in sample_groups.unique():
        block = linear_intensities[sample_groups[sample_groups == group].index]
        per_protein_cv = block.std(axis=1) / block.mean(axis=1)  # base CV formula REQUIRES linear scale
        rows.append({'group': group, 'median_cv_pct': 100 * per_protein_cv.median()})
    return pd.DataFrame(rows)

def geometric_cv_from_log(log_intensities):
    sigma = log_intensities.std(axis=1) * np.log(2)  # convert log2 SD to natural-log SD
    return 100 * np.sqrt(np.expm1(sigma ** 2))  # gCV = sqrt(exp(sigma^2) - 1)

Applying the base formula to log-transformed data compresses CV ~14x (most proteins appear to have CV < 1%) -- meaningless (Brenes 2024). State normalization state, transform, and software params or the CV is uninterpretable: DIA-NN "High precision" mode silently median-normalizes, halving median CV vs "High accuracy". Technical median CV < ~10-20%, biological ~20-40%; a LOWER CV is not automatically better (loose FDR or faulty MS1 extraction produce artificially low CVs).

Missingness Mechanism and Completeness

Goal: Decide how to impute by first deciding why values are missing.

Approach: Diagnose the missingness profile -- left-tail concentration means MNAR (left-censored, abundance-dependent), all-abundance scatter means MCAR -- and filter on completeness before imputing only the shallow remainder.

python
def missingness_profile(log2_intensities, n_bins=20):
    observed = log2_intensities.stack()
    abundance_bins = pd.qcut(observed, n_bins, duplicates='drop')
    present_per_protein = log2_intensities.notna().mean(axis=1)
    mean_abundance = log2_intensities.mean(axis=1)
    return mean_abundance, present_per_protein  # plot present-fraction vs abundance: rising-with-abundance = MNAR

def completeness_filter(log2_intensities, sample_groups, min_valid_frac=0.7):
    keep = pd.Series(False, index=log2_intensities.index)
    for group in sample_groups.unique():
        block = log2_intensities[sample_groups[sample_groups == group].index]
        keep |= block.notna().mean(axis=1) >= min_valid_frac  # valid in >=70% of >=1 condition
    return log2_intensities[keep]

kNN-imputing a genuinely-absent (MNAR) value invents mid-range abundance and KILLS a real present/absent difference; a left-shifted draw (Perseus down-shifted normal, downshift=1.8 SD below the observed mean, width=0.3 of observed SD) on an MCAR gap FABRICATES a false low and inflates a difference. Match the imputer to the mechanism. The imputation mechanics themselves are quantification.

PCA and Batch Detection

Goal: See whether the dominant variance is biology or batch, and flag outlier samples.

Approach: On the normalized survivors, run PCA, color by condition and by batch, and test whether top PCs associate with batch.

python
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from scipy.stats import f_oneway

def pca_batch_check(normalized_log2, sample_info, batch_col='batch'):
    imputed = normalized_log2.apply(lambda r: r.fillna(r.median()), axis=1)  # temporary, for PCA only
    pcs = PCA(n_components=5).fit(StandardScaler().fit_transform(imputed.T))
    coords = pd.DataFrame(pcs.transform(StandardScaler().fit_transform(imputed.T)),
                          columns=[f'PC{i+1}' for i in range(5)], index=normalized_log2.columns).join(sample_info)
    for pc in ['PC1', 'PC2', 'PC3']:
        groups = [coords[coords[batch_col] == b][pc] for b in coords[batch_col].unique()]
        _, p = f_oneway(*groups)
        print(f'{pc} ~ {batch_col}: p={p:.4f}')
    return coords, pcs.explained_variance_ratio_

A sample isolated from its group is a removal/re-run candidate. If batch is PC1, correct it explicitly (ComBat, or include batch in the design matrix downstream) and re-inspect; never let batch be the dominant axis going into differential testing. Visualization of the projection routes to data-visualization/dimensionality-reduction-plots.

Per-Method Failure Modes

Median normalization hides loading failures

Trigger: Normalizing the matrix before inspecting raw per-sample signal. Mechanism: median-centering shifts each sample by a constant to equalize the very statistic that was the symptom of a low load. Symptom: flat, clean boxplots that hide a 3x-low sample now stretched into mid-range. Fix: plot RAW boxplots + ID counts + total signal first; exclude failures; then normalize.

Normalizing with contaminants still in the matrix

Trigger: Contaminant/decoy rows left in before log + normalize. Mechanism: keratin/trypsin/albumin inflate the denominator and shift the median; when their load differs across groups the differential gets normalized into the real proteins. Symptom: spurious fold changes; a contaminant fraction that varies by group. Fix: filter Potential contaminant + Reverse + Only identified by site BEFORE log + normalize.

Show full SKILL.md (1,062 more words)Show less
Wrong imputer for the missingness mechanism

Trigger: kNN on MNAR, or left-shift on MCAR. Mechanism: kNN borrows mid-range neighbors for a value that is low because it is absent; left-shift draws a deep low for a value missing at random. Symptom: killed present/absent calls (kNN-on-MNAR) or inflated false lows (left-shift-on-MCAR). Fix: diagnose left-tail vs all-abundance from the histogram first; mechanics route to quantification.

CV computed on log-transformed data

Trigger: Base CV formula applied after log2. Mechanism: SD/mean is defined for linear intensity; logging compresses it ~14x. Symptom: most proteins appear to have CV < 1%. Fix: compute on linear intensity, or use the geometric-CV formula on logged values; always state transform + normalization + software.

Pearson r on raw intensity

Trigger: Correlating un-logged intensities. Mechanism: a few high-abundance proteins dominate the covariance. Symptom: r = 0.99 while the bulk disagrees. Fix: log2 before correlating; Spearman as a robustness check only.

Bimodal mass-error histogram read as calibration

Trigger: A two-peaked ppm-error distribution. Mechanism: almost always monoisotopic mis-assignment or co-isolation (a search/sample problem), NOT calibration drift; a generous tolerance still IDs the mis-assigned precursors so ID rate looks fine. Symptom: bimodal histogram, normal ID rate. Fix: correct isotope-error tolerance/deisotoping, not recalibration.

Charge-state distribution off the platform baseline

Trigger: The fully-tryptic 2+ fraction drifts from the rolling baseline. Mechanism: a tryptic peptide carries two basic sites (C-terminal K/R + N-terminus) so 2+ dominates; excess 3+/4+ comes from internal basic residues left by missed cleavages, excess 1+ from poor ionization, short peptides, or contaminants. Symptom: raised high-charge fraction (a digestion/chemistry signal) or raised 1+ fraction (an ionization/spray signal). Fix: read the charge distribution together with the missed-cleavage rate (high charge co-moving with missed cleavages = under-digestion) to separate a chemistry problem from a spray problem a raw protein-count drop cannot resolve alone.

TMT ratio compression from co-isolation

Trigger: Isobaric quant with a wide isolation window. Mechanism: near-isobaric co-eluting precursors are co-isolated and add their own reporters across all channels, a uniform pedestal. Symptom: every fold-change compressed toward 1 (a real 10:1 reads ~5:1). Fix: filter isolation interference < 50%; prefer SPS-MS3 (McAlister 2014) / FAIMS / narrow windows; run a TKO or empty-channel control to measure the floor.

Quantitative Thresholds

ThresholdSourceRationale
Mass accuracy (internal/lock) median |err| < 1-3 ppm, single-mode centered 0CONVENTIONwidth approaching MS1 tolerance loses real IDs
iRT RT-fit R^2 > 0.99 (warn below)CONVENTION (mechanism firm)residual is just LC noise on a stable gradient
FWHM alarm on > 20-30% rise vs baseline; >= 8-10 points across FWHM (floor ~5)Kocher 2011peak capacity tracks peptide IDs; points needed for accurate AUC
% MS2 identified: < 20% bad, 20-35% ok, >= 35% greatPTXQC createYaml.Rgeneric pass marks, not a biological ceiling
Missed cleavages: >= 75-85% at 0 MC; flag > 25-30% with >= 1 MCOPERATIONAL (PTXQC-style)porcine trypsin ~78% efficient even ideally
Replicate Pearson r (log2): technical > 0.98, biological 0.90-0.98, floor 0.8CONVENTION (mechanism firm)log2 variance-stabilizes; biological variance is real
Median CV (linear): technical < 10-20%, biological 20-40%CONVENTIONDIA < DDA (no stochastic sampling); lower not always better
Completeness: valid in >= 50-70% of replicates in >= 1 conditionCONVENTIONfilter before imputing
Perseus MNAR imputation: downshift = 1.8 SD, width = 0.3Tyanova 2016deep left tail simulates below-LOD, narrowed so not mistaken for real
TMT channel deviation: investigate > ~2x, flag > ~3-4xCONVENTIONpipetting/labeling vs biology
Isolation interference < 50% PSM filter (< 30% stricter)CONVENTION (PD practice)above ~50% contaminant dominates, ratios uninterpretable
DIA precursor + protein q both <= 0.01, GLOBAL q, PICKED protein estimatorSTANDARDprecursor != protein FDR; route internals to dia-analysis
Levey-Jennings: +/-2 SD warn, +/-3 SD action; QC every 4th-5th injectionBereman 2016~95/99.7% of points under stable normal; catch drift early

Common Errors

Error / symptomCauseSolution
KeyError: 'Mass Error [ppm]'MaxQuant column casing varies by versionmatch case-insensitively (Mass error vs Mass Error); never hard-code
Contaminant rows have True/False, filter keeps allMaxQuant flags with a literal '+', not a booleanfilter col != '+'
CV unexpectedly tiny (< 1%)base CV formula applied to log2 datacompute on linear intensity or use geometric CV
r = 0.99 but samples clearly differPearson on raw (un-logged) intensitylog2 transform before correlating
PCA dominated by injection daybatch effect, not biologycorrect (ComBat / batch in design) and re-inspect; do not proceed
PTXQC "not found" via BiocManagerPTXQC is on CRAN, not Bioconductorinstall.packages('PTXQC')
createReport() errors on a dataframe argit takes a txt-folder path / mzTab / YAML, not dataframespass txt_folder= (the MaxQuant txt/ directory)

References

  • Bielow C, Mastrobuoni G, Kempa S. Proteomics Quality Control: Quality Control Software for MaxQuant Results. J Proteome Res 2016;15(3):777-787.
  • Kovalchik KA, Colborne S, Spencer SEP, et al. RawTools: Rapid and Dynamic Interrogation of Orbitrap Data Files for Mass Spectrometer System Management. J Proteome Res 2019;18(2):700-708.
  • Morgenstern D, Barzilay R, Levin Y. RawBeans: A Simple, Vendor-Independent, Raw-Data Quality-Control Tool. J Proteome Res 2021;20(4):2098-2104.
  • Kockmann T, Panse C. The rawrr R Package: Direct Access to Orbitrap Data and Beyond. J Proteome Res 2021;20(4):2028-2034.
  • Trachsel C, Panse C, Kockmann T, et al. rawDiag: An R Package Supporting Rational LC-MS Method Optimization for Bottom-up Proteomics. J Proteome Res 2018;17(8):2908-2914.
  • Ma ZQ, Polzin KO, Dasari S, et al. QuaMeter: Multivendor Performance Metrics for LC-MS/MS Proteomics Instrumentation. Anal Chem 2012;84(14):5845-5850.
  • Bereman MS, Beri J, Sharma V, et al. An Automated Pipeline to Monitor System Performance in Liquid Chromatography-Tandem Mass Spectrometry Proteomic Experiments. J Proteome Res 2016;15(12):4763-4769.
  • Kocher T, Swart R, Mechtler K. Ultra-High-Pressure RPLC Hyphenated to an LTQ-Orbitrap Velos Reveals a Linear Relation between Peak Capacity and Number of Identified Peptides. Anal Chem 2011;83(7):2699-2704.
  • Tyanova S, Temu T, Sinitcyn P, et al. The Perseus computational platform for comprehensive analysis of (prote)omics data. Nat Methods 2016;13(9):731-740.
  • Brenes AJ. Calculating and Reporting Coefficients of Variation for DIA-Based Proteomics. J Proteome Res 2024;23(12):5274-5278.
  • McAlister GC, Nusinow DP, Jedrychowski MP, et al. MultiNotch MS3 Enables Accurate, Sensitive, and Multiplexed Detection of Differential Expression across Cancer Cell Line Proteomes. Anal Chem 2014;86(14):7150-7158.
  • Huang T, Choi M, Tzouros M, et al. MSstatsTMT: Statistical Detection of Differentially Abundant Proteins in Experiments with Isobaric Labeling and Multiple Mixtures. Mol Cell Proteomics 2020;19(10):1706-1723.
  • Neely BA, Palmblad M, et al. Quality Control in the Mass Spectrometry Proteomics Core: A Practical Primer. J Biomol Tech 2024;35(3).
  • data-import - Load search-engine output and intensity matrices before QC
  • quantification - Normalization and imputation mechanics that QC mandates running AFTER inspection
  • differential-abundance - The moderated statistical test QC gates
  • dia-analysis - DIA q-value/FDR internals behind the protein-count QC
  • data-visualization/dimensionality-reduction-plots - PCA/MDS projection plotting
  • workflows/proteomics-pipeline - End-to-end pipeline placing QC before differential testing

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in proteomics/proteomics-qc of GPTomics/bioSkills.

  • SKILL.md
  • examples/qc_analysis.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Proteomics Proteomics Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Proteomics Proteomics Qc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Proteomics Proteomics Qc this skillGPTomics/bioSkills1.2k1 repos~6.4kAutomated safety check: PassMIT
Tooluniverse Metabolomics Analysiswu-yc/LabClaw1.1k2 repos~5.9kAutomated safety check: PassNone
Bio Spatial Transcriptomics Spatial PreprocessingFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~2kAutomated safety check: PassNone
Exploratory Data Analysisspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: PassMIT
Pyopenmsdavila7/claude-code-templates33k11 repos~1.4kAutomated safety check: PassMIT
Gwas Databasedavila7/claude-code-templates33k10 repos~5kAutomated safety check: PassMIT

Similar skills

  • Analyze metabolomics data including metabolite identification, quantification, pathway analysis, and metabolic flux.

    1.1k GitHub starsUsed in 2 repos~5.9k tokens
    Research & ScienceAuto-check passed
  • Bio Spatial Transcriptomics Spatial Preprocessing

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quality control, filtering, normalization, and feature selection for spatial transcriptomics data.

    3.1k GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Pyopenms

    davila7/claude-code-templates

    Python interface to OpenMS for mass spectrometry data analysis.

    33k GitHub starsUsed in 11 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Gwas Database

    davila7/claude-code-templates

    Query NHGRI-EBI GWAS Catalog for SNP-trait associations. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 10 repos~5k tokens
    Data & AnalyticsAuto-check passed
  • Polars Bio

    ClawBio/ClawBio

    Fast genomic interval operations (overlap, nearest, merge, coverage, cluster, complement, subtract, count-overlaps), multi-format bioinformatics I/O, DataFusion SQL, and pileup on Polars DataFrames…

    1.2k GitHub stars~3.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Proteomics Proteomics Qc

What does Bio Proteomics Proteomics Qc do?

Quality control for bottom-up proteomics across three levels -- instrument/raw-signal (mass accuracy, RT/iRT fit, FWHM, TIC vs injection time, % MS2 identified), identification/run (missed…. Bio Proteomics Proteomics Qc is an agent skill from GPTomics/bioSkills. Quality control for bottom-up proteomics across three levels -- instrument/raw-signal (mass accuracy, RT/iRT fit, FWHM, TIC vs injection time, % MS2 identified), identification/run (missed cleavages, charge states, PTM handling artifacts, contaminants), and experiment/quantitative (replicate correlation on log2, CV on the linear scale, completeness, MNAR-vs-MCAR missingness, PCA/batch, TMT channel balance, DIA q-values).

When should I use Bio Proteomics Proteomics Qc?

Bio Proteomics Proteomics Qc fits situations like: assessing proteomics data quality; diagnosing outlier samples; deciding which samples to exclude before differential testing.

How do I install Bio Proteomics Proteomics Qc in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-proteomics-qc -a claude-code`. Or copy the skill folder (proteomics/proteomics-qc in GPTomics/bioSkills) into .claude/skills/bio-proteomics-proteomics-qc in your project. Claude Code loads it when a task matches its description.

How do I install Bio Proteomics Proteomics Qc in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-proteomics-qc -a codex`. Or copy the skill folder (proteomics/proteomics-qc in GPTomics/bioSkills) into .agents/skills/bio-proteomics-proteomics-qc in your project. Codex loads it when a task matches its description.

Can I use Bio Proteomics Proteomics Qc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-proteomics-proteomics-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-proteomics-proteomics-qc, .gemini/skills/bio-proteomics-proteomics-qc, .github/skills/bio-proteomics-proteomics-qc and .opencode/skills/bio-proteomics-proteomics-qc in your project.

What does Bio Proteomics Proteomics Qc need to run?

Going by SKILL.md and its folder, Bio Proteomics Proteomics Qc needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Proteomics Proteomics Qc access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Proteomics Proteomics Qc safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Proteomics Proteomics Qc use?

Bio Proteomics Proteomics Qc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Proteomics Proteomics Qc use?

About 6.4k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Proteomics Proteomics Qc?

Skills that share tags, products or a category with Bio Proteomics Proteomics Qc: Tooluniverse Metabolomics Analysis (wu-yc/LabClaw, 1.1k stars), Bio Spatial Transcriptomics Spatial Preprocessing (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Exploratory Data Analysis (spacering-net/codeg, 3.9k stars) and Pyopenms (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Proteomics Proteomics Qc?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.