13C Metabolic Flux Analysis
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
MaxQuant + Perseus proteomics pipeline: run MaxQuant for LFQ and SILAC; parse proteinGroups.txt in Python; filter contaminants/decoys; log2 + median-normalize; impute MNAR; t-test with FDR; volcano…
$ npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills maxquant-proteomics --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/proteomics-protein-engineering/maxquant-proteomics .claude/skills/maxquant-proteomics && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "maxquant-proteomics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/maxquant-proteomics into .claude/skills/maxquant-proteomics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "maxquant-proteomics", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/maxquant-proteomicsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills maxquant-proteomics --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/proteomics-protein-engineering/maxquant-proteomics .agents/skills/maxquant-proteomics && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "maxquant-proteomics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/maxquant-proteomics into .agents/skills/maxquant-proteomics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "maxquant-proteomics", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills maxquant-proteomics --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/proteomics-protein-engineering/maxquant-proteomics .cursor/skills/maxquant-proteomics && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "maxquant-proteomics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/maxquant-proteomics into .cursor/skills/maxquant-proteomics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "maxquant-proteomics", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/proteomics-protein-engineering/maxquant-proteomics--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills maxquant-proteomics --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/proteomics-protein-engineering/maxquant-proteomics .gemini/skills/maxquant-proteomics && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "maxquant-proteomics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/maxquant-proteomics into .gemini/skills/maxquant-proteomics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "maxquant-proteomics", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills maxquant-proteomicsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/proteomics-protein-engineering/maxquant-proteomics .github/skills/maxquant-proteomics && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "maxquant-proteomics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/maxquant-proteomics into .github/skills/maxquant-proteomics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "maxquant-proteomics", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills maxquant-proteomics --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/proteomics-protein-engineering/maxquant-proteomics .opencode/skills/maxquant-proteomics && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "maxquant-proteomics" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/maxquant-proteomics into .opencode/skills/maxquant-proteomics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "maxquant-proteomics", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
maxquant-proteomicsMaxQuant + Perseus proteomics pipeline: run MaxQuant for LFQ and SILAC; parse proteinGroups.txt in Python; filter contaminants/decoys; log2 + median-normalize; impute MNAR; t-test with FDR; volcano…
Maxquant Proteomics is an agent skill from jaechang-hits/SciAgent-Skills. MaxQuant + Perseus proteomics pipeline: run MaxQuant for LFQ and SILAC; parse proteinGroups.txt in Python; filter contaminants/decoys; log2 + median-normalize; impute MNAR; t-test with FDR; volcano plot; GO/pathway enrichment. Use Proteome Discoverer for Thermo-native processing; FragPipe/MSFragger for GPU-accelerated DB search.
Its SKILL.md is about 7.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. It works with Python. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
string-db.orgAlso links to:
maxquant.orgdoi.orgmaxquant.netgithub.comgseapy.readthedocs.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Maxquant Proteomics loads about 7.2k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 1,339 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 1,339 words, ~7,249 tokens.
.claude/skills/maxquant-proteomics/SKILL.md (or your agent's skills folder).MaxQuant is the community-standard software for label-free quantification (LFQ) and SILAC proteomics. It performs database search, protein grouping, and intensity-based quantification from raw LC-MS/MS files, producing proteinGroups.txt as the primary output. Downstream statistical analysis — filtering, normalization, imputation, differential abundance testing, and visualization — is performed in Python using pandas, scipy, and matplotlib/seaborn, mirroring the Perseus workflow in a reproducible scripting environment.
proteinGroups.txt) for comparison with published datasetspandas, numpy, scipy, matplotlib, seaborn, statsmodels, gseapy.raw files or mzML-converted files; FASTA protein database (UniProt reviewed + contaminant database)pip install pandas numpy scipy matplotlib seaborn statsmodels gseapy# Install pyMaxQuant for programmatic mqpar.xml configuration
pip install pymaxquantSettle these with the user before writing any analysis code.
decisions:
- id: D1
param: quantificationStrategy
kind: required
source: user
ask: "How were samples quantified - label-free, SILAC, or isobaric tags?"
default: null
- id: D2
param: proteinSequenceDatabase
kind: required
source: user
ask: "Which organism's protein FASTA, and which release, should spectra be searched against?"
default: null
- id: D3
param: digestionEnzyme
kind: required
source: user
ask: "Which protease was used, and how many missed cleavages should be allowed?"
default: "trypsin, up to 2 missed cleavages"
- id: D4
param: modifications
kind: required
source: user
ask: "Which modifications are fixed, and which variable - and is there an enrichment such as phospho to search for?"
default: "fixed carbamidomethyl on cysteine; variable oxidation and N-terminal acetylation"
- id: D5
param: falseDiscoveryRates
kind: required
source: user
ask: "What false-discovery rate should peptide and protein identifications be held to?"
default: "1% at both levels"
- id: D6
param: matchBetweenRuns
kind: required
source: user
depends_on: [D1]
ask: "Should identifications be transferred between runs by retention time, recovering more proteins at the cost of some transferred errors?"
default: "off"
- id: D7
param: lfqMinRatioCount
kind: optional
source: user
depends_on: [D1]
ask: "How many shared peptides must two samples have before their abundances are normalized against each other?"
default: 2
skip_if: "labelled quantification"
- id: D8
param: missingValueImputation
kind: required
source: user
ask: "How should proteins missing from a sample be treated - left missing, or imputed as below the detection limit?"
default: "imputed from a downshifted distribution"
- id: D9
param: differentialThresholds
kind: required
source: user
ask: "What significance and fold-change cuts define a changed protein?"
default: "FDR 0.05, log2 fold change 1.0"D8 is where proteomics differs from RNA-seq and where results most often turn on an unexamined default. A protein absent from every control and present in every treated sample is the strongest possible result or an artefact of detection limits, and the imputation choice decides which one the volcano plot shows. D6 interacts with it: transferred identifications fill in exactly the values that would otherwise be imputed.
import pandas as pd
import numpy as np
# Load MaxQuant output
df = pd.read_csv("combined/txt/proteinGroups.txt", sep="\t", low_memory=False)
print(f"Raw protein groups: {len(df)}")
# Filter contaminants, reverse decoys, only-by-site
mask = (
(df["Potential contaminant"] != "+") &
(df["Reverse"] != "+") &
(df["Only identified by site"] != "+")
)
df = df[mask].copy()
print(f"After filtering: {len(df)} protein groups")
# Extract LFQ intensity columns
lfq_cols = [c for c in df.columns if c.startswith("LFQ intensity ")]
print(f"LFQ columns: {lfq_cols}")
# Log2-transform (0 → NaN)
lfq = df[lfq_cols].replace(0, np.nan)
lfq = np.log2(lfq)
print(f"Valid values per sample:\n{lfq.notna().sum()}")MaxQuant is controlled by an XML parameter file (mqpar.xml). Edit it programmatically to set file paths, enzyme, modifications, and quantification type before running the search.
import xml.etree.ElementTree as ET
def update_mqpar(template_path: str, output_path: str,
raw_files: list[str], fasta_path: str,
experiment_names: list[str]) -> None:
"""Update mqpar.xml with sample-specific file paths."""
tree = ET.parse(template_path)
root = tree.getroot()
# Set raw file paths
file_paths_node = root.find(".//filePaths")
file_paths_node.clear()
for rf in raw_files:
elem = ET.SubElement(file_paths_node, "string")
elem.text = rf
# Set experiment names (maps files to conditions)
experiments_node = root.find(".//experiments")
experiments_node.clear()
for name in experiment_names:
elem = ET.SubElement(experiments_node, "string")
elem.text = name
# Set FASTA database
fasta_node = root.find(".//fastaFiles/FastaFileInfo/fastaFilePath")
fasta_node.text = fasta_path
tree.write(output_path, xml_declaration=True, encoding="utf-8")
print(f"Written: {output_path}")
# Example usage
raw_files = [
r"C:\Data\ctrl_rep1.raw",
r"C:\Data\ctrl_rep2.raw",
r"C:\Data\treat_rep1.raw",
r"C:\Data\treat_rep2.raw",
]
update_mqpar(
template_path="mqpar_template.xml",
output_path="mqpar.xml",
raw_files=raw_files,
fasta_path=r"C:\Databases\human_uniprot_contaminants.fasta",
experiment_names=["ctrl", "ctrl", "treat", "treat"],
)Key mqpar.xml parameters (set in template or edit directly):
<!-- Enzyme and search settings -->
<enzymes>
<string>Trypsin/P</string>
</enzymes>
<maxMissedCleavages>2</maxMissedCleavages>
<variableModifications>
<string>Oxidation (M)</string>
<string>Acetyl (Protein N-term)</string>
</variableModifications>
<fixedModifications>
<string>Carbamidomethyl (C)</string>
</fixedModifications>
<!-- LFQ settings -->
<lfqMode>1</lfqMode> <!-- 1 = LFQ enabled -->
<lfqMinRatioCount>2</lfqMinRatioCount> <!-- minimum peptides for LFQ -->
<matchBetweenRuns>True</matchBetweenRuns>
<!-- FDR thresholds -->
<peptideFdr>0.01</peptideFdr>
<proteinFdr>0.01</proteinFdr>MaxQuant can be run headlessly from the Windows command prompt using the bundled MaxQuantCmd.exe.
REM Windows Command Prompt — run MaxQuant with configured mqpar.xml
REM Adjust path to match your MaxQuant installation directory
set MQ_PATH=C:\Program Files\MaxQuant\bin\MaxQuantCmd.exe
set MQPAR=C:\Projects\proteomics\mqpar.xml
"%MQ_PATH%" "%MQPAR%"
REM For specific workflow steps only (useful for reruns):
REM Step IDs: 0=write tables, 1=feature detection, 7=peptide identification
"%MQ_PATH%" "%MQPAR%" --steps 1,7,11# Cross-platform: run MaxQuant under Wine on Linux/macOS (CI/server use)
wine MaxQuantCmd.exe mqpar.xml
# Monitor progress log
tail -f combined/proc/#runningTimes.txtFilter out reverse decoys, potential contaminants, and proteins only identified by modification site.
import pandas as pd
import numpy as np
def load_protein_groups(path: str) -> pd.DataFrame:
"""Load MaxQuant proteinGroups.txt with quality filters applied."""
df = pd.read_csv(path, sep="\t", low_memory=False)
print(f"Total protein groups: {len(df)}")
# Remove reverse decoys, contaminants, and only-by-site hits
n_before = len(df)
df = df[
(df.get("Reverse", pd.Series("")) != "+") &
(df.get("Potential contaminant", pd.Series("")) != "+") &
(df.get("Only identified by site", pd.Series("")) != "+")
].copy()
print(f"After quality filter: {len(df)} ({n_before - len(df)} removed)")
# Parse gene names (take first entry for multi-gene groups)
df["Gene names"] = df["Gene names"].fillna("Unknown").str.split(";").str[0]
# Set unique index on majority protein ID
df = df.set_index("Majority protein IDs")
return df
# Load output
pg = load_protein_groups("combined/txt/proteinGroups.txt")
# Identify LFQ intensity columns
lfq_cols = [c for c in pg.columns if c.startswith("LFQ intensity ")]
print(f"LFQ samples ({len(lfq_cols)}): {lfq_cols}")
# Output: LFQ samples (6): ['LFQ intensity ctrl_1', 'LFQ intensity ctrl_2', ...]Replace zero intensities with NaN (missing values in MaxQuant are exported as 0), log2-transform, then apply per-sample median centering.
def prepare_lfq_matrix(df: pd.DataFrame, lfq_cols: list[str]) -> pd.DataFrame:
"""Extract, transform, and normalize LFQ intensity matrix."""
# Extract and rename columns (strip 'LFQ intensity ' prefix)
lfq = df[lfq_cols].copy()
lfq.columns = [c.replace("LFQ intensity ", "") for c in lfq_cols]
# Replace 0 with NaN (MaxQuant encodes missing as 0)
lfq = lfq.replace(0, np.nan)
# Log2 transform
lfq = np.log2(lfq)
# Median centering per sample (subtract per-column median of valid values)
col_medians = lfq.median(axis=0)
global_median = col_medians.median()
lfq = lfq.subtract(col_medians, axis=1).add(global_median)
print(f"Matrix shape: {lfq.shape}")
print(f"Missing values per sample:\n{lfq.isna().sum()}")
print(f"Valid values per sample:\n{lfq.notna().sum()}")
return lfq
lfq_matrix = prepare_lfq_matrix(pg, lfq_cols)
# Matrix shape: (3241, 6)
# Missing values per sample: ctrl_1: 421, ctrl_2: 389, ...Missing-not-at-random (MNAR) values arise from proteins below the detection limit. Impute from the low end of the observed intensity distribution — the standard Perseus approach.
def impute_mnar(lfq: pd.DataFrame,
width: float = 0.3,
downshift: float = 1.8,
random_state: int = 42) -> pd.DataFrame:
"""
Impute MNAR missing values from a downshifted Gaussian.
Parameters
----------
width : std of imputation distribution (fraction of sample std)
downshift : downshift in units of sample std below mean
random_state : for reproducibility
"""
rng = np.random.default_rng(random_state)
lfq_imp = lfq.copy()
for col in lfq_imp.columns:
col_data = lfq_imp[col].dropna()
col_mean = col_data.mean()
col_std = col_data.std()
n_missing = lfq_imp[col].isna().sum()
if n_missing > 0:
imputed = rng.normal(
loc=col_mean - downshift * col_std,
scale=width * col_std,
size=n_missing,
)
lfq_imp.loc[lfq_imp[col].isna(), col] = imputed
print(f"Imputed {lfq.isna().sum().sum()} missing values")
return lfq_imp
lfq_imputed = impute_mnar(lfq_matrix)
# Imputed 2847 missing valuesPerform two-sample t-tests for each protein between conditions, then apply Benjamini-Hochberg FDR correction.
from scipy import stats
from statsmodels.stats.multitest import multipletests
def differential_abundance(lfq: pd.DataFrame,
group_a: list[str],
group_b: list[str],
alpha: float = 0.05) -> pd.DataFrame:
"""
Two-sample t-test + BH FDR correction for all proteins.
Parameters
----------
group_a, group_b : sample name lists for each condition
alpha : FDR threshold
"""
results = []
for protein_id, row in lfq.iterrows():
a_vals = row[group_a].dropna().values
b_vals = row[group_b].dropna().values
if len(a_vals) >= 2 and len(b_vals) >= 2:
t_stat, p_val = stats.ttest_ind(a_vals, b_vals, equal_var=False)
log2fc = b_vals.mean() - a_vals.mean()
else:
t_stat, p_val, log2fc = np.nan, np.nan, np.nan
results.append({
"protein_id": protein_id,
"log2FC": log2fc,
"pvalue": p_val,
"t_stat": t_stat,
})
res_df = pd.DataFrame(results).set_index("protein_id")
# BH FDR correction on valid p-values
valid = res_df["pvalue"].notna()
_, padj, _, _ = multipletests(res_df.loc[valid, "pvalue"], method="fdr_bh")
res_df.loc[valid, "padj"] = padj
# Add significance flag
res_df["significant"] = (res_df["padj"] < alpha) & (res_df["pvalue"].notna())
sig_count = res_df["significant"].sum()
print(f"Significant proteins (FDR < {alpha}): {sig_count}")
return res_df.sort_values("padj")
# Define sample groups
group_ctrl = ["ctrl_1", "ctrl_2", "ctrl_3"]
group_treat = ["treat_1", "treat_2", "treat_3"]
results = differential_abundance(lfq_imputed, group_ctrl, group_treat)
print(results[results["significant"]].head(10))
# Significant proteins (FDR < 0.05): 312Plotting: compute the DEG table, then read
skills/data-visualization/omics-plotting/SKILL.mdand follow its "Volcano" recipe on the exported CSV (→figures/volcano_plot.png).
# Volcano input: log2FC vs padj + protein labels; render with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) "Volcano" recipe.
volcano_df = results[["log2FC", "padj"]].copy()
volcano_df["gene"] = pg["Gene names"].reindex(results.index)
volcano_df.dropna(subset=["log2FC", "padj"]).to_csv("volcano_input.csv")
print(f"Volcano input: {int(volcano_df['padj'].notna().sum())} proteins -> volcano_input.csv")
# -> figures/volcano_plot.png via omics-plotting (map log2FC -> log2FoldChange)Run over-representation analysis (ORA) on significantly up- and down-regulated proteins using gseapy's Enrichr API.
import gseapy as gp
def run_enrichment(results: pd.DataFrame,
gene_names: pd.Series,
gene_sets: list[str] | None = None,
top_n: int = 20) -> dict[str, pd.DataFrame]:
"""
ORA enrichment for up- and down-regulated proteins via Enrichr.
Parameters
----------
gene_sets : Enrichr gene set libraries (default: GO BP + KEGG)
top_n : top results to display per direction
"""
if gene_sets is None:
gene_sets = ["GO_Biological_Process_2023", "KEGG_2021_Human"]
results_out = {}
gene_map = gene_names.reindex(results.index).fillna("Unknown")
for direction in ("up", "down"):
sig = results[
(results["significant"]) &
(results["log2FC"] > 0 if direction == "up" else results["log2FC"] < 0)
]
gene_list = gene_map.reindex(sig.index).tolist()
print(f"{direction.capitalize()}-regulated: {len(gene_list)} proteins")
if len(gene_list) < 5:
print(f" Too few proteins for enrichment (n={len(gene_list)}), skipping")
continue
enr = gp.enrichr(
gene_list=gene_list,
gene_sets=gene_sets,
organism="Human",
outdir=f"enrichment_{direction}",
cutoff=0.05,
)
top_results = enr.results.sort_values("Adjusted P-value").head(top_n)
results_out[direction] = top_results
print(top_results[["Term", "Adjusted P-value", "Overlap"]].to_string(index=False))
return results_out
enr_results = run_enrichment(results, pg["Gene names"])| Parameter | Default | Range / Options | Effect |
|---|---|---|---|
matchBetweenRuns | False | True / False | Transfers identifications across runs by retention time matching; increases quantified protein count 10–30% |
lfqMinRatioCount | 2 | 1–5 | Minimum peptide pairs required for LFQ normalization; lower values increase coverage but reduce accuracy |
maxMissedCleavages | 2 | 0–4 | Tryptic missed cleavages allowed; increase for samples with poor digestion |
peptideFdr / proteinFdr | 0.01 | 0.001–0.05 | FDR thresholds for peptide and protein identifications |
MNAR downshift | 1.8 | 1.5–2.5 | Shifts imputation distribution below detection limit in units of column std; larger = more conservative imputation |
MNAR width | 0.3 | 0.1–0.5 | Width of imputed distribution relative to column std |
t-test alpha | 0.05 | 0.01–0.1 | FDR significance threshold for differential abundance |
fc_threshold (volcano) | 1.0 | 0.5–2.0 | log2 fold-change cutoff for "significant" label in volcano plot |
| File | Content |
|---|---|
proteinGroups.txt | Primary output: one row per protein group with LFQ/SILAC intensities, peptide counts, sequence coverage |
peptides.txt | Peptide-level quantification with charge states and modifications |
evidence.txt | Individual MS/MS identifications (one row per peptide-spectrum match) |
msms.txt | Full MS/MS scan data including fragment ions and scores |
summary.txt | Per-raw-file statistics: identifications, MS/MS counts, calibration |
LFQ intensity <sample> columns.iBAQ column.| Perseus step | Python equivalent |
|---|---|
| Filter rows by categorical column | df[df["Reverse"] != "+"] |
| Replace 0 with NaN | df.replace(0, np.nan) |
| Log2 transform | np.log2(df) |
| Median normalization | df.subtract(df.median()).add(global_median) |
| MNAR imputation (normal distribution) | impute_mnar() function above |
| Two-sample t-test | scipy.stats.ttest_ind() + multipletests() |
| Volcano plot | matplotlib.pyplot scatter + threshold lines |
| Hierarchical clustering | seaborn.clustermap() |
When to use: SILAC experiments with H/L or H/M/L labeling instead of LFQ.
import pandas as pd
import numpy as np
# Load proteinGroups.txt for SILAC experiment
df = pd.read_csv("combined/txt/proteinGroups.txt", sep="\t", low_memory=False)
# Filter contaminants and decoys
df = df[(df["Reverse"] != "+") & (df["Potential contaminant"] != "+")].copy()
# Extract H/L ratio columns (log2-transformed)
ratio_cols = [c for c in df.columns if c.startswith("Ratio H/L ") and "normalized" in c.lower()]
if not ratio_cols:
# Fall back to non-normalized
ratio_cols = [c for c in df.columns if c.startswith("Ratio H/L")]
print(f"SILAC ratio columns: {ratio_cols}")
ratios = df[ratio_cols].copy().replace(0, np.nan)
# Log2 transform ratios
log2_ratios = np.log2(ratios)
log2_ratios.columns = [c.replace("Ratio H/L normalized ", "") for c in ratio_cols]
# Summary statistics per sample
print(log2_ratios.describe().round(3))When to use: visualizing patterns across all significant proteins simultaneously. Compute the z-scored matrix below, then read skills/data-visualization/omics-plotting/SKILL.md and follow its "Clustered expression heatmap" recipe (→ figures/heatmap.pdf).
import pandas as pd
def prep_heatmap_matrix(lfq_imputed: pd.DataFrame,
results: pd.DataFrame,
gene_names: pd.Series,
top_n: int = 50,
out_path: str = "heatmap_matrix.csv") -> pd.DataFrame:
"""Z-scored expression matrix of the top significant proteins, for the omics-plotting heatmap."""
sig_proteins = results[results["significant"]].nsmallest(top_n, "padj").index
heatmap_data = lfq_imputed.loc[sig_proteins].copy()
heatmap_data.index = gene_names.reindex(sig_proteins).fillna(sig_proteins)
# Z-score per row
heatmap_z = heatmap_data.subtract(heatmap_data.mean(axis=1), axis=0).divide(
heatmap_data.std(axis=1).replace(0, 1), axis=0
)
heatmap_z.to_csv(out_path)
return heatmap_z
mat = prep_heatmap_matrix(lfq_imputed, results, pg["Gene names"])
print(f"Heatmap matrix: {mat.shape} -> heatmap_matrix.csv")
# Render with the omics-plotting SKILL (`skills/data-visualization/omics-plotting/SKILL.md`) "Clustered expression heatmap" recipe -> figures/heatmap.pdfWhen to use: protein-protein interaction network analysis and enrichment without downloading gene sets locally.
import requests
import pandas as pd
def string_enrichment(gene_list: list[str],
species: int = 9606,
fdr_threshold: float = 0.05) -> pd.DataFrame:
"""Query STRING /enrichment endpoint for GO/KEGG enrichment."""
url = "https://string-db.org/api/json/enrichment"
params = {
"identifiers": "\r".join(gene_list),
"species": species,
"caller_identity": "maxquant_proteomics_skill",
}
response = requests.post(url, data=params)
response.raise_for_status()
enr_df = pd.DataFrame(response.json())
if enr_df.empty:
print("No enrichment results returned")
return enr_df
enr_df = enr_df[enr_df["fdr"].astype(float) < fdr_threshold]
enr_df = enr_df.sort_values("fdr")
print(f"Enriched terms (FDR < {fdr_threshold}): {len(enr_df)}")
print(enr_df[["category", "term", "description", "fdr", "number_of_genes"]].head(15).to_string(index=False))
return enr_df
# Significant up-regulated gene names
up_genes = pg.loc[
results[(results["significant"]) & (results["log2FC"] > 1)].index, "Gene names"
].tolist()
string_enr = string_enrichment(up_genes)When to use: reading and filtering MaxQuant text files with a higher-level API.
# pyMaxQuant provides typed accessors for MaxQuant output files
# Install: pip install pymaxquant
from maxquant.io import read_protein_groups
# Load with built-in contaminant filtering
pg_clean = read_protein_groups(
"combined/txt/proteinGroups.txt",
filter_invalid=True, # removes reverse, contaminant, only-by-site
)
print(f"Loaded {len(pg_clean)} filtered protein groups")
# Access LFQ columns via helper
lfq_df = pg_clean.filter(like="LFQ intensity")
print(f"LFQ matrix: {lfq_df.shape}")| File | Description |
|---|---|
combined/txt/proteinGroups.txt | Main MaxQuant output: protein groups with LFQ intensities, peptide counts, unique peptides, iBAQ |
combined/txt/peptides.txt | Peptide-level quantification with modifications and charge states |
combined/txt/summary.txt | Per-raw-file QC statistics: identification rates, MS/MS counts |
results_differential.csv | Differential abundance table: log2FC, pvalue, padj, significant per protein |
volcano_plot.pdf | Volcano plot with up/down-regulated proteins colored and top proteins labeled |
heatmap.pdf | Hierarchical clustering heatmap of top significant proteins (Z-score normalized) |
enrichment_up/ | gseapy output directory: GO/KEGG enrichment for up-regulated proteins |
enrichment_down/ | gseapy output directory: GO/KEGG enrichment for down-regulated proteins |
| Problem | Cause | Solution |
|---|---|---|
| MaxQuant produces 0 protein identifications | Wrong FASTA database or enzyme settings; raw file path not found | Verify .raw file paths in mqpar.xml are absolute Windows paths; confirm enzyme matches experiment (Trypsin/P vs Trypsin); check summary.txt for identification rate |
| All LFQ intensities are 0 after filtering | matchBetweenRuns off + sparse data, or wrong column selection | Check combined/txt/proteinGroups.txt directly; use pg.filter(like="LFQ intensity") to confirm column names; lower lfqMinRatioCount to 1 |
| Too many missing values after log2 transform | Insufficient replicates, inconsistent sample loading, or undetected peptides | Enable matchBetweenRuns; verify equal protein loading (Bradford/BCA); consider stricter valid-value filter (require 3/3 per group) before imputation |
Memory error loading proteinGroups.txt | File is large (>500 MB for DDA with many samples) | Use pd.read_csv(..., low_memory=False, usecols=[...]) to select only needed columns; or use pd.read_csv(..., chunksize=...) |
| gseapy Enrichr returns empty results | Gene symbols unrecognized or network timeout | Ensure gene list uses HGNC symbols (not UniProt IDs); check internet connectivity; use gp.enrichr(..., timeout=60) |
| Volcano plot: all proteins in "ns" | FDR threshold too stringent or padj not calculated | Verify multipletests returned valid FDR values; try relaxing alpha to 0.1; check sample group assignments are correct |
| MaxQuant run hangs at "Feature detection" | Low memory (MaxQuant needs 4–8 GB RAM per 3–4 raw files) | Process files in smaller batches; increase system RAM; close other applications |
| Imputation inflates false positives | Imputing too aggressively (low downshift) | Increase downshift to 2.0–2.5; alternatively, filter to proteins with ≥ 2 valid values per group before testing |
© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/proteomics-protein-engineering/maxquant-proteomics of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Maxquant Proteomics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Maxquant Proteomics this skilljaechang-hits/SciAgent-Skills | 370 | 1 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| Singlecell Qcxuzhougeng/wisp-science | 1k | — | ~1.6k | Automated safety check: Pass | AGPL-3.0 | |
| Trackplotygidtu/trackplot | 109 | — | ~1.9k | Automated safety check: Pass | BSD-3-Clause | |
| UniProt Database Accessdavila7/claude-code-templates | 32k | 15 repos | ~1.7k | Automated safety check: Pass | MIT |
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
xuzhougeng/wisp-science
A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.
ygidtu/trackplot
Generate sashimi-style genome visualization plots (coverage, line, heatmap, IGV read-by-read, HiC, circRNA, motif) from BAM/bigWig/depth/HiC inputs.
davila7/claude-code-templates
Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.
QING1105/ezST
End-to-end 10x Visium spatial transcriptomics analysis workflow with staged execution and human review gates.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Annotated matrices for single-cell genomics. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
MaxQuant + Perseus proteomics pipeline: run MaxQuant for LFQ and SILAC; parse proteinGroups.txt in Python; filter contaminants/decoys; log2 + median-normalize; impute MNAR; t-test with FDR; volcano…. Maxquant Proteomics is an agent skill from jaechang-hits/SciAgent-Skills.txt in Python; filter contaminants/decoys; log2 + median-normalize; impute MNAR; t-test with FDR; volcano plot; GO/pathway enrichment.
Maxquant Proteomics fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics -a claude-code`. Or copy the skill folder (skills/proteomics-protein-engineering/maxquant-proteomics in jaechang-hits/SciAgent-Skills) into .claude/skills/maxquant-proteomics in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics -a codex`. Or copy the skill folder (skills/proteomics-protein-engineering/maxquant-proteomics in jaechang-hits/SciAgent-Skills) into .agents/skills/maxquant-proteomics in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill maxquant-proteomics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/maxquant-proteomics, .gemini/skills/maxquant-proteomics, .github/skills/maxquant-proteomics and .opencode/skills/maxquant-proteomics in your project.
Going by SKILL.md and its folder, Maxquant Proteomics needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 6 domains. In commands or code: string-db.org; the agent is likely to contact it when it follows the instructions. As links in the text: maxquant.org, doi.org, maxquant.net, github.com and gseapy.readthedocs.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Maxquant Proteomics is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Maxquant Proteomics: 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), Singlecell Qc (xuzhougeng/wisp-science, 1k stars) and Trackplot (ygidtu/trackplot, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 370 GitHub stars. The repository holds 163 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.