Spatial S5 Downstream
QING1105/ezST
Stage 5 of the spatial transcriptomics workflow — neighborhood enrichment and cell-cell communication analysis.
Frames PTM/phosphoproteomics analysis as three stacked inference layers on a biased enrichment - chemistry selection, site localization (FLR), and protein-level-adjusted quantification with…
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-ptm-analysis -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-ptm-analysis --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/proteomics/ptm-analysis .claude/skills/bio-proteomics-ptm-analysis && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-proteomics-ptm-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/ptm-analysis into .claude/skills/bio-proteomics-ptm-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-ptm-analysis", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/proteomics/ptm-analysisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-ptm-analysis -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-ptm-analysis --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/proteomics/ptm-analysis .agents/skills/bio-proteomics-ptm-analysis && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-proteomics-ptm-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/ptm-analysis into .agents/skills/bio-proteomics-ptm-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-ptm-analysis", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-ptm-analysis -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-ptm-analysis --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/proteomics/ptm-analysis .cursor/skills/bio-proteomics-ptm-analysis && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-proteomics-ptm-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/ptm-analysis into .cursor/skills/bio-proteomics-ptm-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-ptm-analysis", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path proteomics/ptm-analysis--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-ptm-analysis -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-ptm-analysis --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/proteomics/ptm-analysis .gemini/skills/bio-proteomics-ptm-analysis && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-proteomics-ptm-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/ptm-analysis into .gemini/skills/bio-proteomics-ptm-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-ptm-analysis", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-proteomics-ptm-analysisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-ptm-analysis -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/proteomics/ptm-analysis .github/skills/bio-proteomics-ptm-analysis && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-proteomics-ptm-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/ptm-analysis into .github/skills/bio-proteomics-ptm-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-ptm-analysis", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-ptm-analysis -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-ptm-analysis --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/proteomics/ptm-analysis .opencode/skills/bio-proteomics-ptm-analysis && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-proteomics-ptm-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/ptm-analysis into .opencode/skills/bio-proteomics-ptm-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-ptm-analysis", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-proteomics-ptm-analysisFrames PTM/phosphoproteomics analysis as three stacked inference layers on a biased enrichment - chemistry selection, site localization (FLR), and protein-level-adjusted quantification with…
Bio Proteomics Ptm Analysis is an agent skill from GPTomics/bioSkills. Frames PTM/phosphoproteomics analysis as three stacked inference layers on a biased enrichment - chemistry selection, site localization (FLR), and protein-level-adjusted quantification with MSstatsPTM - plus kinase-activity and functional triage. Covers MaxQuant Phospho (STY)Sites multiplicity expansion, localization-probability filtering (class I, Ascore, ptmRS, DIA EG.PTMLocalizationProbabilities, DIA-NN PTM.Site.Confidence), false localization rate (LuciPHOr/DeepFLR), motif analysis with experiment-matched…
Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/phospho_analysis.py` and `usage-guide.md`).
It sits in Research & Science, covering Internationalization and Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Proteomics Ptm Analysis loads about 6.9k tokens when it runs. Until then it costs about 259 tokens; SKILL.md has 2,700 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,700 words, ~6,930 tokens.
.claude/skills/bio-proteomics-ptm-analysis/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: MSstatsPTM 2.4+, pandas 2.2+, numpy 1.26+, scipy 1.12+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturespackageVersion('<pkg>') then ?function_name to verify parametersIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Find the regulated phosphosites in my enriched samples" -> Localize each modification, then test whether its abundance change survives subtracting the protein-level change -- because a PTM result is three separate inferences (enrichment, localization, quantification) and each fails silently if the layer below is treated as solved.
MSstatsPTM::groupComparisonPTM() for protein-adjusted site testing (the load-bearing tool)pandas to expand MaxQuant Phospho (STY)Sites multiplicity and filter localization probabilityKSEAapp / PTM-SEA (ssGSEA2.0) for kinase-activity inference from the site fold-changesScope: this skill OWNS enrichment-chemistry framing, site localization and FLR, multiplicity-resolved site quant, protein-level adjustment, motif analysis, and kinase-activity inference. Peptide identification and open/variable-mod search route to peptide-identification; the underlying protein-level (unenriched) quant routes to quantification and differential-abundance; DIA acquisition mechanics route to dia-analysis. OUT OF SCOPE: intact-glycopeptide glycan-composition search (pGlyco3/MSFragger-Glyco) and absolute occupancy from three-ratio SILAC are noted but not implemented here.
log2FC(PTM_observed) = log2FC(occupancy) + log2FC(protein). Without a paired global (unenriched) proteome run on the SAME samples to subtract log2FC(protein), every protein-abundance change masquerades as a regulated site. Because co-regulated proteins move together, the false positives are pathway-coherent and look biologically convincing -- the worst kind of artifact. This is the entire reason MSstatsPTM exists (Kohler 2023).Bottom line: report THREE numbers, not one -- peptide/PSM FDR, per-site localization probability with its threshold, and an empirically estimated global FLR -- and never call a site "regulated" from a phospho-only run without protein-level adjustment.
| Method | Citation | Mechanism / bias | When |
|---|---|---|---|
| TiO2 (MOAC) | Larsen 2005 | Metal-oxide Lewis-acid surface; skews mono-phospho; needs hydroxy-acid additive | General single-method depth; EasyPhos basis |
| Fe(III)-IMAC | Ruprecht 2015 | Chelated Fe3+ coordinates phosphate; multidentate avidity skews MULTI-phospho | Hierarchical/processive signaling; the mono/multi divergence is STRONGEST here vs TiO2 |
| Ti4+/Zr4+-IMAC | Matheron 2014 | Chelated metal ION on immobilized phosphonate (NOT bulk oxide); bias vs TiO2 is SMALL | Modern automated workflows; metal identity matters more than IMAC-vs-MOAC |
| SIMAC (sequential) | Thingholm 2008 | IMAC acidic elution = mono, basic = multi, then TiO2 on mono fraction | Recovering both populations IMAC alone biases |
Naming trap: Ti4+/Zr4+-IMAC (chelated ions) is DIFFERENT chemistry from TiO2/ZrO2 (bulk oxide). Glycolic acid is the modern additive standard (load 80% ACN / 5% TFA / 0.1 M glycolic acid).
| PTM | Reagent / mass | Citation | Headline trap |
|---|---|---|---|
| Ubiquitin (diGly, K-GG, +114.0429) | Anti-K-GG antibody | Xu 2010; Kim 2011 | NOT ubiquitin-specific: K-GG = ubiquitin + NEDD8 (~6% at basal) + ISG15 (rises under interferon). UbiSite (Akimov 2018) is the ubiquitin-specific alternative |
| Ubiquitin alkylation artifact | use chloroacetamide | Nielsen 2008 | Iodoacetamide creates a +114.0429 lysine adduct mimicking ubiquitination; chloroacetamide does not |
| Acetyl-K (+42.0106) | Anti-acetyllysine cocktail | Svinkina 2015 | Isobaric with trimethyl +42.0470 (0.0364 Da, needs high-res); acetyl blocks trypsin -> allow >=4 missed cleavages |
| Glyco N-linked | PNGase F (released) or intact | Riley 2021 | Released loses the glycan; N->D tag +0.984 is isobaric with deamidation -- use PNGase F in H2-18O (+2.988) to disambiguate; N-X-S/T (X!=Pro) sequon is necessary not sufficient |
| Tool | Citation | Mechanism | Note |
|---|---|---|---|
| Ascore | Beausoleil 2006 | Cumulative binomial of site-determining ions; DIFFERENCE between best and 2nd-best localization, peak-depth sweep | Ascore >=19 ~ p 0.01 PAIRWISE per-PSM, NOT a dataset FLR |
| PhosphoRS / ptmRS | Taus 2011 | Per-isomer cumulative binomial, tolerance-aware (correct for high-res), per-site probs sum to 100% | In Proteome Discoverer |
| PTMProphet | Shteynberg (TPP) | EM/Bayesian mixture; per-site probs combinable across PSMs to a global FLR | TPP/FragPipe |
| MaxQuant Localization prob | Cox/Mann (Andromeda) | Normalized posterior on the site (fixed peak depth) | column Localization prob; >=0.75 = class I |
| DIA localization | Bekker-Jensen 2020 | XIC peak-shape correlation substitutes for missing precursor isolation | Spectronaut EG.PTMLocalizationProbabilities; DIA-NN PTM.Site.Confidence |
| DeepFLR | Zong 2023 | Deep-learning spectrum predictor + target-decoy FLR | SOTA direction; DDA + DIA |
| Tool | Citation | Mechanism |
|---|---|---|
| LuciPHOr | Fermin 2013 (MCP) | Decoy localizations on non-modifiable residues; rate decoys win = empirical FLR |
| LuciPHOr2 | Fermin 2015 (Bioinformatics) | Generic-PTM successor (do NOT swap the two journals) |
| Decoy amino-acid FLR | Ramsbottom/Jones 2022 | Add a non-modifiable residue to the candidate set; global FLR ~ decoy-site-hits / target-site-hits, frequency-corrected |
| Tool | Citation | Mechanism | Limitation |
|---|---|---|---|
| KSEA | Casado 2013 | z-score of a kinase's substrate fold-changes | sqrt(m) favors well-annotated kinases; inherits PhosphoSitePlus bias |
| PTM-SEA / PTMsigDB | Krug 2019 | ssGSEA2.0 on site-level +/-7 flanking-sequence signatures | robust to isoform drift; PERT signatures score "looks like EGF stim" |
| RoKAI | Yilmaz 2021 | Network-smooth profiles before z-score so unobserved sites borrow neighbor signal | attacks missingness; feeds KSEA |
Benchmark result (Mueller-Dott 2025): across ~19 methods, simple z-score (KSEA/RoKAI) matched or beat sophisticated methods. Performance is PRIOR-limited, not algorithm-limited; all methods inherit PhosphoSitePlus curation bias toward CK2/CDK1/PKA/MAPK, and the dark kinome is structurally invisible. Spend effort on the substrate prior, not the estimator.
| Scenario | Recommended | Why |
|---|---|---|
| Phospho-only run, want regulated sites | Acquire a PAIRED global proteome -> MSstatsPTM groupComparisonPTM -> require significance in ADJUSTED.Model | Unadjusted site changes are confounded with protein abundance |
| No global proteome available | Report site changes as UNADJUSTED and flag the confound explicitly | Cannot separate occupancy from abundance; do not claim "regulation" |
| Between-method phospho difference | Suspect chemistry (TiO2 vs Fe-IMAC mono/multi bias) BEFORE biology | Enrichment is a confounded filter |
| Multiply-phospho peptides present | Localize per-site (Ascore/ptmRS) AND report empirical global FLR | Peptide FDR != site FDR |
| "Ubiquitination" sites | Confirm chloroacetamide alkylation; treat K-GG as ub + NEDD8 + ISG15; consider UbiSite | Iodoacetamide artifact + NEDD8/ISG15 confound |
| Which kinases moved? | KSEA or PTM-SEA with a curated prior; do not over-interpret dark-kinome silence | Prior-limited; simple z-score suffices |
| Motif logo from the hits | Background = experiment-matched S/T/Y from the identified proteins (NOT whole proteome) | Whole-proteome background rediscovers disordered-region composition bias |
Default when uncertain: localize with the search engine's probability (class I >=0.75), expand MaxQuant multiplicity, run MSstatsPTM with a paired global proteome, and call only ADJUSTED.Model hits regulated.
Goal: Produce a long, multiplicity-resolved, class-I-filtered phosphosite intensity matrix from Phospho (STY)Sites.txt.
Approach: Each site row spreads its quant across Intensity___1/___2/___3 (singly/doubly/triply-phospho forms, THREE underscores); the collapsed base Intensity mixes phospho-states and can fake dephosphorylation. Drop Reverse/contaminant, filter Localization prob, then melt the per-multiplicity columns into rows.
import pandas as pd
import numpy as np
# Filename has a SPACE in the modification name; accept either form.
phospho = pd.read_csv('Phospho (STY)Sites.txt', sep='\t', low_memory=False)
# Newer MaxQuant uses 'Potential contaminant'; older uses 'Contaminant'.
contaminant_col = 'Potential contaminant' if 'Potential contaminant' in phospho.columns else 'Contaminant'
phospho = phospho[(phospho['Reverse'] != '+') & (phospho[contaminant_col] != '+')]
CLASS_I_PROB = 0.75 # Olsen 2006 class-I convention; comparability standard, not a calibrated FLR
phospho = phospho[phospho['Localization prob'] >= CLASS_I_PROB].copy()
gene = phospho['Gene names'].where(phospho['Gene names'].notna(), phospho['Protein'])
phospho['site_id'] = gene.str.split(';').str[0] + '_' + phospho['Amino acid'] + phospho['Position'].astype(int).astype(str)
# Multiplicity columns carry THREE underscores: collapsing them mixes phospho-states.
mult_cols = [c for c in phospho.columns if '___' in c and c.split('___')[-1] in {'1', '2', '3'} and c.startswith('Intensity')]
long = phospho.melt(id_vars=['site_id', 'Amino acid', 'Position', 'Localization prob'], value_vars=mult_cols, var_name='run_multiplicity', value_name='intensity')
long['multiplicity'] = long['run_multiplicity'].str.split('___').str[-1]
long['run'] = long['run_multiplicity'].str.replace(r'___[123]$', '', regex=True).str.replace('Intensity ', '', regex=False)
long = long[long['intensity'] > 0]
long['log2_intensity'] = np.log2(long['intensity'])Goal: Decide whether each site change is real after subtracting the matched protein-abundance change.
Approach: MSstatsPTM carries TWO datasets -- a PTM dataset (enriched) and a PROTEIN dataset (global/unenriched). groupComparisonPTM fits independent linear models to each and returns a list of THREE: PTM.Model (unadjusted), PROTEIN.Model, and ADJUSTED.Model. The adjustment is dFC_adj = dFC_PTM - dFC_protein with SE_adj = sqrt(SE_PTM^2 + SE_protein^2), so adjustment ADDS uncertainty -- a site can be significant unadjusted yet lose significance after adjustment. A confident regulation call requires significance in ADJUSTED.Model.
library(MSstatsPTM)
# Converters are <Tool>toMSstatsPTMFormat and return a list with $PTM and $PROTEIN.
# MaxQtoMSstatsPTMFormat reads the MaxQuant 'evidence.txt' (NOT the Phospho (STY)Sites
# table -- the pandas multiplicity-expansion above is a SEPARATE workflow); the FASTA maps
# peptides back to site coordinates. Supply BOTH the enriched evidence and the global
# proteinGroups; without the protein dataset there is nothing to adjust against.
# Arg-name note: the FASTA argument is `fasta_path` in current MSstatsPTM; older builds may
# differ -- run `?MaxQtoMSstatsPTMFormat` to confirm before relying on it.
input <- MaxQtoMSstatsPTMFormat(
evidence = read.table('evidence.txt', sep = '\t', header = TRUE, quote = ''),
annotation = read.csv('annotation_ptm.csv'),
fasta_path = 'uniprot_human.fasta',
fasta_protein_name = 'uniprot_ac',
proteinGroups = read.table('proteinGroups.txt', sep = '\t', header = TRUE, quote = ''),
annotation_protein = read.csv('annotation_protein.csv'),
mod_id = '\\(Phospho \\(STY\\)\\)',
which_proteinid_ptm = 'Proteins',
use_unmod_peptides = FALSE
)
summarized <- dataSummarizationPTM(input, use_log_file = FALSE)
# LabelFree run: data.type = 'LF' (use 'TMT' for isobaric); contrast.matrix defaults to
# full pairwise. groupComparisonPTM has NO `model` argument -- it always fits independent
# PTM and PROTEIN models, then adjusts.
result <- groupComparisonPTM(summarized, data.type = 'LF')
# Three models; the adjusted one is the deliverable.
adjusted <- result$ADJUSTED.Model
regulated <- adjusted[!is.na(adjusted$adj.pvalue) & adjusted$adj.pvalue < 0.05 & abs(adjusted$log2FC) > 1, ]
# How much of each call was protein-driven: compare PTM.Model vs ADJUSTED.Model.Goal: Find kinase/writer motifs around the modified residue without rediscovering amino-acid composition bias.
Approach: Use the Sequence window (+/-15 residues, 31-mer) MaxQuant already provides, centered on the site. The background MUST be an experiment-matched S/T/Y set drawn from the identified proteins (or a central-residue-preserving shuffle), NOT the whole proteome or IUPAC-random -- those just report the composition of phospho-rich disordered regions. motif-x and MoMo p-values are only valid when the background is built this way.
from collections import Counter
# 'Sequence window' is a 31-mer (+/-15) centered on the modified residue.
WINDOW_HALF = 7 # +/-7 flanking is the standard kinase-motif window
foreground = [w[15 - WINDOW_HALF: 16 + WINDOW_HALF] for w in confident['Sequence window'].dropna() if len(w) >= 31]
# Background: same-residue windows from the matched dataset, NOT the whole proteome.
def position_frequencies(windows):
counts = {i: Counter() for i in range(-WINDOW_HALF, WINDOW_HALF + 1)}
for w in windows:
for offset, aa in zip(range(-WINDOW_HALF, WINDOW_HALF + 1), w):
if aa not in '_X':
counts[offset][aa] += 1
return countsFor a publication-grade enrichment logo, hand the foreground and a matched background to a dedicated tool (motif-x / MoMo) and render with data-visualization/sequence-logos.
The function below is an ILLUSTRATIVE approximation, NOT real Ascore. Real Ascore (Beausoleil 2006) competes the best localization against the second-best, sweeps peak depth 1-10 per 100 Th, and restricts to site-determining ions -- none of which this captures. Use the search engine's own localization probability (MaxQuant Localization prob, ptmRS, PTMProphet) for real work, or pyOpenMS AScore (introspect the exact API before relying on it). The home-grown form is here only to show the binomial intuition.
import numpy as np
from scipy.stats import binom
def illustrative_localization_score(matched_site_ions, total_ions, depth_p=0.04):
'''Binomial intuition only; NOT Ascore (no best-vs-second competition or depth sweep).'''
if total_ions == 0 or matched_site_ions == 0:
return 0.0
p_random = 1 - binom.cdf(matched_site_ions - 1, total_ions, depth_p)
return -10 * np.log10(p_random) if p_random > 0 else 100.0Trigger: Differential testing on a phospho-only run with no paired global proteome. Mechanism: log2FC(PTM_observed) = log2FC(occupancy) + log2FC(protein); the two terms are inseparable. Symptom: Pathway-coherent "regulated sites" that are pure protein-abundance changes (cyclins/histones in cell cycle, stabilized substrates under drug). Fix: Run a matched global proteome and adjust via MSstatsPTM; route the protein-level quant to quantification.
Trigger: Quantifying on base Intensity instead of Intensity___1/___2/___3. Mechanism: The collapsed column mixes singly/doubly/triply-phospho forms of the same site. Symptom: The singly-phospho form dropping as a neighbor gets phosphorylated reads as dephosphorylation. Fix: Expand multiplicity to long form (Perseus "Expand site table" or the melt above) before any stats.
Trigger: Reporting sites at peptide FDR without a localization threshold. Mechanism: Isobaric positional isomers share precursor mass and peptide score; CID/ion-trap neutral loss (-98 Da) starves site-determining ions. Symptom: A 1% peptide FDR result with a much higher true site error. Fix: Filter localization probability (class I >=0.75), report an empirical global FLR (LuciPHOr/DeepFLR), prefer HCD/EThcD.
Trigger: Calling the K-GG proteome "ubiquitination". Mechanism: NEDD8 and ISG15 share the LRLRGG C-terminus and leave the identical +114.0429 remnant; iodoacetamide adds a fourth source. Symptom: Inflated/false ubiquitin sites, worst under interferon (ISG15) or with iodoacetamide. Fix: Chloroacetamide alkylation; treat K-GG as ub+NEDD8+ISG15; use UbiSite for ubiquitin-specific mapping.
Trigger: Whole-proteome or IUPAC-random background. Mechanism: Phosphosites sit in disordered, Ser/Pro/acidic-rich regions; that composition dominates the enrichment. Symptom: "Enriched" proline/serine motifs that are region bias, not kinase preference. Fix: Experiment-matched S/T/Y background or central-residue-preserving shuffle (MoMo default).
Trigger: Naming the top KSEA/atlas kinase as the responsible enzyme. Mechanism: Substrate priors are PhosphoSitePlus-curated (CK2/CDK1/PKA/MAPK heavy); atlas hits are biochemical preference ignoring expression/localization/timing. Symptom: Always-the-usual-suspects kinase lists; dark-kinome activity invisible. Fix: Use a curated prior, report z-scores with their substrate counts, do not infer absence from silence.
| Threshold | Source | Rationale |
|---|---|---|
| Localization prob class I >= 0.75 | Olsen 2006 | Best site holds >3x the posterior of all alternatives (single-phospho); a comparability standard, not a calibrated error rate |
| Class II 0.5-0.75; class III 0.25-0.5 | Olsen 2006 | Partial / poor localization |
| Ascore >= 19 (p~0.01); loose >13 | Beausoleil 2006 | Pairwise per-PSM best-vs-next confidence; NOT a dataset FLR |
| DIA directDIA localization >= 0.99 | Bekker-Jensen 2020 | Stricter than library-based (0.75) to match DDA error rates |
| DIA-NN site matrix 0.90 / 0.99 | DIA-NN docs | phosphosites_90/99.tsv; class-I-equivalent stringency is HIGHER than MaxQuant 0.75 |
| Acetyl missed cleavages >= 4 | -- | Acetyl-K blocks trypsin; pair LysC + trypsin |
| Kinase atlas motif match >= 90th percentile | Johnson 2023 | Strong motif preference, NOT proof the kinase acted |
| Report peptide FDR and site FLR separately | Fermin 2013 | 1% peptide FDR != 1% site FDR; true site error is typically several-fold higher |
| Error / symptom | Cause | Solution |
|---|---|---|
| FileNotFoundError on the sites table | Filename has a SPACE: Phospho (STY)Sites.txt | Accept either spaced or no-space form |
| Apparent dephosphorylation that is not real | Quantified base Intensity, mixing multiplicities | Use Intensity___1/___2/___3 (three underscores) |
KeyError / NaN on Gene names | Column is FASTA-dependent, absent without gene annotation | Guard with .notna() and fall back to Protein |
| All sites "regulated" and pathway-coherent | No protein-level adjustment | Require significance in MSstatsPTM ADJUSTED.Model |
PTM.Q.Value / PhosphoSite not found (DIA-NN) | Those columns do not exist | Use PTM.Site.Confidence and Site.Occupancy.Probabilities |
| False "ubiquitination" sites | Iodoacetamide +114.0429 lysine artifact | Alkylate with chloroacetamide |
| Acetyl confused with trimethyl | +42.0106 vs +42.0470 isobaric at nominal mass | Require high-res MS; check 0.0364 Da split |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in proteomics/ptm-analysis of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Proteomics Ptm Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Proteomics Ptm Analysis this skillGPTomics/bioSkills | 1.2k | 1 repos | ~6.9k | Automated safety check: Pass | MIT | |
| Spatial S5 DownstreamQING1105/ezST | 101 | — | ~513 | Automated safety check: Pass | MIT | |
| Bio Proteomics Ptm AnalysisFreedomIntelligence/OpenClaw-Medical-Skills | 3.1k | 1 repos | ~1.2k | Automated safety check: Pass | None | |
| Proteomics PtmTianGzlab/OmicsClaw | 161 | — | ~989 | Automated safety check: Pass | Apache-2.0 | |
| Bioconductor BandlebioMate-AI/biomate-bioconductor-kb | 804 | — | ~1.3k | Automated safety check: Pass | Custom licence | |
| Bioconductor StatialbioMate-AI/biomate-bioconductor-kb | 804 | — | ~1.5k | Automated safety check: Pass | Custom licence |
QING1105/ezST
Stage 5 of the spatial transcriptomics workflow — neighborhood enrichment and cell-cell communication analysis.
FreedomIntelligence/OpenClaw-Medical-Skills
Post-translational modification analysis including phosphorylation, acetylation, and ubiquitination.
TianGzlab/OmicsClaw
Load when summarising PTM sites (phosphorylation, acetylation, ubiquitination, etc.) from a per-site CSV — site-class assignment (Olsen et al.
bioMate-AI/biomate-bioconductor-kb
The Bandle package enables the analysis and visualisation of differential localisation experiments using mass-spectrometry data.
bioMate-AI/biomate-bioconductor-kb
Statial is a suite of functions for identifying changes in cell state.
jaechang-hits/SciAgent-Skills
Auto-annotate plasmids with features (promoters, terminators, resistance, origins, tags, fluorescent proteins) via BLAST against curated DBs (Addgene, fpbase, SnapGene).
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Frames PTM/phosphoproteomics analysis as three stacked inference layers on a biased enrichment - chemistry selection, site localization (FLR), and protein-level-adjusted quantification with…. Bio Proteomics Ptm Analysis is an agent skill from GPTomics/bioSkills. Frames PTM/phosphoproteomics analysis as three stacked inference layers on a biased enrichment - chemistry selection, site localization (FLR), and protein-level-adjusted quantification with MSstatsPTM - plus kinase-activity and functional triage.
Bio Proteomics Ptm Analysis fits situations like: localizing and quantifying phosphorylation; glycosylation sites from enrichment-based runs and deciding whether an apparent site change is real after subtracting protein abundance.
Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-ptm-analysis -a claude-code`. Or copy the skill folder (proteomics/ptm-analysis in GPTomics/bioSkills) into .claude/skills/bio-proteomics-ptm-analysis in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-ptm-analysis -a codex`. Or copy the skill folder (proteomics/ptm-analysis in GPTomics/bioSkills) into .agents/skills/bio-proteomics-ptm-analysis in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-proteomics-ptm-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-proteomics-ptm-analysis, .gemini/skills/bio-proteomics-ptm-analysis, .github/skills/bio-proteomics-ptm-analysis and .opencode/skills/bio-proteomics-ptm-analysis in your project.
Going by SKILL.md and its folder, Bio Proteomics Ptm Analysis needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Proteomics Ptm Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Proteomics Ptm Analysis: Spatial S5 Downstream (QING1105/ezST, 101 stars), Bio Proteomics Ptm Analysis (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Proteomics Ptm (TianGzlab/OmicsClaw, 161 stars) and Bioconductor Bandle (bioMate-AI/biomate-bioconductor-kb, 804 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.