tangermeme Genomic Model Analysis
jmschrei/tangermeme
Routes agents to the right tangermeme reference for analyzing trained genomic deep learning models, from attributions and motif experiments to variant effects and design.
Analyzes data-independent acquisition (DIA) proteomics by scoring reconstructed fragment-chromatogram peak groups against a decoy null with DIA-NN (library-free directDIA, library-based, or…
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-dia-analysis -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-dia-analysis --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/proteomics/dia-analysis .claude/skills/bio-proteomics-dia-analysis && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-proteomics-dia-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/dia-analysis into .claude/skills/bio-proteomics-dia-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-dia-analysis", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/proteomics/dia-analysisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-dia-analysis -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-dia-analysis --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/proteomics/dia-analysis .agents/skills/bio-proteomics-dia-analysis && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-proteomics-dia-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/dia-analysis into .agents/skills/bio-proteomics-dia-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-dia-analysis", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-dia-analysis -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-dia-analysis --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/proteomics/dia-analysis .cursor/skills/bio-proteomics-dia-analysis && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-proteomics-dia-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/dia-analysis into .cursor/skills/bio-proteomics-dia-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-dia-analysis", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path proteomics/dia-analysis--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-dia-analysis -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-dia-analysis --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/proteomics/dia-analysis .gemini/skills/bio-proteomics-dia-analysis && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-proteomics-dia-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/dia-analysis into .gemini/skills/bio-proteomics-dia-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-dia-analysis", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-proteomics-dia-analysisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-dia-analysis -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/proteomics/dia-analysis .github/skills/bio-proteomics-dia-analysis && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-proteomics-dia-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/dia-analysis into .github/skills/bio-proteomics-dia-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-dia-analysis", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-dia-analysis -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-dia-analysis --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/proteomics/dia-analysis .opencode/skills/bio-proteomics-dia-analysis && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-proteomics-dia-analysis" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/dia-analysis into .opencode/skills/bio-proteomics-dia-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-dia-analysis", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-proteomics-dia-analysisAnalyzes data-independent acquisition (DIA) proteomics by scoring reconstructed fragment-chromatogram peak groups against a decoy null with DIA-NN (library-free directDIA, library-based, or…
Bio Proteomics Dia Analysis is an agent skill from GPTomics/bioSkills. Analyzes data-independent acquisition (DIA) proteomics by scoring reconstructed fragment-chromatogram peak groups against a decoy null with DIA-NN (library-free directDIA, library-based, or deep-learning predicted-library routes), Spectronaut, OpenSWATH, and EncyclopeDIA. Frames the deliverable around q-value LEVEL (precursor/peptide/protein-group) and CONTEXT (run vs experiment-wide/global) rather than a bare "1% FDR", and around the duty-cycle-vs-selectivity acquisition tradeoff (window design, staggered…
Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/diann_analysis.sh` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics, Database schema design and DataFrames. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Proteomics Dia Analysis loads about 5.4k tokens when it runs. Until then it costs about 223 tokens; SKILL.md has 2,313 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,313 words, ~5,386 tokens.
.claude/skills/bio-proteomics-dia-analysis/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: DIA-NN 1.9+, pandas 2.2+, pyarrow 15+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signatures<tool> --version then <tool> --help to confirm flagsIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Identify and quantify proteins from my DIA runs" -> Reconstruct, per candidate peptide, a set of co-eluting fragment extracted-ion chromatograms (XICs) and score whether that peak group is real against a decoy null -- because every wide-isolation-window MS2 is chimeric, so the problem is deconvolution and peak-group scoring, not spectrum matching.
diann --fasta-search for library-free (directDIA) discovery and quantificationdiann --lib predicted.speclib for the deep-learning predicted-library route (the modern default)diann --lib experimental.tsv for an experimental or chromatogram libraryOpenSwathWorkflow + pyprophet when explicit run/experiment/global FDR contexts must be auditableScope: this skill OWNS running the DIA search engine and FILTERING its output at the correct q-value level and context. Building the library (experimental, chromatogram, predicted) -> spectral-libraries. Normalization, MaxLFQ roll-up, and matrix summarization -> quantification. Statistical testing of the protein matrix -> differential-abundance. Loading raw vendor/mzML data -> data-import. OUT OF SCOPE: DDA spectrum-to-peptide matching (peptide-identification); pathway enrichment of the hit list; acquiring the data (the analyst inherits the window design from the core facility).
DIA quantification is peak-group SCORING against a decoy null, not spectrum matching. Gillet 2012 inverted DDA: instead of asking "what peptide is this spectrum", DIA asks, per library peptide, "does a co-eluting peak group of this peptide's expected fragments exist in the chimeric MS2 stream". Decoys are shuffled or reversed peptide queries scored identically; the q-value is the expected fraction of accepted IDs that are decoy-like false peak groups. The count of "proteins found" is therefore a function of the decoy-calibrated threshold, never a quality metric in itself.
"1% FDR" is meaningless without naming the LEVEL and the CONTEXT -- state both. LEVEL = precursor vs peptide vs protein-group (filtering precursors at 1% does NOT give proteins at 1%; control both). CONTEXT = run-specific vs experiment-wide vs global (Rosenberger 2017). Naively filtering N runs at per-run 1% inflates the experiment-wide error: 1% per run accumulates false positives across the union, severe at hundreds-to-thousands of runs. For a cross-run matrix, filter on the GLOBAL protein-group q-value, not the per-run one. The column chosen (Q.Value vs Global.PG.Q.Value) is the decision.
The predicted-library route is now the default recommendation. Library-based search is sensitive but capped by an ill-matched library (wrong organism/tissue/mods silently limits coverage with no error). Library-free directDIA finds sample-specific content but its larger implicit search space can INFLATE IDs if FDR is not controlled across the two-pass process -- worst on wide-window chimeric data. The compromise the field converged on: predict an in-silico library for the whole FASTA digest (DIA-NN built-in predictor, or Prosit/AlphaPeptDeep) and search against THAT, getting directDIA's "no wet-lab library" with library-based's bounded, better-calibrated search.
DDA picks top-N precursors and fragments each in isolation, so every MS2 is nominally one peptide. DIA abandons selection: the quadrupole steps through wide isolation windows (4-25 Th classically, 2 Th on Astral, mobility-gated on timsTOF) and co-fragments EVERY precursor in each window. Consequence: every DIA MS2 is chimeric, a superposition of fragments from all co-isolated precursors. The engine must deconvolve -- reconstruct each candidate's fragment XICs and score the peak group.
Selectivity is set by isolation-window WIDTH. Narrower window = fewer co-isolated precursors = less chimerism = cleaner XICs = fewer false peak groups. But narrower windows mean MORE windows per cycle -> longer duty cycle -> fewer points across each LC peak -> worse quant precision. The central acquisition tradeoff is duty cycle (sampling speed) vs selectivity (window width). Rule of thumb: aim for >= 6 MS2 points across the FWHM of an LC peak for reliable quant. The analyst INHERITS this design and must not pretend all DIA is equivalent:
msconvert --filter "demultiplex optimization=overlap_only") or every tool sees the wide physical window and the selectivity benefit is silently lost. MSX (randomized window combinations) is largely historical.| Tool / method | Citation | Mechanism / role | When |
|---|---|---|---|
| DIA-NN | Demichev 2020 | Deep-NN peak-group scoring + interference correction + QuantUMS quant; library-free, predicted, or library-based | Default for high-throughput, large cohorts, diaPASEF, Astral; free, scriptable |
| Spectronaut | Biognosys (commercial) | directDIA+ pipeline with in-app DL prediction; mature GUI/QC | Regulated/clinical work, polished QC, mixed vendors, when licensed |
| OpenSWATH + PyProphet | Rost 2014; Rosenberger 2017 | Classic peptide-centric extraction + semi-supervised scoring with explicit run/experiment/global q-contexts | When auditable FDR-context control is required; library-based ONLY, needs iRT/RT alignment |
| EncyclopeDIA / Walnut | Searle 2018 | Chromatogram-library search (.dlib/.elib) + GPF; Walnut = library-free mode | Building project-specific chromatogram-library depth on Orbitrap |
| FragPipe (MSFragger-DIA / DIA-Umpire) | -- | Spectrum-centric via pseudo-spectra + Philosopher FDR; IonQuant explicit MBR-FDR | Unified DDA+DIA shop in the MSFragger ecosystem |
| Skyline | MacCoss lab | Targeted/visual peak inspection and demultiplexing; not a discovery engine | Manual peak curation, PRM, library curation -- the microscope, not the engine |
| AlphaDIA | Mann lab | Open transformer-based end-to-end Python; AlphaPeptDeep predictions | Cutting-edge open research, Astral, Python-native pipelines |
| Library build | (route OUT) | Experimental/chromatogram/predicted library construction | -> spectral-libraries |
| Stats on the matrix | (route OUT) | Normalization, roll-up, moderated testing | -> quantification, differential-abundance |
Tool leadership moves fast (Astral, AlphaDIA, DIA-NN releases). Confirm the current recommended engine and version for the specific instrument before committing rather than hard-coding "DIA-NN is best".
| Scenario | Recommended | Why |
|---|---|---|
| Discovery cohort, no wet-lab library | diann --fasta-search --gen-spec-lib (predicted route) then search against it | Bounded, better-calibrated search vs raw directDIA; the modern default |
| Quick single-run discovery, Astral 2-Th data | DIA-NN library-free (directDIA) | Near-non-chimeric MS2 makes directDIA trustworthy |
| Have a deep experimental/chromatogram library | diann --lib library.tsv (no --fasta-search) | Targeted extraction is most sensitive when the library matches |
| Need auditable run/experiment/global FDR for a regulated submission | OpenSWATH + PyProphet | Explicit q-value contexts per Rosenberger 2017 |
| Building chromatogram-library depth for one project on Orbitrap | EncyclopeDIA (GPF) -> .elib -> DIA-NN | Empirical RT and real fragmentation in the project's own LC |
| Large cohort (hundreds-thousands of runs) | DIA-NN --reanalyse, filter on Global.PG.Q.Value | Two-pass global FDR controls cross-run error accumulation |
| Staggered/overlapping acquisition | Demultiplex at conversion FIRST, then any engine | Skipping demux silently keeps wide-window interference |
| PTM / peptidoform-resolved work | DIA-NN --peptidoforms + matched variable mods | Peptidoform-resolved target-decoy scoring |
Default when uncertain: DIA-NN with the predicted-library route (--fasta-search --gen-spec-lib --reanalyse), letting --mass-acc 0 auto-optimize, then filter Q.Value <= 0.01 & PG.Q.Value <= 0.01 per run and Global.PG.Q.Value <= 0.01 for cross-run matrices.
The default route: digest the FASTA in silico, predict a library, and search the DIA data against it in one command. --mass-acc 0 lets DIA-NN auto-optimize tolerances per file (do not hard-code ppm from another instrument). --reanalyse enables the two-pass global FDR / MBR that controls directDIA double-dipping.
diann \
--f sample1.mzML --f sample2.mzML \
--lib "" --fasta uniprot_human.fasta --fasta-search \
--gen-spec-lib --predictor \
--out diann_out/report.parquet \
--out-lib diann_out/report-lib.tsv \
--qvalue 0.01 \
--matrices \
--mass-acc 0 \
--reanalyse --smart-profiling \
--cut K*,R* --missed-cleavages 1 \
--min-pep-len 7 --max-pep-len 30 \
--unimod4 --var-mods 1 --var-mod UniMod:35,15.994915,M \
--threads 8Supply an existing library (experimental, chromatogram-derived, or a previously predicted .speclib). Omit --fasta-search: extraction is targeted to the library content.
diann \
--f sample1.mzML --f sample2.mzML \
--lib spectral_library.tsv \
--out diann_out/report.parquet \
--qvalue 0.01 --matrices \
--mass-acc 0 \
--reanalyse --smart-profiling \
--threads 8DIA-NN 1.9+ writes the main report as Apache Parquet (report.parquet) by default; 2.0 makes it the only default. Matrices stay TSV. Pipelines hard-coding report.tsv silently break or read a stale file -- read parquet. Filter on q-value columns BEFORE pivoting to a matrix.
report.parquet # main report (1.9+ default; was report.tsv pre-1.9)
report.stats.tsv # per-run statistics
report.pg_matrix.tsv # protein-group wide matrix
report.pr_matrix.tsv # precursor wide matrix (verify exact dotting vs installed version)
report.gg_matrix.tsv # gene-group wide matrix
report-lib.tsv # generated library (if --gen-spec-lib)import pandas as pd, numpy as np
report = pd.read_parquet('diann_out/report.parquet') # NOT report.tsv on 1.9+
# Per-run filter: both LEVELS. Add Global.PG.Q.Value for the cross-run matrix.
filt = report[(report['Q.Value'] <= 0.01) &
(report['PG.Q.Value'] <= 0.01) &
(report['Global.PG.Q.Value'] <= 0.01)] # 0.01 = standard 1% FDR
# Pivot to a protein matrix from the filtered long report.
pg = filt.pivot_table(index='Protein.Group', columns='Run', values='PG.MaxLFQ', aggfunc='first')
pg = np.log2(pg.replace(0, np.nan)) # DIA-NN writes 0 for not-quantified; log2(0) = -InfThe *_matrix.tsv files apply an EXTRA 5% run-specific protein FDR (DIA-NN default, --matrix-spec-q), so the matrix protein count can be lower than the report count. A report-vs-matrix mismatch is EXPECTED, not a bug -- do not panic and do not compare the two counts as if they should match.
DIA-NN's built-in predictor covers the common case. For an external predicted library from FragPipe search results, EasyPQP builds it; FragPipe can also emit a DIA-NN-format library directly.
# Real easypqp subcommands: library, convert, insilico-library (NOT a convert --format diann).
easypqp library \
--psmtsv psm.tsv \
--rt_reference irt.tsv \
--peptide_fdr_threshold 0.01 \
--protein_fdr_threshold 0.01 \
--out library.tsv
# FragPipe's DIA workflow can emit a DIA-NN-format library directly -- prefer that when in FragPipe.Trigger: --fasta-search on wide-window (25-Th SWATH) data without two-pass global FDR.
Mechanism: building the library from the same data then quantifying it reuses the data twice; the large implicit search space inflates IDs when decoys are not controlled across both passes.
Symptom: implausibly high protein counts, poor reproducibility across replicates.
Fix: keep --reanalyse (two-pass global FDR) ON and filter on Global.PG.Q.Value; prefer the predicted-library route on chimeric data.
Trigger: library organism/tissue/modifications/gradient do not match the sample. Mechanism: targeted extraction can only find what is in the library; a mismatched library caps coverage with no error. Symptom: low ID counts, conserved-peptide bias (e.g. a human library on a mouse sample finds only conserved peptides). Fix: match the library to the biology (mods especially); regenerate via spectral-libraries or switch to the predicted route.
Trigger: filtering N runs at per-run Q.Value <= 0.01 and unioning.
Mechanism: 1% per run accumulates across the union -> experiment-wide error far above 1%.
Symptom: inflated total protein list; irreproducible "hits" in differential testing.
Fix: filter on Global.PG.Q.Value <= 0.01 (and optionally Lib.PG.Q.Value <= 0.01) for the matrix.
Trigger: treating MBR-transferred quant values as equally confident as directly identified ones.
Mechanism: transferring an ID by RT/m/z/CCS matching can be a false transfer, especially for low-abundance precursors; MBR has its OWN FDR.
Symptom: spurious low-abundance quant filling missing values that should stay missing.
Fix: rely on DIA-NN's global/empirical-library q-values (IonQuant uses an explicit MBR-FDR model); do not disable --reanalyse then trust per-run counts.
Trigger: running staggered/overlapping data through any engine without demultiplexing.
Mechanism: the engine sees the wide physical window, keeping all the interference the staggering was meant to remove.
Symptom: noisy directDIA, poor selectivity on data that should be clean.
Fix: demultiplex at conversion (msconvert --filter "demultiplex optimization=overlap_only") before search.
| Threshold | Source | Rationale |
|---|---|---|
Precursor q-value <= 0.01 | DIA-NN default (--qvalue 0.01) | Standard 1% per-precursor FDR (run context). |
Global.PG.Q.Value <= 0.01 | Rosenberger 2017 | Experiment-wide protein-group FDR for cross-run matrices; run-specific PG q-value is NOT enough for a cohort. |
Matrix run-specific PG filter 0.05 | DIA-NN default (--matrix-spec-q) | Extra 5% run-specific protein FDR applied only when building matrices -> report vs matrix count differs (expected). |
Lib.(PG.)Q.Value <= 0.01 | Demichev recommendation | For very large cohorts, also filter the global library-pass q-values to keep experiment-wide FDR honest. |
Points per peak >= 6 across FWHM | community rule of thumb | Below this, quant precision and peak detection degrade; drives window/cycle design. |
Mass accuracy auto (--mass-acc 0) | DIA-NN | Wrong tolerance silently kills IDs; let DIA-NN auto-calibrate rather than hard-coding ppm from another instrument. |
Missed cleavages 1 | -- | Trypsin standard; raising it expands the search space and the multiple-testing burden. |
| Error / symptom | Cause | Solution |
|---|---|---|
FileNotFoundError: report.tsv or stale data | DIA-NN 1.9+ default is report.parquet, not report.tsv | Read report.parquet (pd.read_parquet); request legacy TSV explicitly only if needed |
| Matrix protein count < report count, looks like data loss | Matrices apply an extra 5% run-specific PG filter | Expected; do not compare the two counts as if equal |
-Inf after log2 of the matrix | DIA-NN writes 0 for not-quantified | Convert 0 -> NaN BEFORE log2/normalization |
| Cohort "hits" do not reproduce | Filtered per-run Q.Value only, not global | Filter Global.PG.Q.Value <= 0.01 for the matrix |
easypqp convert --format diann errors | No such interface; convert/library/insilico-library are the real subcommands | Use easypqp library (with --psmtsv/--rt_reference) or let FragPipe emit a DIA-NN-format library |
KeyError: 'report.pr.matrix.tsv' | Matrix filename dotting varies by version (pr_matrix vs pr.matrix) | ls the output dir after a run and match the installed version's exact names |
| directDIA looks noisy on overlapping-window data | Staggered data not demultiplexed | Demultiplex at conversion before search |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in proteomics/dia-analysis of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Proteomics Dia Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Proteomics Dia Analysis this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.4k | Automated safety check: Pass | MIT | |
| tangermeme Genomic Model Analysisjmschrei/tangermeme | 318 | — | ~1.6k | Automated safety check: Pass | MIT | |
| FlexynesisBIMSBbioinfo/flexynesis | 110 | — | ~2.4k | Automated safety check: Pass | Custom licence | |
| Cellxgene Censusdavila7/claude-code-templates | 33k | 11 repos | ~3.8k | Automated safety check: Pass | MIT | |
| Bulkrna Batch CorrectionTianGzlab/OmicsClaw | 161 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Bulkrna CoexpressionTianGzlab/OmicsClaw | 161 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 |
jmschrei/tangermeme
Routes agents to the right tangermeme reference for analyzing trained genomic deep learning models, from attributions and motif experiments to variant effects and design.
BIMSBbioinfo/flexynesis
Run flexynesis, a deep-learning suite for multi-omics data integration and clinical outcome prediction (drug response, cancer subtyping, survival analysis).
davila7/claude-code-templates
Query CZ CELLxGENE Census (61M+ cells). An agent skill from davila7/claude-code-templates.
TianGzlab/OmicsClaw
Load when correcting batch effects in bulk expression using R sva ComBat or the legacy Python parametric approximation.
TianGzlab/OmicsClaw
Load when discovering bulk gene co-expression modules and hub genes with R WGCNA.
google-deepmind/science-skills
Identify domains, families, and sites in proteins; find all proteins in a family or sharing a domain; explore species distribution for a domain; annotate genomes with protein families and GO terms.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Analyzes data-independent acquisition (DIA) proteomics by scoring reconstructed fragment-chromatogram peak groups against a decoy null with DIA-NN (library-free directDIA, library-based, or…. Bio Proteomics Dia Analysis is an agent skill from GPTomics/bioSkills. Analyzes data-independent acquisition (DIA) proteomics by scoring reconstructed fragment-chromatogram peak groups against a decoy null with DIA-NN (library-free directDIA, library-based, or deep-learning predicted-library routes), Spectronaut, OpenSWATH, and EncyclopeDIA.
Bio Proteomics Dia Analysis fits situations like: identifying and quantifying proteins from DIA mass spectrometry runs and filtering DIA-NN report.parquet/matrix output; tasks that involve Bioinformatics; tasks that involve Database schema design.
Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-dia-analysis -a claude-code`. Or copy the skill folder (proteomics/dia-analysis in GPTomics/bioSkills) into .claude/skills/bio-proteomics-dia-analysis in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-dia-analysis -a codex`. Or copy the skill folder (proteomics/dia-analysis in GPTomics/bioSkills) into .agents/skills/bio-proteomics-dia-analysis in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-proteomics-dia-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-proteomics-dia-analysis, .gemini/skills/bio-proteomics-dia-analysis, .github/skills/bio-proteomics-dia-analysis and .opencode/skills/bio-proteomics-dia-analysis in your project.
Going by SKILL.md and its folder, Bio Proteomics Dia Analysis needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: A Bash shell.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Proteomics Dia Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Proteomics Dia Analysis: tangermeme Genomic Model Analysis (jmschrei/tangermeme, 318 stars), Flexynesis (BIMSBbioinfo/flexynesis, 110 stars), Cellxgene Census (davila7/claude-code-templates, 33k stars) and Bulkrna Batch Correction (TianGzlab/OmicsClaw, 161 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.