Data Table Analysis
NVIDIA-AI-Blueprints/deep-researcher-agent
A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.
Loads mass-spectrometry data into Python/R and strips the search engine's bookkeeping before any number is trusted -- removes decoys (REV/Reverse), contaminants (CON/Potential contaminant)…
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-data-import --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/proteomics/data-import .claude/skills/bio-proteomics-data-import && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-proteomics-data-import" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/data-import into .claude/skills/bio-proteomics-data-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-data-import", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/proteomics/data-importType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-data-import --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/proteomics/data-import .agents/skills/bio-proteomics-data-import && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-proteomics-data-import" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/data-import into .agents/skills/bio-proteomics-data-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-data-import", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-data-import --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/proteomics/data-import .cursor/skills/bio-proteomics-data-import && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-proteomics-data-import" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/data-import into .cursor/skills/bio-proteomics-data-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-data-import", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path proteomics/data-import--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-data-import --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/proteomics/data-import .gemini/skills/bio-proteomics-data-import && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-proteomics-data-import" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/data-import into .gemini/skills/bio-proteomics-data-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-data-import", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-proteomics-data-importInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/proteomics/data-import .github/skills/bio-proteomics-data-import && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-proteomics-data-import" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/data-import into .github/skills/bio-proteomics-data-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-data-import", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-proteomics-data-import --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/proteomics/data-import .opencode/skills/bio-proteomics-data-import && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-proteomics-data-import" agent skill from https://github.com/GPTomics/bioSkills/tree/main/proteomics/data-import into .opencode/skills/bio-proteomics-data-import/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-proteomics-data-import", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-proteomics-data-importLoads mass-spectrometry data into Python/R and strips the search engine's bookkeeping before any number is trusted -- removes decoys (REV/Reverse), contaminants (CON/Potential contaminant)…
Bio Proteomics Data Import is an agent skill from GPTomics/bioSkills. Loads mass-spectrometry data into Python/R and strips the search engine's bookkeeping before any number is trusted -- removes decoys (REV/Reverse), contaminants (CON/Potential contaminant), Only-identified-by-site groups, and resolves semicolon razor/leading protein-ID ambiguity in MaxQuant proteinGroups.txt, DIA-NN report.parquet, and mzML/mzXML. Distinguishes Intensity (raw) vs LFQ intensity (MaxLFQ) vs iBAQ, treats a MaxQuant zero as missing (NaN, not log2(-inf)), and inherits the acquisition mode's…
Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/load_maxquant.py` and `usage-guide.md`).
It sits in Business, Finance & HR, covering Bioinformatics, Accounting and bookkeeping and DataFrames. It works with Python and pandas. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Proteomics Data Import loads about 4.5k tokens when it runs. Until then it costs about 201 tokens; SKILL.md has 1,825 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,825 words, ~4,503 tokens.
.claude/skills/bio-proteomics-data-import/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: pyOpenMS 3.1+, pandas 2.2+, numpy 1.26+, MSnbase 2.28+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturespackageVersion('<pkg>') then ?function_name to verify parametersIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Load my mass spec data into Python" -> Parse spectra or a search-engine table AND immediately enforce two contracts -- which quant column carries real biology, and which rows are search-engine bookkeeping that must be deleted -- because the same proteinGroups.txt yields different conclusions depending on the column read and the rows kept.
pyopenms.MzMLFile().load(path, exp) for raw spectra; pandas.read_csv(sep='\t') for MaxQuant; pandas.read_parquet for DIA-NNSpectra::Spectra() / QFeatures::readQFeatures() for raw and quantified data (MSnbase still works but is in maintenance mode)Scope: this skill owns reading spectra/search outputs into memory, deleting decoy/contaminant/site-only rows, picking the correct quant column, and characterizing missingness. Format conversion (RAW -> mzML) -> peptide-identification. MaxLFQ/TMT reporter quant computation -> quantification. Protein-group parsimony -> protein-inference. Normalization and imputation -> differential-abundance and expression-matrix/normalization. OUT OF SCOPE: statistical testing, batch correction, and the actual imputation step (this skill only diagnoses the missingness so the right imputer is chosen later).
A "data import" is never just file parsing -- it is the moment the acquisition mode's quantitative contract and its missingness structure are inherited. DDA selects the top-N most intense precursors per cycle, and which precursors get picked is partly stochastic and abundance-biased, so the same low-abundance peptide is sampled in run A and missed in run B; this manufactures structured, left-censored MNAR missingness. DIA fragments every precursor in every window every cycle, so its (fewer) missing values are closer to MCAR. The catastrophic error this prevents: imputing a DDA matrix with a mean/KNN method that assumes MCAR, which biases low-abundance proteins upward and manufactures false hits. The mode is born at acquisition and inherited at import; the missingness diagnosis made here dictates which imputation is even legitimate downstream.
The search engine's bookkeeping must be stripped before any number is trusted. A proteinGroups.txt carries decoy rows (Reverse == '+', REV__ prefix in the ID) from the target-decoy FDR machinery, contaminant rows (Potential contaminant == '+', CON__ prefix), and Only-identified-by-site rows (the protein has no unmodified-peptide evidence, only a modified site). Keeping any of these leaks non-biological signal into the intensity matrix and inflates IDs. The catastrophic error: reporting differential abundance on a matrix where decoy or keratin rows survived.
The same proteinGroups.txt yields different biology from different columns, and a zero is not a measurement. Intensity is raw summed precursor signal (not normalized, not comparable across samples for ratios). LFQ intensity is MaxLFQ-normalized and is the column for between-sample comparison. iBAQ is intensity divided by the number of observable tryptic peptides -- a within-sample molar proxy, not a between-sample quant. MaxQuant writes 0 for "not quantified", so log2(0) = -inf; replace 0 -> NaN before any transform. The catastrophic error: log2-transforming raw Intensity (or iBAQ) and reading the ratios as biology.
| Tool / method | Citation | Mechanism / role | When |
|---|---|---|---|
pyOpenMS MzMLFile().load | Chambers 2012 (ProteoWizard lineage) | Loads mzML/mzXML into an MSExperiment in memory; iterate spectra by MS level | Programmatic access to raw peaks, precursor m/z, isolation windows |
pandas read_csv/read_parquet | -- | Tabular ingest of MaxQuant TSV and DIA-NN parquet | All search-engine output tables |
| DIA-NN report | Demichev 2020 | Long-format precursor table; report.parquet is the default (1.9+) and the only default (2.0) | DIA quant; pivot on PG.MaxLFQ after q-filtering |
MaxQuant txt/ outputs | Cox 2014 (MaxLFQ) | proteinGroups.txt (group level), evidence.txt (per-PSM) | DDA label-free / TMT search results |
| Spectra + QFeatures (R) | -- | Current Bioconductor raw + quantified-feature containers; readQFeatures, aggregateFeatures | R pipelines; preferred over MSnbase going forward |
MSnbase readMSData (R) | -- | On-disk raw reading; maintenance mode (route OUT to Spectra/QFeatures) | Legacy R code only |
| ThermoRawFileParser / msconvert | Hulstaert 2020 / Chambers 2012 | RAW -> mzML conversion (route OUT) | File conversion is peptide-identification |
| Scenario | Recommended | Why |
|---|---|---|
| MaxQuant DDA label-free, between-sample comparison | Read LFQ intensity columns from proteinGroups.txt | MaxLFQ-normalized; the only MaxQuant column valid for cross-sample ratios |
| MaxQuant, absolute/molar abundance within one sample | Read iBAQ columns | iBAQ is a within-sample molar proxy; do not use across samples |
| Need raw uncorrected signal for a custom normalization | Read Intensity columns, normalize yourself | Intensity is raw summed precursor area, not comparable as-is |
| DIA-NN output (1.9 or 2.0) | pd.read_parquet('report.parquet'), filter q-values, pivot PG.MaxLFQ | 2.0 dropped the TSV default; q-filter before pivot or low-confidence rows leak in |
| Raw spectra, need peaks/precursor/isolation window | pyOpenMS MzMLFile().load | Programmatic peak and isolation-window access for QC and co-isolation reasoning |
| R-based pipeline, quantified features | QFeatures readQFeatures + aggregateFeatures | Current Bioconductor; MSnbase is maintenance-only |
| Data came from DDA, planning imputation | Diagnose missingness as MNAR -> route to left-censored imputation | DDA top-N sampling makes missingness abundance-dependent |
| Data came from DIA, planning imputation | Treat missingness as closer to MCAR | DIA samples every precursor every cycle |
Default when uncertain: read LFQ intensity (MaxQuant) or PG.MaxLFQ after q-filtering (DIA-NN), strip Reverse/contaminant/site-only rows, set 0 -> NaN, then diagnose missingness before choosing an imputer.
Goal: Parse raw spectra into memory for QC, peak access, and isolation-window reasoning.
Approach: Load into an MSExperiment (filled in place), iterate by MS level; get_peaks() returns a tuple of (mz, intensity) numpy arrays, and getPrecursors() returns a list.
from pyopenms import MSExperiment, MzMLFile
exp = MSExperiment()
MzMLFile().load('sample.mzML', exp) # fills exp in place; returns None
for spectrum in exp:
if spectrum.getMSLevel() == 1:
mz, intensity = spectrum.get_peaks() # tuple of two numpy arrays
elif spectrum.getMSLevel() == 2:
precursor = spectrum.getPrecursors()[0] # getPrecursors returns a list
precursor_mz = precursor.getMZ()
window = precursor.getIsolationWindowLowerOffset() + precursor.getIsolationWindowUpperOffset()Goal: Get a trustworthy log2 intensity matrix with bookkeeping rows removed and missing values represented as NaN.
Approach: Strip Reverse/contaminant/site-only rows, resolve the semicolon protein-ID list to a leading ID, pick LFQ intensity columns, set 0 -> NaN, then log2-transform.
import pandas as pd
import numpy as np
pg = pd.read_csv('proteinGroups.txt', sep='\t', low_memory=False) # mixed-type cols
# Flag columns hold '+' or empty string; all three are proteinGroups-only bookkeeping
mask = (pg.get('Reverse', '') != '+') & (pg.get('Potential contaminant', '') != '+') & (pg.get('Only identified by site', '') != '+')
pg = pg[mask].copy()
# Protein IDs / Majority protein IDs / Gene names are SEMICOLON lists; take the first (leading/razor) entry
pg['leading_protein'] = pg['Protein IDs'].str.split(';').str[0]
pg['leading_gene'] = pg['Gene names'].where(pg['Gene names'].notna(), '').str.split(';').str[0]
lfq_cols = [c for c in pg.columns if c.startswith('LFQ intensity ')] # MaxLFQ-normalized, between-sample comparable
matrix = pg[['leading_protein', 'leading_gene'] + lfq_cols].copy()
matrix[lfq_cols] = matrix[lfq_cols].replace(0, np.nan) # MaxQuant writes 0 for missing; log2(0) = -inf
matrix[lfq_cols] = np.log2(matrix[lfq_cols])Goal: Reshape the long DIA-NN report into a confident protein-by-run matrix.
Approach: Read the parquet (default since 1.9, only default in 2.0), filter precursor- AND protein-group q-values to 1% FDR BEFORE pivoting on PG.MaxLFQ.
import pandas as pd
report = pd.read_parquet('report.parquet') # report.tsv dropped as default in DIA-NN 2.0
report = report[(report['Q.Value'] <= 0.01) & (report['PG.Q.Value'] <= 0.01)] # 1% FDR before quant
matrix = report.pivot_table(index='Protein.Group', columns='Run', values='PG.MaxLFQ', aggfunc='first')Goal: Quantify the missing-value pattern so the legitimate imputation class can be chosen downstream.
Approach: Count NaN per protein and per sample; relate the pattern to acquisition mode (DDA -> structured MNAR; DIA -> closer to MCAR). A correlation between missingness and mean abundance is the MNAR signature.
import numpy as np
def assess_missingness(matrix, sample_cols):
miss_per_protein = matrix[sample_cols].isna().sum(axis=1)
miss_per_sample = matrix[sample_cols].isna().sum(axis=0)
total_pct = 100 * matrix[sample_cols].isna().sum().sum() / matrix[sample_cols].size
mean_abund = matrix[sample_cols].mean(axis=1) # negative corr with missingness => MNAR / left-censored
mnar_corr = mean_abund.corr(miss_per_protein)
return {'per_protein': miss_per_protein, 'per_sample': miss_per_sample, 'total_pct': total_pct, 'abundance_missing_corr': mnar_corr}Trigger: Reading Intensity (raw) or iBAQ when between-sample ratios are intended.
Mechanism: Intensity is un-normalized summed precursor signal; iBAQ is a within-sample molar proxy. Neither is comparable across samples the way LFQ intensity is.
Symptom: Ratios track total loaded protein / sample depth rather than biology; fold changes shift when one sample's loading changes.
Fix: Use LFQ intensity for cross-sample comparison; if computing custom normalization use Intensity and normalize explicitly (expression-matrix/normalization).
Trigger: np.log2 applied directly to a MaxQuant matrix still containing 0.
Mechanism: MaxQuant encodes "not quantified" as 0; log2(0) = -inf, which then propagates into means and tests.
Symptom: -inf values, NaN means, proteins silently dropped or skewed.
Fix: replace(0, np.nan) before any transform; then diagnose missingness.
Trigger: Loading proteinGroups.txt without filtering Reverse / Potential contaminant / Only identified by site.
Mechanism: Decoys exist only for FDR estimation; contaminants are keratin/trypsin/BSA, not the sample; site-only groups have no unmodified-peptide quant evidence. Only identified by site exists only in proteinGroups.txt.
Symptom: Inflated protein counts; a "hit" that is a decoy or keratin.
Fix: Filter all three flag columns; cross-check with REV__/CON__ ID prefixes when joining to peptide tables. Caveat: do not delete CON__ rows blindly if a contaminant (e.g. keratin) is the protein of interest.
Trigger: Treating Protein IDs or Gene names as an atomic single value.
Mechanism: These are semicolon-delimited lists; the first entry is the leading (razor) protein for the group, and Gene names can be blank while protein IDs are present.
Symptom: Merges fail, NaN gene labels, ambiguous identity downstream.
Fix: Split on ; and take the first entry; guard Gene names with .notna(). Group parsimony details -> protein-inference.
Trigger: Reading report.tsv on DIA-NN 2.0, or pivoting before q-filtering.
Mechanism: 2.0 defaults to (and only defaults to) report.parquet; pivoting unfiltered rows includes precursors above 1% FDR.
Symptom: FileNotFoundError on report.tsv; or low-confidence quant inflating the matrix.
Fix: pd.read_parquet('report.parquet'); filter Q.Value <= 0.01 & PG.Q.Value <= 0.01 before pivoting PG.MaxLFQ.
Trigger: Mean/median/KNN imputation on a DDA matrix. Mechanism: DDA missingness is abundance-dependent (left-censored); MCAR imputers fill missing low values with the central tendency, biasing them upward. Symptom: Low-abundance proteins gain false high values; spurious differential hits. Fix: Diagnose the abundance-missingness correlation here; route DDA to left-censored imputation (downshifted-Gaussian / QRILC / MinProb) in differential-abundance; DIA tolerates standard imputers.
| Threshold | Source | Rationale |
|---|---|---|
DIA-NN import filter Q.Value <= 0.01 AND PG.Q.Value <= 0.01 | Demichev 2020; target-decoy convention | Precursor- and protein-group-level 1% FDR enforced before any quant value is used |
| Peptide/protein FDR 1% (q <= 0.01) | Target-decoy convention | Standard ID confidence at both peptide and protein levels |
| MaxQuant zero -> NaN | MaxQuant output convention | 0 encodes "not quantified"; log2(0) = -inf corrupts every transform |
| Min peptides per protein for quant >= 2 | Community quant practice | Single-peptide ("one-hit-wonder") proteins are ID/quant-unreliable |
| Valid-value filter >= 50-70% per group | Modeling choice (document per study) | Caps imputation burden; the exact cutoff is a study decision, not a universal constant |
| Take FIRST semicolon entry as leading protein/gene | MaxQuant proteinGroups convention | The leading/razor protein is the group identifier; trailing entries are shared-peptide members |
| Error / symptom | Cause | Solution |
|---|---|---|
-inf values after log2 | Zeros not converted to NaN | df.replace(0, np.nan) before np.log2 |
FileNotFoundError: report.tsv (DIA-NN 2.0) | TSV no longer the default output | pd.read_parquet('report.parquet') |
KeyError: 'Only identified by site' | That column exists ONLY in proteinGroups.txt | Use df.get('Only identified by site', '') or guard the column lookup |
| Mixed-type / DtypeWarning on MaxQuant load | Wide TSV with mixed column types | pd.read_csv(..., low_memory=False) |
| NaN gene labels break a merge | Gene names is a semicolon list, sometimes blank | .where(notna(), '').str.split(';').str[0] |
| Ratios track loading not biology | Read Intensity (raw) instead of LFQ intensity | Use LFQ intensity for between-sample comparison |
get_peaks() unpacking error | Expecting a 2D array | It returns a tuple (mz, intensity) of two numpy arrays |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in proteomics/data-import of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Proteomics Data Import next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Proteomics Data Import this skillGPTomics/bioSkills | 1.2k | 1 repos | ~4.5k | Automated safety check: Pass | MIT | |
| Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent | 885 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Brokerage Screenshot Trade Extractorguilhermecgs/ir | 179 | — | ~1.2k | Automated safety check: Pass | MPL-2.0 | |
| Candlestick Pattern SignalsHKUDS/Vibe-Trading | 35k | — | ~468 | Automated safety check: Pass | MIT | |
| Tushare Datazillionare/zillionare | 321 | 2 repos | ~2.3k | Automated safety check: Pass | None | |
| Quant Blog Writingzillionare/zillionare | 321 | — | ~895 | Automated safety check: Pass | None |
NVIDIA-AI-Blueprints/deep-researcher-agent
A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.
guilhermecgs/ir
Reads screenshots of brokerage or portfolio transaction tables, checks the rows for duplicates and prints lines you can paste into your operations file by hand.
HKUDS/Vibe-Trading
Detects 15 classic candlestick patterns with vectorized pandas code and combines bullish and bearish scores into a long, short or flat trading signal.
zillionare/zillionare
面向中文自然语言的 Tushare 数据研究技能。用于把“看看这只股票最近怎么样”“帮我查财报趋势”“最近哪个板块最强”“北向资金在买什么”“给我导出一份行情数据”这类请求,转成可执行的数据获取、清洗、对比、筛选、导出与简要分析流程。适用于 A 股、指数、ETF/基金、财务、估值、资金流、公告新闻、板块概念与宏观数据等研究场景。
zillionare/zillionare
撰写文笔精炼、富有深度的量化交易博文,论点清晰、证据确凿、叙事层次更加丰富。适用于量化交易博文、因子研究、回测复盘、数据源排查、市场微观结构、策略原理、风险控制、职业观察、量化人物故事等选题。文章将聚焦具体角度,提供详实的大纲、证据规划及成稿,力求内容兼具思想深度与诚实性,而非单纯口号式宣传;同时,通过人物经历、引言、贡献及行业背景的融入,让文章更具可读性和吸引力。
bex-co/beancount-io
Write or repair a reusable Beangulp importer from a sample bank export, with reviewed golden files and a passing test harness.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
Loads mass-spectrometry data into Python/R and strips the search engine's bookkeeping before any number is trusted -- removes decoys (REV/Reverse), contaminants (CON/Potential contaminant)…. Bio Proteomics Data Import is an agent skill from GPTomics/bioSkills.parquet, and mzML/mzXML.
Bio Proteomics Data Import fits situations like: starting an analysis from raw spectra; A search engine output.
Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a claude-code`. Or copy the skill folder (proteomics/data-import in GPTomics/bioSkills) into .claude/skills/bio-proteomics-data-import in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a codex`. Or copy the skill folder (proteomics/data-import in GPTomics/bioSkills) into .agents/skills/bio-proteomics-data-import in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-proteomics-data-import, .gemini/skills/bio-proteomics-data-import, .github/skills/bio-proteomics-data-import and .opencode/skills/bio-proteomics-data-import in your project.
Going by SKILL.md and its folder, Bio Proteomics Data Import needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Proteomics Data Import is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Proteomics Data Import: Data Table Analysis (NVIDIA-AI-Blueprints/deep-researcher-agent, 885 stars), Brokerage Screenshot Trade Extractor (guilhermecgs/ir, 179 stars), Candlestick Pattern Signals (HKUDS/Vibe-Trading, 35k stars) and Tushare Data (zillionare/zillionare, 321 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.