Agent skill

Bio Proteomics Data Import

by GPTomics in GPTomics/bioSkills

Loads mass-spectrometry data into Python/R and strips the search engine's bookkeeping before any number is trusted -- removes decoys (REV/Reverse), contaminants (CON/Potential contaminant)…

MITAuto-check passedBusiness, Finance & HR

Install Bio Proteomics Data Import

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-proteomics-data-import --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/proteomics/data-import .claude/skills/bio-proteomics-data-import && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-proteomics-data-import
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.5k tokens
SKILL.md length
1,825 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Loads mass-spectrometry data into Python/R and strips the search engine's bookkeeping before any number is trusted -- removes decoys (REV/Reverse), contaminants (CON/Potential contaminant)…

  • Works in 3 steps: A "data import" is never just file… → The search engine's bookkeeping must be… → The same proteinGroups.txt yields…
  • Starting an analysis from raw spectra
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 9 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Bio Proteomics Data Import is an agent skill from GPTomics/bioSkills. Loads mass-spectrometry data into Python/R and strips the search engine's bookkeeping before any number is trusted -- removes decoys (REV/Reverse), contaminants (CON/Potential contaminant), Only-identified-by-site groups, and resolves semicolon razor/leading protein-ID ambiguity in MaxQuant proteinGroups.txt, DIA-NN report.parquet, and mzML/mzXML. Distinguishes Intensity (raw) vs LFQ intensity (MaxLFQ) vs iBAQ, treats a MaxQuant zero as missing (NaN, not log2(-inf)), and inherits the acquisition mode's…

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/load_maxquant.py` and `usage-guide.md`).

It sits in Business, Finance & HR, covering Bioinformatics, Accounting and bookkeeping and DataFrames. It works with Python and pandas. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Starting an analysis from raw spectra
  • A search engine output

Example prompts

  • “Use the bio-proteomics-data-import skill to load mass-spectrometry data into Python/R and strips the search engine's bookkeeping before any number…”
  • “/bio-proteomics-data-import”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. A "data import" is never just file parsing -- it is the moment the acquisition mode's quantitative contract and its missingness structure…
  2. The search engine's bookkeeping must be stripped before any number is trusted. A proteinGroups.txt carries decoy rows (Reverse == '+', REV…
  3. The same proteinGroups.txt yields different biology from different columns, and a zero is not a measurement. Intensity is raw summed…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Proteomics Data Import loads about 4.5k tokens when it runs. Until then it costs about 201 tokens; SKILL.md has 1,825 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~201
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,825 words, ~4,503 tokens.

Download SKILL.mdSave it as .claude/skills/bio-proteomics-data-import/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-proteomics-data-import
description
Loads mass-spectrometry data into Python/R and strips the search engine's bookkeeping before any number is trusted -- removes decoys (REV__/Reverse), contaminants (CON__/Potential contaminant), Only-identified-by-site groups, and resolves semicolon razor/leading protein-ID ambiguity in MaxQuant proteinGroups.txt, DIA-NN report.parquet, and mzML/mzXML. Distinguishes Intensity (raw) vs LFQ intensity (MaxLFQ) vs iBAQ, treats a MaxQuant zero as missing (NaN, not log2(-inf)), and inherits the acquisition mode's missingness contract (DDA MNAR vs DIA MCAR). Use when starting an analysis from raw spectra or a search engine output. Downstream normalization and stats are differential-abundance; reporter-ion/MaxLFQ quant is quantification; protein grouping is protein-inference.
tool_type
mixed
primary_tool
pyOpenMS

Version Compatibility

Reference examples tested with: pyOpenMS 3.1+, pandas 2.2+, numpy 1.26+, MSnbase 2.28+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Mass Spectrometry Data Import -- Inheriting the Acquisition Contract and Stripping the Bookkeeping

"Load my mass spec data into Python" -> Parse spectra or a search-engine table AND immediately enforce two contracts -- which quant column carries real biology, and which rows are search-engine bookkeeping that must be deleted -- because the same proteinGroups.txt yields different conclusions depending on the column read and the rows kept.

  • Python: pyopenms.MzMLFile().load(path, exp) for raw spectra; pandas.read_csv(sep='\t') for MaxQuant; pandas.read_parquet for DIA-NN
  • R: Spectra::Spectra() / QFeatures::readQFeatures() for raw and quantified data (MSnbase still works but is in maintenance mode)

Scope: this skill owns reading spectra/search outputs into memory, deleting decoy/contaminant/site-only rows, picking the correct quant column, and characterizing missingness. Format conversion (RAW -> mzML) -> peptide-identification. MaxLFQ/TMT reporter quant computation -> quantification. Protein-group parsimony -> protein-inference. Normalization and imputation -> differential-abundance and expression-matrix/normalization. OUT OF SCOPE: statistical testing, batch correction, and the actual imputation step (this skill only diagnoses the missingness so the right imputer is chosen later).

The Single Most Important Modern Insight -- Import Is Where Two Contracts Are Read and Enforced

  1. A "data import" is never just file parsing -- it is the moment the acquisition mode's quantitative contract and its missingness structure are inherited. DDA selects the top-N most intense precursors per cycle, and which precursors get picked is partly stochastic and abundance-biased, so the same low-abundance peptide is sampled in run A and missed in run B; this manufactures structured, left-censored MNAR missingness. DIA fragments every precursor in every window every cycle, so its (fewer) missing values are closer to MCAR. The catastrophic error this prevents: imputing a DDA matrix with a mean/KNN method that assumes MCAR, which biases low-abundance proteins upward and manufactures false hits. The mode is born at acquisition and inherited at import; the missingness diagnosis made here dictates which imputation is even legitimate downstream.

  2. The search engine's bookkeeping must be stripped before any number is trusted. A proteinGroups.txt carries decoy rows (Reverse == '+', REV__ prefix in the ID) from the target-decoy FDR machinery, contaminant rows (Potential contaminant == '+', CON__ prefix), and Only-identified-by-site rows (the protein has no unmodified-peptide evidence, only a modified site). Keeping any of these leaks non-biological signal into the intensity matrix and inflates IDs. The catastrophic error: reporting differential abundance on a matrix where decoy or keratin rows survived.

  3. The same proteinGroups.txt yields different biology from different columns, and a zero is not a measurement. Intensity is raw summed precursor signal (not normalized, not comparable across samples for ratios). LFQ intensity is MaxLFQ-normalized and is the column for between-sample comparison. iBAQ is intensity divided by the number of observable tryptic peptides -- a within-sample molar proxy, not a between-sample quant. MaxQuant writes 0 for "not quantified", so log2(0) = -inf; replace 0 -> NaN before any transform. The catastrophic error: log2-transforming raw Intensity (or iBAQ) and reading the ratios as biology.

Tool Taxonomy

Tool / methodCitationMechanism / roleWhen
pyOpenMS MzMLFile().loadChambers 2012 (ProteoWizard lineage)Loads mzML/mzXML into an MSExperiment in memory; iterate spectra by MS levelProgrammatic access to raw peaks, precursor m/z, isolation windows
pandas read_csv/read_parquet--Tabular ingest of MaxQuant TSV and DIA-NN parquetAll search-engine output tables
DIA-NN reportDemichev 2020Long-format precursor table; report.parquet is the default (1.9+) and the only default (2.0)DIA quant; pivot on PG.MaxLFQ after q-filtering
MaxQuant txt/ outputsCox 2014 (MaxLFQ)proteinGroups.txt (group level), evidence.txt (per-PSM)DDA label-free / TMT search results
Spectra + QFeatures (R)--Current Bioconductor raw + quantified-feature containers; readQFeatures, aggregateFeaturesR pipelines; preferred over MSnbase going forward
MSnbase readMSData (R)--On-disk raw reading; maintenance mode (route OUT to Spectra/QFeatures)Legacy R code only
ThermoRawFileParser / msconvertHulstaert 2020 / Chambers 2012RAW -> mzML conversion (route OUT)File conversion is peptide-identification

Decision Tree by Scenario

ScenarioRecommendedWhy
MaxQuant DDA label-free, between-sample comparisonRead LFQ intensity columns from proteinGroups.txtMaxLFQ-normalized; the only MaxQuant column valid for cross-sample ratios
MaxQuant, absolute/molar abundance within one sampleRead iBAQ columnsiBAQ is a within-sample molar proxy; do not use across samples
Need raw uncorrected signal for a custom normalizationRead Intensity columns, normalize yourselfIntensity is raw summed precursor area, not comparable as-is
DIA-NN output (1.9 or 2.0)pd.read_parquet('report.parquet'), filter q-values, pivot PG.MaxLFQ2.0 dropped the TSV default; q-filter before pivot or low-confidence rows leak in
Raw spectra, need peaks/precursor/isolation windowpyOpenMS MzMLFile().loadProgrammatic peak and isolation-window access for QC and co-isolation reasoning
R-based pipeline, quantified featuresQFeatures readQFeatures + aggregateFeaturesCurrent Bioconductor; MSnbase is maintenance-only
Data came from DDA, planning imputationDiagnose missingness as MNAR -> route to left-censored imputationDDA top-N sampling makes missingness abundance-dependent
Data came from DIA, planning imputationTreat missingness as closer to MCARDIA samples every precursor every cycle

Default when uncertain: read LFQ intensity (MaxQuant) or PG.MaxLFQ after q-filtering (DIA-NN), strip Reverse/contaminant/site-only rows, set 0 -> NaN, then diagnose missingness before choosing an imputer.

Loading mzML/mzXML with pyOpenMS

Goal: Parse raw spectra into memory for QC, peak access, and isolation-window reasoning.

Approach: Load into an MSExperiment (filled in place), iterate by MS level; get_peaks() returns a tuple of (mz, intensity) numpy arrays, and getPrecursors() returns a list.

python
from pyopenms import MSExperiment, MzMLFile

exp = MSExperiment()
MzMLFile().load('sample.mzML', exp)  # fills exp in place; returns None

for spectrum in exp:
    if spectrum.getMSLevel() == 1:
        mz, intensity = spectrum.get_peaks()  # tuple of two numpy arrays
    elif spectrum.getMSLevel() == 2:
        precursor = spectrum.getPrecursors()[0]  # getPrecursors returns a list
        precursor_mz = precursor.getMZ()
        window = precursor.getIsolationWindowLowerOffset() + precursor.getIsolationWindowUpperOffset()

Loading and Cleaning MaxQuant proteinGroups.txt

Goal: Get a trustworthy log2 intensity matrix with bookkeeping rows removed and missing values represented as NaN.

Approach: Strip Reverse/contaminant/site-only rows, resolve the semicolon protein-ID list to a leading ID, pick LFQ intensity columns, set 0 -> NaN, then log2-transform.

python
import pandas as pd
import numpy as np

pg = pd.read_csv('proteinGroups.txt', sep='\t', low_memory=False)  # mixed-type cols

# Flag columns hold '+' or empty string; all three are proteinGroups-only bookkeeping
mask = (pg.get('Reverse', '') != '+') & (pg.get('Potential contaminant', '') != '+') & (pg.get('Only identified by site', '') != '+')
pg = pg[mask].copy()

# Protein IDs / Majority protein IDs / Gene names are SEMICOLON lists; take the first (leading/razor) entry
pg['leading_protein'] = pg['Protein IDs'].str.split(';').str[0]
pg['leading_gene'] = pg['Gene names'].where(pg['Gene names'].notna(), '').str.split(';').str[0]

lfq_cols = [c for c in pg.columns if c.startswith('LFQ intensity ')]  # MaxLFQ-normalized, between-sample comparable
matrix = pg[['leading_protein', 'leading_gene'] + lfq_cols].copy()
matrix[lfq_cols] = matrix[lfq_cols].replace(0, np.nan)  # MaxQuant writes 0 for missing; log2(0) = -inf
matrix[lfq_cols] = np.log2(matrix[lfq_cols])

Loading DIA-NN report.parquet

Goal: Reshape the long DIA-NN report into a confident protein-by-run matrix.

Approach: Read the parquet (default since 1.9, only default in 2.0), filter precursor- AND protein-group q-values to 1% FDR BEFORE pivoting on PG.MaxLFQ.

python
import pandas as pd

report = pd.read_parquet('report.parquet')  # report.tsv dropped as default in DIA-NN 2.0
report = report[(report['Q.Value'] <= 0.01) & (report['PG.Q.Value'] <= 0.01)]  # 1% FDR before quant

matrix = report.pivot_table(index='Protein.Group', columns='Run', values='PG.MaxLFQ', aggfunc='first')

Diagnosing the Missingness Contract

Goal: Quantify the missing-value pattern so the legitimate imputation class can be chosen downstream.

Approach: Count NaN per protein and per sample; relate the pattern to acquisition mode (DDA -> structured MNAR; DIA -> closer to MCAR). A correlation between missingness and mean abundance is the MNAR signature.

python
import numpy as np

def assess_missingness(matrix, sample_cols):
    miss_per_protein = matrix[sample_cols].isna().sum(axis=1)
    miss_per_sample = matrix[sample_cols].isna().sum(axis=0)
    total_pct = 100 * matrix[sample_cols].isna().sum().sum() / matrix[sample_cols].size
    mean_abund = matrix[sample_cols].mean(axis=1)  # negative corr with missingness => MNAR / left-censored
    mnar_corr = mean_abund.corr(miss_per_protein)
    return {'per_protein': miss_per_protein, 'per_sample': miss_per_sample, 'total_pct': total_pct, 'abundance_missing_corr': mnar_corr}

Per-Method Failure Modes

MaxQuant wrong quant column

Trigger: Reading Intensity (raw) or iBAQ when between-sample ratios are intended. Mechanism: Intensity is un-normalized summed precursor signal; iBAQ is a within-sample molar proxy. Neither is comparable across samples the way LFQ intensity is. Symptom: Ratios track total loaded protein / sample depth rather than biology; fold changes shift when one sample's loading changes. Fix: Use LFQ intensity for cross-sample comparison; if computing custom normalization use Intensity and normalize explicitly (expression-matrix/normalization).

Show full SKILL.md (724 more words)Show less
Zero treated as a measurement

Trigger: np.log2 applied directly to a MaxQuant matrix still containing 0. Mechanism: MaxQuant encodes "not quantified" as 0; log2(0) = -inf, which then propagates into means and tests. Symptom: -inf values, NaN means, proteins silently dropped or skewed. Fix: replace(0, np.nan) before any transform; then diagnose missingness.

Bookkeeping rows survive

Trigger: Loading proteinGroups.txt without filtering Reverse / Potential contaminant / Only identified by site. Mechanism: Decoys exist only for FDR estimation; contaminants are keratin/trypsin/BSA, not the sample; site-only groups have no unmodified-peptide quant evidence. Only identified by site exists only in proteinGroups.txt. Symptom: Inflated protein counts; a "hit" that is a decoy or keratin. Fix: Filter all three flag columns; cross-check with REV__/CON__ ID prefixes when joining to peptide tables. Caveat: do not delete CON__ rows blindly if a contaminant (e.g. keratin) is the protein of interest.

Razor / leading protein-ID ambiguity ignored

Trigger: Treating Protein IDs or Gene names as an atomic single value. Mechanism: These are semicolon-delimited lists; the first entry is the leading (razor) protein for the group, and Gene names can be blank while protein IDs are present. Symptom: Merges fail, NaN gene labels, ambiguous identity downstream. Fix: Split on ; and take the first entry; guard Gene names with .notna(). Group parsimony details -> protein-inference.

Stale DIA-NN parsing

Trigger: Reading report.tsv on DIA-NN 2.0, or pivoting before q-filtering. Mechanism: 2.0 defaults to (and only defaults to) report.parquet; pivoting unfiltered rows includes precursors above 1% FDR. Symptom: FileNotFoundError on report.tsv; or low-confidence quant inflating the matrix. Fix: pd.read_parquet('report.parquet'); filter Q.Value <= 0.01 & PG.Q.Value <= 0.01 before pivoting PG.MaxLFQ.

MNAR imputed as MCAR

Trigger: Mean/median/KNN imputation on a DDA matrix. Mechanism: DDA missingness is abundance-dependent (left-censored); MCAR imputers fill missing low values with the central tendency, biasing them upward. Symptom: Low-abundance proteins gain false high values; spurious differential hits. Fix: Diagnose the abundance-missingness correlation here; route DDA to left-censored imputation (downshifted-Gaussian / QRILC / MinProb) in differential-abundance; DIA tolerates standard imputers.

Quantitative Thresholds

ThresholdSourceRationale
DIA-NN import filter Q.Value <= 0.01 AND PG.Q.Value <= 0.01Demichev 2020; target-decoy conventionPrecursor- and protein-group-level 1% FDR enforced before any quant value is used
Peptide/protein FDR 1% (q <= 0.01)Target-decoy conventionStandard ID confidence at both peptide and protein levels
MaxQuant zero -> NaNMaxQuant output convention0 encodes "not quantified"; log2(0) = -inf corrupts every transform
Min peptides per protein for quant >= 2Community quant practiceSingle-peptide ("one-hit-wonder") proteins are ID/quant-unreliable
Valid-value filter >= 50-70% per groupModeling choice (document per study)Caps imputation burden; the exact cutoff is a study decision, not a universal constant
Take FIRST semicolon entry as leading protein/geneMaxQuant proteinGroups conventionThe leading/razor protein is the group identifier; trailing entries are shared-peptide members

Common Errors

Error / symptomCauseSolution
-inf values after log2Zeros not converted to NaNdf.replace(0, np.nan) before np.log2
FileNotFoundError: report.tsv (DIA-NN 2.0)TSV no longer the default outputpd.read_parquet('report.parquet')
KeyError: 'Only identified by site'That column exists ONLY in proteinGroups.txtUse df.get('Only identified by site', '') or guard the column lookup
Mixed-type / DtypeWarning on MaxQuant loadWide TSV with mixed column typespd.read_csv(..., low_memory=False)
NaN gene labels break a mergeGene names is a semicolon list, sometimes blank.where(notna(), '').str.split(';').str[0]
Ratios track loading not biologyRead Intensity (raw) instead of LFQ intensityUse LFQ intensity for between-sample comparison
get_peaks() unpacking errorExpecting a 2D arrayIt returns a tuple (mz, intensity) of two numpy arrays

References

  • Cox J, Hein MY, Luber CA, Paron I, Nagaraj N, Mann M. 2014. Accurate proteome-wide label-free quantification by delayed normalization and maximal peptide ratio extraction, termed MaxLFQ. Mol Cell Proteomics 13(9):2513-2526.
  • Demichev V, Messner CB, Vernardis SI, Lilley KS, Ralser M. 2020. DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput. Nat Methods 17(1):41-44.
  • Chambers MC, Maclean B, Burke R, et al. 2012. A cross-platform toolkit for mass spectrometry and proteomics. Nat Biotechnol 30(10):918-920.
  • Hulstaert N, Shofstahl J, Sachsenberg T, et al. 2020. ThermoRawFileParser: modular, scalable, and cross-platform RAW file conversion. J Proteome Res 19(1):537-542.
  • peptide-identification - search raw spectra and convert vendor RAW to mzML
  • quantification - compute MaxLFQ and TMT reporter-ion quantities from imported data
  • protein-inference - resolve protein-group parsimony and razor assignment
  • differential-abundance - normalize, impute (per the missingness diagnosis), and test
  • proteomics-qc - assess run-level identification and quant quality
  • dia-analysis - run DIA-NN to produce the report this skill imports
  • expression-matrix/normalization - general intensity-matrix normalization patterns
  • workflows/proteomics-pipeline - end-to-end pipeline that begins with this import step

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in proteomics/data-import of GPTomics/bioSkills.

  • SKILL.md
  • examples/load_maxquant.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Proteomics Data Import next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Proteomics Data Import compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Proteomics Data Import this skillGPTomics/bioSkills1.2k1 repos~4.5kAutomated safety check: PassMIT
Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent885—~2.5kAutomated safety check: PassApache-2.0
Brokerage Screenshot Trade Extractorguilhermecgs/ir179—~1.2kAutomated safety check: PassMPL-2.0
Candlestick Pattern SignalsHKUDS/Vibe-Trading35k—~468Automated safety check: PassMIT
Tushare Datazillionare/zillionare3212 repos~2.3kAutomated safety check: PassNone
Quant Blog Writingzillionare/zillionare321—~895Automated safety check: PassNone

Similar skills

  • Data Table Analysis

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.

    885 GitHub stars~2.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Reads screenshots of brokerage or portfolio transaction tables, checks the rows for duplicates and prints lines you can paste into your operations file by hand.

    179 GitHub stars~1.2k tokensUpdated 3 mo ago
    Business, Finance & HRAuto-check passed
  • Candlestick Pattern Signals

    HKUDS/Vibe-Trading

    Detects 15 classic candlestick patterns with vectorized pandas code and combines bullish and bearish scores into a long, short or flat trading signal.

    35k GitHub stars~468 tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Tushare Data

    zillionare/zillionare

    面向中文自然语言的 Tushare 数据研究技能。用于把“看看这只股票最近怎么样”“帮我查财报趋势”“最近哪个板块最强”“北向资金在买什么”“给我导出一份行情数据”这类请求,转成可执行的数据获取、清洗、对比、筛选、导出与简要分析流程。适用于 A 股、指数、ETF/基金、财务、估值、资金流、公告新闻、板块概念与宏观数据等研究场景。

    321 GitHub starsUsed in 2 repos~2.3k tokens
    Business, Finance & HRAuto-check passed
  • Quant Blog Writing

    zillionare/zillionare

    撰写文笔精炼、富有深度的量化交易博文,论点清晰、证据确凿、叙事层次更加丰富。适用于量化交易博文、因子研究、回测复盘、数据源排查、市场微观结构、策略原理、风险控制、职业观察、量化人物故事等选题。文章将聚焦具体角度,提供详实的大纲、证据规划及成稿,力求内容兼具思想深度与诚实性,而非单纯口号式宣传;同时,通过人物经历、引言、贡献及行业背景的融入,让文章更具可读性和吸引力。

    321 GitHub stars~895 tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Beancount Importer Author

    bex-co/beancount-io

    Write or repair a reusable Beangulp importer from a sample bank export, with reviewed golden files and a passing test harness.

    296 GitHub stars~2k tokensUpdated today
    Business, Finance & HRAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Proteomics Data Import

What does Bio Proteomics Data Import do?

Loads mass-spectrometry data into Python/R and strips the search engine's bookkeeping before any number is trusted -- removes decoys (REV/Reverse), contaminants (CON/Potential contaminant)…. Bio Proteomics Data Import is an agent skill from GPTomics/bioSkills.parquet, and mzML/mzXML.

When should I use Bio Proteomics Data Import?

Bio Proteomics Data Import fits situations like: starting an analysis from raw spectra; A search engine output.

How do I install Bio Proteomics Data Import in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a claude-code`. Or copy the skill folder (proteomics/data-import in GPTomics/bioSkills) into .claude/skills/bio-proteomics-data-import in your project. Claude Code loads it when a task matches its description.

How do I install Bio Proteomics Data Import in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a codex`. Or copy the skill folder (proteomics/data-import in GPTomics/bioSkills) into .agents/skills/bio-proteomics-data-import in your project. Codex loads it when a task matches its description.

Can I use Bio Proteomics Data Import in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-proteomics-data-import -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-proteomics-data-import, .gemini/skills/bio-proteomics-data-import, .github/skills/bio-proteomics-data-import and .opencode/skills/bio-proteomics-data-import in your project.

What does Bio Proteomics Data Import need to run?

Going by SKILL.md and its folder, Bio Proteomics Data Import needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Proteomics Data Import access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Proteomics Data Import safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Proteomics Data Import use?

Bio Proteomics Data Import is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Proteomics Data Import use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Proteomics Data Import?

Skills that share tags, products or a category with Bio Proteomics Data Import: Data Table Analysis (NVIDIA-AI-Blueprints/deep-researcher-agent, 885 stars), Brokerage Screenshot Trade Extractor (guilhermecgs/ir, 179 stars), Candlestick Pattern Signals (HKUDS/Vibe-Trading, 35k stars) and Tushare Data (zillionare/zillionare, 321 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Proteomics Data Import?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.