Agent skill

Genotex Benchmark Guide

by wentorai in wentorai/research-plugins

Benchmark for LLM agents on gene expression data analysis. An agent skill from wentorai/research-plugins.

MITAuto-check passedData & Analytics

Install Genotex Benchmark Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill genotex-benchmark-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins genotex-benchmark-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/domains/biomedical/genotex-benchmark-guide .claude/skills/genotex-benchmark-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
genotex-benchmark-guide
GitHub stars
298
Used in
1 other repo
Token cost
~956 tokens
SKILL.md length
106 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Benchmark for LLM agents on gene expression data analysis. An agent skill from wentorai/research-plugins.

  • Works in 5 steps: Agent evaluation: Test bioinformatics… → Method comparison: Compare LLM agents on… → Benchmark development: Extend with new… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, Benchmark Structure, Usage and Running Evaluations, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Genotex Benchmark Guide is an agent skill from wentorai/research-plugins. Benchmark for LLM agents on gene expression data analysis

Its SKILL.md is about 960 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Bioinformatics and Data analysis. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve Data analysis

Example prompts

  • “/genotex-benchmark-guide”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Agent evaluation: Test bioinformatics agents on real tasks
  2. Method comparison: Compare LLM agents on genomics
  3. Benchmark development: Extend with new GEO datasets
  4. Teaching: Standard tasks for bioinformatics education
  5. Tool development: Test new analysis pipelines

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • ncbi.nlm.nih.gov
    • mlcb.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Genotex Benchmark Guide loads about 956 tokens when it runs. Until then it costs about 20 tokens; SKILL.md has 106 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~20
When it runs · the whole SKILL.md, loaded when a task matches
~956

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 106 words, ~956 tokens.

Download SKILL.mdSave it as .claude/skills/genotex-benchmark-guide/SKILL.md (or your agent's skills folder).
name
genotex-benchmark-guide
description
Benchmark for LLM agents on gene expression data analysis

GenoTEX Benchmark Guide

Overview

GenoTEX is a benchmark for evaluating LLM-based agents on gene expression data analysis tasks. It provides curated datasets from GEO (Gene Expression Omnibus) with ground-truth analysis pipelines, testing agents on data preprocessing, differential expression, enrichment analysis, and biological interpretation. Published at MLCB 2025 as an oral presentation.

Benchmark Structure

GenoTEX Benchmark
├── Data Collection
│   └── Curated GEO datasets with ground truth
├── Task Categories
│   ├── Data preprocessing (QC, normalization)
│   ├── Differential expression analysis
│   ├── Gene set enrichment analysis
│   ├── Clustering and classification
│   └── Biological interpretation
├── Evaluation
│   ├── Code correctness (executes without error)
│   ├── Statistical validity (appropriate tests)
│   ├── Result accuracy (vs ground truth)
│   └── Interpretation quality (biological insight)
└── Baselines
    ├── GPT-4 agent
    ├── Claude agent
    └── Domain-specific fine-tuned models

Usage

python
from genotex import GenoTEXBenchmark

bench = GenoTEXBenchmark()

# List available tasks
tasks = bench.list_tasks()
for task in tasks[:5]:
    print(f"Task: {task.id}")
    print(f"  Dataset: {task.geo_accession}")
    print(f"  Category: {task.category}")
    print(f"  Difficulty: {task.difficulty}")

# Get a specific task
task = bench.get_task("GSE12345_DEG")
print(f"Description: {task.description}")
print(f"Input files: {task.input_files}")
print(f"Expected output: {task.expected_output_type}")

Running Evaluations

python
# Evaluate an agent on GenoTEX
from genotex import evaluate_agent

results = evaluate_agent(
    agent_fn=my_agent_function,
    tasks="all",            # or specific task IDs
    timeout_per_task=300,   # seconds
)

print(f"Tasks completed: {results.completed}/{results.total}")
print(f"Code correctness: {results.code_correct_rate:.1%}")
print(f"Statistical validity: {results.stats_valid_rate:.1%}")
print(f"Result accuracy: {results.accuracy:.3f}")

Task Examples

python
# Example: Differential Expression Analysis
task = {
    "id": "GSE12345_DEG",
    "description": "Identify differentially expressed genes "
                   "between treatment and control groups in "
                   "this RNA-seq dataset.",
    "input": "GSE12345_counts.csv",  # Raw count matrix
    "metadata": "GSE12345_metadata.csv",  # Sample info
    "expected": {
        "method": "DESeq2 or limma-voom",
        "output": "DEG table with log2FC, p-value, adj.p",
        "ground_truth": "GSE12345_deg_truth.csv",
    },
}

# Example: Gene Set Enrichment
task = {
    "id": "GSE12345_GSEA",
    "description": "Perform gene set enrichment analysis on "
                   "the DEGs and identify enriched pathways.",
    "input": "GSE12345_deg_results.csv",
    "expected": {
        "method": "fgsea, clusterProfiler, or enrichR",
        "output": "Enriched pathways with NES and FDR",
    },
}

Use Cases

  1. Agent evaluation: Test bioinformatics agents on real tasks
  2. Method comparison: Compare LLM agents on genomics
  3. Benchmark development: Extend with new GEO datasets
  4. Teaching: Standard tasks for bioinformatics education
  5. Tool development: Test new analysis pipelines

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/domains/biomedical/genotex-benchmark-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Genotex Benchmark Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Genotex Benchmark Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Genotex Benchmark Guide this skillwentorai/research-plugins2981 repos~956Automated safety check: PassMIT
Exploratory Data Analysisspacering-net/codeg3.9k14 repos~3.6kAutomated safety check: PassMIT
Pyopenmsdavila7/claude-code-templates32k11 repos~1.4kAutomated safety check: PassMIT
Gwas Databasedavila7/claude-code-templates32k10 repos~5kAutomated safety check: PassMIT
Bioconductor BiomartbioMate-AI/biomate-bioconductor-kb804—~4.5kAutomated safety check: PassCustom licence
Bio Data Visualization Manhattan Qq LocuszoomGPTomics/bioSkills1.2k2 repos~4.3kAutomated safety check: PassMIT

Similar skills

  • Exploratory Data Analysis

    spacering-net/codeg

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    3.9k GitHub starsUsed in 14 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Pyopenms

    davila7/claude-code-templates

    Python interface to OpenMS for mass spectrometry data analysis.

    32k GitHub starsUsed in 11 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Gwas Database

    davila7/claude-code-templates

    Query NHGRI-EBI GWAS Catalog for SNP-trait associations. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~5k tokens
    Data & AnalyticsAuto-check passed
  • Bioconductor Biomart

    bioMate-AI/biomate-bioconductor-kb

    In recent years a wealth of biological data has become available in public data repositories.

    804 GitHub stars~4.5k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Build Manhattan, Miami, QQ, and locuszoom-style regional plots from GWAS, TWAS, PWAS, and QTL summary statistics with correct genomic-inflation diagnostics, multi-trait overlays, lead-SNP labeling…

    1.2k GitHub starsUsed in 2 repos~4.3k tokens
    Data & AnalyticsAuto-check passed
  • Heatmap Beautifier

    aipoch/medical-research-skills

    Professional beautification tool for gene expression heatmaps, automatically adds clustering trees, color annotation tracks, and intelligently optimizes label layout.

    2k GitHub stars~3.5k tokensUpdated 23 days ago
    Data & AnalyticsAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Genotex Benchmark Guide

What does Genotex Benchmark Guide do?

Benchmark for LLM agents on gene expression data analysis. An agent skill from wentorai/research-plugins. Genotex Benchmark Guide is an agent skill from wentorai/research-plugins.

When should I use Genotex Benchmark Guide?

Genotex Benchmark Guide fits situations like: tasks that involve Bioinformatics; tasks that involve Data analysis.

How do I install Genotex Benchmark Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill genotex-benchmark-guide -a claude-code`. Or copy the skill folder (skills/domains/biomedical/genotex-benchmark-guide in wentorai/research-plugins) into .claude/skills/genotex-benchmark-guide in your project. Claude Code loads it when a task matches its description.

How do I install Genotex Benchmark Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill genotex-benchmark-guide -a codex`. Or copy the skill folder (skills/domains/biomedical/genotex-benchmark-guide in wentorai/research-plugins) into .agents/skills/genotex-benchmark-guide in your project. Codex loads it when a task matches its description.

Can I use Genotex Benchmark Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill genotex-benchmark-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/genotex-benchmark-guide, .gemini/skills/genotex-benchmark-guide, .github/skills/genotex-benchmark-guide and .opencode/skills/genotex-benchmark-guide in your project.

What does Genotex Benchmark Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: Genotex Benchmark Guide is instructions for the agent only. Our summary lists: Python 3.

Does Genotex Benchmark Guide access the network?

SKILL.md names 3 domains. As links in the text: github.com, ncbi.nlm.nih.gov and mlcb.github.io. This is read from the text; nothing was executed.

Is Genotex Benchmark Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Genotex Benchmark Guide use?

Genotex Benchmark Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Genotex Benchmark Guide use?

About 956 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Genotex Benchmark Guide?

Skills that share tags, products or a category with Genotex Benchmark Guide: Exploratory Data Analysis (spacering-net/codeg, 3.9k stars), Pyopenms (davila7/claude-code-templates, 32k stars), Gwas Database (davila7/claude-code-templates, 32k stars) and Bioconductor Biomart (bioMate-AI/biomate-bioconductor-kb, 804 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Genotex Benchmark Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.