Agent skill

Experiment Lab

by Citrus-bit in Citrus-bit/Anaxa

A skill your agent uses whenever the user wants reproducible CS/AI experiments, model evaluation, regression/classification/clustering analyses, bioinformatics workflows, QC, differential…

MITAuto-check passedResearch & Science

Install Experiment Lab

skills CLI
$ npx skills add Citrus-bit/Anaxa --skill experiment-lab -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Citrus-bit/Anaxa experiment-lab --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Citrus-bit/Anaxa.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/public/experiment-lab .claude/skills/experiment-lab && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-lab
GitHub stars
120
Token cost
~2.1k tokens
SKILL.md length
915 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses whenever the user wants reproducible CS/AI experiments, model evaluation, regression/classification/clustering analyses, bioinformatics workflows, QC, differential…

  • The user wants reproducible CS/AI experiments
  • SKILL.md covers Use This Skill For, Core Rules, Iterative Experiment Loop and CS/AI Routing, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Model evaluation

What it does

Experiment Lab is an agent skill from Citrus-bit/Anaxa. Use this skill whenever the user wants reproducible CS/AI experiments, model evaluation, regression/classification/clustering analyses, bioinformatics workflows, QC, differential expression, single-cell starter analysis, or manuscript-ready experiment bundles with scientific figures. Also use it when the user asks for heatmaps, ROC/PR curves, volcano plots, UMAP/PCA projections, or wants experimental results packaged for a paper or report.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics, Data visualization and Machine learning. It works with UMAP and Python. The repository describes itself as: Anaxa 是一个面向科研工作流的开源智能体系统。它不是单纯的聊天机器人,也不是无人监管的自动发论文机器,而是把文献检索、证据审计、实验执行、论文写作、同行评审式检查和最终产物打包放进同一个可追踪的研究生命周期中。 The licence is MIT.

When your agent uses it

  • The user wants reproducible CS/AI experiments
  • Model evaluation
  • Regression/classification/clustering analyses
  • Bioinformatics workflows

Example prompts

  • “/experiment-lab”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d57c708. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Lab loads about 2.1k tokens when it runs. Until then it costs about 115 tokens; SKILL.md has 915 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Citrus-bit/Anaxa at commit d57c708, republished under its MIT licence (© Citrus-bit). 915 words, ~2,055 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-lab/SKILL.md (or your agent's skills folder).
name
experiment-lab
description
Use this skill whenever the user wants reproducible CS/AI experiments, model evaluation, regression/classification/clustering analyses, bioinformatics workflows, QC, differential expression, single-cell starter analysis, or manuscript-ready experiment bundles with scientific figures. Also use it when the user asks for heatmaps, ROC/PR curves, volcano plots, UMAP/PCA projections, or wants experimental results packaged for a paper or report.

Experiment Lab Skill

This skill governs MedrixFlow's experiment workflow. It does not replace execution tools. It tells the agent when to use the structured experiment pipeline and how to keep results defensible.

Use This Skill For

  • CS/AI experiments on local structured data
  • Regression, classification, clustering, PCA/UMAP, and simple ablation tasks
  • Empirical social-science or public-health data analysis when paired with /mnt/skills/public/empirical-research-methods/SKILL.md
  • Bioinformatics workflows on local expression data, metadata tables, h5ad, or 10x matrix bundles
  • QC, differential expression, enrichment attempts, marker-gene summaries
  • Requests for manuscript-grade figures or experiment bundles for a report

Core Rules

  • Do not invent datasets, metrics, plots, baselines, enrichments, or scientific conclusions.
  • Exception: when thread context enables Synthetic Experiment Mode, personal experiment data may be simulated from explicit assumptions to complete the manuscript workflow. Public literature, DOI, baseline, leaderboard, benchmark, license, and dataset-version facts must still be real and verifiable.
  • For benchmark-driven work, use dataset_benchmark_discovery first to map candidate datasets, leaderboards, metrics, access limits, and baseline/SOTA hints.
  • If the task requires actual experimental output, prefer the experiment_lab tool instead of ad hoc reasoning.
  • Ask for missing inputs before running a workflow when the absence would invalidate the result:
    • no dataset
    • no target column for supervised CS/AI work
    • no metadata/group labels for requested differential analysis
  • Keep literature evidence separate from experimental evidence. Use academic_research only for related work or method framing.
  • Keep candidate benchmark evidence separate from executed experiment evidence. A dataset map can justify a plan, but only experiment_lab or attached result files can support manuscript result claims.
  • Default to Python-first execution. Do not ask the user to choose Python or R unless they explicitly request R output.
  • For empirical methods such as DID, IV, RDD, PSM/IPW, synthetic control, DML, causal forest, target-trial emulation, TMLE, survival, Table 1, event study, robustness, heterogeneity, or mechanism analysis, first load /mnt/skills/public/empirical-research-methods/SKILL.md and follow its identification gates.

Iterative Experiment Loop

For code-tuning, model-training, ablation, or autonomous experiment requests, use an autoresearch-style loop when the environment supports it:

  • Establish a baseline before changing code or hyperparameters.
  • Define one primary metric and direction up front, such as lower validation loss or higher AUROC.
  • Keep the evaluation harness, dataset split, and comparison budget fixed unless the user explicitly asks to change them.
  • Change one coherent idea per trial so the result remains attributable.
  • Record each trial with commit or run id, metric, memory/runtime if available, status (keep, discard, or crash), and a short description.
  • Keep a change only when it improves the primary metric or clearly simplifies the system without hurting the metric.
  • Treat crashes, OOMs, timeout, missing metrics, or incompatible dependencies as failed trials; summarize the failure and move on.
  • Do not start an indefinite or overnight loop unless the user explicitly asks for a long-running autonomous run.

CS/AI Routing

  • Supervised tabular prediction: run regression or classification through experiment_lab
  • Diagnostics:
    • binary classification -> confusion matrix, ROC, PR, feature importance if available
    • regression -> predicted-vs-actual, residual distribution, feature importance if available
  • Unsupervised tasks:
    • clustering -> cluster projection + silhouette
    • dimensionality reduction -> PCA/UMAP embedding
Show full SKILL.md (422 more words)Show less

Empirical Research Routing

For empirical social-science and public-health datasets, the empirical-research-methods skill defines the method contract. Use it to decide whether the requested work is descriptive, predictive, or causal, then pass method details into experiment_lab metadata:

  • empirical_method: did, staggered_did, iv, rdd, psm, synthetic_control, dml, causal_forest, target_trial, tmle, survival, or regression
  • estimand: ATE, ATT, LATE, CATE, event-time effect, risk difference, hazard ratio, or predictive metric
  • outcome, treatment, unit_id, time, treatment_time, instrument, running_variable, cutoff, covariates, fixed_effects, cluster
  • required_outputs: Table 1, main results, identification graphic, robustness, heterogeneity, mechanism, and reproducibility ledger

If the current backend cannot execute the requested causal estimator, run only the closest safe descriptive/regression workflow, record the requested method as not executed, and avoid causal claims.

Bioinformatics Routing

  • Bulk expression:
    • sample QC
    • sample PCA
    • top-variable-gene heatmap
    • differential analysis when metadata allows
    • enrichment attempt when significant genes and dependencies allow
  • Single-cell starter:
    • cell QC
    • PCA/UMAP embedding
    • clustering
    • marker summary
    • marker heatmap or violin plot

Expected Deliverables

When using the structured pipeline, prefer returning a concise summary plus artifacts. The artifact bundle should usually include:

  • experiment_contract.json
  • experiment_plan.md
  • methods.md
  • results.md
  • metrics.json
  • baseline_results.json
  • ablation_results.json
  • robustness_results.json
  • error_analysis.md
  • claim_support_matrix.json
  • Synthetic Experiment Mode also requires simulated_experiment_contract.json, simulation_assumptions.json, and synthetic_results.json.
  • Synthetic-mode claim maps should use supported_by_simulation, not supported_by_experiment, for simulated personal experimental outputs.
  • figure_manifest.json
  • figures/
  • tables/

If the experiment is linked to an academic project, also retain:

  • paper_ready_results.md
  • evidence.json

Tool Contract

Use experiment_lab when the user needs execution, metrics, or figure generation. Pass:

  • topic
  • dataset_paths
  • domain when clear
  • analysis_type when explicit
  • target_column for supervised CS/AI work
  • metadata_path, sample_id_column, and group_column for bulk expression tasks when available
  • linked_academic_project_id when the experiment should feed a formal report
  • synthetic_data_mode=true when Synthetic Experiment Mode is enabled
  • metadata with empirical method, estimand, variables, identification assumptions, and required outputs for empirical research tasks; in Synthetic Experiment Mode, include synthetic_data_mode=true and simulation assumptions when known

MATLAB Routing

Use matlab_execution only when the user explicitly needs MATLAB and command-line MATLAB is available. Generate a complete .m script, run it via batch mode, and write all figures, .mat, .csv, and logs into the provided output directory. Do not claim MATLAB GUI control; GUI automation requires a separate desktop integration outside this skill.

MATLAB results still need the same evidence discipline:

  • save numeric outputs in machine-readable files
  • summarize the method and parameters
  • link manuscript claims to generated artifacts
  • mark missing ablations or robustness checks as limitations

Final Response Pattern

  • Say what analysis ran
  • Report the main metrics or biological outputs
  • Mention important fallbacks or skipped stages
  • Point to the generated artifact bundle instead of dumping long raw output into chat

© Citrus-bit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/public/experiment-lab of Citrus-bit/Anaxa.

Open the folder on GitHubat commit d57c708

Compare with similar skills

Experiment Lab next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Lab compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Lab this skillCitrus-bit/Anaxa120—~2.1kAutomated safety check: PassMIT
deepTools NGS Toolkitdavila7/claude-code-templates32k13 repos~4.5kAutomated safety check: PassMIT
Gtars Genomic Interval Toolkitdavila7/claude-code-templates32k12 repos~1.9kAutomated safety check: PassMIT
FBA Flux Analyzeraiming-lab/AutoResearchClaw15k—~2.3kAutomated safety check: PassMIT
ScanpyK-Dense-AI/scientific-agent-skills48k1 repos~5.1kAutomated safety check: PassBSD-3-Clause
Bio Phylo Tree ManipulationGPTomics/bioSkills1.2k1 repos~5kAutomated safety check: PassMIT

Similar skills

  • deepTools NGS Toolkit

    davila7/claude-code-templates

    Guides use of deepTools on sequencing data: BAM to bigWig conversion, QC, sample correlation, and heatmaps or profiles around TSS and peaks for ChIP-seq, RNA-seq and ATAC-seq.

    32k GitHub starsUsed in 13 repos~4.5k tokens
    Research & ScienceAuto-check passed
  • Gtars Genomic Interval Toolkit

    davila7/claude-code-templates

    Works with genomic intervals using gtars, a Rust toolkit with Python bindings: overlap detection, coverage tracks, tokenization for ML models and reference sequences.

    32k GitHub starsUsed in 12 repos~1.9k tokens
    Research & ScienceAuto-check passed
  • FBA Flux Analyzer

    aiming-lab/AutoResearchClaw

    Turns raw flux balance analysis output and a COBRApy model into gene essentiality maps, phenotypic phase planes, flux sampling results, pathway summaries and secretion predictions.

    15k GitHub stars~2.3k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Scanpy

    K-Dense-AI/scientific-agent-skills

    Performs Scanpy single-cell RNA-seq QC, normalization, HVG selection, PCA/UMAP/t-SNE, clustering, exploratory marker ranking, pseudobulk preparation, visualization, and Seurat or…

    48k GitHub starsUsed in 1 repo~5.1k tokens
    Research & ScienceAuto-check passed
  • Bio Phylo Tree Manipulation

    GPTomics/bioSkills

    Edit phylogenetic tree structure with Biopython Bio.Phylo, and treat rooting as a separate statistical inference rather than a display choice.

    1.2k GitHub starsUsed in 1 repo~5k tokens
    Research & ScienceAuto-check passed
  • Bio Single Cell Clustering

    GPTomics/bioSkills

    Dimensionality reduction and graph-based clustering for single-cell RNA-seq with Scanpy (Python) and Seurat (R).

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Research & ScienceAuto-check passed

More from Citrus-bit/Anaxa

All 15 skills in this repo
  • Nature Figure

    Citrus-bit/Anaxa

    Submission-grade Nature/high-impact journal figure workflow for Python or R.

    120 GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Chart Visualization

    Citrus-bit/Anaxa

    A skill your agent uses whenever the user wants a chart, graph, plot, dashboard visual, or asks to visualize structured numbers, trends, comparisons, proportions, distributions, correlations…

    120 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Claude To Medrixflow

    Citrus-bit/Anaxa

    Interact with MedrixFlow AI agent platform via its HTTP API.

    120 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Fireworks Tech Graph

    Citrus-bit/Anaxa

    A skill your agent uses when the user wants to create any technical diagram - architecture, data flow, flowchart, sequence, agent/memory, or concept map - and export as SVG+PNG.

    120 GitHub stars~5.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Bootstrap

    Citrus-bit/Anaxa

    Generate a personalized SOUL.md through a warm, adaptive onboarding conversation.

    120 GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Deep Research

    Citrus-bit/Anaxa

    A skill your agent uses for general web research that needs current online information, multiple source angles, and synthesis, when no more specific research skill applies.

    120 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Experiment Lab

What does Experiment Lab do?

A skill your agent uses whenever the user wants reproducible CS/AI experiments, model evaluation, regression/classification/clustering analyses, bioinformatics workflows, QC, differential…. Experiment Lab is an agent skill from Citrus-bit/Anaxa. Use this skill whenever the user wants reproducible CS/AI experiments, model evaluation, regression/classification/clustering analyses, bioinformatics workflows, QC, differential expression, single-cell starter analysis, or manuscript-ready experiment bundles with scientific figures.

When should I use Experiment Lab?

Experiment Lab fits situations like: the user wants reproducible CS/AI experiments; model evaluation; regression/classification/clustering analyses; bioinformatics workflows.

How do I install Experiment Lab in Claude Code?

Run `npx skills add Citrus-bit/Anaxa --skill experiment-lab -a claude-code`. Or copy the skill folder (skills/public/experiment-lab in Citrus-bit/Anaxa) into .claude/skills/experiment-lab in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Lab in Codex?

Run `npx skills add Citrus-bit/Anaxa --skill experiment-lab -a codex`. Or copy the skill folder (skills/public/experiment-lab in Citrus-bit/Anaxa) into .agents/skills/experiment-lab in your project. Codex loads it when a task matches its description.

Can I use Experiment Lab in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Citrus-bit/Anaxa --skill experiment-lab -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-lab, .gemini/skills/experiment-lab, .github/skills/experiment-lab and .opencode/skills/experiment-lab in your project.

What does Experiment Lab need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiment Lab is instructions for the agent only. Our summary lists: Python 3.

Does Experiment Lab access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment Lab safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Lab use?

Experiment Lab is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Lab use?

About 2.1k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Experiment Lab?

Skills that share tags, products or a category with Experiment Lab: deepTools NGS Toolkit (davila7/claude-code-templates, 32k stars), Gtars Genomic Interval Toolkit (davila7/claude-code-templates, 32k stars), FBA Flux Analyzer (aiming-lab/AutoResearchClaw, 15k stars) and Scanpy (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Lab?

Citrus-bit (a GitHub user) maintains it in Citrus-bit/Anaxa, which has 120 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on September 7, 2026.

Source: Citrus-bit/Anaxa on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.