Agent skill

Arboreto Grn Inference

by jaechang-hits in jaechang-hits/SciAgent-Skills

GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest).

BSD-3-ClauseAuto-check passedResearch & Science

Install Arboreto Grn Inference

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills arboreto-grn-inference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/arboreto-grn-inference .claude/skills/arboreto-grn-inference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
arboreto-grn-inference
GitHub stars
374
Used in
2 other repos
Token cost
~5.3k tokens
SKILL.md length
1,350 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest).

  • Works in 7 steps: Load Expression Data → Load Transcription Factor List → Configure Dask Client (Optional) → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Pre-flight Interview, plus 10 more sections
  • Calls pip

What it does

Arboreto Grn Inference is an agent skill from jaechang-hits/SciAgent-Skills. GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest). Load matrix, filter by TFs, infer TF-target-importance links, save network. Dask-parallelized to single-cell scale. Core SCENIC component.

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics, DataFrames and Machine learning. It works with Dask. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.

When your agent uses it

  • Tasks that involve Bioinformatics
  • Tasks that involve DataFrames
  • Tasks that involve Machine learning

Example prompts

  • “/arboreto-grn-inference”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Load Expression Data
  2. Load Transcription Factor List
  3. Configure Dask Client (Optional)
  4. Run GRN Inference
  5. Filter Results
  6. Save Network
  7. Visualize Results (Optional)

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • pypi.org
    • distributed.dask.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Arboreto Grn Inference loads about 5.3k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 1,350 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 1,350 words, ~5,283 tokens.

Download SKILL.mdSave it as .claude/skills/arboreto-grn-inference/SKILL.md (or your agent's skills folder).
name
arboreto-grn-inference
description
GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest). Load matrix, filter by TFs, infer TF-target-importance links, save network. Dask-parallelized to single-cell scale. Core SCENIC component.
license
BSD-3-Clause

Arboreto GRN Inference

Overview

Arboreto infers gene regulatory networks (GRNs) from gene expression data using parallelized tree-based regression. For each target gene, it trains a regression model with all other genes (or a specified TF list) as features and emits TF-target-importance triplets. It provides two interchangeable algorithms -- GRNBoost2 (gradient boosting, fast) and GENIE3 (Random Forest, classic) -- sharing identical input/output formats. Computation is Dask-parallelized, scaling from laptop cores to HPC clusters.

When to Use

  • Inferring transcription factor-to-target gene regulatory relationships from bulk RNA-seq expression data
  • Building gene regulatory networks from single-cell RNA-seq count matrices (cells as rows, genes as columns)
  • Generating the adjacency matrix (Step 1) of the pySCENIC regulatory analysis pipeline
  • Comparing regulatory network structure across experimental conditions (e.g., control vs treatment)
  • Producing consensus regulatory networks by running inference across multiple random seeds
  • Validating GRN results by comparing GRNBoost2 and GENIE3 outputs on the same dataset
  • For downstream regulon identification and activity scoring, use arboreto output with pySCENIC
  • For single-cell preprocessing (QC, normalization, clustering) before GRN inference, use scanpy-scrna-seq

Prerequisites

  • Python packages: arboreto, pandas, numpy, dask, distributed, scikit-learn, scipy
  • Data requirements: Gene expression matrix (genes as columns, observations as rows) in TSV/CSV; optionally a TF list file (one gene name per line)
  • Environment: Python 3.8+; optional networkx, matplotlib for visualization
bash
pip install arboreto distributed networkx matplotlib

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: algorithm
    kind: required
    source: user
    ask: "Infer the network with gradient boosting (faster) or random forests (the original GENIE3 formulation)?"
    default: "GRNBoost2"

  - id: D2
    param: expressionMatrix
    kind: required
    source: upstream
    ask: "Which processed expression matrix, and which cells or samples within it, should the network be inferred from?"
    default: null

  - id: D3
    param: regulatorList
    kind: required
    source: literature
    ask: "Restrict candidate regulators to a curated transcription-factor list, or let every gene be a candidate regulator?"
    default: "a species-matched TF list"

  - id: D4
    param: edgeThreshold
    kind: required
    source: user
    ask: "The output ranks every regulator-target pair - how many of the top edges should be kept as the network?"
    default: null

  - id: D5
    param: seed
    kind: never_ask
    source: data
    reason: "Fixes the draw for reproducibility; does not change what the data supports"
    default: 0

  - id: D6
    param: daskClient
    kind: never_ask
    source: data
    reason: "Execution backend; affects runtime only"
    default: "local"

D3 is the decision that separates a usable network from a hairball. Leaving regulators unrestricted lets any correlated gene appear as a regulator, and the result still returns a fully populated ranked table. D4 has no safe default: the algorithm scores all pairs, so where the list is cut is the user's call about the network they intend to interpret.

Quick Start

Complete GRN inference in a single block. The if __name__ == '__main__': guard is required because Dask spawns worker processes via multiprocessing.

python
import pandas as pd
from arboreto.algo import grnboost2
from arboreto.utils import load_tf_names

if __name__ == '__main__':
    # Load expression matrix (observations x genes)
    expression_matrix = pd.read_csv('expression_data.tsv', sep='\t')
    tf_names = load_tf_names('tf_list.txt')  # optional TF filter

    # Infer GRN (uses all local cores by default)
    network = grnboost2(expression_data=expression_matrix,
                        tf_names=tf_names, seed=777)

    # Filter top links and save
    top_network = network[network['importance'] > 1.0]
    top_network.to_csv('grn_output.tsv', sep='\t', index=False, header=False)
    print(f"Inferred {len(network)} links, kept {len(top_network)} above threshold")
    # Example: Inferred 185432 links, kept 12876 above threshold

Workflow

Step 1: Load Expression Data

Arboreto accepts a pandas DataFrame (recommended) or NumPy array. Rows are observations (cells or samples), columns are genes. Gene names must be column headers.

python
import pandas as pd

# From TSV (genes as columns, observations as rows)
expression_matrix = pd.read_csv('expression_data.tsv', sep='\t')
print(f"Shape: {expression_matrix.shape}  "
      f"({expression_matrix.shape[0]} observations x {expression_matrix.shape[1]} genes)")
# Example: Shape: (5000, 18654)  (5000 observations x 18654 genes)

# From AnnData (e.g., after scanpy preprocessing)
import anndata as ad
adata = ad.read_h5ad('preprocessed.h5ad')
expression_matrix = pd.DataFrame(
    adata.X.toarray() if hasattr(adata.X, 'toarray') else adata.X,
    columns=adata.var_names.tolist()
)
print(f"Converted AnnData: {expression_matrix.shape}")
# Example: Converted AnnData: (5000, 18654)
Step 2: Load Transcription Factor List

Providing a TF list restricts regulators to known transcription factors, reducing computation time and improving biological relevance. If omitted, all genes are treated as potential regulators.

python
from arboreto.utils import load_tf_names

# From file (one TF name per line)
tf_names = load_tf_names('human_tfs.txt')
print(f"Loaded {len(tf_names)} transcription factors")
# Example: Loaded 1639 transcription factors

# Or define directly
tf_names = ['MYC', 'TP53', 'SOX2', 'NANOG', 'POU5F1']

# Verify TFs exist in expression matrix columns
tf_in_data = [tf for tf in tf_names if tf in expression_matrix.columns]
print(f"TFs found in expression data: {len(tf_in_data)}/{len(tf_names)}")
# Example: TFs found in expression data: 1583/1639
Step 3: Configure Dask Client (Optional)

By default arboreto creates an internal Dask client using all local cores. Create an explicit client for resource control, monitoring, or cluster deployment.

python
from distributed import LocalCluster, Client

# Custom local client with resource limits
local_cluster = LocalCluster(
    n_workers=8,
    threads_per_worker=1,   # avoid GIL contention in scikit-learn
    memory_limit='4GB'       # per worker
)
client = Client(local_cluster)
print(f"Dashboard: {client.dashboard_link}")
# Example: Dashboard: http://127.0.0.1:8787/status
Step 4: Run GRN Inference

Call grnboost2() (recommended) or genie3(). Both share the same signature and output format.

python
from arboreto.algo import grnboost2

if __name__ == '__main__':
    network = grnboost2(
        expression_data=expression_matrix,
        tf_names=tf_names,         # 'all' if no TF list
        client_or_address=client,  # omit to use default local scheduler
        seed=777,                  # for reproducibility
        verbose=True               # print progress
    )
    print(f"Inferred {len(network)} regulatory links")
    print(network.head())
    # Example:
    #       TF  target  importance
    # 0   MYC   CDK4       3.214
    # 1   MYC   CCND1      2.871
    # 2  TP53  CDKN1A      2.654
Step 5: Filter Results

Raw output contains links for every TF-target pair with non-zero importance. Filter to retain high-confidence regulatory interactions.

python
# Strategy 1: Importance threshold
threshold = 1.0
filtered = network[network['importance'] > threshold]
print(f"Threshold {threshold}: {len(filtered)} links "
      f"({len(filtered)/len(network)*100:.1f}% retained)")
# Example: Threshold 1.0: 12876 links (6.9% retained)

# Strategy 2: Top N links per target gene
top_n = 10
top_per_target = (network.groupby('target')
                  .apply(lambda g: g.nlargest(top_n, 'importance'))
                  .reset_index(drop=True))
print(f"Top {top_n} per target: {len(top_per_target)} links")
# Example: Top 10 per target: 186540 links
Step 6: Save Network

Save as a TSV file with three columns: TF, target, importance.

python
# Save full network
network.to_csv('full_network.tsv', sep='\t', index=False)
print(f"Saved full network: {len(network)} links")

# Save filtered network (without header for pySCENIC compatibility)
filtered.to_csv('filtered_network.tsv', sep='\t', index=False, header=False)
print(f"Saved filtered network: {len(filtered)} links")

# Clean up Dask client if explicitly created
client.close()
local_cluster.close()
Step 7: Visualize Results (Optional)

Plot the top regulatory interactions as a directed network graph.

python
import networkx as nx
import matplotlib.pyplot as plt

# Build directed graph from top links
top_links = network.nlargest(50, 'importance')
G = nx.from_pandas_edgelist(
    top_links, source='TF', target='target',
    edge_attr='importance', create_using=nx.DiGraph()
)

# Draw network
fig, ax = plt.subplots(figsize=(12, 10))
pos = nx.spring_layout(G, k=2, seed=42)
nx.draw_networkx_nodes(G, pos, node_size=300, node_color='lightblue', ax=ax)
nx.draw_networkx_labels(G, pos, font_size=7, ax=ax)
edges = nx.draw_networkx_edges(
    G, pos, edge_color=[G[u][v]['importance'] for u, v in G.edges()],
    edge_cmap=plt.cm.Reds, width=1.5, arrows=True, ax=ax
)
plt.colorbar(edges, ax=ax, label='Importance')
ax.set_title('Top 50 Regulatory Interactions')
plt.tight_layout()
plt.savefig('grn_network.png', dpi=300, bbox_inches='tight')
print("Saved grn_network.png")

Key Parameters

ParameterDefaultRange / OptionsEffect
expression_data(required)DataFrame or ndarrayExpression matrix, observations x genes
tf_names'all'list of str or 'all'Restrict regulators to known TFs; 'all' uses every gene
gene_namesNonelist of strRequired when expression_data is a NumPy array
client_or_address'local'Client, str, or 'local'Dask client instance or scheduler address
seedNoneintRandom seed for reproducibility; always set for replicable results
verboseFalseboolPrint progress messages during inference

Key Concepts

Algorithm Selection Guide

Both algorithms follow the same multiple-regression strategy (train one model per target gene, extract feature importances) but differ in the underlying regressor.

FeatureGRNBoost2GENIE3
MethodStochastic gradient boosting with early stoppingRandom Forest (or ExtraTrees)
SpeedFast -- optimized for large datasetsSlower -- higher per-model cost
MemoryLower (early stopping limits tree depth)Higher (full forests per target)
Best forDefault choice; 10k+ observations, single-cell scaleComparison with published GENIE3 results; validation
Importfrom arboreto.algo import grnboost2from arboreto.algo import genie3
Output formatTF, target, importance (identical)TF, target, importance (identical)

Decision rule: Start with GRNBoost2. Use GENIE3 only when reproducing published GENIE3 analyses or as an independent validation of GRNBoost2 results.

Output Format

The network DataFrame has three columns, sorted by descending importance:

ColumnTypeDescription
TFstrTranscription factor (regulator) gene name
targetstrTarget gene name
importancefloatRegulatory importance score (higher = stronger predicted regulation)

Importance scores are not p-values. They represent feature importance from the tree-based model. Filter by absolute threshold or top-N per target for downstream analysis.

The if __name__ == '__main__': Requirement

Dask uses Python's multiprocessing module to spawn worker processes. Without the main guard, each spawned process re-executes the script, leading to infinite process spawning. Always wrap arboreto calls in if __name__ == '__main__': when running as a script. This is not needed inside Jupyter notebooks.

Common Recipes

Recipe: Distributed Computing on a Cluster

When to use: datasets too large for a single machine (e.g., 50k+ cells, 20k+ genes). Requires a Dask distributed scheduler running on the cluster.

bash
# On head node: start scheduler
dask-scheduler
# Output: Scheduler at tcp://10.118.224.134:8786

# On each compute node: start workers
dask-worker tcp://10.118.224.134:8786 --nprocs 4 --nthreads 1 --memory-limit 16GB
python
from distributed import Client
from arboreto.algo import grnboost2
import pandas as pd

if __name__ == '__main__':
    client = Client('tcp://10.118.224.134:8786')
    print(f"Connected to cluster: {client.scheduler_info()['workers'].__len__()} workers")
    # Example: Connected to cluster: 16 workers

    expression_data = pd.read_csv('large_scrna.tsv', sep='\t')
    tf_names = pd.read_csv('tf_list.txt', header=None)[0].tolist()

    network = grnboost2(
        expression_data=expression_data,
        tf_names=tf_names,
        client_or_address=client,
        seed=42, verbose=True
    )
    network.to_csv('cluster_grn.tsv', sep='\t', index=False)
    print(f"Cluster inference complete: {len(network)} links")
    client.close()
Show full SKILL.md (550 more words)Show less
Recipe: Consensus Network from Multiple Seeds

When to use: improve robustness by aggregating networks from multiple random seeds. Links appearing consistently across runs are more reliable.

python
import pandas as pd
from distributed import LocalCluster, Client
from arboreto.algo import grnboost2

if __name__ == '__main__':
    client = Client(LocalCluster(n_workers=8, threads_per_worker=1))
    expression_data = pd.read_csv('expression_data.tsv', sep='\t')
    tf_names = pd.read_csv('tf_list.txt', header=None)[0].tolist()

    seeds = [42, 123, 456, 789, 1001]
    all_networks = []

    for seed in seeds:
        net = grnboost2(expression_data=expression_data,
                        tf_names=tf_names,
                        client_or_address=client, seed=seed)
        net['seed'] = seed
        all_networks.append(net)
        print(f"Seed {seed}: {len(net)} links")

    combined = pd.concat(all_networks, ignore_index=True)

    # Consensus: mean importance across seeds, keep links found in 3+ runs
    consensus = (combined.groupby(['TF', 'target'])
                 .agg(mean_importance=('importance', 'mean'),
                      n_seeds=('seed', 'nunique'))
                 .reset_index())
    consensus = consensus[consensus['n_seeds'] >= 3]
    consensus = consensus.sort_values('mean_importance', ascending=False)
    consensus.to_csv('consensus_network.tsv', sep='\t', index=False)
    print(f"Consensus network: {len(consensus)} links (found in 3+ of {len(seeds)} runs)")
    # Example: Consensus network: 28453 links (found in 3+ of 5 runs)

    client.close()
Recipe: pySCENIC Integration

When to use: downstream regulatory analysis after arboreto GRN inference. Arboreto produces the adjacency matrix (Step 1 of SCENIC); pySCENIC performs regulon identification (Step 2) and activity scoring (Step 3).

python
from arboreto.algo import grnboost2
from arboreto.utils import load_tf_names
import pandas as pd

if __name__ == '__main__':
    # Step 1 (arboreto): GRN inference
    expression_matrix = pd.read_csv('scrna_expression.tsv', sep='\t')
    tf_names = load_tf_names('allTFs_hg38.txt')

    adjacencies = grnboost2(
        expression_data=expression_matrix,
        tf_names=tf_names, seed=777
    )
    adjacencies.to_csv('adjacencies.tsv', sep='\t', index=False, header=False)
    print(f"SCENIC Step 1 complete: {len(adjacencies)} adjacencies")
    # Example: SCENIC Step 1 complete: 234567 adjacencies

    # Steps 2-3 use pySCENIC CLI (requires pyscenic installation):
    # pyscenic ctx adjacencies.tsv hg38_ranking.feather \
    #     --annotations_fname motifs-v10nr.hgnc-m0.001-o0.0.tbl \
    #     --output regulons.csv
    #
    # pyscenic aucell scrna_expression.tsv regulons.csv \
    #     --output auc_matrix.csv
    print("Next: run pyscenic ctx (Step 2) and pyscenic aucell (Step 3)")

Expected Outputs

  • full_network.tsv -- complete TF-target-importance table; columns: TF, target, importance
  • filtered_network.tsv -- high-confidence subset after threshold or top-N filtering
  • consensus_network.tsv -- aggregated network from multi-seed runs; columns: TF, target, mean_importance, n_seeds
  • adjacencies.tsv -- arboreto output formatted for pySCENIC input (no header)
  • grn_network.png -- directed graph visualization of top regulatory interactions

Troubleshooting

ProblemCauseSolution
RuntimeError: freeze_support() or infinite process spawningMissing if __name__ == '__main__': guardWrap all arboreto calls in main guard; not needed in Jupyter
MemoryError during inferenceExpression matrix too large for available RAMFilter low-variance genes first, reduce to top 5-10k genes, or use distributed cluster
Empty or near-empty networkTF names do not match expression matrix column namesVerify TF names overlap with expression_matrix.columns; check case sensitivity
Very slow inferenceUsing GENIE3 on large dataset, or too few Dask workersSwitch to GRNBoost2; create explicit Dask client with more workers
ModuleNotFoundError: No module named 'arboreto'Package not installedpip install arboreto
All importance scores are 0 or identicalExpression matrix has no variance (e.g., all zeros after filtering)Check data quality; ensure matrix contains non-zero expression values with variance
TypeError with NumPy array inputMissing gene_names parameterProvide gene_names= list matching the column count of the array

Bundled Resources

This entry is fully self-contained with no references/ directory.

Original reference file dispositions:

  • references/algorithms.md (139 lines) -- Algorithm comparison, parameter details, custom regressor kwargs. Consolidated into Key Concepts (Algorithm Selection Guide table) and Key Parameters. Custom regressor kwargs (regressor_type, regressor_kwargs) omitted: advanced tuning rarely needed; users can pass scikit-learn kwargs directly per the arboreto API docs.
  • references/basic_inference.md (152 lines) -- Input formats, TF loading, basic workflow, output format. Fully consolidated into Workflow Steps 1-2, Key Concepts (Output Format table), and Quick Start.
  • references/distributed_computing.md (242 lines) -- Local client, cluster setup, monitoring, performance tips. Consolidated into Workflow Step 3 and Recipe: Distributed Computing. Omitted: Dask dashboard panel descriptions (monitoring UI detail beyond scope), detailed network/storage tuning (cluster-admin concern).
  • scripts/basic_grn_inference.py (98 lines) -- CLI wrapper with argparse. Core load-infer-save logic absorbed into Workflow Steps 1-6 and Quick Start. Argparse/CLI boilerplate stripped.

Narrative use-case dispositions (from original Common Use Cases):

  • Single-cell RNA-seq analysis -- absorbed into When to Use bullets and Workflow Steps 1/4
  • Bulk RNA-seq with TF filtering -- absorbed into Workflow Step 2 and Quick Start
  • Comparative analysis (multiple conditions) -- omitted as standalone recipe: trivial loop over conditions, pattern shown in Consensus Network recipe
  • scanpy-scrna-seq -- upstream single-cell preprocessing (QC, normalization, HVG selection) before GRN inference
  • scvi-tools-single-cell -- deep generative models for batch correction and imputation prior to network inference
  • pyscenic (planned) -- downstream regulon identification and cellular activity scoring using arboreto adjacencies
  • networkx-graph-analysis -- graph-theoretic analysis of inferred regulatory networks (centrality, communities)

References

  • Arboreto GitHub repository -- source code, README, examples
  • Arboreto PyPI -- installation and version history
  • Moerman et al. (2019) "GRNBoost2 and Arboreto: efficient and scalable inference of gene regulatory networks" Bioinformatics 35(12):2159-2161 -- algorithm paper
  • Aibar et al. (2017) "SCENIC: single-cell regulatory network inference and clustering" Nature Methods 14:1083-1086 -- SCENIC pipeline paper
  • Dask distributed documentation -- Dask client, cluster setup, dashboard

© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/genomics-bioinformatics/arboreto-grn-inference of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Arboreto Grn Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Arboreto Grn Inference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Arboreto Grn Inference this skilljaechang-hits/SciAgent-Skills3742 repos~5.3kAutomated safety check: PassBSD-3-Clause
ArboretoK-Dense-AI/scientific-agent-skills48k1 repos~2.7kAutomated safety check: PassBSD-3-Clause
Bio Expression Matrix Sparse HandlingGPTomics/bioSkills1.2k1 repos~5.6kAutomated safety check: PassMIT
Gtars Genomic Interval Toolkitdavila7/claude-code-templates33k11 repos~1.9kAutomated safety check: PassMIT
Bulkrna Batch CorrectionTianGzlab/OmicsClaw161—~1.2kAutomated safety check: PassApache-2.0
Bulkrna CoexpressionTianGzlab/OmicsClaw161—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Arboreto

    K-Dense-AI/scientific-agent-skills

    Infers candidate gene regulatory networks from bulk or single-cell expression data using AertsLab Arboreto GRNBoost2 and GENIE3.

    48k GitHub starsUsed in 1 repo~2.7k tokens
    Research & ScienceAuto-check passed
  • Stores and operates on sparse expression matrices for single-cell and large bulk RNA-seq, covering dgCMatrix/dgRMatrix/dgTMatrix when-each-is-fast, the dgCMatrix (CSC, R) <- CSR (Python) implicit…

    1.2k GitHub starsUsed in 1 repo~5.6k tokens
    Research & ScienceAuto-check passed
  • Gtars Genomic Interval Toolkit

    davila7/claude-code-templates

    Works with genomic intervals using gtars, a Rust toolkit with Python bindings: overlap detection, coverage tracks, tokenization for ML models and reference sequences.

    33k GitHub starsUsed in 11 repos~1.9k tokens
    Research & ScienceAuto-check passed
  • Bulkrna Batch Correction

    TianGzlab/OmicsClaw

    Load when correcting batch effects in bulk expression using R sva ComBat or the legacy Python parametric approximation.

    161 GitHub stars~1.2k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Bulkrna Coexpression

    TianGzlab/OmicsClaw

    Load when discovering bulk gene co-expression modules and hub genes with R WGCNA.

    161 GitHub stars~1.3k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Bio Spatial Transcriptomics Spatial Preprocessing

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quality control, filtering, normalization, and feature selection for spatial transcriptomics data.

    3.1k GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Works with

Questions about Arboreto Grn Inference

What does Arboreto Grn Inference do?

GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest). Arboreto Grn Inference is an agent skill from jaechang-hits/SciAgent-Skills. GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest).

When should I use Arboreto Grn Inference?

Arboreto Grn Inference fits situations like: tasks that involve Bioinformatics; tasks that involve DataFrames; tasks that involve Machine learning.

How do I install Arboreto Grn Inference in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/arboreto-grn-inference in jaechang-hits/SciAgent-Skills) into .claude/skills/arboreto-grn-inference in your project. Claude Code loads it when a task matches its description.

How do I install Arboreto Grn Inference in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/arboreto-grn-inference in jaechang-hits/SciAgent-Skills) into .agents/skills/arboreto-grn-inference in your project. Codex loads it when a task matches its description.

Can I use Arboreto Grn Inference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/arboreto-grn-inference, .gemini/skills/arboreto-grn-inference, .github/skills/arboreto-grn-inference and .opencode/skills/arboreto-grn-inference in your project.

What does Arboreto Grn Inference need to run?

Going by SKILL.md and its folder, Arboreto Grn Inference needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Arboreto Grn Inference access the network?

SKILL.md names 3 domains. As links in the text: github.com, pypi.org and distributed.dask.org. This is read from the text; nothing was executed.

Is Arboreto Grn Inference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Arboreto Grn Inference use?

Arboreto Grn Inference is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Arboreto Grn Inference use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Arboreto Grn Inference?

Skills that share tags, products or a category with Arboreto Grn Inference: Arboreto (K-Dense-AI/scientific-agent-skills, 48k stars), Bio Expression Matrix Sparse Handling (GPTomics/bioSkills, 1.2k stars), Gtars Genomic Interval Toolkit (davila7/claude-code-templates, 33k stars) and Bulkrna Batch Correction (TianGzlab/OmicsClaw, 161 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Arboreto Grn Inference?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.