Arboreto
K-Dense-AI/scientific-agent-skills
Infers candidate gene regulatory networks from bulk or single-cell expression data using AertsLab Arboreto GRNBoost2 and GENIE3.
GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest).
$ npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills arboreto-grn-inference --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/arboreto-grn-inference .claude/skills/arboreto-grn-inference && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "arboreto-grn-inference" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/arboreto-grn-inference into .claude/skills/arboreto-grn-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "arboreto-grn-inference", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/arboreto-grn-inferenceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills arboreto-grn-inference --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/arboreto-grn-inference .agents/skills/arboreto-grn-inference && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "arboreto-grn-inference" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/arboreto-grn-inference into .agents/skills/arboreto-grn-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "arboreto-grn-inference", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills arboreto-grn-inference --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/arboreto-grn-inference .cursor/skills/arboreto-grn-inference && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "arboreto-grn-inference" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/arboreto-grn-inference into .cursor/skills/arboreto-grn-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "arboreto-grn-inference", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/arboreto-grn-inference--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills arboreto-grn-inference --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/arboreto-grn-inference .gemini/skills/arboreto-grn-inference && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "arboreto-grn-inference" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/arboreto-grn-inference into .gemini/skills/arboreto-grn-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "arboreto-grn-inference", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills arboreto-grn-inferenceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/arboreto-grn-inference .github/skills/arboreto-grn-inference && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "arboreto-grn-inference" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/arboreto-grn-inference into .github/skills/arboreto-grn-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "arboreto-grn-inference", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills arboreto-grn-inference --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/arboreto-grn-inference .opencode/skills/arboreto-grn-inference && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "arboreto-grn-inference" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/arboreto-grn-inference into .opencode/skills/arboreto-grn-inference/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "arboreto-grn-inference", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
arboreto-grn-inferenceGRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest).
Arboreto Grn Inference is an agent skill from jaechang-hits/SciAgent-Skills. GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest). Load matrix, filter by TFs, infer TF-target-importance links, save network. Dask-parallelized to single-cell scale. Core SCENIC component.
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics, DataFrames and Machine learning. It works with Dask. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.compypi.orgdistributed.dask.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Arboreto Grn Inference loads about 5.3k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 1,350 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 1,350 words, ~5,283 tokens.
.claude/skills/arboreto-grn-inference/SKILL.md (or your agent's skills folder).Arboreto infers gene regulatory networks (GRNs) from gene expression data using parallelized tree-based regression. For each target gene, it trains a regression model with all other genes (or a specified TF list) as features and emits TF-target-importance triplets. It provides two interchangeable algorithms -- GRNBoost2 (gradient boosting, fast) and GENIE3 (Random Forest, classic) -- sharing identical input/output formats. Computation is Dask-parallelized, scaling from laptop cores to HPC clusters.
arboreto, pandas, numpy, dask, distributed, scikit-learn, scipynetworkx, matplotlib for visualizationpip install arboreto distributed networkx matplotlibSettle these with the user before writing any analysis code.
decisions:
- id: D1
param: algorithm
kind: required
source: user
ask: "Infer the network with gradient boosting (faster) or random forests (the original GENIE3 formulation)?"
default: "GRNBoost2"
- id: D2
param: expressionMatrix
kind: required
source: upstream
ask: "Which processed expression matrix, and which cells or samples within it, should the network be inferred from?"
default: null
- id: D3
param: regulatorList
kind: required
source: literature
ask: "Restrict candidate regulators to a curated transcription-factor list, or let every gene be a candidate regulator?"
default: "a species-matched TF list"
- id: D4
param: edgeThreshold
kind: required
source: user
ask: "The output ranks every regulator-target pair - how many of the top edges should be kept as the network?"
default: null
- id: D5
param: seed
kind: never_ask
source: data
reason: "Fixes the draw for reproducibility; does not change what the data supports"
default: 0
- id: D6
param: daskClient
kind: never_ask
source: data
reason: "Execution backend; affects runtime only"
default: "local"D3 is the decision that separates a usable network from a hairball. Leaving regulators unrestricted lets any correlated gene appear as a regulator, and the result still returns a fully populated ranked table. D4 has no safe default: the algorithm scores all pairs, so where the list is cut is the user's call about the network they intend to interpret.
Complete GRN inference in a single block. The if __name__ == '__main__': guard is required because Dask spawns worker processes via multiprocessing.
import pandas as pd
from arboreto.algo import grnboost2
from arboreto.utils import load_tf_names
if __name__ == '__main__':
# Load expression matrix (observations x genes)
expression_matrix = pd.read_csv('expression_data.tsv', sep='\t')
tf_names = load_tf_names('tf_list.txt') # optional TF filter
# Infer GRN (uses all local cores by default)
network = grnboost2(expression_data=expression_matrix,
tf_names=tf_names, seed=777)
# Filter top links and save
top_network = network[network['importance'] > 1.0]
top_network.to_csv('grn_output.tsv', sep='\t', index=False, header=False)
print(f"Inferred {len(network)} links, kept {len(top_network)} above threshold")
# Example: Inferred 185432 links, kept 12876 above thresholdArboreto accepts a pandas DataFrame (recommended) or NumPy array. Rows are observations (cells or samples), columns are genes. Gene names must be column headers.
import pandas as pd
# From TSV (genes as columns, observations as rows)
expression_matrix = pd.read_csv('expression_data.tsv', sep='\t')
print(f"Shape: {expression_matrix.shape} "
f"({expression_matrix.shape[0]} observations x {expression_matrix.shape[1]} genes)")
# Example: Shape: (5000, 18654) (5000 observations x 18654 genes)
# From AnnData (e.g., after scanpy preprocessing)
import anndata as ad
adata = ad.read_h5ad('preprocessed.h5ad')
expression_matrix = pd.DataFrame(
adata.X.toarray() if hasattr(adata.X, 'toarray') else adata.X,
columns=adata.var_names.tolist()
)
print(f"Converted AnnData: {expression_matrix.shape}")
# Example: Converted AnnData: (5000, 18654)Providing a TF list restricts regulators to known transcription factors, reducing computation time and improving biological relevance. If omitted, all genes are treated as potential regulators.
from arboreto.utils import load_tf_names
# From file (one TF name per line)
tf_names = load_tf_names('human_tfs.txt')
print(f"Loaded {len(tf_names)} transcription factors")
# Example: Loaded 1639 transcription factors
# Or define directly
tf_names = ['MYC', 'TP53', 'SOX2', 'NANOG', 'POU5F1']
# Verify TFs exist in expression matrix columns
tf_in_data = [tf for tf in tf_names if tf in expression_matrix.columns]
print(f"TFs found in expression data: {len(tf_in_data)}/{len(tf_names)}")
# Example: TFs found in expression data: 1583/1639By default arboreto creates an internal Dask client using all local cores. Create an explicit client for resource control, monitoring, or cluster deployment.
from distributed import LocalCluster, Client
# Custom local client with resource limits
local_cluster = LocalCluster(
n_workers=8,
threads_per_worker=1, # avoid GIL contention in scikit-learn
memory_limit='4GB' # per worker
)
client = Client(local_cluster)
print(f"Dashboard: {client.dashboard_link}")
# Example: Dashboard: http://127.0.0.1:8787/statusCall grnboost2() (recommended) or genie3(). Both share the same signature and output format.
from arboreto.algo import grnboost2
if __name__ == '__main__':
network = grnboost2(
expression_data=expression_matrix,
tf_names=tf_names, # 'all' if no TF list
client_or_address=client, # omit to use default local scheduler
seed=777, # for reproducibility
verbose=True # print progress
)
print(f"Inferred {len(network)} regulatory links")
print(network.head())
# Example:
# TF target importance
# 0 MYC CDK4 3.214
# 1 MYC CCND1 2.871
# 2 TP53 CDKN1A 2.654Raw output contains links for every TF-target pair with non-zero importance. Filter to retain high-confidence regulatory interactions.
# Strategy 1: Importance threshold
threshold = 1.0
filtered = network[network['importance'] > threshold]
print(f"Threshold {threshold}: {len(filtered)} links "
f"({len(filtered)/len(network)*100:.1f}% retained)")
# Example: Threshold 1.0: 12876 links (6.9% retained)
# Strategy 2: Top N links per target gene
top_n = 10
top_per_target = (network.groupby('target')
.apply(lambda g: g.nlargest(top_n, 'importance'))
.reset_index(drop=True))
print(f"Top {top_n} per target: {len(top_per_target)} links")
# Example: Top 10 per target: 186540 linksSave as a TSV file with three columns: TF, target, importance.
# Save full network
network.to_csv('full_network.tsv', sep='\t', index=False)
print(f"Saved full network: {len(network)} links")
# Save filtered network (without header for pySCENIC compatibility)
filtered.to_csv('filtered_network.tsv', sep='\t', index=False, header=False)
print(f"Saved filtered network: {len(filtered)} links")
# Clean up Dask client if explicitly created
client.close()
local_cluster.close()Plot the top regulatory interactions as a directed network graph.
import networkx as nx
import matplotlib.pyplot as plt
# Build directed graph from top links
top_links = network.nlargest(50, 'importance')
G = nx.from_pandas_edgelist(
top_links, source='TF', target='target',
edge_attr='importance', create_using=nx.DiGraph()
)
# Draw network
fig, ax = plt.subplots(figsize=(12, 10))
pos = nx.spring_layout(G, k=2, seed=42)
nx.draw_networkx_nodes(G, pos, node_size=300, node_color='lightblue', ax=ax)
nx.draw_networkx_labels(G, pos, font_size=7, ax=ax)
edges = nx.draw_networkx_edges(
G, pos, edge_color=[G[u][v]['importance'] for u, v in G.edges()],
edge_cmap=plt.cm.Reds, width=1.5, arrows=True, ax=ax
)
plt.colorbar(edges, ax=ax, label='Importance')
ax.set_title('Top 50 Regulatory Interactions')
plt.tight_layout()
plt.savefig('grn_network.png', dpi=300, bbox_inches='tight')
print("Saved grn_network.png")| Parameter | Default | Range / Options | Effect |
|---|---|---|---|
expression_data | (required) | DataFrame or ndarray | Expression matrix, observations x genes |
tf_names | 'all' | list of str or 'all' | Restrict regulators to known TFs; 'all' uses every gene |
gene_names | None | list of str | Required when expression_data is a NumPy array |
client_or_address | 'local' | Client, str, or 'local' | Dask client instance or scheduler address |
seed | None | int | Random seed for reproducibility; always set for replicable results |
verbose | False | bool | Print progress messages during inference |
Both algorithms follow the same multiple-regression strategy (train one model per target gene, extract feature importances) but differ in the underlying regressor.
| Feature | GRNBoost2 | GENIE3 |
|---|---|---|
| Method | Stochastic gradient boosting with early stopping | Random Forest (or ExtraTrees) |
| Speed | Fast -- optimized for large datasets | Slower -- higher per-model cost |
| Memory | Lower (early stopping limits tree depth) | Higher (full forests per target) |
| Best for | Default choice; 10k+ observations, single-cell scale | Comparison with published GENIE3 results; validation |
| Import | from arboreto.algo import grnboost2 | from arboreto.algo import genie3 |
| Output format | TF, target, importance (identical) | TF, target, importance (identical) |
Decision rule: Start with GRNBoost2. Use GENIE3 only when reproducing published GENIE3 analyses or as an independent validation of GRNBoost2 results.
The network DataFrame has three columns, sorted by descending importance:
| Column | Type | Description |
|---|---|---|
TF | str | Transcription factor (regulator) gene name |
target | str | Target gene name |
importance | float | Regulatory importance score (higher = stronger predicted regulation) |
Importance scores are not p-values. They represent feature importance from the tree-based model. Filter by absolute threshold or top-N per target for downstream analysis.
if __name__ == '__main__': RequirementDask uses Python's multiprocessing module to spawn worker processes. Without the main guard, each spawned process re-executes the script, leading to infinite process spawning. Always wrap arboreto calls in if __name__ == '__main__': when running as a script. This is not needed inside Jupyter notebooks.
When to use: datasets too large for a single machine (e.g., 50k+ cells, 20k+ genes). Requires a Dask distributed scheduler running on the cluster.
# On head node: start scheduler
dask-scheduler
# Output: Scheduler at tcp://10.118.224.134:8786
# On each compute node: start workers
dask-worker tcp://10.118.224.134:8786 --nprocs 4 --nthreads 1 --memory-limit 16GBfrom distributed import Client
from arboreto.algo import grnboost2
import pandas as pd
if __name__ == '__main__':
client = Client('tcp://10.118.224.134:8786')
print(f"Connected to cluster: {client.scheduler_info()['workers'].__len__()} workers")
# Example: Connected to cluster: 16 workers
expression_data = pd.read_csv('large_scrna.tsv', sep='\t')
tf_names = pd.read_csv('tf_list.txt', header=None)[0].tolist()
network = grnboost2(
expression_data=expression_data,
tf_names=tf_names,
client_or_address=client,
seed=42, verbose=True
)
network.to_csv('cluster_grn.tsv', sep='\t', index=False)
print(f"Cluster inference complete: {len(network)} links")
client.close()When to use: improve robustness by aggregating networks from multiple random seeds. Links appearing consistently across runs are more reliable.
import pandas as pd
from distributed import LocalCluster, Client
from arboreto.algo import grnboost2
if __name__ == '__main__':
client = Client(LocalCluster(n_workers=8, threads_per_worker=1))
expression_data = pd.read_csv('expression_data.tsv', sep='\t')
tf_names = pd.read_csv('tf_list.txt', header=None)[0].tolist()
seeds = [42, 123, 456, 789, 1001]
all_networks = []
for seed in seeds:
net = grnboost2(expression_data=expression_data,
tf_names=tf_names,
client_or_address=client, seed=seed)
net['seed'] = seed
all_networks.append(net)
print(f"Seed {seed}: {len(net)} links")
combined = pd.concat(all_networks, ignore_index=True)
# Consensus: mean importance across seeds, keep links found in 3+ runs
consensus = (combined.groupby(['TF', 'target'])
.agg(mean_importance=('importance', 'mean'),
n_seeds=('seed', 'nunique'))
.reset_index())
consensus = consensus[consensus['n_seeds'] >= 3]
consensus = consensus.sort_values('mean_importance', ascending=False)
consensus.to_csv('consensus_network.tsv', sep='\t', index=False)
print(f"Consensus network: {len(consensus)} links (found in 3+ of {len(seeds)} runs)")
# Example: Consensus network: 28453 links (found in 3+ of 5 runs)
client.close()When to use: downstream regulatory analysis after arboreto GRN inference. Arboreto produces the adjacency matrix (Step 1 of SCENIC); pySCENIC performs regulon identification (Step 2) and activity scoring (Step 3).
from arboreto.algo import grnboost2
from arboreto.utils import load_tf_names
import pandas as pd
if __name__ == '__main__':
# Step 1 (arboreto): GRN inference
expression_matrix = pd.read_csv('scrna_expression.tsv', sep='\t')
tf_names = load_tf_names('allTFs_hg38.txt')
adjacencies = grnboost2(
expression_data=expression_matrix,
tf_names=tf_names, seed=777
)
adjacencies.to_csv('adjacencies.tsv', sep='\t', index=False, header=False)
print(f"SCENIC Step 1 complete: {len(adjacencies)} adjacencies")
# Example: SCENIC Step 1 complete: 234567 adjacencies
# Steps 2-3 use pySCENIC CLI (requires pyscenic installation):
# pyscenic ctx adjacencies.tsv hg38_ranking.feather \
# --annotations_fname motifs-v10nr.hgnc-m0.001-o0.0.tbl \
# --output regulons.csv
#
# pyscenic aucell scrna_expression.tsv regulons.csv \
# --output auc_matrix.csv
print("Next: run pyscenic ctx (Step 2) and pyscenic aucell (Step 3)")full_network.tsv -- complete TF-target-importance table; columns: TF, target, importancefiltered_network.tsv -- high-confidence subset after threshold or top-N filteringconsensus_network.tsv -- aggregated network from multi-seed runs; columns: TF, target, mean_importance, n_seedsadjacencies.tsv -- arboreto output formatted for pySCENIC input (no header)grn_network.png -- directed graph visualization of top regulatory interactions| Problem | Cause | Solution |
|---|---|---|
RuntimeError: freeze_support() or infinite process spawning | Missing if __name__ == '__main__': guard | Wrap all arboreto calls in main guard; not needed in Jupyter |
MemoryError during inference | Expression matrix too large for available RAM | Filter low-variance genes first, reduce to top 5-10k genes, or use distributed cluster |
| Empty or near-empty network | TF names do not match expression matrix column names | Verify TF names overlap with expression_matrix.columns; check case sensitivity |
| Very slow inference | Using GENIE3 on large dataset, or too few Dask workers | Switch to GRNBoost2; create explicit Dask client with more workers |
ModuleNotFoundError: No module named 'arboreto' | Package not installed | pip install arboreto |
| All importance scores are 0 or identical | Expression matrix has no variance (e.g., all zeros after filtering) | Check data quality; ensure matrix contains non-zero expression values with variance |
TypeError with NumPy array input | Missing gene_names parameter | Provide gene_names= list matching the column count of the array |
This entry is fully self-contained with no references/ directory.
Original reference file dispositions:
regressor_type, regressor_kwargs) omitted: advanced tuning rarely needed; users can pass scikit-learn kwargs directly per the arboreto API docs.Narrative use-case dispositions (from original Common Use Cases):
© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/arboreto-grn-inference of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Arboreto Grn Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Arboreto Grn Inference this skilljaechang-hits/SciAgent-Skills | 374 | 2 repos | ~5.3k | Automated safety check: Pass | BSD-3-Clause | |
| ArboretoK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~2.7k | Automated safety check: Pass | BSD-3-Clause | |
| Bio Expression Matrix Sparse HandlingGPTomics/bioSkills | 1.2k | 1 repos | ~5.6k | Automated safety check: Pass | MIT | |
| Gtars Genomic Interval Toolkitdavila7/claude-code-templates | 33k | 11 repos | ~1.9k | Automated safety check: Pass | MIT | |
| Bulkrna Batch CorrectionTianGzlab/OmicsClaw | 161 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Bulkrna CoexpressionTianGzlab/OmicsClaw | 161 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 |
K-Dense-AI/scientific-agent-skills
Infers candidate gene regulatory networks from bulk or single-cell expression data using AertsLab Arboreto GRNBoost2 and GENIE3.
GPTomics/bioSkills
Stores and operates on sparse expression matrices for single-cell and large bulk RNA-seq, covering dgCMatrix/dgRMatrix/dgTMatrix when-each-is-fast, the dgCMatrix (CSC, R) <- CSR (Python) implicit…
davila7/claude-code-templates
Works with genomic intervals using gtars, a Rust toolkit with Python bindings: overlap detection, coverage tracks, tokenization for ML models and reference sequences.
TianGzlab/OmicsClaw
Load when correcting batch effects in bulk expression using R sva ComBat or the legacy Python parametric approximation.
TianGzlab/OmicsClaw
Load when discovering bulk gene co-expression modules and hub genes with R WGCNA.
FreedomIntelligence/OpenClaw-Medical-Skills
Quality control, filtering, normalization, and feature selection for spatial transcriptomics data.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Works with
Categories
GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest). Arboreto Grn Inference is an agent skill from jaechang-hits/SciAgent-Skills. GRN inference from expression via GRNBoost2 (gradient boosting) or GENIE3 (Random Forest).
Arboreto Grn Inference fits situations like: tasks that involve Bioinformatics; tasks that involve DataFrames; tasks that involve Machine learning.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/arboreto-grn-inference in jaechang-hits/SciAgent-Skills) into .claude/skills/arboreto-grn-inference in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/arboreto-grn-inference in jaechang-hits/SciAgent-Skills) into .agents/skills/arboreto-grn-inference in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill arboreto-grn-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/arboreto-grn-inference, .gemini/skills/arboreto-grn-inference, .github/skills/arboreto-grn-inference and .opencode/skills/arboreto-grn-inference in your project.
Going by SKILL.md and its folder, Arboreto Grn Inference needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 3 domains. As links in the text: github.com, pypi.org and distributed.dask.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Arboreto Grn Inference is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Arboreto Grn Inference: Arboreto (K-Dense-AI/scientific-agent-skills, 48k stars), Bio Expression Matrix Sparse Handling (GPTomics/bioSkills, 1.2k stars), Gtars Genomic Interval Toolkit (davila7/claude-code-templates, 33k stars) and Bulkrna Batch Correction (TianGzlab/OmicsClaw, 161 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.