Bulkrna Geneid Mapping
TianGzlab/OmicsClaw
Load when converting Ensembl, Entrez or symbol IDs in a bulk RNA count matrix using an explicit mapping or a small human demo reference.
NCBI Gene via E-utilities: curated records across 1M+ taxa. An agent skill from jaechang-hits/SciAgent-Skills.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gene-database -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gene-database --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gene-database .claude/skills/gene-database && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gene-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gene-database into .claude/skills/gene-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gene-database", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gene-databaseType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gene-database -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gene-database --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gene-database .agents/skills/gene-database && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gene-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gene-database into .agents/skills/gene-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gene-database", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gene-database -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gene-database --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gene-database .cursor/skills/gene-database && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gene-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gene-database into .cursor/skills/gene-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gene-database", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/databases/gene-database--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gene-database -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gene-database --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gene-database .gemini/skills/gene-database && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gene-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gene-database into .gemini/skills/gene-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gene-database", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills gene-databaseInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill gene-database -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gene-database .github/skills/gene-database && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gene-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gene-database into .github/skills/gene-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gene-database", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill gene-database -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills gene-database --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/databases/gene-database .opencode/skills/gene-database && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gene-database" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/databases/gene-database into .opencode/skills/gene-database/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gene-database", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gene-databaseNCBI Gene via E-utilities: curated records across 1M+ taxa. An agent skill from jaechang-hits/SciAgent-Skills.
Gene Database is an agent skill from jaechang-hits/SciAgent-Skills. NCBI Gene via E-utilities: curated records across 1M+ taxa. Official symbols, aliases, RefSeq IDs, summaries, coordinates, GO, interactions. Use for gene ID resolution and cross-species function queries. For sequences use Ensembl; for expression use geo-database.
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. It works with NCBI and Ensembl. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC0-1.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
eutils.ncbi.nlm.nih.govapi.ncbi.nlm.nih.govAlso links to:
ncbi.nlm.nih.govFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Gene Database loads about 4.4k tokens when it runs. Until then it costs about 69 tokens; SKILL.md has 902 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC0-1.0 licence (© jaechang-hits). 902 words, ~4,428 tokens.
.claude/skills/gene-database/SKILL.md (or your agent's skills folder).NCBI Gene is the authoritative curated database for gene-centric information, covering 1M+ genes across hundreds of thousands of taxa. Each gene record includes the official symbol, aliases, full name, functional summary, genomic coordinates (GRCh38/GRCh37), RefSeq accessions, GO annotations, interaction partners, and links to related databases. Access is free via E-utilities REST API (no API key required, though recommended).
gene_gene_homolog retired with HomoloGene in 2019)geo-database; for variant annotations use clinvar-database or ensembl-databaserequests, xml.etree.ElementTree (stdlib), pandas (optional)email parameter)pip install requests pandasimport requests
EMAIL = "your@email.com"
BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
def gene_search(query, retmax=5):
r = requests.get(f"{BASE}/esearch.fcgi",
params={"db": "gene", "term": query,
"retmax": retmax, "retmode": "json", "email": EMAIL})
r.raise_for_status()
return r.json()["esearchresult"]["idlist"]
# Find human BRCA1 gene ID
ids = gene_search("BRCA1[sym] AND Homo sapiens[orgn]")
print(f"Gene IDs for BRCA1: {ids}") # → ['672']Use ESearch with field tags for precise queries.
import requests
EMAIL = "your@email.com"
BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
# Exact symbol match for human gene
r = requests.get(f"{BASE}/esearch.fcgi",
params={"db": "gene", "email": EMAIL, "retmode": "json",
"term": "TP53[sym] AND Homo sapiens[orgn] AND alive[prop]"})
ids = r.json()["esearchresult"]["idlist"]
print(f"TP53 Gene ID: {ids}") # → ['7157']# Search by function keyword
r = requests.get(f"{BASE}/esearch.fcgi",
params={"db": "gene", "email": EMAIL, "retmode": "json",
"term": "CRISPR[title] AND Homo sapiens[orgn]", "retmax": 5})
ids = r.json()["esearchresult"]["idlist"]
print(f"CRISPR-related gene IDs: {ids}")Retrieve key metadata fields for a list of Gene IDs.
import requests
EMAIL = "your@email.com"
BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
def esummary_gene(gene_ids):
r = requests.post(f"{BASE}/esummary.fcgi",
data={"db": "gene", "id": ",".join(gene_ids),
"retmode": "json", "email": EMAIL})
r.raise_for_status()
return r.json()["result"]
result = esummary_gene(["672", "675", "7157"]) # BRCA1, BRCA2, TP53
for uid in result.get("uids", []):
g = result[uid]
print(f"\n{g.get('name')} (ID {uid})")
print(f" Official symbol : {g.get('nomenclaturesymbol', g.get('name'))}")
print(f" Chr location : {g.get('maplocation')}")
print(f" Summary (first 100): {g.get('summary', '')[:100]}...")
print(f" Aliases: {g.get('otheraliases', 'none')}")Retrieve the complete gene record in XML for RefSeq accessions, GO terms, and interaction data.
import requests
import xml.etree.ElementTree as ET
EMAIL = "your@email.com"
BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
def efetch_gene_xml(gene_id):
r = requests.get(f"{BASE}/efetch.fcgi",
params={"db": "gene", "id": gene_id,
"rettype": "gene_table", "retmode": "text", "email": EMAIL})
r.raise_for_status()
return r.text
# Get gene table (tab-delimited overview)
table = efetch_gene_xml("672")
print(table[:500])# XML for RefSeq accession extraction
r = requests.get(f"{BASE}/efetch.fcgi",
params={"db": "gene", "id": "672",
"rettype": "xml", "retmode": "xml", "email": EMAIL})
root = ET.fromstring(r.text)
# Extract RefSeq mRNA accessions
for ref in root.iter("Gene-commentary"):
acc = ref.find("Gene-commentary_accession")
ver = ref.find("Gene-commentary_version")
typ = ref.find("Gene-commentary_type")
if acc is not None and acc.text and acc.text.startswith("NM_"):
print(f"RefSeq mRNA: {acc.text}.{ver.text if ver is not None else ''}")Map a list of gene symbols to NCBI Gene IDs efficiently.
import requests, time
EMAIL = "your@email.com"
BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
def symbols_to_ids(symbols, organism="Homo sapiens"):
"""Map gene symbols to NCBI Gene IDs. Returns dict {symbol: gene_id}."""
mapping = {}
for sym in symbols:
r = requests.get(f"{BASE}/esearch.fcgi",
params={"db": "gene", "email": EMAIL, "retmode": "json",
"term": f"{sym}[sym] AND {organism}[orgn] AND alive[prop]"})
ids = r.json()["esearchresult"]["idlist"]
mapping[sym] = ids[0] if ids else None
time.sleep(0.1)
return mapping
genes = ["EGFR", "KRAS", "BRAF", "PIK3CA", "PTEN"]
id_map = symbols_to_ids(genes)
for sym, gid in id_map.items():
print(f"{sym:10s} → Gene ID {gid}")Parse GO terms from the gene XML record.
import requests
import xml.etree.ElementTree as ET
EMAIL = "your@email.com"
BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
r = requests.get(f"{BASE}/efetch.fcgi",
params={"db": "gene", "id": "7157",
"rettype": "xml", "retmode": "xml", "email": EMAIL})
root = ET.fromstring(r.text)
# Extract GO annotations
go_terms = []
for ref in root.iter("Gene-commentary"):
heading = ref.find("Gene-commentary_heading")
label = ref.find("Gene-commentary_label")
if heading is not None and "Gene Ontology" in heading.text:
if label is not None:
go_terms.append(label.text)
print(f"TP53 GO terms ({len(go_terms)} found):")
for term in go_terms[:10]:
print(f" {term}")Find orthologs across species. Note: the legacy E-utilities link gene_gene_homolog was retired with HomoloGene in 2019 — the modern path is the NCBI Datasets v2 REST API, which exposes a dedicated orthologs endpoint.
import requests, time
DATASETS_BASE = "https://api.ncbi.nlm.nih.gov/datasets/v2"
def get_orthologs(gene_id, taxon_filter=None):
"""Return ortholog Gene reports for a given NCBI Gene ID.
taxon_filter: optional tax_id (int) or list of tax_ids to narrow species."""
params = {}
if taxon_filter is not None:
# tax_ids: human=9606, mouse=10090, rat=10116, zebrafish=7955, fly=7227
ids = taxon_filter if isinstance(taxon_filter, (list, tuple)) else [taxon_filter]
params["taxon_filter"] = [str(t) for t in ids]
r = requests.get(f"{DATASETS_BASE}/gene/id/{gene_id}/orthologs",
params=params, timeout=30)
r.raise_for_status()
return r.json().get("reports", [])
# Mouse ortholog of human TP53 (Gene ID 7157)
reports = get_orthologs("7157", taxon_filter=10090)
for rep in reports[:5]:
g = rep.get("gene", {})
print(f" {g.get('symbol'):8s} (tax {g.get('tax_id')}, gene_id {g.get('gene_id')}): "
f"{g.get('description', '')[:60]}")
# Expect: Trp53 (tax 10090, gene_id 22059): transformation related protein 53
time.sleep(0.34)
# All orthologs (every species in the orthology group)
all_orthologs = get_orthologs("7157")
print(f"\nTotal TP53 orthologs across species: {len(all_orthologs)}")NCBI Gene IDs are integers assigned per gene per organism (e.g., human TP53 = 7157). These are distinct from HGNC IDs (e.g., HGNC:11998) and Ensembl IDs (ENSG00000141510). Many downstream NCBI databases (ClinVar, dbSNP, GEO) use NCBI Gene IDs internally.
alive[prop] FilterNCBI Gene records for discontinued genes have status=discontinued. Always add AND alive[prop] to symbol queries to exclude retired entries and avoid retrieving stale data.
Goal: For a list of gene symbols, retrieve Gene ID, official name, chromosomal location, and description.
import requests, time, pandas as pd
EMAIL = "your@email.com"
BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
def search_gene(sym, organism="Homo sapiens"):
r = requests.get(f"{BASE}/esearch.fcgi",
params={"db": "gene", "email": EMAIL, "retmode": "json",
"term": f"{sym}[sym] AND {organism}[orgn] AND alive[prop]"})
ids = r.json()["esearchresult"]["idlist"]
return ids[0] if ids else None
def batch_summary(gene_ids):
r = requests.post(f"{BASE}/esummary.fcgi",
data={"db": "gene", "id": ",".join(gene_ids),
"retmode": "json", "email": EMAIL})
return r.json()["result"]
symbols = ["BRCA1", "BRCA2", "TP53", "EGFR", "MYC", "KRAS", "PTEN"]
# Step 1: Symbol → Gene ID
id_map = {}
for sym in symbols:
gid = search_gene(sym)
id_map[sym] = gid
time.sleep(0.12)
# Step 2: Batch summary
valid_ids = [v for v in id_map.values() if v]
result = batch_summary(valid_ids)
rows = []
sym_to_id = {v: k for k, v in id_map.items() if v}
for uid in result.get("uids", []):
g = result[uid]
rows.append({
"symbol": sym_to_id.get(uid, g.get("name")),
"gene_id": uid,
"full_name": g.get("description"),
"chr_location": g.get("maplocation"),
"summary": g.get("summary", "")[:200],
})
df = pd.DataFrame(rows)
df.to_csv("gene_annotations.csv", index=False)
print(df[["symbol", "gene_id", "full_name", "chr_location"]].to_string(index=False))Goal: Retrieve all human genes associated with a biological keyword from the NCBI Gene summary field.
import requests, time, pandas as pd
EMAIL = "your@email.com"
BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
keyword = "DNA mismatch repair"
r = requests.get(f"{BASE}/esearch.fcgi",
params={"db": "gene", "email": EMAIL, "retmode": "json",
"retmax": 50,
"term": f"{keyword}[title/abstract] AND Homo sapiens[orgn] AND alive[prop]"})
ids = r.json()["esearchresult"]["idlist"]
print(f"Found {len(ids)} genes related to '{keyword}'")
# Fetch summaries
r2 = requests.post(f"{BASE}/esummary.fcgi",
data={"db": "gene", "id": ",".join(ids), "retmode": "json", "email": EMAIL})
result = r2.json()["result"]
rows = []
for uid in result.get("uids", []):
g = result[uid]
rows.append({"gene_id": uid, "symbol": g.get("name"),
"description": g.get("description"),
"location": g.get("maplocation")})
df = pd.DataFrame(rows)
print(df.to_string(index=False))
df.to_csv(f"{keyword.replace(' ', '_')}_genes.csv", index=False)| Parameter | Module | Default | Range / Options | Effect |
|---|---|---|---|---|
retmax | ESearch | 20 | 1–10000 | Max records returned |
retmode | ESearch/ESummary | "xml" | "json", "xml" | Response format |
rettype | EFetch | depends | "xml", "gene_table", "text" | Record format for full fetch |
[sym] field tag | ESearch | — | gene symbol | Match exact official symbol only |
[orgn] field tag | ESearch | — | organism name or tax ID | Filter by taxonomy |
alive[prop] | ESearch | — | boolean flag | Exclude discontinued gene records |
Always add alive[prop]: Discontinued gene records remain in the database. Without this filter, symbol searches may return outdated records.
Use Gene IDs in pipelines: Downstream NCBI databases (ClinVar, dbSNP, GEO) accept Gene IDs; avoid re-searching by symbol in each call.
Use ESummary for metadata, EFetch for full records: ESummary returns JSON with all common fields; EFetch XML is needed only for RefSeq accessions, GO terms, or interaction links.
Register for a free API key: Triple your rate limit (3 → 10 req/s) at https://www.ncbi.nlm.nih.gov/account/. Pass as api_key parameter.
Batch with ESummary: POST up to 200 Gene IDs per call to ESummary instead of querying one at a time.
When to use: Get the canonical mRNA accession for a protein-coding gene.
import requests, re
EMAIL = "your@email.com"
GENE_ID = "672" # BRCA1
r = requests.get(
"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi",
params={"db": "gene", "id": GENE_ID, "rettype": "gene_table",
"retmode": "text", "email": EMAIL}
)
nm_accessions = re.findall(r"NM_\d+\.\d+", r.text)
print(f"RefSeq mRNA accessions: {list(set(nm_accessions))}")When to use: Resolve legacy/alias symbols to the current official NCBI symbol.
import requests
EMAIL = "your@email.com"
# P53 is an alias for TP53
r = requests.get(
"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi",
params={"db": "gene", "email": EMAIL, "retmode": "json",
"term": "p53[sym] AND Homo sapiens[orgn] AND alive[prop]"}
)
ids = r.json()["esearchresult"]["idlist"]
r2 = requests.post("https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi",
data={"db": "gene", "id": ",".join(ids[:1]),
"retmode": "json", "email": EMAIL})
g = r2.json()["result"][ids[0]]
print(f"Official symbol : {g.get('nomenclaturesymbol', g.get('name'))}")
print(f"Other aliases : {g.get('otheraliases')}")
print(f"Designations : {g.get('otherdesignations', '')[:100]}")When to use: Get all protein-coding genes on a specific human chromosome.
import requests
EMAIL = "your@email.com"
r = requests.get(
"https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi",
params={"db": "gene", "email": EMAIL, "retmode": "json", "retmax": 5,
"term": "17[chr] AND Homo sapiens[orgn] AND protein coding[filter] AND alive[prop]"}
)
result = r.json()["esearchresult"]
print(f"Protein-coding genes on chr17: {result['count']} total")
print(f"Sample IDs: {result['idlist']}")| Problem | Cause | Solution |
|---|---|---|
Empty idlist for known symbol | Symbol is an alias, not the official term | Use [gene name] or [title] field tag; check aliases via ESummary |
| Wrong species returned | Missing organism filter | Add AND Homo sapiens[orgn] or target tax ID (9606[taxid]) |
| Discontinued gene returned | Missing alive[prop] filter | Append AND alive[prop] to all symbol queries |
HTTP 429 rate limit | Too many requests | Add time.sleep(0.35) between calls; use NCBI API key |
ESummary missing uids key | All IDs invalid/absent | Check id values are valid integers, not empty strings |
| XML parse error | Malformed XML for rare genes | Wrap ET.fromstring in try/except; retry with rettype=text |
| Empty ortholog list from ELink | Legacy linkname=gene_gene_homolog retired with HomoloGene in 2019 | Use NCBI Datasets v2 /gene/id/{gene_id}/orthologs instead (Query 6) |
geo-database — Gene Expression Omnibus for retrieving expression data linked to genes found hereclinvar-database — Clinical variant data indexed by NCBI Gene IDsensembl-database — Complementary gene annotations with VEP and comparative genomicsbiopython-molecular-biology — Biopython Entrez module wraps E-utilities with typed return values© jaechang-hits, CC0-1.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/databases/gene-database of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Gene Database next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Gene Database this skilljaechang-hits/SciAgent-Skills | 371 | 1 repos | ~4.4k | Automated safety check: Pass | CC0-1.0 | |
| Bulkrna Geneid MappingTianGzlab/OmicsClaw | 161 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Tooluniverse Phylogeneticswu-yc/LabClaw | 1.1k | 2 repos | ~4.2k | Automated safety check: Pass | None | |
| Bio Ensembl RESTGPTomics/bioSkills | 1.2k | 2 repos | ~3.6k | Automated safety check: Pass | MIT | |
| Ggetaipoch/medical-research-skills | 2k | — | ~816 | Automated safety check: Pass | MIT | |
| Bioconductor BiomartbioMate-AI/biomate-bioconductor-kb | 804 | — | ~4.5k | Automated safety check: Pass | Custom licence |
TianGzlab/OmicsClaw
Load when converting Ensembl, Entrez or symbol IDs in a bulk RNA count matrix using an explicit mapping or a small human demo reference.
wu-yc/LabClaw
Production-ready phylogenetics and sequence analysis skill for alignment processing, tree analysis, and evolutionary metrics.
GPTomics/bioSkills
Query the Ensembl REST API for gene/transcript/protein lookup, sequence retrieval, comparative genomics (Compara), variant effect prediction (VEP), regulatory features, and cross-species…
aipoch/medical-research-skills
Unified CLI/Python interface for querying genomic, proteomic, structure, and expression data across 20+ bioinformatics databases; use when you need fast, scriptable retrieval by gene/protein IDs or…
bioMate-AI/biomate-bioconductor-kb
In recent years a wealth of biological data has become available in public data repositories.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
NCBI Gene via E-utilities: curated records across 1M+ taxa. An agent skill from jaechang-hits/SciAgent-Skills. Gene Database is an agent skill from jaechang-hits/SciAgent-Skills. NCBI Gene via E-utilities: curated records across 1M+ taxa.
Gene Database fits situations like: gene ID resolution and cross-species function queries; tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill gene-database -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/databases/gene-database in jaechang-hits/SciAgent-Skills) into .claude/skills/gene-database in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill gene-database -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/databases/gene-database in jaechang-hits/SciAgent-Skills) into .agents/skills/gene-database in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill gene-database -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gene-database, .gemini/skills/gene-database, .github/skills/gene-database and .opencode/skills/gene-database in your project.
Going by SKILL.md and its folder, Gene Database needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 3 domains. In commands or code: eutils.ncbi.nlm.nih.gov and api.ncbi.nlm.nih.gov; the agent is likely to contact these when it follows the instructions. As links in the text: ncbi.nlm.nih.gov. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Gene Database is published under the CC0-1.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Gene Database: Bulkrna Geneid Mapping (TianGzlab/OmicsClaw, 161 stars), Tooluniverse Phylogenetics (wu-yc/LabClaw, 1.1k stars), Bio Ensembl REST (GPTomics/bioSkills, 1.2k stars) and Gget (aipoch/medical-research-skills, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 371 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.