Agent skill

Bioinformatics

by ZimoLiao in ZimoLiao/scholaraio

A skill your agent uses when working on bioinformatics workflows such as alignment, variant calling, phylogenetics, or protein-structure analysis, especially across BLAST, minimap2, samtools…

MITAuto-check passedResearch & Science

Install Bioinformatics

skills CLI
$ npx skills add ZimoLiao/scholaraio --skill bioinformatics -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZimoLiao/scholaraio bioinformatics --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZimoLiao/scholaraio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/bioinformatics .claude/skills/bioinformatics && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bioinformatics
GitHub stars
577
Token cost
~1.4k tokens
SKILL.md length
321 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when working on bioinformatics workflows such as alignment, variant calling, phylogenetics, or protein-structure analysis, especially across BLAST, minimap2, samtools…

  • Works in 5 steps: 先判断当前任务属于哪类 → 再用 toolref show bioinformatics ... 或… → 不要把“生信工具链”当一个大黑箱查 → …
  • Working on bioinformatics workflows such as alignment
  • SKILL.md covers Agent 默认协议(toolref-first,…, 前置条件, 何时使用 and Toolref 优先, plus 6 more sections
  • Calls pip and conda

What it does

Bioinformatics is an agent skill from ZimoLiao/scholaraio. Use when working on bioinformatics workflows such as alignment, variant calling, phylogenetics, or protein-structure analysis, especially across BLAST, minimap2, samtools, bcftools, MAFFT, IQ-TREE, or ESMFold.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: Scholar All-In-One: A research infrastructure for AI agents. The licence is MIT.

When your agent uses it

  • Working on bioinformatics workflows such as alignment
  • Variant calling
  • Protein-structure analysis
  • Especially across BLAST

Example prompts

  • “/bioinformatics”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. 先判断当前任务属于哪类
  2. 再用 toolref show bioinformatics ... 或 search --program 查对应程序
  3. 不要把“生信工具链”当一个大黑箱查
  4. 如果某个子工具当前 toolref 覆盖不全,agent 应先回退该工具的官方手册或 README,再继续任务
  5. 不要让普通用户自己补齐某个子工具的 toolref

What it can do on your machine

Read from SKILL.md and the folder at commit 777628b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bioinformatics loads about 1.4k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 321 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ZimoLiao/scholaraio at commit 777628b, republished under its MIT licence (© ZimoLiao). 321 words, ~1,350 tokens.

Download SKILL.mdSave it as .claude/skills/bioinformatics/SKILL.md (or your agent's skills folder).
name
bioinformatics
description
Use when working on bioinformatics workflows such as alignment, variant calling, phylogenetics, or protein-structure analysis, especially across BLAST, minimap2, samtools, bcftools, MAFFT, IQ-TREE, or ESMFold.

生物信息学分析

用生物信息学工具链做序列比对、变异检测、系统发育和蛋白质结构分析。

本 skill 故意保持轻量:

  • 它负责告诉 agent 哪类问题该用哪类工具、标准分析链路是什么、哪些生信规范不能忽略
  • 它不承担各个命令行工具的完整手册职责
  • 具体 CLI 选项、子命令、输入输出细节统一去查 scholaraio toolref

Agent 默认协议(toolref-first, toolchain-aware)

Bioinformatics 不是单一程序,而是一组工具链。agent 必须先判断自己在用哪一个子工具,再决定怎么查。

默认顺序:

  1. 先判断当前任务属于哪类:
    • 同源搜索:BLAST
    • 组装序列比对:minimap2
    • BAM/SAM 处理:samtools
    • 变异调用:bcftools
    • 多序列比对:MAFFT
    • 建树:IQ-TREE
    • 结构预测:ESMFold
  2. 再用 toolref show bioinformatics <program> ... 或 search --program <program> 查对应程序
  3. 不要把“生信工具链”当一个大黑箱查
  4. 如果某个子工具当前 toolref 覆盖不全,agent 应先回退该工具的官方手册或 README,再继续任务
  5. 不要让普通用户自己补齐某个子工具的 toolref

这意味着:

  • bioinformatics skill 负责先分流,再选工具
  • toolref 负责各子工具的接口细节
  • 当前覆盖不全时,复杂度应由 agent 吸收,而不是由用户承担

前置条件

bash
# 核心工具(conda bioconda 频道)
conda install -c bioconda minimap2 mafft iqtree bcftools samtools blast

# Python 库
pip install biopython py3Dmol pycirclize toytree matplotlib seaborn pandas

# 蛋白质结构预测(需 GPU)
pip install fair-esm

# 数据获取
pip install ncbi-datasets-cli

验证:minimap2 --version、samtools --version、blastn -version 均应正常输出。

何时使用

适合:

  • 序列相似性搜索、参考比对、变异检测、系统发育树构建
  • 蛋白质结构预测与突变位点解释

不适合:

  • 没有明确数据类型就盲选工具
  • 把生信流程当“黑箱一键按钮”,不检查质量控制和统计假设

Toolref 优先

当 agent 不确定子命令、选项、参数含义时,先查 toolref。

常用查法:

bash
scholaraio toolref show bioinformatics samtools sort
scholaraio toolref show bioinformatics bcftools manual
scholaraio toolref show bioinformatics minimap2 manual
scholaraio toolref show bioinformatics blast blastn
scholaraio toolref search bioinformatics bootstrap tree --program iqtree

推荐习惯:

  • 在决定工具前先确认数据类型:组装序列、短读段、蛋白序列、树推断
  • 写命令前先查对应手册页,而不是靠记忆拼接参数
  • 报告结果时带上阈值、模型和置信度,而不是只给一张图

如果遇到覆盖缺口:

  • 先回退到对应子工具的官方手册
  • 在回答里明确指出是哪个子工具存在 toolref 覆盖不足
  • 不要让用户为了当前分析去维护 toolref

核心工具链

工具功能何时用
BLAST序列相似性搜索查找同源序列、注释未知基因
minimap2序列比对组装序列/长读段 vs 参考基因组
BWA-MEM2短读段比对Illumina 短读段 vs 参考基因组
samtoolsBAM/SAM 操作排序、索引、统计
bcftools变异检测SNP/InDel calling
MAFFT多序列比对建树前的全局比对
IQ-TREE最大似然系统发育建进化树(支持 bootstrap)
FastTree快速近似建树大规模序列(>1000 条)
ESMFold蛋白质结构预测AI 蛋白质折叠(用 A100 GPU)
BioPython通用生物信息学PDB 解析、序列操作、Entrez 查询
工具选择规范
场景正确工具常见错误
组装基因组 vs 参考minimap2用 BWA(BWA 是短读段工具)
短读段 vs 参考BWA-MEM2用 minimap2(不够精确)
建进化树IQ-TREE (ML)用 NJ(邻接法太粗糙)
蛋白质结构ESMFold 或 PDB 实验结构盲目信任预测不看 pLDDT

工作流模板

变异分析

建议流程:

  1. 明确数据类型和参考序列
  2. 选择正确比对工具
  3. 用 samtools / bcftools 做排序、索引和变异调用
  4. 做变异注释与功能解释
  5. 与论文或数据库中的关键突变结论交叉验证
系统发育分析

建议流程:

  1. 明确序列集合和研究问题
  2. 多序列比对
  3. 用 ML 方法建树
  4. 检查 bootstrap 支持度
  5. 讨论拓扑、分支和可能的趋同进化,而不是只贴树图
蛋白质结构预测

建议流程:

  1. 优先检查是否已有实验结构
  2. 无实验结构时再做预测
  3. 报告 pLDDT 或其他置信度
  4. 将结构解释和突变、生物功能联系起来
重点查询点
  • minimap2 的预设与输入输出格式
  • samtools sort/view/index 的正确用法
  • bcftools 变异调用链路
  • iqtree 的 bootstrap 与模型参数
  • blastn 的输出格式和阈值

这些细节优先查 toolref。

可视化

系统发育树
python
import toytree
import matplotlib.pyplot as plt

tree = toytree.tree("tree.treefile")
canvas, axes, marks = tree.draw(
    width=600, height=800,
    tip_labels_align=True,
    node_sizes=[0 if not n.is_leaf() else 8 for n in tree.treenode.traverse()],
)
# 按类群着色需自定义 node_colors
Circos 基因组图
python
from pycirclize import Circos

circos = Circos(sectors={"genome": genome_length})
sector = circos.sectors[0]

# 轨道 1: 基因注释
track1 = sector.add_track((90, 95))
# 轨道 2: 变异密度
track2 = sector.add_track((80, 88))
# 轨道 3: GC 含量
track3 = sector.add_track((70, 78))

circos.savefig("circos.png", dpi=300)
3D 蛋白质结构
python
import py3Dmol

view = py3Dmol.view(width=800, height=600)
view.addModel(pdb_string, "pdb")

# 卡通表示 + 突变位点高亮
view.setStyle({"cartoon": {"color": "spectrum"}})
# 突变残基显示为球棍
view.addStyle({"resi": [484, 501, 681]},
              {"stick": {"colorscheme": "redCarbon"}})
view.zoomTo()
view.show()
突变景观图(Lollipop plot)
python
import matplotlib.pyplot as plt

fig, ax = plt.subplots(figsize=(15, 4))
# x 轴: 蛋白位置
# y 轴: 携带该突变的变异株数量
# 颜色: 按功能域着色(NTD, RBD, S1/S2, S2)
ax.vlines(positions, 0, counts, colors=domain_colors, linewidth=1.5)
ax.scatter(positions, counts, c=domain_colors, s=30, zorder=5)
# 标注关键突变
for pos, name in key_mutations:
    ax.annotate(name, (pos, counts[pos]), fontsize=8, rotation=45)

GPU 使用(ESMFold)

序列长度VRAM 需求A100 40GB
< 400 aa~10 GB单 GPU
400-800 aa~15-20 GB单 GPU
800-1200 aa~25-35 GB单 GPU
> 1200 aa> 40 GB需拆分域或多 GPU

建议:预测蛋白质域(200-500 aa)而非全长(可能超出 VRAM)。

科学规范

检查项正确做法常见错误
比对工具组装用 minimap2,短读段用 BWA混用
建树方法ML (IQ-TREE) + bootstrap ≥1000用 NJ 不做 bootstrap
替换模型DNA: GTR+G4,蛋白: LG+G4不做模型选择
结构预测报告 pLDDT,与实验结构对比盲目信任预测
突变命名标准命名(N501Y 而非"501位天冬酰胺突变")非标准命名
趋同进化区分趋同进化和共祖混淆
E-valueBLAST 结果按 e-value 过滤不设阈值

Agent 行为准则

  • 不要先想命令,先判断数据类型和科学问题
  • 不要混用工具适用场景,尤其是 minimap2 / BWA / BLAST
  • 不要只给图,不给阈值、支持度、置信度和过滤标准
  • 不要把结构预测或树拓扑当事实,必须解释其可信边界

© ZimoLiao, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/bioinformatics of ZimoLiao/scholaraio.

Open the folder on GitHubat commit 777628b

Compare with similar skills

Bioinformatics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bioinformatics compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bioinformatics this skillZimoLiao/scholaraio577—~1.4kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from ZimoLiao/scholaraio

All 43 skills in this repo
  • Document

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to create or inspect DOCX, PPTX, or XLSX files, generate a downloadable Office deliverable, or verify its structure and layout warnings with scholaraio…

    577 GitHub stars~674 tokensUpdated 14 days ago
    Auto-check passed
  • Academic Writing

    ZimoLiao/scholaraio

    A skill your agent uses when the user needs help choosing or organizing an academic-writing workflow by deliverable, stage, or format, including review articles, guided reading, paper sections, PPT…

    577 GitHub stars~827 tokensUpdated 14 days ago
    Auto-check passed
  • Arxiv

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to browse arXiv preprints, search arXiv directly, fetch a PDF by arXiv ID or URL, or send a preprint into the ScholarAIO ingest pipeline.

    577 GitHub stars~799 tokensUpdated 14 days ago
    Auto-check passed
  • Citation Check

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to verify citations in AI-generated or human-written text against the local knowledge base and catch hallucinated, wrong, or missing references.

    577 GitHub stars~454 tokensUpdated 14 days ago
    Auto-check passed
  • Draw

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants diagrams, flowcharts, architecture visuals, data relationships, timelines, concept maps, Mermaid, Graphviz, drawio, or polished paper figures generated…

    577 GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check: notes
  • Explore

    ZimoLiao/scholaraio

    A skill your agent uses when the user wants to survey a journal or field, fetch papers from OpenAlex, cluster topics, build exploration embeddings, or search named explore libraries under…

    577 GitHub stars~755 tokensUpdated 14 days ago
    Auto-check passed

Questions about Bioinformatics

What does Bioinformatics do?

A skill your agent uses when working on bioinformatics workflows such as alignment, variant calling, phylogenetics, or protein-structure analysis, especially across BLAST, minimap2, samtools…. Bioinformatics is an agent skill from ZimoLiao/scholaraio. Use when working on bioinformatics workflows such as alignment, variant calling, phylogenetics, or protein-structure analysis, especially across BLAST, minimap2, samtools, bcftools, MAFFT, IQ-TREE, or ESMFold.

When should I use Bioinformatics?

Bioinformatics fits situations like: working on bioinformatics workflows such as alignment; variant calling; protein-structure analysis; especially across BLAST.

How do I install Bioinformatics in Claude Code?

Run `npx skills add ZimoLiao/scholaraio --skill bioinformatics -a claude-code`. Or copy the skill folder (.claude/skills/bioinformatics in ZimoLiao/scholaraio) into .claude/skills/bioinformatics in your project. Claude Code loads it when a task matches its description.

How do I install Bioinformatics in Codex?

Run `npx skills add ZimoLiao/scholaraio --skill bioinformatics -a codex`. Or copy the skill folder (.claude/skills/bioinformatics in ZimoLiao/scholaraio) into .agents/skills/bioinformatics in your project. Codex loads it when a task matches its description.

Can I use Bioinformatics in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZimoLiao/scholaraio --skill bioinformatics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bioinformatics, .gemini/skills/bioinformatics, .github/skills/bioinformatics and .opencode/skills/bioinformatics in your project.

What does Bioinformatics need to run?

Going by SKILL.md and its folder, Bioinformatics needs the command-line tools its instructions call (pip and conda). Our summary lists: Python 3.

Does Bioinformatics access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bioinformatics safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bioinformatics use?

Bioinformatics is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bioinformatics use?

About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bioinformatics?

Skills that share tags, products or a category with Bioinformatics: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bioinformatics?

ZimoLiao (a GitHub user) maintains it in ZimoLiao/scholaraio, which has 577 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 25, 2026.

Source: ZimoLiao/scholaraio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.