Agent skill

Biopython

by aipoch in aipoch/medical-research-skills

A comprehensive toolbox for computational molecular biology; use it when you need programmatic sequence/structure parsing, batch bioinformatics pipelines, or automated NCBI/BLAST workflows.

MITAuto-check passedResearch & Science

Install Biopython

skills CLI
$ npx skills add aipoch/medical-research-skills --skill biopython -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills biopython --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Data Analysis/biopython' .claude/skills/biopython && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
biopython
GitHub stars
2k
Token cost
~1.7k tokens
SKILL.md length
427 words
Files
9 (incl. references)
Skills in repo
567
Repo updated
First seen
Licence
MIT

At a glance

A comprehensive toolbox for computational molecular biology; use it when you need programmatic sequence/structure parsing, batch bioinformatics pipelines, or automated NCBI/BLAST workflows.

  • Works in 5 steps: parses a FASTA file, → computes GC fraction, → runs a remote BLAST (optional), → …
  • You need programmatic sequence/structure parsing
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 1 more section
  • Calls python; needs NCBI_API_KEY

What it does

Biopython is an agent skill from aipoch/medical-research-skills. A comprehensive toolbox for computational molecular biology; use it when you need programmatic sequence/structure parsing, batch bioinformatics pipelines, or automated NCBI/BLAST workflows.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `biopython_audit_result_v1.json`, `references/advanced.md` and `references/alignment.md`).

It sits in Research & Science, covering Bioinformatics. It works with NCBI and Biopython. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • You need programmatic sequence/structure parsing
  • Batch bioinformatics pipelines
  • Automated NCBI/BLAST workflows

Example prompts

  • “/biopython”

Requirements

  • Python 3
  • A credential in NCBI_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. parses a FASTA file,
  2. computes GC fraction,
  3. runs a remote BLAST (optional),
  4. fetches the top hit from NCBI,
  5. prints basic results.

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NCBI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Biopython loads about 1.7k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 50 tokens; SKILL.md has 427 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~23k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 427 words, ~1,652 tokens.

Download SKILL.mdSave it as .claude/skills/biopython/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
biopython
description
A comprehensive toolbox for computational molecular biology; use it when you need programmatic sequence/structure parsing, batch bioinformatics pipelines, or automated NCBI/BLAST workflows.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

When to Use

Use this skill when you need to:

  • Batch-process DNA/RNA/protein sequences (translation, reverse complement, statistics) as part of a custom pipeline.
  • Parse, validate, convert, or stream large bioinformatics files (FASTA/FASTQ/GenBank/PDB/mmCIF) without loading everything into memory.
  • Programmatically query and download records from NCBI (GenBank, PubMed, Gene, Protein) via Bio.Entrez, respecting rate limits.
  • Automate BLAST searches (web or local) and parse results to extract top hits and metadata.
  • Build or manipulate phylogenetic trees from alignments or distance matrices (e.g., NJ trees) for downstream analysis.

Note: For quick one-off queries, tools like gget may be more convenient; for multi-service API aggregation, bioservices may be a better fit.

Key Features

  • Sequence objects and utilities: Bio.Seq, Bio.SeqRecord, Bio.SeqUtils (GC fraction, molecular weight, translation, etc.).
  • File I/O and format conversion: Bio.SeqIO, Bio.AlignIO for FASTA/FASTQ/GenBank and alignment formats.
  • NCBI access: Bio.Entrez for esearch, efetch, elink, and structured parsing via Entrez.read.
  • BLAST: Bio.Blast.NCBIWWW for remote BLAST and Bio.Blast.NCBIXML for XML parsing.
  • Structural bioinformatics: Bio.PDB for PDB/mmCIF parsing, hierarchy traversal, and geometry calculations.
  • Phylogenetics: Bio.Phylo and Bio.Phylo.TreeConstruction for tree I/O, distances, and construction.

Reference guides (if present in this repository) can be consulted for deeper module-specific patterns:

  • references/sequence_io.md
  • references/alignment.md
  • references/databases.md
  • references/blast.md
  • references/structure.md
  • references/phylogenetics.md
  • references/advanced.md

Dependencies

  • Python >= 3.8 (Biopython 1.85 supports Python 3)
  • biopython==1.85
  • numpy>=1.20 (required by Biopython)

Install:

bash
python -m pip install "biopython==1.85" "numpy>=1.20"

Example Usage

A complete, runnable example that:

  1. parses a FASTA file,
  2. computes GC fraction,
  3. runs a remote BLAST (optional),
  4. fetches the top hit from NCBI,
  5. prints basic results.

Create example_biopython_pipeline.py:

python
from __future__ import annotations

import os
import time
from typing import Optional

from Bio import Entrez, SeqIO
from Bio.SeqUtils import gc_fraction

# Optional BLAST (remote). Comment out if you do not want network calls.
from Bio.Blast import NCBIWWW, NCBIXML


def configure_entrez() -> None:
    """
    NCBI requires an email. An API key increases rate limits.
    Set these via environment variables to avoid hardcoding secrets.
    """
    email = os.environ.get("NCBI_EMAIL")
    if not email:
        raise RuntimeError("Set NCBI_EMAIL env var (required by NCBI). Example: export NCBI_EMAIL='you@org.org'")
    Entrez.email = email

    api_key = os.environ.get("NCBI_API_KEY")
    if api_key:
        Entrez.api_key = api_key


def read_first_fasta_record(path: str):
    with open(path, "r", encoding="utf-8") as handle:
        return next(SeqIO.parse(handle, "fasta"))


def blast_top_accession(sequence: str, program: str = "blastn", database: str = "nt") -> Optional[str]:
    """
    Remote BLAST can be slow and rate-limited. For large-scale BLAST, prefer local BLAST+.
    """
    result_handle = NCBIWWW.qblast(program, database, sequence)
    blast_record = NCBIXML.read(result_handle)

    if not blast_record.alignments:
        return None

    # Many BLAST titles include multiple identifiers; accession is usually available directly.
    return blast_record.alignments[0].accession


def fetch_fasta_by_accession(accession: str) -> str:
    with Entrez.efetch(db="nucleotide", id=accession, rettype="fasta", retmode="text") as handle:
        return handle.read()


def main() -> None:
    configure_entrez()

    record = read_first_fasta_record("input.fasta")
    seq = record.seq

    print(f"ID: {record.id}")
    print(f"Length: {len(seq)}")
    print(f"GC fraction: {gc_fraction(seq):.2%}")

    # Be polite to NCBI services in batch workflows.
    time.sleep(0.34)

    top_acc = blast_top_accession(str(seq))
    if not top_acc:
        print("No BLAST hits found.")
        return

    print(f"Top BLAST accession: {top_acc}")

    time.sleep(0.34)
    fasta_text = fetch_fasta_by_accession(top_acc)
    print("Top hit FASTA:")
    print(fasta_text)


if __name__ == "__main__":
    main()

Run:

bash
export NCBI_EMAIL="your.email@example.com"
# export NCBI_API_KEY="your_ncbi_api_key"  # optional
python example_biopython_pipeline.py

Provide an input.fasta in the same directory, e.g.:

text
>demo
ATCGATCGATCGATCGATCG
Show full SKILL.md (170 more words)Show less

Implementation Details

  • Streaming I/O for large datasets: Prefer iterator-based parsing (SeqIO.parse) to avoid loading entire files into memory. Use SeqIO.read only when exactly one record is expected.
  • Entrez configuration and rate limits:
    • Always set Entrez.email (NCBI requirement).
    • Optionally set Entrez.api_key to increase request limits.
    • In batch jobs, add delays (e.g., time.sleep(0.34) as a conservative baseline) and implement retries for transient HTTP failures.
  • BLAST considerations:
    • NCBIWWW.qblast(...) is convenient but can be slow and is not ideal for high-throughput workloads.
    • Parse results with NCBIXML.read(...) (single record) or NCBIXML.parse(...) (multiple records).
    • Filter hits by HSP metrics (e-value, identity) by iterating alignment.hsps.
  • Sequence statistics and transformations:
    • Use Bio.SeqUtils.gc_fraction(seq) for GC fraction (returns 0–1).
    • Use seq.translate(table=...) with the correct genetic code table for reproducibility.
  • Structure parsing (if used):
    • Use Bio.PDB.PDBParser(QUIET=True) to suppress warnings when appropriate.
    • Navigate the SMCRA hierarchy (Structure → Model → Chain → Residue → Atom) for robust traversal and geometry calculations.
  • Reproducibility:
    • Record key parameters (file formats, translation table, BLAST program/database, e-value thresholds, NCBI query terms).
    • Cache downloaded records when iterating to avoid repeated network calls.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in scientific-skills/Data Analysis/biopython of aipoch/medical-research-skills.

  • SKILL.md
  • biopython_audit_result_v1.json
  • references/advanced.md
  • references/alignment.md
  • references/blast.md
  • references/databases.md
  • references/phylogenetics.md
  • references/sequence_io.md
  • references/structure.md

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Biopython next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Biopython compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Biopython this skillaipoch/medical-research-skills2k—~1.7kAutomated safety check: PassMIT
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
Bio Write SequencesGPTomics/bioSkills1.2k3 repos~2.1kAutomated safety check: PassMIT
Biopythondavila7/claude-code-templates32k13 repos~3.4kAutomated safety check: PassMIT
BiopythonK-Dense-AI/scientific-agent-skills48k1 repos~4.3kAutomated safety check: NotesMIT
Biopythonlamm-mit/scienceclaw244—~3.9kAutomated safety check: PassApache-2.0

Similar skills

  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 13 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Biopython

    K-Dense-AI/scientific-agent-skills

    Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).

    48k GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    lamm-mit/scienceclaw

    Computational molecular biology library (sequence I/O, alignment, phylogenetics).

    244 GitHub stars~3.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Production-ready phylogenetics and sequence analysis skill for alignment processing, tree analysis, and evolutionary metrics.

    1.1k GitHub starsUsed in 2 repos~4.2k tokens
    Research & ScienceAuto-check passed

More from aipoch/medical-research-skills

All 567 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 20 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 20 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 20 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 20 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 20 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 20 days ago
    Auto-check passed

Works with

Questions about Biopython

What does Biopython do?

A comprehensive toolbox for computational molecular biology; use it when you need programmatic sequence/structure parsing, batch bioinformatics pipelines, or automated NCBI/BLAST workflows. Biopython is an agent skill from aipoch/medical-research-skills. A comprehensive toolbox for computational molecular biology; use it when you need programmatic sequence/structure parsing, batch bioinformatics pipelines, or automated NCBI/BLAST workflows.

When should I use Biopython?

Biopython fits situations like: you need programmatic sequence/structure parsing; batch bioinformatics pipelines; automated NCBI/BLAST workflows.

How do I install Biopython in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill biopython -a claude-code`. Or copy the skill folder (scientific-skills/Data Analysis/biopython in aipoch/medical-research-skills) into .claude/skills/biopython in your project. Claude Code loads it when a task matches its description.

How do I install Biopython in Codex?

Run `npx skills add aipoch/medical-research-skills --skill biopython -a codex`. Or copy the skill folder (scientific-skills/Data Analysis/biopython in aipoch/medical-research-skills) into .agents/skills/biopython in your project. Codex loads it when a task matches its description.

Can I use Biopython in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill biopython -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/biopython, .gemini/skills/biopython, .github/skills/biopython and .opencode/skills/biopython in your project.

What does Biopython need to run?

Going by SKILL.md and its folder, Biopython needs the command-line tools its instructions call (python) and credentials named NCBI_API_KEY. Our summary lists: Python 3; A credential in NCBI_API_KEY.

Does Biopython access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Biopython safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Biopython use?

Biopython is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Biopython use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 22k tokens, read only when the agent opens those files.

What are the alternatives to Biopython?

Skills that share tags, products or a category with Biopython: Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Bio Write Sequences (GPTomics/bioSkills, 1.2k stars), Biopython (davila7/claude-code-templates, 32k stars) and Biopython (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Biopython?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,973 GitHub stars. The repository holds 567 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.