Biopython Bioinformatics
aiming-lab/AutoResearchClaw
Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.
Biopython sequence analysis: parse FASTA/FASTQ/GenBank/GFF (SeqIO), NCBI Entrez (esearch/efetch/elink), remote/local BLAST, pairwise/MSA alignment (PairwiseAligner, MUSCLE/ClustalW), phylogenetic…
$ npx skills add jaechang-hits/SciAgent-Skills --skill biopython-sequence-analysis -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills biopython-sequence-analysis --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/biopython-sequence-analysis .claude/skills/biopython-sequence-analysis && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "biopython-sequence-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/biopython-sequence-analysis into .claude/skills/biopython-sequence-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "biopython-sequence-analysis", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/biopython-sequence-analysisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill biopython-sequence-analysis -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills biopython-sequence-analysis --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/genomics-bioinformatics/biopython-sequence-analysis .agents/skills/biopython-sequence-analysis && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "biopython-sequence-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/biopython-sequence-analysis into .agents/skills/biopython-sequence-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "biopython-sequence-analysis", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill biopython-sequence-analysis -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills biopython-sequence-analysis --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/genomics-bioinformatics/biopython-sequence-analysis .cursor/skills/biopython-sequence-analysis && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "biopython-sequence-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/biopython-sequence-analysis into .cursor/skills/biopython-sequence-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "biopython-sequence-analysis", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/genomics-bioinformatics/biopython-sequence-analysis--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill biopython-sequence-analysis -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills biopython-sequence-analysis --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/genomics-bioinformatics/biopython-sequence-analysis .gemini/skills/biopython-sequence-analysis && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "biopython-sequence-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/biopython-sequence-analysis into .gemini/skills/biopython-sequence-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "biopython-sequence-analysis", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills biopython-sequence-analysisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill biopython-sequence-analysis -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/genomics-bioinformatics/biopython-sequence-analysis .github/skills/biopython-sequence-analysis && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "biopython-sequence-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/biopython-sequence-analysis into .github/skills/biopython-sequence-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "biopython-sequence-analysis", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill biopython-sequence-analysis -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills biopython-sequence-analysis --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/genomics-bioinformatics/biopython-sequence-analysis .opencode/skills/biopython-sequence-analysis && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "biopython-sequence-analysis" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/genomics-bioinformatics/biopython-sequence-analysis into .opencode/skills/biopython-sequence-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "biopython-sequence-analysis", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
biopython-sequence-analysisBiopython sequence analysis: parse FASTA/FASTQ/GenBank/GFF (SeqIO), NCBI Entrez (esearch/efetch/elink), remote/local BLAST, pairwise/MSA alignment (PairwiseAligner, MUSCLE/ClustalW), phylogenetic…
Biopython Sequence Analysis is an agent skill from jaechang-hits/SciAgent-Skills. Biopython sequence analysis: parse FASTA/FASTQ/GenBank/GFF (SeqIO), NCBI Entrez (esearch/efetch/elink), remote/local BLAST, pairwise/MSA alignment (PairwiseAligner, MUSCLE/ClustalW), phylogenetic trees (Phylo). Use for gene family studies, phylogenomics, comparative genomics, NCBI pipelines. For PCR/restriction/cloning use biopython-molecular-biology; for SAM/BAM use pysam.
Its SKILL.md is about 8.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics. It works with NCBI, Biopython and pysam. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipcondaFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
ncbi.nlm.nih.govAlso links to:
biopython.orggithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Biopython Sequence Analysis loads about 8.5k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 1,536 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 1,536 words, ~8,477 tokens.
.claude/skills/biopython-sequence-analysis/SKILL.md (or your agent's skills folder).Biopython provides a comprehensive suite of modules for sequence-centric bioinformatics: reading and writing every major biological file format (FASTA, FASTQ, GenBank, GFF), querying NCBI databases programmatically, running BLAST searches and parsing results, aligning sequences pairwise or in multiple-sequence alignments, and building and visualizing phylogenetic trees. This skill focuses on analysis workflows — from NCBI data retrieval through alignment to phylogenetic inference.
For PCR primer design, restriction enzyme digestion, cloning simulation, protein structure analysis (Bio.PDB), and molecular weight/Tm calculations, see biopython-molecular-biology.
nt or nr, filter significant hits, and fetch their full sequencesSeqIO.index() for random-access retrieval without loading all sequences into RAMbiopython, numpy, matplotlibBio.Align.Applications wrappers)Entrez.email before any E-utilities call; obtain a free API key at https://www.ncbi.nlm.nih.gov/account/ for 10 req/s (default is 3 req/s)conda install -c bioconda blast) for offline searchesCheck before installing: The tool may already be available in the current environment (e.g., inside a
pixi/condaenv). Runcommand -v pythonfirst and skip the install commands below if it returns a path. When running inside a pixi project, invoke the tool viapixi run pythonrather than barepython.
pip install biopython numpy matplotlib
conda install -c bioconda blast # optional, for local BLASTfrom Bio import SeqIO, Entrez
from Bio.SeqUtils import gc_fraction
# Fetch a GenBank record and display basic stats
Entrez.email = "your.email@example.com"
handle = Entrez.efetch(db="nucleotide", id="NM_007294", rettype="gb", retmode="text")
record = SeqIO.read(handle, "genbank")
handle.close()
print(f"ID: {record.id}")
print(f"Length: {len(record.seq)} bp")
print(f"GC content: {gc_fraction(record.seq)*100:.1f}%")
print(f"Features: {len(record.features)}")
print(f"First feature: {record.features[0].type} at {record.features[0].location}")
# ID: NM_007294.4
# Length: 7207 bp
# GC content: 47.3%
# Features: 9Bio.SeqIO reads and writes every major sequence format (FASTA, FASTQ, GenBank, EMBL, PHYLIP, Nexus). SeqIO.parse() returns an iterator of SeqRecord objects; SeqIO.index() builds an on-disk or in-memory dictionary for large files.
from Bio import SeqIO
# Parse FASTA and FASTQ
fasta_records = list(SeqIO.parse("sequences.fasta", "fasta"))
print(f"FASTA: {len(fasta_records)} sequences")
# FASTQ: access quality scores
for rec in SeqIO.parse("reads.fastq", "fastq"):
quals = rec.letter_annotations["phred_quality"]
avg_q = sum(quals) / len(quals)
print(f" {rec.id}: {len(rec.seq)} bp, mean Q={avg_q:.1f}")
break # show one example
# Parse GenBank with feature access
for rec in SeqIO.parse("chromosome.gb", "genbank"):
cdss = [f for f in rec.features if f.type == "CDS"]
print(f"{rec.id}: {len(cdss)} CDS features")
for cds in cdss[:3]:
gene = cds.qualifiers.get("gene", ["unknown"])[0]
print(f" {gene}: {cds.location}")from Bio import SeqIO
# SeqIO.index() for random access without loading all records
# Useful for large reference FASTA files (genomes, nr database subsets)
idx = SeqIO.index("large_genome.fasta", "fasta")
print(f"Index contains {len(idx)} sequences")
# Retrieve specific sequences by ID in O(1)
target = idx["chr1"]
print(f"chr1: {len(target.seq):,} bp")
region = target.seq[1_000_000:1_001_000]
print(f"Region [1M-1M+1kb]: {region[:60]}...")
# Format conversion: GenBank → FASTA in one call
n = SeqIO.convert("annotation.gb", "genbank", "sequences.fasta", "fasta")
print(f"Converted {n} records to FASTA")Bio.Seq represents a biological sequence with in-place operations. Bio.SeqRecord wraps a Seq with an ID, description, and a dictionary of feature annotations. Features use FeatureLocation with strand (+1, -1).
from Bio.Seq import Seq
from Bio.SeqRecord import SeqRecord
from Bio.SeqFeature import SeqFeature, FeatureLocation
from Bio.SeqUtils import gc_fraction, MeltingTemp
dna = Seq("ATGAAACCCGGGTTTTAA")
print(f"GC content: {gc_fraction(dna)*100:.1f}%")
print(f"Reverse complement: {dna.reverse_complement()}")
print(f"Transcript: {dna.transcribe()}")
print(f"Translation: {dna.translate()}")
print(f"Translation (to stop): {dna.translate(to_stop=True)}")
# Tm calculation for a short primer
primer = Seq("ATGAAACCCGGG")
tm = MeltingTemp.Tm_Wallace(primer)
print(f"Tm (Wallace): {tm:.1f}°C")from Bio.Seq import Seq
from Bio.SeqRecord import SeqRecord
from Bio.SeqFeature import SeqFeature, FeatureLocation
# Build a SeqRecord with annotated features
gene_seq = Seq("ATGAAACCCGGGTTTTAAATCGATCG" * 10)
record = SeqRecord(gene_seq, id="MY_GENE_001", description="synthetic example gene")
# Annotate a CDS feature
cds = SeqFeature(
FeatureLocation(0, 18, strand=+1),
type="CDS",
qualifiers={"gene": ["myGene"], "product": ["hypothetical protein"]}
)
record.features.append(cds)
# Extract and translate the feature
cds_seq = cds.location.extract(record.seq)
protein = cds_seq.translate(to_stop=True)
print(f"CDS: {cds_seq}")
print(f"Protein: {protein}")
# Slice record preserving feature annotations
sub = record[0:60]
print(f"Subrecord: {len(sub.seq)} bp, {len(sub.features)} features")Bio.Entrez wraps the NCBI E-utilities API: esearch finds records matching a query, efetch retrieves full records, elink finds related records across databases, and esummary returns document summaries. Always set Entrez.email and respect the 3 req/s rate limit (10 req/s with API key).
from Bio import Entrez, SeqIO
import time
Entrez.email = "your.email@example.com"
# Entrez.api_key = "YOUR_API_KEY" # for 10 req/s
# esearch: find IDs matching a query
handle = Entrez.esearch(db="nucleotide", term="BRCA1[Gene] AND Homo sapiens[Organism] AND mRNA[Filter]", retmax=10)
search_results = Entrez.read(handle)
handle.close()
print(f"Total hits: {search_results['Count']}")
ids = search_results["IdList"]
print(f"Retrieved IDs: {ids}")
# efetch: download full GenBank records
for acc_id in ids[:3]:
handle = Entrez.efetch(db="nucleotide", id=acc_id, rettype="gb", retmode="text")
record = SeqIO.read(handle, "genbank")
handle.close()
print(f" {record.id}: {len(record.seq)} bp — {record.description[:60]}")
time.sleep(0.4) # stay within rate limitfrom Bio import Entrez
import time
Entrez.email = "your.email@example.com"
# elink: cross-database links (e.g., PubMed article → related nucleotide sequences)
handle = Entrez.elink(dbfrom="pubmed", db="nucleotide", id="29087512")
link_results = Entrez.read(handle)
handle.close()
linked_ids = []
for linkset in link_results:
for db_links in linkset.get("LinkSetDb", []):
if db_links["DbTo"] == "nucleotide":
linked_ids = [lnk["Id"] for lnk in db_links["Link"]]
break
print(f"Nucleotide sequences linked to PubMed 29087512: {len(linked_ids)}")
print(f"First IDs: {linked_ids[:5]}")
# esummary: lightweight metadata without downloading full records
if linked_ids:
handle = Entrez.esummary(db="nucleotide", id=",".join(linked_ids[:5]))
summaries = Entrez.read(handle)
handle.close()
for doc in summaries:
print(f" {doc['AccessionVersion']}: {doc['Title'][:70]}")Bio.Blast.NCBIWWW.qblast() submits queries to NCBI BLAST servers and returns XML handles. Bio.Blast.NCBIXML.parse() (or .read() for single queries) yields Blast objects with alignments and hsps. For large-scale searches, use local BLAST+ via subprocess.
from Bio.Blast import NCBIWWW, NCBIXML
from Bio.Seq import Seq
# Remote BLASTP against Swiss-Prot (small, reviewed database)
query = Seq("MTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSY")
result_handle = NCBIWWW.qblast("blastp", "swissprot", str(query), hitlist_size=10)
# Parse results
blast_record = NCBIXML.read(result_handle)
print(f"Query: {blast_record.query}")
print(f"Database: {blast_record.database}")
print(f"Hits: {len(blast_record.alignments)}")
E_VALUE_THRESH = 1e-5
for alignment in blast_record.alignments[:5]:
for hsp in alignment.hsps:
if hsp.expect < E_VALUE_THRESH:
identity_pct = hsp.identities / hsp.align_length * 100
print(f"\n Hit: {alignment.title[:70]}")
print(f" E-value: {hsp.expect:.2e}, Identity: {identity_pct:.1f}%, Score: {hsp.score}")
print(f" Query: {hsp.query[:60]}")
print(f" Match: {hsp.match[:60]}")
print(f" Sbjct: {hsp.sbjct[:60]}")import subprocess
from Bio.Blast import NCBIXML
from Bio import SeqIO
# Local BLAST+ via subprocess (faster for large batches)
# Requires: makeblastdb and blastp installed (conda install -c bioconda blast)
# Step 1: Build a local database from a FASTA file
subprocess.run(
["makeblastdb", "-in", "ref_proteins.fasta", "-dbtype", "prot", "-out", "ref_db"],
check=True
)
# Step 2: Run blastp against local database
result = subprocess.run(
["blastp", "-query", "query.fasta", "-db", "ref_db",
"-outfmt", "5", # XML output for NCBIXML parsing
"-evalue", "1e-5",
"-num_threads", "4",
"-out", "blast_results.xml"],
check=True
)
# Step 3: Parse XML results
with open("blast_results.xml") as fh:
for blast_rec in NCBIXML.parse(fh):
print(f"Query: {blast_rec.query_id}")
for aln in blast_rec.alignments[:3]:
hsp = aln.hsps[0]
print(f" Hit: {aln.hit_id}, E={hsp.expect:.2e}, Id={hsp.identities}/{hsp.align_length}")Bio.Phylo reads Newick, Nexus, PhyloXML, and NeXML formats. Trees are represented as Clade objects with nested children. Key methods: find_clades() (generator over nodes), common_ancestor(), distance(), root_with_outgroup(), and prune(). Visualization uses matplotlib.
from Bio import Phylo
import io
# Parse a Newick tree from a string
newick = "((Homo_sapiens:0.01, Pan_troglodytes:0.012):0.08, (Mus_musculus:0.15, Rattus_norvegicus:0.14):0.12, Drosophila_melanogaster:0.85);"
tree = Phylo.read(io.StringIO(newick), "newick")
print(f"Number of terminals: {tree.count_terminals()}")
print(f"Total branch length: {tree.total_branch_length():.3f}")
print(f"Terminals: {[t.name for t in tree.get_terminals()]}")
# Root with outgroup and compute distances
tree.root_with_outgroup("Drosophila_melanogaster")
human = tree.find_any("Homo_sapiens")
chimp = tree.find_any("Pan_troglodytes")
mouse = tree.find_any("Mus_musculus")
print(f"\nDistance Human–Chimp: {tree.distance(human, chimp):.4f}")
print(f"Distance Human–Mouse: {tree.distance(human, mouse):.4f}")
# Traverse all internal nodes
for clade in tree.find_clades(order="level"):
if not clade.is_terminal():
children = [c.name or "internal" for c in clade.clades]
print(f" Internal node → {children}")from Bio import Phylo
import matplotlib.pyplot as plt
import io
newick = "((Homo_sapiens:0.01, Pan_troglodytes:0.012):0.08, (Mus_musculus:0.15, Rattus_norvegicus:0.14):0.12, Drosophila_melanogaster:0.85);"
tree = Phylo.read(io.StringIO(newick), "newick")
tree.root_with_outgroup("Drosophila_melanogaster")
tree.ladderize() # sort clades by size for a clean layout
# Annotate bootstrap support (if present in node labels)
for clade in tree.find_clades():
if clade.confidence is not None:
clade.name = f"{clade.confidence:.0f}"
fig, ax = plt.subplots(figsize=(8, 5))
Phylo.draw(tree, axes=ax, do_show=False, show_confidence=True)
ax.set_title("Primate + Outgroup Phylogeny")
plt.tight_layout()
plt.savefig("phylogeny.png", dpi=150, bbox_inches="tight")
print("Saved phylogeny.png")Bio.Align.PairwiseAligner replaces the legacy pairwise2 module (deprecated since Biopython 1.80). It supports local and global alignment, gap open/extend penalties, and any substitution matrix from Bio.Align.substitution_matrices.
from Bio.Align import PairwiseAligner, substitution_matrices
# Global protein alignment with BLOSUM62
aligner = PairwiseAligner()
aligner.mode = "global"
aligner.substitution_matrix = substitution_matrices.load("BLOSUM62")
aligner.open_gap_score = -11
aligner.extend_gap_score = -1
seq1 = "MTEYKLVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSY"
seq2 = "MTEYKLVVVGAVGVGKSALTIQLIQNHFVDEYDPTIEDSY" # G12V variant
alignments = aligner.align(seq1, seq2)
best = alignments[0]
print(f"Score: {best.score:.1f}")
print(f"Number of alignments: {aligner.score(seq1, seq2):.0f} (score only, faster)")
print(best) # formatted alignment string
identity = sum(a == b for a, b in zip(*best) if a != "-") / best.shape[1]
print(f"Identity: {identity*100:.1f}%")from Bio.Align import PairwiseAligner, substitution_matrices
# Local DNA alignment
aligner = PairwiseAligner()
aligner.mode = "local"
aligner.match_score = 2
aligner.mismatch_score = -1
aligner.open_gap_score = -5
aligner.extend_gap_score = -0.5
query = "ACGTACGTACGT"
subject = "TTTTACGTACGTACGTTTTT"
alignments = list(aligner.align(query, subject))
print(f"Local alignments found: {len(alignments)}")
best = alignments[0]
print(f"Score: {best.score}")
print(f"Aligned region in subject: [{best.coordinates[1][0]}:{best.coordinates[1][-1]}]")
print(best)
# Batch pairwise identity matrix
seqs = ["ACGTACGT", "ACGTATGT", "TTGTACGT", "ACGTACGG"]
n = len(seqs)
aligner2 = PairwiseAligner(mode="global", match_score=1, mismatch_score=-1)
print("\nPairwise identity matrix:")
for i in range(n):
row = []
for j in range(n):
score = aligner2.score(seqs[i], seqs[j])
max_len = max(len(seqs[i]), len(seqs[j]))
row.append(f"{score/max_len:.2f}")
print(" " + " ".join(row))SeqRecord stores sequence metadata and a list of SeqFeature objects. Features use zero-based, half-open FeatureLocation(start, end, strand) — the same convention as Python slicing. Use feature.location.extract(record.seq) to obtain the strand-correct subsequence. Compound locations (e.g., spliced exons) use CompoundLocation.
from Bio import SeqIO
# Extract all CDS protein translations from a GenBank record
for rec in SeqIO.parse("gene.gb", "genbank"):
for feat in rec.features:
if feat.type == "CDS":
gene = feat.qualifiers.get("gene", ["?"])[0]
product = feat.qualifiers.get("product", ["unknown"])[0]
cds_nt = feat.location.extract(rec.seq)
protein = cds_nt.translate(to_stop=True)
print(f"{gene} ({product}): {len(protein)} aa — {protein[:10]}...")A Clade is both a node and the subtree rooted at that node. Terminal clades (leaves) have .name set; internal clades may have .confidence (bootstrap) and .branch_length. Use tree.find_clades() with terminal=True/False to filter. The root is tree.root (also a Clade). Phylo.read() returns a Tree wrapper; tree.root gives the root Clade.
Goal: Download all BRCA1 orthologs from Vertebrata, align with MUSCLE, and build a neighbor-joining tree.
import subprocess
import time
from Bio import Entrez, SeqIO, Phylo, AlignIO
from Bio.Align import MultipleSeqAlignment
import matplotlib.pyplot as plt
Entrez.email = "your.email@example.com"
# Step 1: Search NCBI Protein for BRCA1 orthologs
handle = Entrez.esearch(
db="protein",
term="BRCA1[Gene] AND Vertebrata[Organism] AND RefSeq[Filter]",
retmax=20
)
results = Entrez.read(handle); handle.close()
ids = results["IdList"]
print(f"Found {results['Count']} sequences, using {len(ids)}")
# Step 2: Fetch sequences in FASTA format
handle = Entrez.efetch(db="protein", id=",".join(ids), rettype="fasta", retmode="text")
with open("brca1_orthologs.fasta", "w") as fh:
fh.write(handle.read())
handle.close()
print(f"Saved {len(ids)} sequences to brca1_orthologs.fasta")
time.sleep(1)
# Step 3: Align with MUSCLE (must be installed: conda install -c bioconda muscle)
subprocess.run(
["muscle", "-align", "brca1_orthologs.fasta", "-output", "brca1_aligned.fasta"],
check=True
)
alignment = AlignIO.read("brca1_aligned.fasta", "fasta")
print(f"Alignment: {len(alignment)} sequences × {alignment.get_alignment_length()} columns")
# Step 4: Build a neighbor-joining tree using Bio.Phylo + distance matrix
from Bio.Phylo.TreeConstruction import DistanceCalculator, DistanceTreeConstructor
calculator = DistanceCalculator("identity")
dm = calculator.get_distance(alignment)
constructor = DistanceTreeConstructor(calculator, method="nj")
tree = constructor.build_tree(alignment)
# Step 5: Root and save
tree.root_at_midpoint()
tree.ladderize()
Phylo.write(tree, "brca1_nj.nwk", "newick")
print("Saved NJ tree to brca1_nj.nwk")
# Step 6: Visualize
fig, ax = plt.subplots(figsize=(10, 8))
Phylo.draw(tree, axes=ax, do_show=False)
ax.set_title("BRCA1 Ortholog NJ Tree")
plt.tight_layout()
plt.savefig("brca1_tree.png", dpi=150, bbox_inches="tight")
print("Saved brca1_tree.png")Goal: Take a query FASTA, BLAST against NCBI nr, filter significant hits, retrieve full GenBank records for the top matches, and write a summary CSV.
import csv
import time
from Bio import SeqIO, Entrez
from Bio.Blast import NCBIWWW, NCBIXML
Entrez.email = "your.email@example.com"
# Step 1: Load query sequence
query_record = SeqIO.read("query.fasta", "fasta")
print(f"Query: {query_record.id} ({len(query_record.seq)} aa)")
# Step 2: Remote BLAST against nr protein database
print("Submitting BLAST job (may take 1-5 minutes)...")
result_handle = NCBIWWW.qblast(
"blastp", "nr", str(query_record.seq),
hitlist_size=20,
expect=1e-5,
word_size=6,
matrix_name="BLOSUM62"
)
# Step 3: Parse results and filter
E_THRESH = 1e-5
MIN_IDENTITY = 50.0
top_hits = []
blast_record = NCBIXML.read(result_handle)
for alignment in blast_record.alignments:
for hsp in alignment.hsps:
if hsp.expect > E_THRESH:
continue
identity_pct = hsp.identities / hsp.align_length * 100
if identity_pct < MIN_IDENTITY:
continue
accession = alignment.accession
top_hits.append({
"accession": accession,
"title": alignment.title[:80],
"length": alignment.length,
"score": hsp.score,
"evalue": hsp.expect,
"identity_pct": round(identity_pct, 1),
"coverage": round(hsp.align_length / blast_record.query_length * 100, 1),
})
print(f"Filtered to {len(top_hits)} significant hits")
# Step 4: Fetch full GenBank records for top 5 hits
accessions = [h["accession"] for h in top_hits[:5]]
time.sleep(1)
handle = Entrez.efetch(db="protein", id=",".join(accessions), rettype="gb", retmode="text")
gb_records = list(SeqIO.parse(handle, "genbank"))
handle.close()
SeqIO.write(gb_records, "top_blast_hits.gb", "genbank")
print(f"Saved {len(gb_records)} full GenBank records to top_blast_hits.gb")
# Step 5: Write CSV summary
with open("blast_summary.csv", "w", newline="") as fh:
writer = csv.DictWriter(fh, fieldnames=top_hits[0].keys())
writer.writeheader()
writer.writerows(top_hits)
print(f"Saved blast_summary.csv ({len(top_hits)} hits)")Goal: Read a FASTQ file, filter reads by mean quality and length, and output a FASTA for downstream alignment.
from Bio import SeqIO
from Bio.SeqUtils import gc_fraction
input_fastq = "raw_reads.fastq"
output_fasta = "filtered_reads.fasta"
MIN_QUAL = 20 # mean Phred quality
MIN_LEN = 100 # minimum read length
MAX_LEN = 300 # maximum read length
def passes_qc(record, min_q=MIN_QUAL, min_l=MIN_LEN, max_l=MAX_LEN):
quals = record.letter_annotations["phred_quality"]
mean_q = sum(quals) / len(quals)
length = len(record.seq)
return mean_q >= min_q and min_l <= length <= max_l
kept = 0
total = 0
with open(output_fasta, "w") as out_fh:
for record in SeqIO.parse(input_fastq, "fastq"):
total += 1
if passes_qc(record):
# Convert FASTQ → FASTA (drops quality scores)
SeqIO.write(record, out_fh, "fasta")
kept += 1
print(f"Total reads: {total:,}")
print(f"Passed QC: {kept:,} ({kept/total*100:.1f}%)")
print(f"Saved to: {output_fasta}")| Parameter | Module / Function | Default | Range / Options | Effect |
|---|---|---|---|---|
retmax | Entrez.esearch() | 20 | 1–10000 | Max IDs returned per search; use history server for >10000 |
hitlist_size | NCBIWWW.qblast() | 50 | 1–5000 | Max BLAST alignments returned |
expect | NCBIWWW.qblast() | 10.0 | 1e-100–1000 | E-value cutoff for BLAST hit reporting |
matrix_name | NCBIWWW.qblast() | "BLOSUM62" | "BLOSUM45", "BLOSUM80", "PAM250" | Substitution matrix; BLOSUM62 for general use |
mode | PairwiseAligner | "global" | "global", "local" | Needleman-Wunsch (global) vs Smith-Waterman (local) |
open_gap_score | PairwiseAligner | -1 | Negative float | Penalty for opening a gap; increase magnitude to penalize more |
extend_gap_score | PairwiseAligner | 0 | Negative float | Per-residue gap extension penalty |
method | DistanceTreeConstructor | "nj" | "nj", "upgma" | NJ (unrooted) vs UPGMA (rooted, assumes clock) |
format | Phylo.read/write() | required | "newick", "nexus", "phyloxml" | Tree file format |
rettype | Entrez.efetch() | required | "fasta", "gb", "xml" | Record format returned; pair with matching SeqIO format string |
Always set Entrez.email and respect rate limits. NCBI blocks IPs that exceed 3 req/s without a key. Add time.sleep(0.4) between requests in loops, or use Entrez.api_key for 10 req/s.
from Bio import Entrez
import time
Entrez.email = "your.email@institution.edu"
Entrez.api_key = "YOUR_API_KEY" # from https://www.ncbi.nlm.nih.gov/account/
# In any batch loop: time.sleep(0.11) # ~9 req/s with keyUse SeqIO.index() for large FASTA files instead of list(SeqIO.parse(...)). Loading all records into a list consumes O(N) memory; index() reads only the byte offsets and fetches records on demand.
# Bad for large files: loads everything into RAM
# records = {r.id: r for r in SeqIO.parse("genome.fasta", "fasta")}
# Good: on-disk index, O(1) access
from Bio import SeqIO
idx = SeqIO.index("genome.fasta", "fasta")
rec = idx["chr22"] # fetches only this recordUse PairwiseAligner not the legacy pairwise2. Bio.pairwise2 is deprecated since Biopython 1.80 and will be removed. PairwiseAligner is faster, supports substitution matrices directly, and returns Alignment objects with coordinate arrays.
Close Entrez handles immediately after reading. Handles are HTTP connections; leaving them open risks timeouts and resource exhaustion.
handle = Entrez.efetch(db="nucleotide", id="NM_007294", rettype="gb", retmode="text")
record = SeqIO.read(handle, "genbank")
handle.close() # do not skip thisPrefer local BLAST for batch searches. Remote NCBIWWW.qblast() is suitable for ad hoc queries but can queue for minutes on NCBI servers. For screening >100 sequences, build a local BLAST+ database with makeblastdb and call blastp/blastn via subprocess with -outfmt 5 (XML) for NCBIXML parsing.
Root trees before measuring distances. Phylo.distance() measures the sum of branch lengths along the path between two nodes. On an unrooted tree, the path is still unique, but midpoint-rooting or outgroup-rooting makes biological sense for visualizations and clade assertions.
Strip alignment gaps before building SeqRecord collections. BLAST and alignment results may include gap characters (-). Seq operations like .translate() will raise errors on gapped sequences; strip with seq.replace("-", "") or use ungap().
When to use: Downloading more than 500 records from NCBI — avoids URL length limits and keeps search results server-side.
from Bio import Entrez, SeqIO
import time
Entrez.email = "your.email@example.com"
# Search and store results on NCBI history server
handle = Entrez.esearch(db="nucleotide", term="16S rRNA[Gene] AND Bacteria[Organism]",
retmax=500, usehistory="y")
results = Entrez.read(handle); handle.close()
webenv = results["WebEnv"]
query_key = results["QueryKey"]
count = int(results["Count"])
print(f"Total: {count} sequences")
# Fetch in batches of 200
batch_size = 200
all_records = []
for start in range(0, min(count, 1000), batch_size):
handle = Entrez.efetch(
db="nucleotide", rettype="fasta", retmode="text",
retstart=start, retmax=batch_size,
webenv=webenv, query_key=query_key
)
batch = list(SeqIO.parse(handle, "fasta"))
handle.close()
all_records.extend(batch)
print(f" Fetched {len(all_records)}/{min(count, 1000)}")
time.sleep(0.5)
SeqIO.write(all_records, "16S_sequences.fasta", "fasta")
print(f"Saved {len(all_records)} sequences")When to use: Extract gene/CDS sequences from a GFF3 annotation paired with a reference FASTA.
from Bio import SeqIO
# Load genome FASTA into indexed dict
genome = SeqIO.to_dict(SeqIO.parse("genome.fasta", "fasta"))
# Parse GFF3 manually (Biopython does not have a native GFF3 parser;
# use the gffutils package for complex queries, or parse directly for simple cases)
genes = []
with open("annotation.gff3") as fh:
for line in fh:
if line.startswith("#") or not line.strip():
continue
fields = line.strip().split("\t")
if len(fields) < 9 or fields[2] != "CDS":
continue
chrom, _, feat_type, start, end, _, strand, _, attrs = fields
start, end = int(start) - 1, int(end) # GFF3 is 1-based
if chrom not in genome:
continue
seq = genome[chrom].seq[start:end]
if strand == "-":
seq = seq.reverse_complement()
attr_dict = dict(kv.split("=") for kv in attrs.strip().split(";") if "=" in kv)
gene_id = attr_dict.get("gene_id", attr_dict.get("ID", "unknown"))
genes.append((gene_id, seq, strand))
print(f"Extracted {len(genes)} CDS features")
for gid, seq, strand in genes[:3]:
print(f" {gid} ({strand}): {len(seq)} bp — protein: {seq.translate(to_stop=True)[:8]}...")When to use: Quickly assess sequence diversity within an alignment file (FASTA, PHYLIP, Clustal) and identify outlier sequences.
from Bio import AlignIO
import numpy as np
alignment = AlignIO.read("aligned_sequences.fasta", "fasta")
n = len(alignment)
names = [rec.id for rec in alignment]
length = alignment.get_alignment_length()
# Build identity matrix
identity_matrix = np.zeros((n, n))
for i in range(n):
for j in range(n):
if i == j:
identity_matrix[i, j] = 1.0
continue
matches = sum(
a == b and a != "-"
for a, b in zip(str(alignment[i].seq), str(alignment[j].seq))
)
aligned_cols = sum(a != "-" and b != "-"
for a, b in zip(str(alignment[i].seq), str(alignment[j].seq)))
identity_matrix[i, j] = matches / aligned_cols if aligned_cols else 0.0
print("Pairwise identity matrix:")
print(f"{'':20s} " + " ".join(f"{n[:8]:>8s}" for n in names))
for i, name in enumerate(names):
row = " ".join(f"{identity_matrix[i,j]*100:8.1f}" for j in range(n))
print(f"{name[:20]:20s} {row}")
# Find most divergent pair
min_id = np.min(identity_matrix[identity_matrix > 0])
idx = np.unravel_index(np.argmin(np.where(identity_matrix > 0, identity_matrix, 1)), identity_matrix.shape)
print(f"\nMost divergent pair: {names[idx[0]]} vs {names[idx[1]]} ({min_id*100:.1f}%)")| Problem | Cause | Solution |
|---|---|---|
urllib.error.HTTPError: 429 Too Many Requests | Exceeding NCBI rate limit (3 req/s) | Add time.sleep(0.4) between calls, or register a free API key at https://www.ncbi.nlm.nih.gov/account/ for 10 req/s |
RuntimeError: Too many requests were made without pausing | NCBI detects rapid-fire requests | Set Entrez.email (required), reduce loop frequency, batch IDs into comma-joined strings in a single efetch call |
Bio.Application.ApplicationError: blastall not found | Legacy blastall replaced by BLAST+ | Replace NcbiblastpCommandline with subprocess.run(["blastp", ...]) and BLAST+ tools |
AttributeError: 'PairwiseAligner' object has no attribute 'align' | Biopython < 1.78 | Upgrade: pip install --upgrade biopython (PairwiseAligner requires ≥ 1.72; stable from 1.78) |
Bio.pairwise2 deprecation warning | Using the legacy pairwise2 API | Replace with from Bio.Align import PairwiseAligner (removed in Biopython 1.84+) |
ValueError: Sequence contains letters not in the alphabet | Non-standard characters (e.g., N, ambiguity codes) in sequence before translate() | Strip or replace ambiguous bases; use seq.translate(table=1) which handles ambiguous codons |
Empty BLAST results / StopIteration from NCBIXML.read() | Empty or malformed XML; network timeout; query too short | Check query sequence length (>10 aa recommended); switch from .read() to .parse() and check blast_record.alignments length |
Phylo.draw() hangs or produces blank plot | Missing matplotlib backend | Call import matplotlib; matplotlib.use("Agg") before importing Phylo for headless environments |
TreeConstruction distance matrix dimension mismatch | Alignment contains duplicate IDs | Deduplicate SeqRecord IDs before constructing the alignment: {r.id: r for r in records}.values() |
Bio.pairwise2 module© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/genomics-bioinformatics/biopython-sequence-analysis of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Biopython Sequence Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Biopython Sequence Analysis this skilljaechang-hits/SciAgent-Skills | 370 | 1 repos | ~8.5k | Automated safety check: Pass | BSD-3-Clause | |
| Biopython Bioinformaticsaiming-lab/AutoResearchClaw | 15k | — | ~810 | Automated safety check: Pass | MIT | |
| Bio Write SequencesGPTomics/bioSkills | 1.2k | 3 repos | ~2.1k | Automated safety check: Pass | MIT | |
| Biopythondavila7/claude-code-templates | 32k | 13 repos | ~3.4k | Automated safety check: Pass | MIT | |
| BiopythonK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~4.3k | Automated safety check: Notes | MIT | |
| Biopythonlamm-mit/scienceclaw | 244 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 |
aiming-lab/AutoResearchClaw
Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
davila7/claude-code-templates
Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.
K-Dense-AI/scientific-agent-skills
Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).
lamm-mit/scienceclaw
Computational molecular biology library (sequence I/O, alignment, phylogenetics).
wu-yc/LabClaw
Production-ready phylogenetics and sequence analysis skill for alignment processing, tree analysis, and evolutionary metrics.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
Biopython sequence analysis: parse FASTA/FASTQ/GenBank/GFF (SeqIO), NCBI Entrez (esearch/efetch/elink), remote/local BLAST, pairwise/MSA alignment (PairwiseAligner, MUSCLE/ClustalW), phylogenetic…. Biopython Sequence Analysis is an agent skill from jaechang-hits/SciAgent-Skills. Biopython sequence analysis: parse FASTA/FASTQ/GenBank/GFF (SeqIO), NCBI Entrez (esearch/efetch/elink), remote/local BLAST, pairwise/MSA alignment (PairwiseAligner, MUSCLE/ClustalW), phylogenetic trees (Phylo).
Biopython Sequence Analysis fits situations like: gene family studies; comparative genomics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill biopython-sequence-analysis -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/biopython-sequence-analysis in jaechang-hits/SciAgent-Skills) into .claude/skills/biopython-sequence-analysis in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill biopython-sequence-analysis -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/biopython-sequence-analysis in jaechang-hits/SciAgent-Skills) into .agents/skills/biopython-sequence-analysis in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill biopython-sequence-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/biopython-sequence-analysis, .gemini/skills/biopython-sequence-analysis, .github/skills/biopython-sequence-analysis and .opencode/skills/biopython-sequence-analysis in your project.
Going by SKILL.md and its folder, Biopython Sequence Analysis needs the command-line tools its instructions call (pip and conda). Our summary lists: Python 3; A credential in YOUR_API_KEY.
SKILL.md names 3 domains. In commands or code: ncbi.nlm.nih.gov; the agent is likely to contact it when it follows the instructions. As links in the text: biopython.org and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Biopython Sequence Analysis is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.5k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Biopython Sequence Analysis: Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Bio Write Sequences (GPTomics/bioSkills, 1.2k stars), Biopython (davila7/claude-code-templates, 32k stars) and Biopython (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 370 GitHub stars. The repository holds 165 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.