Agent skill

Protein Sequence Msa

by google-deepmind in google-deepmind/science-skills

Performs multiple sequence alignment of proteins with EBI Clustal Omega.

Apache-2.0Auto-check: notesLegal & Compliance

Install Protein Sequence Msa

skills CLI
$ npx skills add google-deepmind/science-skills --skill protein-sequence-msa -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install google-deepmind/science-skills protein-sequence-msa --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/google-deepmind/science-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/protein_sequence_msa .claude/skills/protein-sequence-msa && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
protein-sequence-msa
GitHub stars
3.2k
Used in
1 other repo
Token cost
~1.5k tokens
SKILL.md length
647 words
Files
3 (incl. scripts, references)
Skills in repo
40
Repo updated
First seen
Licence
Apache-2.0

At a glance

Performs multiple sequence alignment of proteins with EBI Clustal Omega.

  • Works in 4 steps: uv: Read the uv skill and follow its… → User Notification: If… → .env file: Make sure the .env file… → …
  • You need to align multiple sequences to assess similarity
  • SKILL.md covers Prerequisites, Core Rules, Goal and Instructions, plus 1 more section
  • Runs Python scripts from its folder

What it does

Protein Sequence Msa is an agent skill from google-deepmind/science-skills. Performs multiple sequence alignment of proteins with EBI Clustal Omega. Use when you need to align multiple sequences to assess similarity, domain conservation, or key residue conservation. Supports up to 4000 sequences and a maximum file size of 4 MB. Do not use to search for homologous proteins in a database (use MMseqs2, BLAST), align non-protein sequences (DNA, RNA), perform structural alignment (use Foldseek, PyMOL), or if you only have a single sequence.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `scripts/msa_align.py`).

It sits in Legal & Compliance. The repository describes itself as: GDM Science Skills to speed up agentic scientific workflows with better grounding and higher token efficiency. Integrate insights from AlphaGenome, AFDB, UniProt and 30+ other… The licence is Apache-2.0.

When your agent uses it

  • You need to align multiple sequences to assess similarity
  • Domain conservation
  • Key residue conservation
  • Search for homologous proteins in a database (use MMseqs2

Example prompts

  • “Use the protein-sequence-msa skill to perform multiple sequence alignment of proteins with EBI Clustal Omega”
  • “/protein-sequence-msa”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. uv: Read the uv skill and follow its Setup instructions to ensure
  2. User Notification: If .licenses/protein_sequence_msa_LICENSE.txt does
  3. .env file: Make sure the .env file exists in your home directory.
  4. USER_EMAIL: Required by the wrapper script for Clustal Omega job

What it can do on your machine

Read from SKILL.md and the folder at commit 6883275. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ebi.ac.uk

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Protein Sequence Msa loads about 1.5k tokens when it runs, and up to ~2.7k if it reads all its reference files. Until then it costs about 122 tokens; SKILL.md has 647 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:23
    3.  **`.env` file**: Make sure the `.env` file exists in your home directory.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from google-deepmind/science-skills at commit 6883275, republished under its Apache-2.0 licence (© google-deepmind). 647 words, ~1,529 tokens.

Download SKILL.mdSave it as .claude/skills/protein-sequence-msa/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
protein-sequence-msa
description
Performs multiple sequence alignment of proteins with EBI Clustal Omega. Use when you need to align multiple sequences to assess similarity, domain conservation, or key residue conservation. Supports up to 4000 sequences and a maximum file size of 4 MB. Do not use to search for homologous proteins in a database (use MMseqs2, BLAST), align non-protein sequences (DNA, RNA), perform structural alignment (use Foldseek, PyMOL), or if you only have a single sequence.

Prerequisites

  1. uv: Read the uv skill and follow its Setup instructions to ensure uv is installed and on PATH.
  2. User Notification: If .licenses/protein_sequence_msa_LICENSE.txt does not already exist in the workspace root directory then (1) prominently notify the user to check the terms at https://www.ebi.ac.uk/jdispatcher/msa/clustalo and https://www.ebi.ac.uk/about/terms-of-use/, then (2) create the file recording the notification text and timestamp.
  3. .env file: Make sure the .env file exists in your home directory. Create one if it does not exist.
  4. USER_EMAIL: Required by the wrapper script for Clustal Omega job tracking (recommended by the EBI). You MUST use the safe credentials protocol in the credentials skill to check for and request this credential if this skill looks relevant to the user's request.

Core Rules

  • Use the Wrapper: ALWAYS execute the alignment using scripts/msa_align.py rather than writing your own curl or custom Python requests. The script automatically enforces the required rate limit to respect EBI's Terms of Use.
  • Notification: If this skill is used, ensure this is mentioned in the output.
  • Always state the method: Every report must clearly state that the alignment was performed using EBI Clustal Omega.
  • No Hallucinations: Do NOT invent alignments or conservation metrics. Report only what is present in the alignment file.

Goal

Take a file containing multiple protein sequences in FASTA format, perform multiple sequence alignment using the EBI Clustal Omega API, save the resulting alignment locally for future programmatic analysis, and interpret the results towards addressing the user's specific research objective (e.g., assessing similarity, identifying conserved domains, or analyzing key residues).

Instructions

  1. Prepare Input File: The input must be a plain text file containing two or more protein sequences in FASTA format. Each sequence header must start with a > symbol. Example:

    >Sequence_1_Name
    MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQ
    QRLIFAGKQLEDGRTLSDYNIQKESTLHLVLRLRGG
    >Sequence_2_Name
    MQIFVKTLTGKTITLEVEPSDTIENVKAKIQDKEGIPPDQ
    QRLIFAGKQLEDGRTLSDYNIQKESTLHLVLRLRGG
  2. Execute Alignment: Run the alignment script:

    bash
    uv run scripts/msa_align.py <INPUT_FASTA> -o <OUTPUT_FILE>

    Always specify the output file with -o or --output.

  3. Interpret and Report Results: Analyze the Clustal Omega alignment by selecting metrics and mapping strategies aligned with the research objective. Note that while Clustal Omega produces a Global Alignment, pairwise metrics can be extracted to evaluate specific relationships within the set:

    • Identity Metric Options: The choice of denominator determines how insertions/deletions (gaps) affect the final percentage. Select the most appropriate calculation based on the biological context:
      • Pairwise - Sequence Coverage: (Identical Residue Matches) / (Length of Shorter Sequence). Use when determining if a specific domain or fragment is fully preserved within a larger protein. This ignores gaps in the longer sequence, focusing purely on the "content" of the shorter one.
      • Pairwise - Global Identity: (Identical Residue Matches) / (Total Alignment Columns). Use when comparing full-length sequences of similar expected length. This is the most conservative metric; it penalizes for all gaps (indels) introduced by any sequence in the MSA.
      • Pairwise - Overlap Identity: (Identical Residue Matches) / (Total Alignment Columns - Terminal Gaps). Use when comparing a fragment to a full-length protein or when sequences have long unaligned "tails." This focuses on similarity only where the sequences physically overlap.
      • Multisequence - Conservation Index: (Fully Conserved Columns) / (Total Alignment Columns). Use for quantifying the percentage of residues that are 100% identical across the entire alignment set. This identifies the core evolutionary signature of the protein family.
    • Feature Mapping: Leverage known biological data from specific sequences to ground the analysis:
      • Knowledge Gathering: Identify relevant known sites or regions (e.g., catalytic residues, binding motifs) from your input or via external tools.
      • Coordinate Projection: Map these features onto the corresponding Column Indices of the alignment.
      • Targeted Discussion: Use these columns to drive the assessment:
        • Local Conservation: Analyze if the known functional residues are invariant across the set.
        • Region-Specific Metrics: Calculate identity/similarity specifically within the mapped functional regions rather than the whole sequence.
        • Goal Contribution: Discuss how this data contributes to your goal, e.g. using conservation to corroborate a prediction or divergence to reject a functional hypothesis.
Show full SKILL.md (5 more words)Show less

References

© google-deepmind, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in skills/protein_sequence_msa of google-deepmind/science-skills.

  • SKILL.md
  • references/citation.bib
  • scripts/msa_align.py

Open the folder on GitHubat commit 6883275

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in google-deepmind/science-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Protein Sequence Msa next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Protein Sequence Msa compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Protein Sequence Msa this skillgoogle-deepmind/science-skills3.2k1 repos~1.5kAutomated safety check: NotesApache-2.0
Paper to Chinese Patent DrafterYuan1z0825/nature-skills46k1 repos~1.1kAutomated safety check: PassApache-2.0
C15tc15t/c15t1.9k1 repos~1.6kAutomated safety check: PassApache-2.0
Contract Reviewevolsb/claude-legal-skill4611 repos~3.6kAutomated safety check: PassMIT
Legal Clinic Client Intakeanthropics/claude-for-legal9.6k3 repos~3.2kAutomated safety check: PassApache-2.0
Paper To Cn Patentsnipp-zha/Paper-to-patent-Skill1051 repos~959Automated safety check: PassNone

Similar skills

  • Paper to Chinese Patent Drafter

    Yuan1z0825/nature-skills

    Drafts Chinese invention patent applications and technical disclosures from research papers or inventor materials, tying each claim feature to source evidence.

    46k GitHub starsUsed in 1 repo~1.1k tokens
    Legal & ComplianceAuto-check passed
  • C15t

    c15t/c15t

    Work with c15t consent management docs, APIs, and integrations for Next.js, React, and JavaScript.

    1.9k GitHub starsUsed in 1 repo~1.6k tokens
    Legal & ComplianceAuto-check passed
  • Contract Review

    evolsb/claude-legal-skill

    Review legal contracts, NDAs, employment agreements, SaaS terms, and M&A documents.

    461 GitHub starsUsed in 1 repo~3.6k tokens
    Legal & ComplianceAuto-check passed
  • Legal Clinic Client Intake

    anthropics/claude-for-legal

    Official

    Structures a legal clinic client intake interview and produces a case summary with cross-area issue spotting, conflict flags and triage classification.

    9.6k GitHub starsUsed in 3 repos~3.2k tokens
    Legal & ComplianceAuto-check passed
  • Paper To Cn Patent

    snipp-zha/Paper-to-patent-Skill

    Convert scientific papers, theses, technical reports, source code, figures, or research manuscripts into evidence-grounded Chinese invention patent drafts.

    105 GitHub starsUsed in 1 repo~959 tokens
    Legal & ComplianceAuto-check passed
  • Commercial Legal Pl

    apiotrowski-afk/commercial-legal-pl

    Skill do analizy i tworzenia umów według polskiego prawa, ze szczególnym uwzględnieniem umów B2B, IP i IT (body leasing, NDA, wdrożenia, SaaS, przeniesienie praw autorskich, ugody).

    176 GitHub starsUsed in 1 repo~4.1k tokens
    Legal & ComplianceAuto-check passed

More from google-deepmind/science-skills

All 40 skills in this repo
  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Auto-check: notes
  • Chembl Database

    google-deepmind/science-skills

    Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures.

    3.2k GitHub starsUsed in 2 repos~2.9k tokens
    Auto-check passed
  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Auto-check: notes
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Auto-check: notes

Questions about Protein Sequence Msa

What does Protein Sequence Msa do?

Performs multiple sequence alignment of proteins with EBI Clustal Omega. Protein Sequence Msa is an agent skill from google-deepmind/science-skills. Performs multiple sequence alignment of proteins with EBI Clustal Omega.

When should I use Protein Sequence Msa?

Protein Sequence Msa fits situations like: you need to align multiple sequences to assess similarity; domain conservation; key residue conservation; search for homologous proteins in a database (use MMseqs2.

How do I install Protein Sequence Msa in Claude Code?

Run `npx skills add google-deepmind/science-skills --skill protein-sequence-msa -a claude-code`. Or copy the skill folder (skills/protein_sequence_msa in google-deepmind/science-skills) into .claude/skills/protein-sequence-msa in your project. Claude Code loads it when a task matches its description.

How do I install Protein Sequence Msa in Codex?

Run `npx skills add google-deepmind/science-skills --skill protein-sequence-msa -a codex`. Or copy the skill folder (skills/protein_sequence_msa in google-deepmind/science-skills) into .agents/skills/protein-sequence-msa in your project. Codex loads it when a task matches its description.

Can I use Protein Sequence Msa in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add google-deepmind/science-skills --skill protein-sequence-msa -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/protein-sequence-msa, .gemini/skills/protein-sequence-msa, .github/skills/protein-sequence-msa and .opencode/skills/protein-sequence-msa in your project.

What does Protein Sequence Msa need to run?

Going by SKILL.md and its folder, Protein Sequence Msa needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Protein Sequence Msa access the network?

SKILL.md names 1 domain. As links in the text: ebi.ac.uk. This is read from the text; nothing was executed.

Is Protein Sequence Msa safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Protein Sequence Msa use?

Protein Sequence Msa is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Protein Sequence Msa use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Protein Sequence Msa?

Skills that share tags, products or a category with Protein Sequence Msa: Paper to Chinese Patent Drafter (Yuan1z0825/nature-skills, 46k stars), C15t (c15t/c15t, 1.9k stars), Contract Review (evolsb/claude-legal-skill, 461 stars) and Legal Clinic Client Intake (anthropics/claude-for-legal, 9.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Protein Sequence Msa?

google-deepmind (a GitHub organization) maintains it in google-deepmind/science-skills, which has 3,216 GitHub stars. The repository holds 40 skills in this directory. The repository was last updated on September 15, 2026.

Source: google-deepmind/science-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.