Agent skill

Uniprot Protein Retrieval

by InternScience in InternScience/scp

Retrieve protein sequences and functional information from UniProt database by protein name, enabling protein analysis and bioinformatics workflows.

MITAuto-check passedResearch & Science

Install Uniprot Protein Retrieval

skills CLI
$ npx skills add InternScience/scp --skill uniprot-protein-retrieval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install InternScience/scp uniprot-protein-retrieval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/InternScience/scp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/uniprot-protein-retrieval .claude/skills/uniprot-protein-retrieval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
uniprot-protein-retrieval
GitHub stars
170
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
333 words
Files
1
Skills in repo
73
Repo updated
First seen
Licence
MIT

At a glance

Retrieve protein sequences and functional information from UniProt database by protein name, enabling protein analysis and bioinformatics workflows.

  • Works in 2 steps: MCP Server Definition → Protein Sequence Retrieval Workflow
  • Tasks that involve Protein structure and design
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Bioinformatics

What it does

Uniprot Protein Retrieval is an agent skill from InternScience/scp. Retrieve protein sequences and functional information from UniProt database by protein name, enabling protein analysis and bioinformatics workflows.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Protein structure and design and Bioinformatics. It works with UniProt. The licence is MIT.

When your agent uses it

  • Tasks that involve Protein structure and design
  • Tasks that involve Bioinformatics

Example prompts

  • “/uniprot-protein-retrieval”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. MCP Server Definition
  2. Protein Sequence Retrieval Workflow

What it can do on your machine

Read from SKILL.md and the folder at commit cea5398. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Uniprot Protein Retrieval loads about 2k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 333 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from InternScience/scp at commit cea5398, republished under its MIT licence (© InternScience). 333 words, ~2,025 tokens.

Download SKILL.mdSave it as .claude/skills/uniprot-protein-retrieval/SKILL.md (or your agent's skills folder).
name
uniprot-protein-retrieval
description
Retrieve protein sequences and functional information from UniProt database by protein name, enabling protein analysis and bioinformatics workflows.
license
MIT license
metadata.skill-author
PJLab

UniProt Protein Sequence Retrieval

Usage

1. MCP Server Definition

Use the standard MCP client pattern for Origene-UniProt server.

2. Protein Sequence Retrieval Workflow

This workflow retrieves protein sequences and associated information from the UniProt database using protein names or identifiers.

Workflow Steps:

  1. Query by Protein Name - Search UniProt using common protein names
  2. Retrieve Sequence Data - Get amino acid sequence and metadata

Implementation:

python
from mcp.client.streamable_http import streamablehttp_client
from mcp import ClientSession
import json

class OrigeneClient:
    def __init__(self, server_url: str):
        self.server_url = server_url
        self.session = None

    async def connect(self):
        try:
            self.transport = streamablehttp_client(
                url=self.server_url,
                headers={"SCP-HUB-API-KEY": "<your-api-key>"}
            )
            self.read, self.write, self.get_session_id = await self.transport.__aenter__()
            self.session_ctx = ClientSession(self.read, self.write)
            self.session = await self.session_ctx.__aenter__()
            await self.session.initialize()
            print("✓ Connected to Origene-UniProt")
            return True
        except Exception as e:
            print(f"✗ Connection failed: {e}")
            return False

    async def disconnect(self):
        try:
            if self.session:
                await self.session_ctx.__aexit__(None, None, None)
            if hasattr(self, 'transport'):
                await self.transport.__aexit__(None, None, None)
            print("✓ Disconnected")
        except Exception as e:
            print(f"✗ Disconnect error: {e}")

    def parse_result(self, result):
        try:
            if hasattr(result, 'content') and result.content:
                content = result.content[0]
                if hasattr(content, 'text'):
                    return json.loads(content.text)
            return str(result)
        except Exception as e:
            return {"error": f"Parse error: {e}", "raw": str(result)}

## Initialize client
client = OrigeneClient("https://scp.intern-ai.org.cn/api/v1/mcp/10/Origene-UniProt")
if not await client.connect():
    print("Connection failed")
    return

## Step 1: Retrieve protein sequence by name
protein_name = "insulin"  # Can be common name, gene symbol, or UniProt ID

result = await client.session.call_tool(
    "get_protein_sequence_by_name",
    arguments={
        "protein_name": protein_name
    }
)

result_data = client.parse_result(result)

## Display results
print(f"\nProtein: {protein_name}")
print("=" * 80)

if "sequence" in result_data:
    sequence = result_data["sequence"]
    print(f"Amino Acid Sequence ({len(sequence)} residues):")
    print(sequence)

    # Format sequence in blocks of 60
    print("\nFormatted Sequence:")
    for i in range(0, len(sequence), 60):
        position = i + 1
        block = sequence[i:i+60]
        print(f"{position:6d} {block}")

if "uniprot_id" in result_data:
    print(f"\nUniProt ID: {result_data['uniprot_id']}")

if "protein_names" in result_data:
    print(f"Protein Names: {result_data['protein_names']}")

if "organism" in result_data:
    print(f"Organism: {result_data['organism']}")

if "function" in result_data:
    print(f"Function: {result_data['function'][:200]}...")

await client.disconnect()
Extended Example: Multiple Protein Retrieval
python
## Retrieve multiple proteins
protein_list = ["p53", "BRCA1", "insulin", "hemoglobin"]

sequences = {}
for protein in protein_list:
    result = await client.session.call_tool(
        "get_protein_sequence_by_name",
        arguments={"protein_name": protein}
    )
    data = client.parse_result(result)

    if "sequence" in data:
        sequences[protein] = {
            "sequence": data["sequence"],
            "length": len(data["sequence"]),
            "uniprot_id": data.get("uniprot_id", "N/A")
        }

## Display summary
print("\nProtein Sequence Summary:")
print(f"{'Protein':<15} {'UniProt ID':<12} {'Length':<10}")
print("-" * 40)
for name, info in sequences.items():
    print(f"{name:<15} {info['uniprot_id']:<12} {info['length']:<10}")
Tool Description

Origene-UniProt Server:

  • get_protein_sequence_by_name: Retrieve protein sequence from UniProt database
    • Args:
      • protein_name (str): Protein common name, gene symbol, or UniProt ID
    • Returns:
      • sequence (str): Amino acid sequence (one-letter code)
      • uniprot_id (str): UniProt accession number
      • protein_names (str): Official and alternative protein names
      • organism (str): Source organism
      • function (str): Protein function description
      • length (int): Sequence length in residues
      • mass (float): Molecular mass (Da)
Input/Output

Input:

  • protein_name: Protein identifier (flexible format)
    • Examples: "insulin", "P53", "BRCA1", "P01308"
    • Supports: common names, gene symbols, UniProt IDs

Output:

  • Protein sequence and comprehensive metadata
  • Ready for downstream analysis (alignment, structure prediction, etc.)
Supported Query Types
  1. Common Names: "insulin", "hemoglobin", "actin"
  2. Gene Symbols: "TP53", "BRCA1", "EGFR"
  3. UniProt IDs: "P01308", "P04637"
  4. Protein Families: "kinase", "protease" (returns multiple entries)
Applications

Use retrieved sequences for:

  • Protein alignment and homology analysis
  • Structure prediction (AlphaFold, ESM Fold)
  • Primer design for cloning
  • Antibody epitope mapping
  • Conservation analysis
  • Mutation impact assessment
  • Phylogenetic studies
Integration with Other Workflows

Combine with:

  1. Protein BLAST → Find homologs
  2. InterProScan → Identify domains
  3. AlphaFold → Predict 3D structure
  4. STRING → Find protein interactions
  5. OpenTargets → Link to diseases
Example: Complete Protein Analysis Pipeline
python
## 1. Retrieve sequence
result = await uniprot_client.session.call_tool(
    "get_protein_sequence_by_name",
    arguments={"protein_name": "BRCA1"}
)
sequence = uniprot_client.parse_result(result)["sequence"]

## 2. Find similar proteins (BLAST)
result = await biotools_client.session.call_tool(
    "blast_search",
    arguments={
        "sequence": sequence,
        "evalue": 1e-10,
        "max_hits": 20
    }
)
homologs = biotools_client.parse_result(result)

## 3. Identify domains (InterProScan)
result = await biotools_client.session.call_tool(
    "interproscan_analyze",
    arguments={
        "sequence": sequence,
        "databases": ["Pfam", "SMART"]
    }
)
domains = biotools_client.parse_result(result)

## 4. Get disease associations (OpenTargets)
result = await opentargets_client.session.call_tool(
    "get_target_associated_diseases",
    arguments={"gene_symbol": "BRCA1"}
)
diseases = opentargets_client.parse_result(result)

print(f"Complete analysis for BRCA1:")
print(f"- Sequence length: {len(sequence)} amino acids")
print(f"- Homologs found: {len(homologs)}")
print(f"- Functional domains: {len(domains)}")
print(f"- Associated diseases: {len(diseases)}")
Error Handling

Common issues:

  • Protein not found: Check spelling, try alternative names or UniProt ID
  • Multiple matches: Use more specific identifier (UniProt ID preferred)
  • No sequence available: Some entries may lack sequence data
  • Network timeout: Retry with exponential backoff
Data Quality Notes
  • UniProt is manually curated (Swiss-Prot) and computationally annotated (TrEMBL)
  • Sequence quality: Swiss-Prot entries are highly reliable
  • Updates: UniProt is updated regularly; sequences may change
  • Isoforms: Multiple isoforms may exist; canonical sequence is returned by default

© InternScience, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/uniprot-protein-retrieval of InternScience/scp.

Open the folder on GitHubat commit cea5398

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in InternScience/scp, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Uniprot Protein Retrieval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Uniprot Protein Retrieval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Uniprot Protein Retrieval this skillInternScience/scp1701 repos~2kAutomated safety check: PassMIT
Ggetdavila7/claude-code-templates33k10 repos~6.3kAutomated safety check: PassMIT
Tooluniverseynulihao/AgentSkillOS6182 repos~2.5kAutomated safety check: PassNone
Ggetaipoch/medical-research-skills1.9k—~816Automated safety check: PassMIT
Uniprot Protein Databasejaechang-hits/SciAgent-Skills3741 repos~3.4kAutomated safety check: PassCC-BY-4.0
Alphafold Database Fetch And Analyzegoogle-deepmind/science-skills3.2k2 repos~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 10 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • Tooluniverse

    ynulihao/AgentSkillOS

    A skill your agent uses when working with scientific research tools and workflows across bioinformatics, cheminformatics, genomics, structural biology, proteomics, and drug discovery.

    618 GitHub starsUsed in 2 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Gget

    aipoch/medical-research-skills

    Unified CLI/Python interface for querying genomic, proteomic, structure, and expression data across 20+ bioinformatics databases; use when you need fast, scriptable retrieval by gene/protein IDs or…

    1.9k GitHub stars~816 tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Uniprot Protein Database

    jaechang-hits/SciAgent-Skills

    Query UniProt REST API: search by gene/protein name, fetch FASTA, map IDs (Ensembl, PDB, RefSeq), access Swiss-Prot annotations.

    374 GitHub starsUsed in 1 repo~3.4k tokens
    Research & ScienceAuto-check passed
  • Alphafold Database Fetch And Analyze

    google-deepmind/science-skills

    Retrieve and analyze AlphaFold predicted structures for a protein.

    3.2k GitHub starsUsed in 2 repos~1.2k tokens
    Research & ScienceAuto-check passed
  • Biomarker Database Analysis

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    A skill your agent uses when a researcher needs to query biomedical databases for biomarker discovery, build target profiles from UniProt/Open Targets/STRING, rank biomarker candidates by evidence…

    274 GitHub stars~1.1k tokensUpdated 9 days ago
    Research & ScienceAuto-check passed

More from InternScience/scp

All 73 skills in this repo
  • Given an rsID, query multiple databases (dbSNP, FAVOR, GWAS Catalog, ClinVar, gnomAD, PharmGKB, ClinGen) for comprehensive annotation.

    170 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Calculate atmospheric parameters including Coriolis parameter, geostrophic wind, heat index, potential temperature, and dewpoint for meteorology and climate science.

    170 GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Biomedical Web Search

    InternScience/scp

    Search biomedical literature and web content using Tavily search engine for research and clinical information.

    170 GitHub starsUsed in 1 repo~598 tokens
    Auto-check passed
  • Calculate buoyancy forces and acceleration for fluid mechanics and hydrodynamics analysis.

    170 GitHub starsUsed in 1 repo~540 tokens
    Auto-check passed
  • Capacitance Calculation

    InternScience/scp

    Calculate electrical capacitance from geometric parameters and dielectric properties for circuit design.

    170 GitHub starsUsed in 1 repo~537 tokens
    Auto-check passed
  • Chembl Molecule Search

    InternScience/scp

    Search ChEMBL database for molecule information by name to retrieve bioactivity data and chemical structures.

    170 GitHub starsUsed in 1 repo~757 tokens
    Auto-check passed

Works with

Questions about Uniprot Protein Retrieval

What does Uniprot Protein Retrieval do?

Retrieve protein sequences and functional information from UniProt database by protein name, enabling protein analysis and bioinformatics workflows. Uniprot Protein Retrieval is an agent skill from InternScience/scp. Retrieve protein sequences and functional information from UniProt database by protein name, enabling protein analysis and bioinformatics workflows.

When should I use Uniprot Protein Retrieval?

Uniprot Protein Retrieval fits situations like: tasks that involve Protein structure and design; tasks that involve Bioinformatics.

How do I install Uniprot Protein Retrieval in Claude Code?

Run `npx skills add InternScience/scp --skill uniprot-protein-retrieval -a claude-code`. Or copy the skill folder (skills/uniprot-protein-retrieval in InternScience/scp) into .claude/skills/uniprot-protein-retrieval in your project. Claude Code loads it when a task matches its description.

How do I install Uniprot Protein Retrieval in Codex?

Run `npx skills add InternScience/scp --skill uniprot-protein-retrieval -a codex`. Or copy the skill folder (skills/uniprot-protein-retrieval in InternScience/scp) into .agents/skills/uniprot-protein-retrieval in your project. Codex loads it when a task matches its description.

Can I use Uniprot Protein Retrieval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add InternScience/scp --skill uniprot-protein-retrieval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/uniprot-protein-retrieval, .gemini/skills/uniprot-protein-retrieval, .github/skills/uniprot-protein-retrieval and .opencode/skills/uniprot-protein-retrieval in your project.

What does Uniprot Protein Retrieval need to run?

SKILL.md names no scripts, command-line tools or credentials: Uniprot Protein Retrieval is instructions for the agent only. Our summary lists: Python 3.

Does Uniprot Protein Retrieval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Uniprot Protein Retrieval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Uniprot Protein Retrieval use?

Uniprot Protein Retrieval is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Uniprot Protein Retrieval use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Uniprot Protein Retrieval?

Skills that share tags, products or a category with Uniprot Protein Retrieval: Gget (davila7/claude-code-templates, 33k stars), Tooluniverse (ynulihao/AgentSkillOS, 618 stars), Gget (aipoch/medical-research-skills, 1.9k stars) and Uniprot Protein Database (jaechang-hits/SciAgent-Skills, 374 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Uniprot Protein Retrieval?

InternScience (a GitHub organization) maintains it in InternScience/scp, which has 170 GitHub stars. The repository holds 73 skills in this directory. The repository was last updated on June 3, 2026.

Source: InternScience/scp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.