Agent skill

Comprehensive Protein Analysis

by InternScience in InternScience/scp

Comprehensive protein analysis combining InterProScan domain identification with BLAST similarity search to provide complete functional and evolutionary annotation.

MITAuto-check passedResearch & Science

Install Comprehensive Protein Analysis

skills CLI
$ npx skills add InternScience/scp --skill comprehensive-protein-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install InternScience/scp comprehensive-protein-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/InternScience/scp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/comprehensive-protein-analysis .claude/skills/comprehensive-protein-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
comprehensive-protein-analysis
GitHub stars
169
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
436 words
Files
1
Skills in repo
73
Repo updated
First seen
Licence
MIT

At a glance

Comprehensive protein analysis combining InterProScan domain identification with BLAST similarity search to provide complete functional and evolutionary annotation.

  • Works in 2 steps: MCP Server Definition → Comprehensive Protein Analysis Workflow
  • Tasks that involve Vector databases
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Comprehensive Protein Analysis is an agent skill from InternScience/scp. Comprehensive protein analysis combining InterProScan domain identification with BLAST similarity search to provide complete functional and evolutionary annotation.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Vector databases. The licence is MIT.

When your agent uses it

  • Tasks that involve Vector databases

Example prompts

  • “/comprehensive-protein-analysis”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. MCP Server Definition
  2. Comprehensive Protein Analysis Workflow

What it can do on your machine

Read from SKILL.md and the folder at commit cea5398. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Comprehensive Protein Analysis loads about 2k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 436 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from InternScience/scp at commit cea5398, republished under its MIT licence (© InternScience). 436 words, ~2,021 tokens.

Download SKILL.mdSave it as .claude/skills/comprehensive-protein-analysis/SKILL.md (or your agent's skills folder).
name
comprehensive-protein-analysis
description
Comprehensive protein analysis combining InterProScan domain identification with BLAST similarity search to provide complete functional and evolutionary annotation.
license
MIT license
metadata.skill-author
PJLab

Comprehensive Protein Analysis

Usage

1. MCP Server Definition

Use the same BioInfoToolsClient class as defined in the protein-blast-search skill.

2. Comprehensive Protein Analysis Workflow

This workflow combines InterProScan domain analysis with BLAST similarity search to provide a complete functional and evolutionary annotation of a protein sequence.

Workflow Steps:

  1. Validate Input - Check protein sequence format
  2. Run InterProScan - Identify functional domains and GO terms
  3. Run BLAST Search - Find similar sequences and homologs
  4. Integrate Results - Combine domain and homology information for comprehensive annotation

Implementation:

python
from datetime import timedelta

## Initialize client
client = BioInfoToolsClient(
    "https://scp.intern-ai.org.cn/api/v1/mcp/17/BioInfo-Tools",
    "<your-api-key>"
)

if not await client.connect():
    print("connection failed")
    exit()

## Input: Protein sequence to analyze
protein_sequence = """
MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN
"""

sequence_id = "INS_HUMAN"

## Step 1, 2 & 3: Run comprehensive analysis (InterProScan + BLAST)
result = await client.session.call_tool(
    "analyze_protein",
    arguments={
        "sequence": protein_sequence.strip(),
        "sequence_id": sequence_id,
        "databases": ["Pfam"],     # InterProScan databases
        "evalue": 1e-5,             # BLAST E-value threshold (more stringent)
        "max_hits": 10              # BLAST max hits
    },
    read_timeout_seconds=timedelta(seconds=1200)  # Allow up to 20 minutes
)

## Step 4: Parse and display comprehensive results
result_data = client.parse_result(result)

print(f"{'='*80}")
print(f"Comprehensive Protein Analysis: {sequence_id}")
print(f"{'='*80}\n")

# InterProScan Results
ips_result = result_data.get("interproscan", {})
if ips_result.get("success"):
    ips_data = ips_result.get("results", {})
    domains = ips_data.get('domains', [])
    go_terms = ips_data.get('go_terms', [])

    print("=== DOMAIN ANALYSIS (InterProScan) ===")
    print(f"Execution time: {ips_result.get('time_seconds', '?')} seconds")
    print(f"Domains found: {len(domains)}")
    print(f"GO annotations: {len(go_terms)}\n")

    if domains:
        print("Functional Domains:")
        for domain in domains:
            print(f"  • {domain.get('name', 'N/A')} ({domain.get('database', 'N/A')})")
            if domain.get('description'):
                print(f"    Description: {domain.get('description')}")
            locations = domain.get('locations', [])
            if locations:
                loc = locations[0]
                print(f"    Position: {loc.get('start')}-{loc.get('end')} aa")
        print()

    if go_terms:
        print("Gene Ontology Annotations:")
        for go in go_terms[:5]:  # Show top 5
            print(f"  • {go.get('id', 'N/A')}: {go.get('name', 'N/A')}")
            print(f"    Category: {go.get('category', 'N/A')}")
        if len(go_terms) > 5:
            print(f"  ... and {len(go_terms) - 5} more")
        print()
else:
    print(f"❌ InterProScan failed: {ips_result.get('error', 'Unknown')}\n")

# BLAST Results
blast_result = result_data.get("blast", {})
if blast_result.get("success"):
    hits = blast_result.get('hits', [])

    print("=== HOMOLOGY SEARCH (BLAST) ===")
    print(f"Execution time: {blast_result.get('time_seconds', '?')} seconds")
    print(f"Similar sequences found: {blast_result.get('total_hits', 0)}")
    print(f"E-value threshold: {1e-5}\n")

    if hits:
        print("Top Homologous Proteins:")
        for i, hit in enumerate(hits[:5], 1):
            print(f"  {i}. {hit['uniprot_id']} - {hit.get('organism', 'N/A')}")
            print(f"     Description: {hit['description']}")
            print(f"     Identity: {hit['identity_percent']:.1f}%, E-value: {hit['evalue']:.2e}")
        if len(hits) > 5:
            print(f"  ... and {len(hits) - 5} more matches")
        print()
    else:
        print("No significant homologs found (E-value threshold may be too stringent)\n")
else:
    print(f"❌ BLAST failed: {blast_result.get('error', 'Unknown')}\n")

# Summary
print("=== FUNCTIONAL SUMMARY ===")
if domains:
    print(f"Protein Family: {domains[0].get('name', 'Unknown')}")
if hits:
    most_similar = hits[0]
    print(f"Most Similar Protein: {most_similar['uniprot_id']} ({most_similar['identity_percent']:.1f}% identity)")
    print(f"Organism: {most_similar.get('organism', 'Unknown')}")
print(f"{'='*80}")

await client.disconnect()
Tool Descriptions

BioInfo-Tools Server:

  • analyze_protein: Comprehensive protein analysis combining InterProScan and BLAST
    • Args:
      • sequence (str): Protein sequence in amino acid single-letter code
      • sequence_id (str, optional): Identifier for the query sequence
      • databases (list, optional): InterProScan databases (default: ["Pfam"])
      • evalue (float, optional): BLAST E-value threshold (default: 0.01)
      • max_hits (int, optional): Maximum BLAST hits (default: 10)
    • Returns:
      • interproscan (dict): InterProScan analysis results
        • success (bool): Whether InterProScan completed
        • results (dict): Domains and GO terms
        • time_seconds (float): Execution time
      • blast (dict): BLAST search results
        • success (bool): Whether BLAST completed
        • hits (list): Similar proteins
        • total_hits (int): Number of matches
        • time_seconds (float): Execution time
Input/Output

Input:

  • sequence: Protein sequence (amino acid single-letter code)
  • sequence_id: Optional identifier for the query
  • databases: List of InterProScan databases to query
  • evalue: BLAST E-value threshold (lower = more stringent)
  • max_hits: Maximum number of BLAST hits to return

Output:

  • InterProScan Results:
    • Functional domains with positions
    • Protein family classifications
    • Gene Ontology annotations
  • BLAST Results:
    • Homologous proteins across species
    • Sequence identity and alignment statistics
    • Evolutionary relationships
Show full SKILL.md (191 more words)Show less
Analysis Strategy

This comprehensive approach provides:

  1. Structural Information (InterProScan):

    • Domain architecture and organization
    • Functional motifs and active sites
    • Protein family membership
  2. Evolutionary Context (BLAST):

    • Homologs in other species
    • Sequence conservation patterns
    • Potential orthologs and paralogs
  3. Functional Prediction:

    • Combining domain and homology information
    • GO term annotations for molecular function
    • Biological process involvement
Performance Notes
  • Total execution time: 2-20 minutes depending on sequence length
    • InterProScan: 30 seconds to 15 minutes
    • BLAST: 10-90 seconds
    • Both run sequentially in this workflow
  • Timeout recommendation: Set to at least 1200 seconds (20 minutes)
  • E-value tuning: Use lower E-values (e.g., 1e-10) for highly conserved proteins, higher (e.g., 0.01) for divergent families
Use Cases
  • Complete functional annotation of unknown proteins
  • Validate predicted protein functions
  • Study protein evolution and conservation
  • Identify potential drug targets
  • Annotate proteomes and genome sequences
  • Compare protein function across species
Interpretation Tips
  • High domain coverage + high homology: Well-characterized protein with known function
  • Domains but no homologs: Novel protein with conserved domains, function can be inferred from domains
  • Homologs but no domains: May need more sensitive domain detection or represents a novel fold
  • Neither domains nor homologs: Potentially novel protein, may require experimental characterization

© InternScience, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/comprehensive-protein-analysis of InternScience/scp.

Open the folder on GitHubat commit cea5398

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in InternScience/scp, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Comprehensive Protein Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Comprehensive Protein Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Comprehensive Protein Analysis this skillInternScience/scp1691 repos~2kAutomated safety check: PassMIT
Drugbank Databasedavila7/claude-code-templates32k10 repos~2.3kAutomated safety check: PassMIT
Scholar RAGjoshzyj/open-scholar-skill168—~7.4kAutomated safety check: NotesCustom licence
Protein Sequence Similarity Searchgoogle-deepmind/science-skills3.2k1 repos~2.7kAutomated safety check: NotesApache-2.0
Uniprot Databasegoogle-deepmind/science-skills3.2k1 repos~3.1kAutomated safety check: PassApache-2.0
Medical Vector Searchaipoch/medical-research-skills2k—~2kAutomated safety check: PassMIT

Similar skills

  • Drugbank Database

    davila7/claude-code-templates

    Access and analyze comprehensive drug information from the DrugBank database including drug properties, interactions, targets, pathways, chemical structures, and pharmacology data.

    32k GitHub starsUsed in 10 repos~2.3k tokens
    Research & ScienceAuto-check passed
  • Scholar RAG

    joshzyj/open-scholar-skill

    Build and query a local vector database + GraphRAG over your entire reference library (Zotero or a PDF folder) for literature review.

    168 GitHub stars~7.4k tokensUpdated 20 days ago
    Research & ScienceAuto-check: notes
  • Protein Sequence Similarity Search

    google-deepmind/science-skills

    Searches for homologous protein sequences using MMseqs2 (fast, default) or BLAST (comprehensive, fallback).

    3.2k GitHub starsUsed in 1 repo~2.7k tokens
    Research & ScienceAuto-check: notes
  • Uniprot Database

    google-deepmind/science-skills

    Access protein metadata, function, taxonomy, and sequences across UniProtKB, UniParc, and UniRef.

    3.2k GitHub starsUsed in 1 repo~3.1k tokens
    Research & ScienceAuto-check passed
  • Medical Vector Search

    aipoch/medical-research-skills

    Vector database retrieval and evidence-based answering for medical research topics.

    2k GitHub stars~2k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed
  • Geniml

    aipoch/medical-research-skills

    Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run…

    2k GitHub stars~1.9k tokensUpdated 21 days ago
    Research & ScienceAuto-check passed

More from InternScience/scp

All 73 skills in this repo
  • Given an rsID, query multiple databases (dbSNP, FAVOR, GWAS Catalog, ClinVar, gnomAD, PharmGKB, ClinGen) for comprehensive annotation.

    169 GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Drugsda Esmfold

    InternScience/scp

    Use ESMFold model to predict 3D structure of the input protein sequence.

    169 GitHub starsUsed in 2 repos~721 tokens
    Auto-check passed
  • Drugsda Prosst

    InternScience/scp

    Given a protein sequence and its structure, employ ProSST model to predict mutation effects and obtain the top-k mutated sequences.

    169 GitHub starsUsed in 2 repos~949 tokens
    Auto-check passed
  • Calculate atmospheric parameters including Coriolis parameter, geostrophic wind, heat index, potential temperature, and dewpoint for meteorology and climate science.

    169 GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Biomedical Web Search

    InternScience/scp

    Search biomedical literature and web content using Tavily search engine for research and clinical information.

    169 GitHub starsUsed in 1 repo~598 tokens
    Auto-check passed
  • Calculate buoyancy forces and acceleration for fluid mechanics and hydrodynamics analysis.

    169 GitHub starsUsed in 1 repo~540 tokens
    Auto-check passed

Questions about Comprehensive Protein Analysis

What does Comprehensive Protein Analysis do?

Comprehensive protein analysis combining InterProScan domain identification with BLAST similarity search to provide complete functional and evolutionary annotation. Comprehensive Protein Analysis is an agent skill from InternScience/scp. Comprehensive protein analysis combining InterProScan domain identification with BLAST similarity search to provide complete functional and evolutionary annotation.

When should I use Comprehensive Protein Analysis?

Comprehensive Protein Analysis fits situations like: tasks that involve Vector databases.

How do I install Comprehensive Protein Analysis in Claude Code?

Run `npx skills add InternScience/scp --skill comprehensive-protein-analysis -a claude-code`. Or copy the skill folder (skills/comprehensive-protein-analysis in InternScience/scp) into .claude/skills/comprehensive-protein-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Comprehensive Protein Analysis in Codex?

Run `npx skills add InternScience/scp --skill comprehensive-protein-analysis -a codex`. Or copy the skill folder (skills/comprehensive-protein-analysis in InternScience/scp) into .agents/skills/comprehensive-protein-analysis in your project. Codex loads it when a task matches its description.

Can I use Comprehensive Protein Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add InternScience/scp --skill comprehensive-protein-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/comprehensive-protein-analysis, .gemini/skills/comprehensive-protein-analysis, .github/skills/comprehensive-protein-analysis and .opencode/skills/comprehensive-protein-analysis in your project.

What does Comprehensive Protein Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Comprehensive Protein Analysis is instructions for the agent only. Our summary lists: Python 3.

Does Comprehensive Protein Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Comprehensive Protein Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Comprehensive Protein Analysis use?

Comprehensive Protein Analysis is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Comprehensive Protein Analysis use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Comprehensive Protein Analysis?

Skills that share tags, products or a category with Comprehensive Protein Analysis: Drugbank Database (davila7/claude-code-templates, 32k stars), Scholar RAG (joshzyj/open-scholar-skill, 168 stars), Protein Sequence Similarity Search (google-deepmind/science-skills, 3.2k stars) and Uniprot Database (google-deepmind/science-skills, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Comprehensive Protein Analysis?

InternScience (a GitHub organization) maintains it in InternScience/scp, which has 169 GitHub stars. The repository holds 73 skills in this directory. The repository was last updated on June 3, 2026.

Source: InternScience/scp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.