Agent skill

Analyze Fasta

by ClawBio in ClawBio/ClawBio

Analyze a single FASTA file (nucleotide or protein), compute sequence-level metrics (GC, ORFs, MW, pI, GRAVY, secondary-structure fractions) with Biopython, and write a Markdown report plus…

MITAuto-check passedResearch & Science

Install Analyze Fasta

skills CLI
$ npx skills add ClawBio/ClawBio --skill analyze-fasta -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio analyze-fasta --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analyze-fasta .claude/skills/analyze-fasta && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyze-fasta
GitHub stars
1.2k
Used in
1 other repo
Token cost
~3.4k tokens
SKILL.md length
1,201 words
Files
6
Skills in repo
104
Repo updated
First seen
Licence
MIT

At a glance

Analyze a single FASTA file (nucleotide or protein), compute sequence-level metrics (GC, ORFs, MW, pI, GRAVY, secondary-structure fractions) with Biopython, and write a Markdown report plus…

  • Works in 3 steps: Auto-detect sequence type: nucleotide vs… → Nucleotide metrics: length, GC% / AT%,… → Protein metrics: length, MW, isoelectric…
  • Tasks that involve Bioinformatics
  • SKILL.md covers Trigger, Why This Exists, Core Capabilities and Scope, plus 15 more sections
  • Runs Python scripts from its folder; calls python

What it does

Analyze Fasta is an agent skill from ClawBio/ClawBio. Analyze a single FASTA file (nucleotide or protein), compute sequence-level metrics (GC, ORFs, MW, pI, GRAVY, secondary-structure fractions) with Biopython, and write a Markdown report plus structured JSON for downstream chaining.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `INTENTS.json`, `analyze_fasta.py` and `tests/test_analyze_fasta.py`).

It sits in Research & Science, covering Bioinformatics. It works with Biopython. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/analyze-fasta”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Auto-detect sequence type: nucleotide vs protein (>=85% ACGTUN ratio threshold over the first 500 chars).
  2. Nucleotide metrics: length, GC% / AT%, base and dinucleotide composition, ORF discovery (>=100 aa), N50 across multi-record FASTAs, MW.
  3. Protein metrics: length, MW, isoelectric point (pI), instability index, GRAVY (hydrophobicity), aromaticity, charged/aromatic residue %…

What it can do on your machine

Read from SKILL.md and the folder at commit 5e045e3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • doi.org
    • biopython.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyze Fasta loads about 3.4k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 1,201 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit 5e045e3, republished under its MIT licence (© ClawBio). 1,201 words, ~3,394 tokens.

Download SKILL.mdSave it as .claude/skills/analyze-fasta/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
analyze-fasta
description
Analyze a single FASTA file (nucleotide or protein), compute sequence-level metrics (GC, ORFs, MW, pI, GRAVY, secondary-structure fractions) with Biopython, and write a Markdown report plus structured JSON for downstream chaining.
license
MIT
metadata.version
0.1.0
metadata.author
Santiago Rodriguez Salinas
metadata.domain
genomics
metadata.tags
fasta, biopython, sequence-analysis, gc-content, orf, protein-properties, isoelectric-point, gravy

🧬 analyze-fasta

You are analyze-fasta, a specialised ClawBio agent for single-FASTA inspection. Your role is to take a FASTA file (nucleotide or protein), auto-detect its type, compute the standard set of sequence-level metrics with Biopython, and produce a structured report that downstream skills can chain to.

Trigger

Fire this skill when the user says any of:

  • "analyze this fasta"
  • "analiza este fasta"
  • "what's the GC content of this sequence"
  • "find ORFs in this sequence"
  • "compute pI / isoelectric point of this protein"
  • "GRAVY index"
  • "protein properties from this fasta"
  • "summarise this fasta"
  • "describe this sequence"

Do NOT fire when:

  • The user has FASTQ reads — route to seq-wrangler (alignment QC).
  • The user has a VCF — route to variant-annotation or clinical-variant-reporter.
  • The user wants comparison between two FASTA — route to genome-compare.
  • The user wants 3D structure prediction — route to struct-predictor.

Why This Exists

  • Without it: Users open Biopython interactively, copy boilerplate to compute GC / ProtParam metrics, and hand-format a report. Common values get computed inconsistently across notebooks.
  • With it: One command turns a FASTA into a Markdown report + JSON suitable for orchestration. Detection of nucleotide vs protein is automatic. ORFs, GC%, MW, pI, GRAVY, secondary-structure fractions, dinucleotide counts, and N50 all come out at once.
  • Why ClawBio: Output is structured (result.json) so the bio-orchestrator can chain analyze-fasta → variant-annotation, struct-predictor, or pubmed-summariser without reparsing prose.

Core Capabilities

  1. Auto-detect sequence type: nucleotide vs protein (>=85% ACGTUN ratio threshold over the first 500 chars).
  2. Nucleotide metrics: length, GC% / AT%, base and dinucleotide composition, ORF discovery (>=100 aa), N50 across multi-record FASTAs, MW.
  3. Protein metrics: length, MW, isoelectric point (pI), instability index, GRAVY (hydrophobicity), aromaticity, charged/aromatic residue %, secondary-structure fractions (helix/turn/sheet), AA composition.

Scope

One skill, one task. This skill describes a single FASTA file. It does not align, blast, fold, compare, or annotate. If the user wants any of those, the skill should refuse and route elsewhere.

Input Formats

FormatExtensionRequired FieldsExample
FASTA (nucleotide).fasta, .fa, .fna>header line + ACGTUN sequenceexample_data/demo_nucleotide.fasta
FASTA (protein).fasta, .fa, .faa>header line + amino-acid sequenceexample_data/demo_protein.fasta

Workflow

When the user asks for FASTA analysis:

  1. Validate (prescriptive): file exists; at least one record; first record >=10 chars; <=50% Ns. Any failure → exit 1 with explicit message. Never write a partial report.
  2. Detect type (prescriptive): nucleotide if >=85% of first 500 chars are in ACGTUNacgtun, else protein.
  3. Compute metrics per record (prescriptive): use Biopython gc_fraction, molecular_weight, ProteinAnalysis. Round consistently (GC to 2 dp, MW to 1 dp, pI to 2 dp).
  4. Generate (prescriptive): write result.json (full structured data), report.md (human-readable), report.html (visual), and reproducibility/{commands.sh,environment.yml,checksums.sha256,run.json}.
  5. Interpret (flexible — agent layer): the LLM may add a short biological narrative on top of the report (likely organism class from GC, predicted protein family from pI/GRAVY) but must not modify the numeric metrics.

CLI Reference

bash
# Standard usage (ClawBio convention)
python skills/analyze-fasta/analyze_fasta.py \
  --input <fasta_file> --output <report_dir>

# Demo mode (uses bundled synthetic nucleotide FASTA)
python skills/analyze-fasta/analyze_fasta.py --demo --output /tmp/analyze_fasta_demo

# Via ClawBio runner
python clawbio.py run analyze-fasta --input <fasta_file> --output <dir>
python clawbio.py run analyze-fasta --demo

# Legacy modes (backward compat with the original TP1 release)
python skills/analyze-fasta/analyze_fasta.py <file.fasta> --json
python skills/analyze-fasta/analyze_fasta.py <file.fasta> --html out.html

Demo

bash
python clawbio.py run analyze-fasta --demo

Expected output: a report.md with summary metrics for the bundled ~720 bp synthetic nucleotide (GC ~50%, 1 ORF detected, AA composition table) plus the matching result.json and reproducibility/ bundle.

Algorithm / Methodology

So an LLM agent can apply the same logic without the script:

  1. Sequence type detection: count chars in first 500 of the first record that match [ACGTUNacgtun]. Ratio >= 0.85 → nucleotide, else protein. (No silent fallback; if ambiguous, document in result.json.)
  2. Nucleotide GC: gc = (G + C) / (A + T + G + C + N) * 100. Use Biopython gc_fraction to match the production behaviour.
  3. ORF discovery: scan all 3 forward frames for ATG ... [TAA|TAG|TGA]. Keep ORFs with length_bp >= 300 (>= 100 aa).
  4. N50: sort lengths descending; cumulative sum until it reaches half of the total. Length at that point is N50.
  5. Protein metrics: Biopython ProteinAnalysis. Strip X and * before instantiating to avoid ProtParam errors.
  6. Secondary-structure fractions: ProtParam secondary_structure_fraction() → (helix, turn, sheet); convert to percent.

Key thresholds:

  • Min sequence length: 10 chars (source: arbitrary lower bound to reject empty/garbage input).
  • Max N ratio: 50% (source: arbitrary; below this Biopython metrics become unreliable).
  • ORF min length: 300 bp / 100 aa (source: standard convention for naive ORF finders, avoids spurious short ORFs).
  • Sequence-type detection threshold: 85% (source: heuristic that handles common ambiguity codes without misclassifying short proteins).

Example Queries

  • "Analyze sample.fasta"
  • "Analiza este FASTA, decime el GC y los ORFs"
  • "What's the molecular weight of this protein?"
  • "Compute pI of the FASTA in /tmp/x.fa"

Example Output

markdown
# analyze-fasta Report

**Input file:** `demo_nucleotide.fasta`
**Analysis date:** 2026-05-05 12:00:00
**Sequence type:** `nucleotide`
**Total sequences:** 1

## Summary

| Metric | Value |
|---|---|
| total_sequences | 1 |
| total_residues | 720 |
| min_length | 720 |
| max_length | 720 |
| avg_length | 720.0 |
| n50 | 720 |
| avg_gc_content | 50.42 |
| total_orfs | 1 |

## Per-sequence metrics

### 1. synthetic_demo_orf

- **Description:** synthetic_demo_orf | Synthetic E. coli-like ORF
- **Length:** 720 bp
- **GC content:** 50.42%
- **AT content:** 49.58%
- **ORFs (>=100 aa):** 1

---

_ClawBio is a research and educational tool. It is not a medical device and does not provide clinical diagnoses. Consult a healthcare professional before making any medical decisions._

Output Structure

<output_dir>/
├── report.md              # Primary markdown report
├── report.html            # Standalone visual report
├── result.json            # Machine-readable results
└── reproducibility/
    ├── commands.sh        # Portable replay command ($CLAWBIO_ROOT / $OUTPUT_DIR)
    ├── environment.yml    # Conda recipe (biopython)
    ├── checksums.sha256   # SHA-256 of every output file
    └── run.json           # Run metadata (versions, timestamps, input size)
Show full SKILL.md (488 more words)Show less

Dependencies

Required:

  • biopython >= 1.80; sequence parsing, ProtParam, gc_fraction, molecular_weight.

Optional:

  • None beyond Biopython. The skill is intentionally lean; stdlib + Biopython.

Gotchas

  • The model will want to claim "this is gene X / from organism Y" from GC content alone. Do not. GC is a weak signal — many taxa overlap. State GC as a number; if the user asks for a guess, frame it explicitly as "consistent with" rather than "this is".
  • The model will treat ORFs >100 aa as proof of coding. Do not. The ORF finder is naive: forward strand only, no reading-frame validation against known annotations, no Kozak / Shine-Dalgarno check. Frame ORFs as candidates, never confirmed.
  • The model will silently re-interpret a sequence with many Ns as a real result. Do not. The script aborts with >50% Ns; the agent must not bypass that with a "best-effort" fallback. Surface the failure to the user.
  • The model will mix nucleotide and protein metrics if a multi-record FASTA mixes types. The skill detects type from the first record only. If the FASTA mixes nucleotides and proteins, ask the user to split the file rather than reporting hybrid metrics.
  • The model will use the script's HTML output as the primary deliverable. Use report.md for chaining; the HTML is a courtesy for human inspection only.

Safety

  • Local-first: no network calls; everything runs against the local file.
  • Disclaimer: every report.md includes the standard ClawBio research-tool disclaimer.
  • Audit trail: every run writes reproducibility/run.json with timestamps, Python and Biopython versions, and input file size.
  • No hallucinated science: thresholds (GC, ORF, N ratio) are documented in this SKILL.md; the agent must not invent new ones.

Agent Boundary

The agent (LLM) decides whether to fire this skill, may add a short biological-context paragraph on top of the report, and may suggest follow-up skills (struct-predictor, variant-annotation, pubmed-summariser). The skill (Python) executes the metrics and writes the artefacts. The agent must NOT recompute metrics, override thresholds, or fabricate organism-of-origin claims.

Integration with Bio Orchestrator

Trigger conditions: the orchestrator routes here when the input is a single .fasta/.fa/.fna/.faa file or the query mentions gc content, orfs, pi, gravy, or protein properties.

Chaining partners:

  • struct-predictor: take a single protein record from the input FASTA and predict structure.
  • variant-annotation: out of scope here, but the user often asks for variant context after sequence inspection.
  • pubmed-summariser: useful when the FASTA header contains a gene/organism name that the user wants literature for.

Output is JSON + Markdown with stable keys, so it composes cleanly into pipelines.

Maintenance

  • Review cadence: re-evaluate quarterly or when Biopython releases a major version.
  • Staleness signals: Biopython API breaks (ProteinAnalysis signature changes), or ORF heuristics receive a community-standard upgrade (e.g., GeneMark-style probabilistic finders).
  • Deprecation: archive to skills/_deprecated/analyze-fasta/ only if a more capable single-FASTA skill (e.g., one wrapping seqkit stats) replaces it across the catalog.

Citations

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/analyze-fasta of ClawBio/ClawBio.

  • SKILL.md
  • INTENTS.json
  • analyze_fasta.py
  • example_data/demo_nucleotide.fasta
  • example_data/demo_protein.fasta
  • tests/test_analyze_fasta.py

Open the folder on GitHubat commit 5e045e3

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ClawBio/ClawBio, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Analyze Fasta next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyze Fasta compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyze Fasta this skillClawBio/ClawBio1.2k1 repos~3.4kAutomated safety check: PassMIT
Bio Alignment IoGPTomics/bioSkills1.2k3 repos~4.9kAutomated safety check: PassMIT
Biopython Bioinformaticsaiming-lab/AutoResearchClaw15k—~810Automated safety check: PassMIT
Bio Write SequencesGPTomics/bioSkills1.2k3 repos~2.1kAutomated safety check: PassMIT
Biopythondavila7/claude-code-templates32k13 repos~3.4kAutomated safety check: PassMIT
Ggetdavila7/claude-code-templates32k11 repos~6.3kAutomated safety check: PassMIT

Similar skills

  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Research & ScienceAuto-check passed
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 13 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Gget

    davila7/claude-code-templates

    CLI/Python toolkit for rapid bioinformatics queries. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~6.3k tokens
    Research & ScienceAuto-check passed
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Research & ScienceAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Works with

Questions about Analyze Fasta

What does Analyze Fasta do?

Analyze a single FASTA file (nucleotide or protein), compute sequence-level metrics (GC, ORFs, MW, pI, GRAVY, secondary-structure fractions) with Biopython, and write a Markdown report plus…. Analyze Fasta is an agent skill from ClawBio/ClawBio. Analyze a single FASTA file (nucleotide or protein), compute sequence-level metrics (GC, ORFs, MW, pI, GRAVY, secondary-structure fractions) with Biopython, and write a Markdown report plus structured JSON for downstream chaining.

When should I use Analyze Fasta?

Analyze Fasta fits situations like: tasks that involve Bioinformatics.

How do I install Analyze Fasta in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill analyze-fasta -a claude-code`. Or copy the skill folder (skills/analyze-fasta in ClawBio/ClawBio) into .claude/skills/analyze-fasta in your project. Claude Code loads it when a task matches its description.

How do I install Analyze Fasta in Codex?

Run `npx skills add ClawBio/ClawBio --skill analyze-fasta -a codex`. Or copy the skill folder (skills/analyze-fasta in ClawBio/ClawBio) into .agents/skills/analyze-fasta in your project. Codex loads it when a task matches its description.

Can I use Analyze Fasta in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill analyze-fasta -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-fasta, .gemini/skills/analyze-fasta, .github/skills/analyze-fasta and .opencode/skills/analyze-fasta in your project.

What does Analyze Fasta need to run?

Going by SKILL.md and its folder, Analyze Fasta needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Analyze Fasta access the network?

SKILL.md names 2 domains. As links in the text: doi.org and biopython.org. This is read from the text; nothing was executed.

Is Analyze Fasta safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Analyze Fasta use?

Analyze Fasta is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyze Fasta use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyze Fasta?

Skills that share tags, products or a category with Analyze Fasta: Bio Alignment Io (GPTomics/bioSkills, 1.2k stars), Biopython Bioinformatics (aiming-lab/AutoResearchClaw, 15k stars), Bio Write Sequences (GPTomics/bioSkills, 1.2k stars) and Biopython (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyze Fasta?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 7, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.