Agent skill

Gi Annotation

by ClawBio in ClawBio/ClawBio

Predict gene and transcript structure (intervals, exons, strand) from a DNA sequence using the Genomic Intelligence DNA Annotation model, via the hosted /v1/tasks/annotation/predict API.

MITAuto-check: notesResearch & Science

Install Gi Annotation

skills CLI
$ npx skills add ClawBio/ClawBio --skill gi-annotation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio gi-annotation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gi-annotation .claude/skills/gi-annotation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gi-annotation
GitHub stars
1.2k
Token cost
~2.1k tokens
SKILL.md length
730 words
Files
6
Skills in repo
104
Repo updated
First seen
Licence
MIT

At a glance

Predict gene and transcript structure (intervals, exons, strand) from a DNA sequence using the Genomic Intelligence DNA Annotation model, via the hosted /v1/tasks/annotation/predict API.

  • Works in 4 steps: Parse: single-record FASTA. → Submit async: POST… → Poll: stream progress (percent, message)… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Trigger, Why This Exists, API Backed and Workflow, plus 7 more sections
  • Runs Python scripts from its folder; calls python; reaches api.genomicintelligence.ai; needs GI_API_KEY

What it does

Gi Annotation is an agent skill from ClawBio/ClawBio. Predict gene and transcript structure (intervals, exons, strand) from a DNA sequence using the Genomic Intelligence DNA Annotation model, via the hosted /v1/tasks/annotation/predict API. Submitted asynchronously — the pipeline takes ~20 s for ~20 kbp.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `api.py`, `gi_annotation.py` and `tests/__init__.py`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/gi-annotation”

Requirements

  • Python 3
  • A credential in GI_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Parse: single-record FASTA.
  2. Submit async: POST /v1/tasks/annotation/predict with Prefer: respond-async → 202 + job_id.
  3. Poll: stream progress (percent, message) until terminal.
  4. Render: report.md (transcripts table) + result.json (full response) + reproducibility/.

What it can do on your machine

Read from SKILL.md and the folder at commit 5e045e3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.genomicintelligence.ai

    Also links to:

    • genomicintelligence.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gi Annotation loads about 2.1k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 730 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:143
    cp .env.example .env
  • NoteMentions a .env fileSKILL.md:144
    set -a && source .env && set +a

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit 5e045e3, republished under its MIT licence (© ClawBio). 730 words, ~2,088 tokens.

Download SKILL.mdSave it as .claude/skills/gi-annotation/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
gi-annotation
description
Predict gene and transcript structure (intervals, exons, strand) from a DNA sequence using the Genomic Intelligence DNA Annotation model, via the hosted /v1/tasks/annotation/predict API. Submitted asynchronously — the pipeline takes ~20 s for ~20 kbp.
license
MIT
metadata.author
ClawBio + Genomic Intelligence
metadata.domain
genomics
metadata.tags
genomics, annotation, gene-prediction, transcript-prediction, gene-structure, dna-lm, gi-api
metadata.version
0.1.0

📜 gi-annotation

You are gi-annotation, a ClawBio agent that calls the Genomic Intelligence DNA annotation pipeline. Given a genomic region, it predicts gene boundaries → intervals → transcripts, all from sequence alone (no external annotation database).

⚠️ Remote inference — opt-in required. Unlike most ClawBio skills, this skill uploads your FASTA sequence to the hosted Genomic Intelligence API at https://api.genomicintelligence.ai. The same models also run interactively at https://genomicintelligence.ai. Do not submit identifiable patient data without an appropriate data-use agreement. Key setup: see Authentication below.

Trigger

Fire this skill when the user says any of:

  • "annotate this DNA sequence"
  • "predict genes / transcripts in this region"
  • "what genes are encoded here?" (from sequence, not coordinates)
  • "de novo gene prediction"
  • "gi-annotation"

Do NOT fire when:

  • The user has a VCF and wants variant consequences → variant-annotation (VEP)
  • The user wants known gene records by coordinate → external NCBI / Ensembl lookup

Why This Exists

  • Without it: Running AUGUSTUS / Helixer locally requires species models + dependency setup.
  • With it: One CLI call → predicted transcript structures, in ~20 s for ~20 kbp.
  • Why ClawBio: Hosted private weights (ModernBERT-based) plus ClawBio's reproducibility bundle and progress streaming for long jobs.

API Backed

POST https://api.genomicintelligence.ai/v1/tasks/annotation/predict with Prefer: respond-async. The API accepts either delivery mode on every task; this skill always submits async because the pipeline is long-running. The pipeline streams progress through GET /v1/tasks/jobs/{job_id} (typically: load → gene-boundaries → gene-intervals → transcripts).

Contract note. The Genomic Intelligence API publishes one operation per task, each with its own request schema: per-task minLength/maxLength on sequence, and a typed, closed options object (an unknown option key is a 422 validation_failed, not a silent ignore). The bounds quoted in this file are the published ones, but the authority is always the served schema: GET https://api.genomicintelligence.ai/v1/openapi.json.

Workflow

  1. Parse: single-record FASTA.
  2. Submit async: POST /v1/tasks/annotation/predict with Prefer: respond-async → 202 + job_id.
  3. Poll: stream progress (percent, message) until terminal.
  4. Render: report.md (transcripts table) + result.json (full response) + reproducibility/.

CLI Reference

bash
# Demo — bundled TP53 region (~20 s)
python skills/gi-annotation/gi_annotation.py --demo --output /tmp/gi-annotation-demo

# Your own FASTA
python skills/gi-annotation/gi_annotation.py --input my_region.fa --output report_dir

# Via ClawBio runner
python clawbio.py run gi-annotation --demo

Authentication

The skill requires a Genomic Intelligence partner key in GI_API_KEY. Resolution order:

  1. --api-key <value> CLI flag (explicit override).
  2. GI_API_KEY environment variable.
  3. Otherwise: the skill raises a RuntimeError pointing here.
Quick start — ClawBio hackathon key

A shared hackathon-tier key ships in .env.example at the repo root (opt-in only). Caps are per-key and are not published as a fixed number — read RateLimit-Limit / RateLimit-Remaining on any /v1/tasks/ response for the live allowance. The runner keeps them for you: they are in result.json under rate_limit, and a 429 names them on the error line. From wherever the ClawBio files live on your machine:

bash
# Repo root (git clone) — or ~/.claude/plugins/cache/clawbio/clawbio/<version>/ for plugin installs
cp .env.example .env
set -a && source .env && set +a
Production / heavier use

Request an individual key at contact@genomicintelligence.ai, then:

bash
export GI_API_KEY=gi_yourkeyhere
Show full SKILL.md (302 more words)Show less

Demo

bash
python clawbio.py run gi-annotation --demo

Bundled fixture is the TP53 locus (19 kbp). Expect several transcripts — TP53 has multiple annotated isoforms — and a wall time of roughly 20 s.

Gotchas

  • Always submitted async. This skill sends Prefer: respond-async and polls, so it never returns a synchronous response — though the API itself serves annotation synchronously when the header is omitted. Delivery mode is a per-request choice, not a property of the task.
  • Length bounds are 1,000–500,000 bp, published as minLength / maxLength on AnnotationPredictRequest and counted after whitespace is stripped. Both ends are a 422 validation_failed (over-max is not a 413 — 413 is the separate 16 MiB raw-body cap). The skill rejects either locally before spending a request. The 1,000 bp floor is the highest of the six tasks: the gene finder needs a region, not a single exon.
  • Long input is normal. The model handles tens-to-hundreds of kbp up to the 500 kbp cap; longer regions take proportionally more time. Annotation reports bio_spec.context_window_bp: null — there is no sliding-window regime caveat here, unlike promoter / splice / enhancer / chromatin.
  • First-call cold-start. The annotation pipeline is the heaviest GI model — first request after a cold service takes ~30+ s; subsequent calls are warm.
  • The model is trained on human + a few other vertebrates. Bacterial / fungal / plant predictions are out of distribution.
  • Hackathon key is shared. Async jobs count toward concurrent caps too — under heavy hackathon load, you may queue.

Output Structure

output_dir/
├── report.md
├── result.json
└── reproducibility/
    ├── command.sh
    └── environment.json

Integration with Bio Orchestrator

Routes here on: "annotate sequence", "predict genes", "gene structure", "de novo annotation".

Chains with: gi-promoter (validate predicted TSSes), gi-splice (cross-check predicted exon boundaries against splice-site calls), gi-expression (predict expression for each predicted transcript by extracting its TSS-centered window).

Safety

Research and development use. Not for clinical or diagnostic decisions. Predicted gene structures are model outputs, not curated reference annotations — for clinical interpretation, anchor to RefSeq / Ensembl.

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/gi-annotation of ClawBio/ClawBio.

  • SKILL.md
  • api.py
  • example_data/annotation_tp53.fa
  • gi_annotation.py
  • tests/__init__.py
  • tests/test_gi_annotation.py

Open the folder on GitHubat commit 5e045e3

Compare with similar skills

Gi Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gi Annotation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gi Annotation this skillClawBio/ClawBio1.2k—~2.1kAutomated safety check: NotesMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k3 repos~3.4kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
MFA Pipeline Orchestratoraiming-lab/AutoResearchClaw15k—~923Automated safety check: PassMIT

Similar skills

  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 3 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Singlecell Qc

    xuzhougeng/wisp-science

    A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.

    1k GitHub stars~1.6k tokensUpdated today
    Research & ScienceAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated today
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Questions about Gi Annotation

What does Gi Annotation do?

Predict gene and transcript structure (intervals, exons, strand) from a DNA sequence using the Genomic Intelligence DNA Annotation model, via the hosted /v1/tasks/annotation/predict API. Gi Annotation is an agent skill from ClawBio/ClawBio. Predict gene and transcript structure (intervals, exons, strand) from a DNA sequence using the Genomic Intelligence DNA Annotation model, via the hosted /v1/tasks/annotation/predict API.

When should I use Gi Annotation?

Gi Annotation fits situations like: tasks that involve Bioinformatics.

How do I install Gi Annotation in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill gi-annotation -a claude-code`. Or copy the skill folder (skills/gi-annotation in ClawBio/ClawBio) into .claude/skills/gi-annotation in your project. Claude Code loads it when a task matches its description.

How do I install Gi Annotation in Codex?

Run `npx skills add ClawBio/ClawBio --skill gi-annotation -a codex`. Or copy the skill folder (skills/gi-annotation in ClawBio/ClawBio) into .agents/skills/gi-annotation in your project. Codex loads it when a task matches its description.

Can I use Gi Annotation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill gi-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gi-annotation, .gemini/skills/gi-annotation, .github/skills/gi-annotation and .opencode/skills/gi-annotation in your project.

What does Gi Annotation need to run?

Going by SKILL.md and its folder, Gi Annotation needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named GI_API_KEY. Our summary lists: Python 3; A credential in GI_API_KEY.

Does Gi Annotation access the network?

SKILL.md names 2 domains. In commands or code: api.genomicintelligence.ai; the agent is likely to contact it when it follows the instructions. As links in the text: genomicintelligence.ai. This is read from the text; nothing was executed.

Is Gi Annotation safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Gi Annotation use?

Gi Annotation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gi Annotation use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gi Annotation?

Skills that share tags, products or a category with Gi Annotation: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars) and Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gi Annotation?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 7, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.