Agent skill

Polars Bio

by ClawBio in ClawBio/ClawBio

Fast genomic interval operations (overlap, nearest, merge, coverage, cluster, complement, subtract, count-overlaps), multi-format bioinformatics I/O, DataFusion SQL, and pileup on Polars DataFrames…

Apache-2.0Auto-check passedData & Analytics

Install Polars Bio

skills CLI
$ npx skills add ClawBio/ClawBio --skill polars-bio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio polars-bio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/polars-bio .claude/skills/polars-bio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
polars-bio
GitHub stars
1.2k
Token cost
~3.4k tokens
SKILL.md length
1,035 words
Files
24 (incl. references)
Skills in repo
104
Repo updated
First seen
Licence
Apache-2.0

At a glance

Fast genomic interval operations (overlap, nearest, merge, coverage, cluster, complement, subtract, count-overlaps), multi-format bioinformatics I/O, DataFusion SQL, and pileup on Polars DataFrames…

  • Works in 4 steps: Interval operations: overlap, nearest,… → Multi-format I/O: read/scan BED, VCF,… → DataFusion SQL: register a file as table… → …
  • Tasks that involve DataFrames
  • SKILL.md covers Trigger, Why This Exists, Core Capabilities and Scope, plus 11 more sections
  • Runs Python scripts from its folder; calls python

What it does

Polars Bio is an agent skill from ClawBio/ClawBio. Fast genomic interval operations (overlap, nearest, merge, coverage, cluster, complement, subtract, count-overlaps), multi-format bioinformatics I/O, DataFusion SQL, and pileup on Polars DataFrames via polars-bio. A scalable bioframe/bedtools alternative.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 25 other files, including reference files (for example `examples/generate_fixtures.py`, `polars_bio_runner.py` and `references/configuration.md`).

It sits in Data & Analytics, covering DataFrames and Bioinformatics. It works with Polars and SQL. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve DataFrames
  • Tasks that involve Bioinformatics

Example prompts

  • “/polars-bio”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Interval operations: overlap, nearest, merge, coverage, cluster, complement,
  2. Multi-format I/O: read/scan BED, VCF, VCF Zarr, GFF, GTF, FASTA, FASTQ, BAM,
  3. DataFusion SQL: register a file as table t and run SQL.
  4. Pileup: per-base read depth (mosdepth-compatible) from an indexed BAM.

What it can do on your machine

Read from SKILL.md and the folder at commit dece754. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • doi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Polars Bio loads about 3.4k tokens when it runs, and up to ~6.9k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 1,035 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit dece754, republished under its Apache-2.0 licence (© ClawBio). 1,035 words, ~3,379 tokens.

Download SKILL.mdSave it as .claude/skills/polars-bio/SKILL.md (or your agent's skills folder). This skill also uses 23 other files; get the full folder from GitHub.
name
polars-bio
description
Fast genomic interval operations (overlap, nearest, merge, coverage, cluster, complement, subtract, count-overlaps), multi-format bioinformatics I/O, DataFusion SQL, and pileup on Polars DataFrames via polars-bio. A scalable bioframe/bedtools alternative.
license
Apache-2.0
metadata.version
0.1.0
metadata.author
ClawBio (adapted from K-Dense scientific-agent-skills/polars-bio)
metadata.domain
genomics
metadata.tags
genomic-intervals, interval-arithmetic, bioframe-alternative, file-io, datafusion-sql, polars

🐻 polars-bio

You are polars-bio, a ClawBio agent for fast genomic interval arithmetic and bioinformatics file I/O on Polars DataFrames. You dispatch the polars_bio_runner.py CLI; the library does the compute.

Trigger

Fire this skill when the user says any of:

  • "find overlapping intervals", "intersect these BED files", "overlap a.bed b.bed"
  • "nearest interval / nearest feature", "merge overlapping intervals", "cluster intervals"
  • "interval coverage", "complement / gaps between intervals", "subtract intervals", "count overlaps"
  • "bioframe alternative", "faster than bedtools/pyranges", "interval arithmetic"
  • "read/scan a BED/VCF/GFF/GTF/FASTA/FASTQ/BAM/BigWig/BigBed", "inspect file schema"
  • "run SQL on a VCF/BED", "DataFusion SQL on genomic files"
  • "per-base depth / pileup from a BAM", "polars-bio"

Do NOT fire when:

  • The user wants variant annotation / pathogenicity → variant-annotation, vcf-annotator, clinical-variant-reporter.
  • The user wants a phylogenetic tree / distance matrix → fastreer, phylogenetics-builder.
  • The user wants multi-sample QC aggregation → multiqc-reporter.
  • The user wants variant calling from FASTQ/BAM → nfcore-sarek-wrapper.

Why This Exists

ClawBio has variant/VCF skills and a phylogenetics tool, but no fast, DataFrame-native interval-operations engine.

  • Without it: users hand-roll overlaps in pandas/bioframe or shell out to bedtools, with no reproducible ClawBio report.
  • With it: the full interval-op set plus multi-format I/O and SQL, streaming and cloud-native, with a report + JSON + figure bundle.
  • Why ClawBio: grounded in a peer-reviewed, benchmarked library — not guesswork.

Performance (attributed to the polars-bio docs/paper, not invented): 6–38× faster than bioframe on interval benchmarks; streaming throughput ~20–28M rows/s; substantially faster VCF parsing; ~20× less memory than vanilla Polars on GFF reads. See references/polars_primer.md.

Core Capabilities

  1. Interval operations: overlap, nearest, merge, coverage, cluster, complement, subtract, count-overlaps.
  2. Multi-format I/O: read/scan BED, VCF, VCF Zarr, GFF, GTF, FASTA, FASTQ, BAM, CRAM, SAM, Pairs, BigWig, BigBed; --describe for schema-only inspection (VCF/VCF Zarr/BAM/CRAM/SAM).
  3. DataFusion SQL: register a file as table t and run SQL.
  4. Pileup: per-base read depth (mosdepth-compatible) from an indexed BAM.

Scope

One library, one cohesive surface. This skill wraps polars-bio operations and nothing else. Annotation, calling, QC, and phylogenetics live in other skills.

Polars & the Python ecosystem

polars-bio extends Polars (a Rust-backed, Apache Arrow-native DataFrame library) with genomics. The stack:

Polars (LazyFrame/DataFrame)  ->  Apache Arrow (columnar memory)
   ->  Apache DataFusion (query/SQL engine)  ->  datafusion-bio (BED/VCF/BAM/... readers)

Genomic interval work stays inside the same DataFrame pipeline as the rest of a Python analysis — no pandas/bedtools round-trips. Interop: .to_pandas(), pyarrow hand-off, and output_type="polars.DataFrame" for eager results. Full primer (Polars vs pandas, neighbors bioframe/pyranges1/pybedtools/GenomicRanges, Rust backends ruranges/ superintervals): references/polars_primer.md.

Input Formats

Canonical list of what this skill accepts (reader functions and parameters are detailed in references/file_io.md).

FormatExtensionNotes
BED.bed>=4 columns required (chrom,start,end,name); interval ops + io/sql
VCF.vcf/.vcf.gzio/sql; --describe lists INFO/FORMAT fields
VCF Zarr.zarr dirio/sql; array-native variant store
GFF / GTF.gff3/.gtfannotations; io/sql
FASTA / FASTQ.fasta/.fastqsequences; io/sql
BAM.bam (+.bai)io/sql/pileup; index required
CRAM.cramio/pileup; needs --reference FASTA
SAM.samtext alignments; io/sql
Pairs.pairsHi-C contacts; io/sql
BigWig / BigBed.bw/.bbsignal / interval tracks; io/sql

Workflow

  1. Validate: confirm the subcommand and that required inputs exist (BED >=4 cols; BAM has a .bai).
  2. Run the CLI subcommand with --output <dir> (see CLI Reference).
  3. Open the generated figure.png and read report.md; summarize the row counts and schema for the user.
  4. DEMO FALLBACK (mandatory): if the user has no file, do NOT refuse — run --demo immediately ("I'll run a demo on synthetic BED data so you can see it").

CLI Reference

bash
# Interval operations (BED in, report/json/figure/table out)
python skills/polars-bio/polars_bio_runner.py overlap  --a a.bed --b b.bed --output <dir>
python skills/polars-bio/polars_bio_runner.py nearest  --a a.bed --b b.bed --k 1 --output <dir>
python skills/polars-bio/polars_bio_runner.py merge     --a a.bed --output <dir>
python skills/polars-bio/polars_bio_runner.py coverage  --a a.bed --b b.bed --output <dir>
python skills/polars-bio/polars_bio_runner.py cluster    --a a.bed --output <dir>
python skills/polars-bio/polars_bio_runner.py complement --a a.bed --output <dir>
python skills/polars-bio/polars_bio_runner.py subtract   --a a.bed --b b.bed --output <dir>
python skills/polars-bio/polars_bio_runner.py count-overlaps --a a.bed --b b.bed --output <dir>

# File I/O (schema + head); --describe for schema-only inspection
python skills/polars-bio/polars_bio_runner.py io --input s.vcf --format vcf --output <dir>
python skills/polars-bio/polars_bio_runner.py io --input s.vcf --format vcf --describe --output <dir>

# DataFusion SQL (file registered as table `t`)
python skills/polars-bio/polars_bio_runner.py sql --input s.vcf --query "SELECT chrom,start FROM t" --output <dir>

# Pileup (indexed BAM)
python skills/polars-bio/polars_bio_runner.py pileup --input aln.bam --min-mapping-quality 20 --output <dir>

# Demo (synthetic BED overlap)
python skills/polars-bio/polars_bio_runner.py --demo --output /tmp/polars_bio_demo

Global flags: --one-based (output 1-based closed coords; default is 0-based half-open, BED-native), --genome <chrom-sizes> (bound complement gaps), --output (required). The coordinate flag sets the output representation only — it does not change how inputs are parsed, and interval results are identical either way.

Example Output

--demo runs overlap on the bundled synthetic BED sets and writes report.md, result.json, figure.png, and result.csv.

result.json (actual):

json
{
  "skill": "polars-bio",
  "subcommand": "overlap",
  "params": { "k": 1, "zero_based": true },
  "polars_bio_version": "<runtime-detected>",
  "output_rows": 5,
  "output_schema": {
    "chrom_1": "String", "start_1": "UInt32", "end_1": "UInt32", "name_1": "String",
    "chrom_2": "String", "start_2": "UInt32", "end_2": "UInt32", "name_2": "String"
  },
  "figure": "figure.png",
  "report": "report.md"
}

report.md (excerpt):

markdown
# polars-bio — overlap

**polars-bio version:** <runtime-detected>
**Output rows:** 5

## Output schema
| Column | Type |
|--------|------|
| `chrom_1` | String |
| `start_1` | UInt32 |
Show full SKILL.md (438 more words)Show less

Gotchas

  1. BED needs >=4 columns. The polars-bio BED reader silently returns 0 rows for 3-column BED. Demo files are BED6. Always include a name column.
  2. The coordinate flag is output representation, not input parsing. The reader already knows BED is 0-based half-open on disk. --one-based only changes how output coordinates are displayed (1-based closed shifts each start +1); it does not change which intervals overlap — pairings are identical in both modes. The CLI defaults to 0-based half-open so BED round-trips (merge/complement/ subtract/cluster) come back BED-native. io and sql honor the same default. The runner records the actually-installed polars-bio version in result.json (no hardcoded version anywhere).
  3. complement needs contig bounds. Without --genome, trailing gaps span to i64::MAX (not genomically meaningful); the runner emits a caveat in report.md and stderr. Pass --genome <chrom-sizes> (chrom<TAB>size per line) for bounded gaps.
  4. Operations return LazyFrames. The CLI requests eager output (output_type="polars.DataFrame"); if you call the library directly, remember .collect().
  5. Probe-build order matters. For two-input ops the first DataFrame is probed against the second — pass the larger set first for speed.
  6. expand and sort_bedframe are not exposed as functions in the current polars-bio Python API, so they are intentionally not subcommands. Use Polars expressions for padding/sorting if needed.
  7. BAM needs a .bai index for io/sql/pileup; the runner errors clearly if missing (samtools index aln.bam). CRAM needs a reference_path.
  8. Nested columns can't be CSV. GFF/GTF/VCF outputs with list/struct columns are written as result.ndjson instead of result.csv automatically.
  9. INT32 position limit (~2.1 Gb) per contig — fine for known genomes.

Safety

ClawBio is a research and educational tool. It is not a medical device and does not provide clinical diagnoses. Consult a healthcare professional before making any medical decisions. Genomic data is processed locally; cloud paths use your own SDK credentials only when an s3:///gs:///az:// URI is accessed.

Agent Boundary

The agent dispatches the subcommand, explains parameters, and interprets the report. The skill (polars_bio_runner.py) executes the computation via polars-bio. The agent does not invent thresholds, schemas, or benchmark numbers — those come from the library and references/.

Chaining Partners

  • vcf-annotator / variant-annotation: annotate variants that interval ops select.
  • multiqc-reporter: aggregate QC alongside coverage/pileup outputs.
  • fastreer / phylogenetics-builder: downstream phylogenetics on selected regions.
  • nfcore-sarek-wrapper: upstream calling that produces the VCFs/BAMs analyzed here.

Maintenance

  • Review cadence: per polars-bio minor release.
  • Staleness signals: a new polars-bio release changing the interval-op set, reader formats, or coordinate behavior; a new describe_*/register_* function.
  • Deprecation: retire if polars-bio is abandoned or superseded in ClawBio.

Citation

Wiewiórka M, Khamutou P, Zbysiński M, Gambin T. polars-bio — fast, scalable, and out-of-core operations on large genomic interval datasets. Bioinformatics, 2025, 41(12):btaf640. https://doi.org/10.1093/bioinformatics/btaf640

© ClawBio, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 23 other files (references) in skills/polars-bio of ClawBio/ClawBio.

  • SKILL.md
  • examples/demo.bam
  • examples/demo.bam.bai
  • examples/demo.bb
  • examples/demo.bw
  • examples/demo.fasta
  • examples/demo.fastq
  • examples/demo.gff3
  • examples/demo.gtf
  • examples/demo.pairs
  • examples/demo.sam
  • examples/demo.vcf
  • examples/demo_a.bed
  • examples/demo_b.bed
  • examples/generate_fixtures.py
  • polars_bio_runner.py
  • references/configuration.md
  • references/file_io.md
  • references/interval_operations.md
  • … and 5 more

Open the folder on GitHubat commit dece754

Compare with similar skills

Polars Bio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Polars Bio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Polars Bio this skillClawBio/ClawBio1.2k—~3.4kAutomated safety check: PassApache-2.0
Polars BioK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: NotesApache-2.0
Transforming Dataancoleman/ai-design-components526—~3kAutomated safety check: PassMIT
Analyzing Dataastronomer/agents451—~1.3kAutomated safety check: PassApache-2.0
Bio Genome Intervals Gtf Gff HandlingGPTomics/bioSkills1.2k1 repos~4.6kAutomated safety check: PassMIT
Sc GrnTianGzlab/OmicsClaw161—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Polars Bio

    K-Dense-AI/scientific-agent-skills

    Performs genomic interval overlap, nearest, merge, coverage, complement and subtraction on Polars DataFrames, and reads or writes BED, VCF, BCF, BAM, CRAM, GFF, GTF, FASTA and FASTQ data.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Data & AnalyticsAuto-check: notes
  • Transforming Data

    ancoleman/ai-design-components

    Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).

    526 GitHub stars~3k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed
  • Analyzing Data

    astronomer/agents

    Queries the data warehouse with SQL and answers business questions about data.

    451 GitHub stars~1.3k tokensUpdated 2 days ago
    DatabasesAuto-check passed
  • Parses, queries, converts, and extracts from GTF and GFF3 gene-model annotation files - walking the gene/transcript/exon/CDS hierarchy with gffutils (queryable SQLite DB), converting formats and…

    1.2k GitHub starsUsed in 1 repo~4.6k tokens
    Data & AnalyticsAuto-check passed
  • Sc Grn

    TianGzlab/OmicsClaw

    Load when inferring TF → target gene regulatory networks on a normalised scRNA AnnData via pySCENIC (GRNBoost2 + cisTarget + AUCell) or correlation-based GRN fallback (when arboreto is unavailable…

    161 GitHub stars~1.7k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Chdb Datastore

    vemetric/vemetric

    A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

    395 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.7k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated yesterday
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Works with

Questions about Polars Bio

What does Polars Bio do?

Fast genomic interval operations (overlap, nearest, merge, coverage, cluster, complement, subtract, count-overlaps), multi-format bioinformatics I/O, DataFusion SQL, and pileup on Polars DataFrames…. Polars Bio is an agent skill from ClawBio/ClawBio. Fast genomic interval operations (overlap, nearest, merge, coverage, cluster, complement, subtract, count-overlaps), multi-format bioinformatics I/O, DataFusion SQL, and pileup on Polars DataFrames via polars-bio.

When should I use Polars Bio?

Polars Bio fits situations like: tasks that involve DataFrames; tasks that involve Bioinformatics.

How do I install Polars Bio in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill polars-bio -a claude-code`. Or copy the skill folder (skills/polars-bio in ClawBio/ClawBio) into .claude/skills/polars-bio in your project. Claude Code loads it when a task matches its description.

How do I install Polars Bio in Codex?

Run `npx skills add ClawBio/ClawBio --skill polars-bio -a codex`. Or copy the skill folder (skills/polars-bio in ClawBio/ClawBio) into .agents/skills/polars-bio in your project. Codex loads it when a task matches its description.

Can I use Polars Bio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill polars-bio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/polars-bio, .gemini/skills/polars-bio, .github/skills/polars-bio and .opencode/skills/polars-bio in your project.

What does Polars Bio need to run?

Going by SKILL.md and its folder, Polars Bio needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Polars Bio access the network?

SKILL.md names 1 domain. As links in the text: doi.org. This is read from the text; nothing was executed.

Is Polars Bio safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Polars Bio use?

Polars Bio is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Polars Bio use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.5k tokens, read only when the agent opens those files.

What are the alternatives to Polars Bio?

Skills that share tags, products or a category with Polars Bio: Polars Bio (K-Dense-AI/scientific-agent-skills, 48k stars), Transforming Data (ancoleman/ai-design-components, 526 stars), Analyzing Data (astronomer/agents, 451 stars) and Bio Genome Intervals Gtf Gff Handling (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Polars Bio?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 8, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.