Agent skill

Bio Genome Intervals Gtf Gff Handling

by GPTomics in GPTomics/bioSkills

Parses, queries, converts, and extracts from GTF and GFF3 gene-model annotation files - walking the gene/transcript/exon/CDS hierarchy with gffutils (queryable SQLite DB), converting formats and…

MITAuto-check passedData & Analytics

Install Bio Genome Intervals Gtf Gff Handling

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-genome-intervals-gtf-gff-handling -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-genome-intervals-gtf-gff-handling --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/genome-intervals/gtf-gff-handling .claude/skills/bio-genome-intervals-gtf-gff-handling && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-genome-intervals-gtf-gff-handling
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.6k tokens
SKILL.md length
1,969 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Parses, queries, converts, and extracts from GTF and GFF3 gene-model annotation files - walking the gene/transcript/exon/CDS hierarchy with gffutils (queryable SQLite DB), converting formats and…

  • Works in 3 steps: The coordinate conversion is asymmetric.… → The all-zero count matrix.… → phase is not frame, and the stop codon…
  • Extracting features
  • SKILL.md covers Version Compatibility, The Single Most Important…, Tool Taxonomy and Decision Tree by Scenario, plus 9 more sections
  • Runs Shell and Python scripts from its folder; calls pip

What it does

Bio Genome Intervals Gtf Gff Handling is an agent skill from GPTomics/bioSkills. Parses, queries, converts, and extracts from GTF and GFF3 gene-model annotation files - walking the gene/transcript/exon/CDS hierarchy with gffutils (queryable SQLite DB), converting formats and extracting transcript/CDS/protein FASTA with gffread, slurping to dataframes with gtfparse/pyranges, and sanitizing malformed files with AGAT. Covers the 1-based-inclusive vs 0-based BED coordinate conversion (start-1 only), deriving implicit features (introns/UTRs/TSS), phase-not-frame, the stop-codon-in-or-out-of-CDS…

Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/gffread_convert.sh`, `examples/parse_gtf.py` and `usage-guide.md`).

It sits in Data & Analytics, covering Bioinformatics and DataFrames. It works with SQLite and pandas. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Extracting features
  • Sequences from an annotation
  • Converting GTF<-GFF3
  • Traversing the gene tree

Example prompts

  • “/bio-genome-intervals-gtf-gff-handling”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The coordinate conversion is asymmetric. GTF/GFF3 are 1-based fully inclusive [start, end]; BED (and pyranges-internal) are 0-based…
  2. The all-zero count matrix. featureCounts/htseq-count match a read to a feature by string equality on the chromosome name, so chr1 != 1 !=…
  3. phase is not frame, and the stop codon is a 3-bp ghost. Phase (column 8) is the strand-aware count of bases to trim from the segment's…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Genome Intervals Gtf Gff Handling loads about 4.6k tokens when it runs. Until then it costs about 221 tokens; SKILL.md has 1,969 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~221
When it runs · the whole SKILL.md, loaded when a task matches
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,969 words, ~4,598 tokens.

Download SKILL.mdSave it as .claude/skills/bio-genome-intervals-gtf-gff-handling/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-genome-intervals-gtf-gff-handling
description
Parses, queries, converts, and extracts from GTF and GFF3 gene-model annotation files - walking the gene/transcript/exon/CDS hierarchy with gffutils (queryable SQLite DB), converting formats and extracting transcript/CDS/protein FASTA with gffread, slurping to dataframes with gtfparse/pyranges, and sanitizing malformed files with AGAT. Covers the 1-based-inclusive vs 0-based BED coordinate conversion (start-1 only), deriving implicit features (introns/UTRs/TSS), phase-not-frame, the stop-codon-in-or-out-of-CDS convention, and the chr1-vs-1 seqid and gene-ID-version mismatches that silently produce all-zero count matrices and dropped joins. Use when extracting features or sequences from an annotation, converting GTF<->GFF3 or GTF->BED, traversing the gene tree, or diagnosing a coordinate/provenance mismatch upstream of counting or DE.
tool_type
mixed
primary_tool
gffutils

Version Compatibility

Reference examples tested with: gffutils 0.13+, gffread 0.12+, gtfparse 2.x, pyranges 0.1+ (or 1.0+ - see note), AGAT 1.4+.

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures

Two version landmines specific to this skill: (1) gtfparse changed its return type - older releases returned a pandas DataFrame, gtfparse >=2.x returns a polars DataFrame by default; pass result_type='pandas' before chaining pandas idioms (.copy(), boolean masks). (2) pyranges has a major-version API split - pyranges 0.x and the 1.0 rewrite differ in method names and attribute access; check import pyranges; pyranges.__version__ before pasting code. gffutils stores 1-based coordinates while pyranges stores 0-based - their start fields differ by one by design. If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt rather than retrying.

GTF/GFF Handling

"Pull these features (or their sequences) out of my annotation, convert it, or find out why my counts are wrong." -> Treat the file as a serialized gene-model tree: walk gene->transcript->exon/CDS, derive implicit features, and reconcile coordinate and namespace conventions before trusting any number.

  • CLI: gffread in.gtf -T -o out.gtf (convert), gffread -w tx.fa -g genome.fa in.gtf (FASTA), agat_convert_sp_gxf2gxf.pl (sanitize)
  • Python: gffutils.create_db(...) then db.children(gene, featuretype='exon') (tree query); gtfparse.read_gtf(..., result_type='pandas') / pyranges.read_gtf(...) (dataframe)

The Single Most Important Modern Insight -- A GTF/GFF3 Is a Serialized Gene-Model Tree, Not a Table of Intervals

Almost every painful bug here comes from the file looking like a CSV while behaving like a tree, or from a coordinate/provenance mismatch the tools never warn about - and every one of these failures is silent: nothing throws, the wrong answer just propagates. Three load-bearing facts the tutorials skip:

  1. The coordinate conversion is asymmetric. GTF/GFF3 are 1-based fully inclusive [start, end]; BED (and pyranges-internal) are 0-based half-open [start-1, end). Convert to BED by subtracting 1 from the start only - the end is unchanged (the inclusive 1-based end and the exclusive 0-based end are the same integer). Doing start-1 AND end-1 shifts the feature one base left and is the classic over-correction: invisible in coverage/overlap, catastrophic in CDS translation (a one-base frameshift garbles the protein). gffutils keeps 1-based, pyranges stores 0-based, so their start fields differ by one correctly - never "fix" that discrepancy.

  2. The all-zero count matrix. featureCounts/htseq-count match a read to a feature by string equality on the chromosome name, so chr1 != 1 != NC_000001.11 produces a perfectly well-formed matrix of zeros with no error or warning - the only signal is ~0% assigned in the summary. Same bug one altitude up: gene-ID version suffixes (ENSG00000223972.5 vs ENSG00000223972) silently drop rows on an annotation join. Audit every cross-file key (chromosomes between BAM/GTF/FASTA, gene IDs between GTF/count-matrix/annotation) by set intersection, never by eye, before any count or join.

  3. phase is not frame, and the stop codon is a 3-bp ghost. Phase (column 8) is the strand-aware count of bases to trim from the segment's transcriptional 5' end to reach the next codon (0/1/2) - recompute it on any CDS edit (AGAT/gffread do; hand-editing coordinates without fixing phase frameshifts the translation). GTF (Ensembl/GENCODE) excludes the stop codon from CDS; GenBank/GFF3 often include it - so a CDS length off by exactly 3 nt (or a protein +/-1 stop) between two sources is a convention mismatch, not a bug.

Tool Taxonomy

ToolRoleMechanismWhen
gffutilsQueryable gene-tree DB (Python)builds a SQLite DB; children/parents/region traverse the hierarchy; keeps 1-based coordswalk gene->transcript->exon/CDS, derive introns, query by ID/coordinate
gffreadConverter + sequence extractor (CLI)fast C++, genome-aware; one binaryGTF<->GFF3, extract transcript/CDS/protein FASTA, region filter
pyrangesVectorized interval engine (Python)PyRanges/pandas-like; stores 0-based half-openoverlap joins, set ops, dataframe-native interval work
gtfparseGTF -> dataframe (Python)one call explodes column 9 into attribute columnsquick column/filter work; NOT hierarchy-aware (flat table)
AGATGFF/GTF sanitizer (Perl CLI)reconstructs the full tree; adds missing features, fixes IDs/phase, deflates attributesa malformed/non-standard file - run FIRST, before parsing

Decision Tree by Scenario

ScenarioRecommendedWhy
Walk the gene/transcript/exon hierarchy, derive intronsgffutils create_db + children/parentsthe hierarchy is the point; flat parsers lose it
Convert GTF<->GFF3 or extract transcript/CDS/protein FASTAgffread (-T, -w/-x/-y -g)genome-aware, knows the stop-codon convention
Quick column/filter on a clean modern GTFgtfparse (result_type='pandas')one-call dataframe; verify return type first
Overlap/set ops, large in-memory interval joinspyrangesvectorized; route arithmetic -> interval-arithmetic
File malformed: no ##gff-version, missing gene/exon lines, dup IDs, mixed conventionsAGAT agat_convert_sp_gxf2gxf.pl firstsanitize once vs writing a brittle parser around it
GTF -> BED for bedtoolsstart-1, end unchanged (-> bed-file-basics)the off-by-one boundary is where it bites
Counts came out all-zero or DE join dropped rowsintersect seqid / gene-ID namespacesstring-equality match; no error is emitted
Counting reads per gene/feature-> rna-quantification/featurecounts-countingthe seqid/strand landmines live there; set -s from chemistry
Judge whether the annotation itself is sound-> genome-annotation/annotation-qcthis skill operates on the file, not its quality

Walk the Gene Tree and Derive Introns (gffutils)

Goal: Traverse gene -> transcript -> exon and reconstruct features (introns) that the file does not store explicitly.

Approach: Build a SQLite DB once (disabling gene/transcript inference when those lines already exist, for a ~100x speedup), then query children ordered by position and synthesize introns from the exon gaps.

python
import gffutils

# disable_infer_* is GTF-only and applies when gene/transcript lines ALREADY exist (modern GENCODE/Ensembl) -> ~100x faster
db = gffutils.create_db('annotation.gtf', 'annotation.db', force=True,
                        disable_infer_genes=True, disable_infer_transcripts=True,
                        merge_strategy='create_unique')

gene = db['ENSG00000141510']                                    # gffutils returns 1-based coords (raw record)
for tx in db.children(gene, featuretype=['mRNA', 'transcript'], order_by='start'):
    exons = list(db.children(tx, featuretype='exon', order_by='start'))
    introns = list(db.interfeatures(exons, new_featuretype='intron'))   # introns are not stored - derived from exon gaps
    print(tx.id, len(exons), 'exons', len(introns), 'introns')

A modern GTF without disable_infer_* triggers the slow inference/merge machinery; an older minimal GTF lacking gene/transcript lines needs inference ON so gffutils reconstructs the envelopes. Match the flag to the file. introns, UTRs (exon - CDS), and TSS are derived, not stored - never infer biological absence from a missing feature line.

Convert Formats and Extract Sequences (gffread)

gffread is genome-aware and respects the stop-codon convention, so it is the safe path for sequence extraction (naive coordinate math is not).

bash
gffread annotation.gff3 -T -o annotation.gtf            # GFF3 -> GTF2 (default output is GFF3)
gffread -w transcripts.fa -g genome.fa annotation.gtf   # spliced exon (mature transcript) FASTA
gffread -x cds.fa         -g genome.fa annotation.gtf   # spliced CDS nucleotide FASTA
gffread -y proteins.fa    -g genome.fa annotation.gtf   # translated-CDS protein FASTA
gffread annotation.gtf -C -o coding.gtf                 # keep only coding transcripts

-g needs the genome FASTA (gffread auto-creates the .fai). -w/-x/-y splice the segments per transcript, so they handle multi-exon models correctly - do not concatenate exon FASTAs by hand.

Convert GTF to BED with the Right Coordinate Shift

Goal: Emit a BED of a chosen feature type for bedtools, without the off-by-one frameshift.

Approach: Parse to a pandas frame, filter to the feature type, subtract 1 from the start only, leave the end untouched.

python
import gtfparse

df = gtfparse.read_gtf('annotation.gtf', result_type='pandas')   # gtfparse >=2.x defaults to POLARS - force pandas
genes = df[df['feature'] == 'gene'].copy()
genes['start'] = genes['start'] - 1                              # 1-based inclusive -> 0-based half-open: START ONLY
bed = genes[['seqname', 'start', 'end', 'gene_id', 'score', 'strand']]
bed.to_csv('genes.bed', sep='\t', header=False, index=False)

For TSS/promoter derivation (strand-aware: + strand TSS = start, - strand TSS = end), route to proximity-operations - the promoter window is an imposed definition, not an annotated feature.

Sanitize a Malformed File First (AGAT)

When a file lacks ##gff-version 3, has non-Sequence-Ontology types, is missing gene/exon/UTR lines, has duplicate IDs, or mixes conventions, sanitize it once rather than coding around it:

bash
agat_convert_sp_gxf2gxf.pl -g messy.gff3 -o clean.gff3   # adds missing ID/Parent + features, fixes dup IDs, recomputes phase, sorts
agat_convert_sp_gff2gtf.pl -g clean.gff3 -o clean.gtf    # GFF3 -> GTF (collapses level1->gene, level2->transcript)

AGAT makes decisions (which convention to standardize to, how to derive missing features) - usually a feature, but when a source's exact encoding must be preserved (e.g. auditing a submission), inspect what it changed rather than trusting blindly.

Per-Method Failure Modes

Over-correcting the coordinate conversion

Trigger: subtracting 1 from both start and end when converting to BED. Mechanism: only the start representation differs; the inclusive 1-based end equals the exclusive 0-based end. Symptom: every feature shifted one base left; invisible in coverage, frameshifts CDS translation. Fix: start-1, end unchanged.

Comparing coordinates across gffutils and pyranges

Trigger: asserting equality on start fields from both libraries in one script. Mechanism: gffutils keeps 1-based, pyranges stores 0-based. Symptom: an off-by-one that looks like a bug; "fixing" it introduces a real error. Fix: confirm each library's convention; expect the difference.

Show full SKILL.md (765 more words)Show less
All-zero count matrix (seqid mismatch)

Trigger: BAM aligned to chr1, GTF annotated with 1. Mechanism: counters match reads to features by chromosome-name string equality. Symptom: well-formed matrix of zeros, no error; ~0% assigned in the summary. Fix: intersect the BAM @SQ/idxstats chromosome set with the GTF column-1 set programmatically; remap one namespace, re-confirm.

Dropped rows on a gene-ID join

Trigger: count matrix keyed ENSG... joined to annotation keyed ENSG....5. Mechanism: exact string match on a versioned vs unversioned ID. Symptom: join returns a dataframe but rows vanish / annotation is NA. Fix: strip .\d+$ on both sides for matching; keep the version in the stored annotation for provenance.

CDS length off by exactly 3

Trigger: comparing CDS/protein length across two sources, or translating after a convention-flipping conversion. Mechanism: GTF excludes the stop codon from CDS; GenBank/GFF3 often include it. Symptom: length differs by 3 nt / 1 aa; protein does/does not end in *. Fix: suspect the convention before debugging code; extract CDS with gffread/AGAT, which know it.

Editing CDS coordinates without recomputing phase

Trigger: trimming/merging/lifting CDS coordinates, leaving column 8 as-is. Mechanism: phase is a static integer; the chain of per-segment phases depends on cumulative coding length. Symptom: downstream translation (gffread -y, table2asn, EMBL) frameshifts or rejects. Fix: treat a CDS edit + phase recompute as one atomic operation; let AGAT/gffread recompute.

gffutils pathologically slow on a modern GTF

Trigger: create_db on a GENCODE/Ensembl GTF without the infer flags. Mechanism: gffutils infers gene/transcript envelopes and runs the merge machinery. Symptom: create_db hangs for many minutes. Fix: disable_infer_genes=True, disable_infer_transcripts=True when those lines already exist (~100x faster).

Quantitative Thresholds

Convention / thresholdSourceRationale
GTF/GFF3 1-based inclusive; convert to BED with start-1, end unchangedUCSC/SO format specsthe inclusive 1-based end == the exclusive 0-based end; over-correcting both ends frameshifts CDS
CDS length differs by exactly 3 nt between sourcesGTF vs GenBank/GFF3 stop-codon conventionGTF (Ensembl/GENCODE) excludes the stop from CDS; GenBank/GFF3 often include it
disable_infer_* -> ~100x create_db speedupgffutils docsinference/merge machinery is skipped when gene/transcript lines already exist
seqid intersection required before countingfeatureCounts/htseq string-equality matchnon-overlapping chromosome names -> all-zero matrix with no error
Strip .\d+$ from gene IDs on both sides before a joinEnsembl/GENCODE/RefSeq versioned accessionsversion suffix tracks model revision; mismatch drops rows silently
featureCounts default -s 0 (unstranded) vs htseq-count -s yes (stranded)tool defaults (Liao 2014; Anders 2015)switching tools changes the counting model; set -s from library chemistry, not the default

Common Errors

Error / symptomCauseSolution
All genes count zeroseqid mismatch (chr1 vs 1 vs NC_...)intersect BAM and GTF chromosome sets; remap one namespace
Counts low and flip when -s changeswrong strandedness (featureCounts vs htseq defaults differ)set -s from the library prep chemistry; verify assignment rate
Join drops rows / NA annotationgene-ID version suffix (ENSG....5 vs ENSG...)strip .\d+$ on both sides for matching
Biotype filter returns emptyattribute key differs by source: GENCODE uses gene_type, Ensembl/RefSeq use gene_biotypecheck the actual key (it travels with the chr1-vs-1 provenance split); query the present key
Translated protein is garbageover-corrected coordinate (start-1 AND end-1)subtract 1 from start only
CDS/protein off by 3 nt / 1 aastop-codon-in-or-out-of-CDS conventionextract with gffread/AGAT; do not debug coordinate math
gtfparse pandas idioms raise AttributeErrorgtfparse >=2.x returns polarspass result_type='pandas'
pyranges AttributeError0.x vs 1.0 API mismatchcheck pyranges.__version__; use matching method names
gffutils create_db hangsinfer machinery on a modern GTFset disable_infer_genes=True, disable_infer_transcripts=True
gffread -w/-x/-y errorsmissing or unindexed genome FASTApass -g genome.fa (gffread creates the .fai)

References

  • Pertea G, Pertea M. 2020. GFF Utilities: GffRead and GffCompare. F1000Research 9:304.
  • Stovner EB, Saetrom P. 2020. PyRanges: efficient comparison of genomic intervals in Python. Bioinformatics 36:918-919.
  • Dale R. gffutils: GFF and GTF file manipulation and interconversion. Software, https://github.com/daler/gffutils (no journal publication).
  • Dainat J. AGAT: Another Gff Analysis Toolkit to handle annotations in any GTF/GFF format. Zenodo. doi:10.5281/zenodo.3552717.
  • Rubinsteyn A, et al. gtfparse: parsing tools for GTF (gene transfer format) files. Software, https://github.com/openvax/gtfparse (no journal publication).
  • Liao Y, Smyth GK, Shi W. 2014. featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30:923-930.
  • Anders S, Pyl PT, Huber W. 2015. HTSeq - a Python framework to work with high-throughput sequencing data. Bioinformatics 31:166-169.
  • bed-file-basics - BED format and the coordinate conversion this skill feeds into
  • interval-arithmetic - Set operations on the features extracted here
  • proximity-operations - Strand-aware TSS/promoter derivation from extracted features
  • rna-quantification/featurecounts-counting - Consumes the GTF/GFF features; the seqid/strand landmines live there
  • genome-annotation/functional-annotation - Downstream of feature/sequence extraction from the annotation
  • genome-annotation/annotation-qc - Judges whether the annotation this skill parses is sound
  • differential-expression/de-results - Map gene coordinates back to DE results

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in genome-intervals/gtf-gff-handling of GPTomics/bioSkills.

  • SKILL.md
  • examples/gffread_convert.sh
  • examples/parse_gtf.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Genome Intervals Gtf Gff Handling next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Genome Intervals Gtf Gff Handling compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Genome Intervals Gtf Gff Handling this skillGPTomics/bioSkills1.2k1 repos~4.6kAutomated safety check: PassMIT
Polars BioClawBio/ClawBio1.2k—~3.4kAutomated safety check: PassApache-2.0
Polars BioK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: NotesApache-2.0
Pydeseqaipoch/medical-research-skills1.9k—~1.8kAutomated safety check: PassMIT
Sc GrnTianGzlab/OmicsClaw161—~1.7kAutomated safety check: PassApache-2.0
Chdb Datastorevemetric/vemetric3952 repos~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Polars Bio

    ClawBio/ClawBio

    Fast genomic interval operations (overlap, nearest, merge, coverage, cluster, complement, subtract, count-overlaps), multi-format bioinformatics I/O, DataFusion SQL, and pileup on Polars DataFrames…

    1.2k GitHub stars~3.4k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Polars Bio

    K-Dense-AI/scientific-agent-skills

    Performs genomic interval overlap, nearest, merge, coverage, complement and subtraction on Polars DataFrames, and reads or writes BED, VCF, BCF, BAM, CRAM, GFF, GTF, FASTA and FASTQ data.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Data & AnalyticsAuto-check: notes
  • Pydeseq

    aipoch/medical-research-skills

    Differential gene expression analysis for bulk RNA-seq count matrices using a DESeq2-like workflow in Python; use when you need Wald tests, FDR correction, and optional LFC shrinkage for…

    1.9k GitHub stars~1.8k tokensUpdated 24 days ago
    Data & AnalyticsAuto-check passed
  • Sc Grn

    TianGzlab/OmicsClaw

    Load when inferring TF → target gene regulatory networks on a normalised scRNA AnnData via pySCENIC (GRNBoost2 + cisTarget + AUCell) or correlation-based GRN fallback (when arboreto is unavailable…

    161 GitHub stars~1.7k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Chdb Datastore

    vemetric/vemetric

    A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

    395 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • CSV Data Summarizer

    coffeefuelbump/csv-data-summarizer-claude-skill

    Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.

    468 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Genome Intervals Gtf Gff Handling

What does Bio Genome Intervals Gtf Gff Handling do?

Parses, queries, converts, and extracts from GTF and GFF3 gene-model annotation files - walking the gene/transcript/exon/CDS hierarchy with gffutils (queryable SQLite DB), converting formats and…. Bio Genome Intervals Gtf Gff Handling is an agent skill from GPTomics/bioSkills. Parses, queries, converts, and extracts from GTF and GFF3 gene-model annotation files - walking the gene/transcript/exon/CDS hierarchy with gffutils (queryable SQLite DB), converting formats and extracting transcript/CDS/protein FASTA with gffread, slurping to dataframes with gtfparse/pyranges, and sanitizing malformed files with AGAT.

When should I use Bio Genome Intervals Gtf Gff Handling?

Bio Genome Intervals Gtf Gff Handling fits situations like: extracting features; sequences from an annotation; converting GTF<-GFF3; traversing the gene tree.

How do I install Bio Genome Intervals Gtf Gff Handling in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-intervals-gtf-gff-handling -a claude-code`. Or copy the skill folder (genome-intervals/gtf-gff-handling in GPTomics/bioSkills) into .claude/skills/bio-genome-intervals-gtf-gff-handling in your project. Claude Code loads it when a task matches its description.

How do I install Bio Genome Intervals Gtf Gff Handling in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-genome-intervals-gtf-gff-handling -a codex`. Or copy the skill folder (genome-intervals/gtf-gff-handling in GPTomics/bioSkills) into .agents/skills/bio-genome-intervals-gtf-gff-handling in your project. Codex loads it when a task matches its description.

Can I use Bio Genome Intervals Gtf Gff Handling in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-genome-intervals-gtf-gff-handling -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-genome-intervals-gtf-gff-handling, .gemini/skills/bio-genome-intervals-gtf-gff-handling, .github/skills/bio-genome-intervals-gtf-gff-handling and .opencode/skills/bio-genome-intervals-gtf-gff-handling in your project.

What does Bio Genome Intervals Gtf Gff Handling need to run?

Going by SKILL.md and its folder, Bio Genome Intervals Gtf Gff Handling needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Genome Intervals Gtf Gff Handling access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Bio Genome Intervals Gtf Gff Handling safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Genome Intervals Gtf Gff Handling use?

Bio Genome Intervals Gtf Gff Handling is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Genome Intervals Gtf Gff Handling use?

About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Genome Intervals Gtf Gff Handling?

Skills that share tags, products or a category with Bio Genome Intervals Gtf Gff Handling: Polars Bio (ClawBio/ClawBio, 1.2k stars), Polars Bio (K-Dense-AI/scientific-agent-skills, 48k stars), Pydeseq (aipoch/medical-research-skills, 1.9k stars) and Sc Grn (TianGzlab/OmicsClaw, 161 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Genome Intervals Gtf Gff Handling?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.