Agent skill

Bio Rna Quantification Tximport Workflow

by GPTomics in GPTomics/bioSkills

Import transcript-level quantifications from Salmon/kallisto/RSEM into R for gene-level analysis with DESeq2/edgeR using tximport or tximeta.

MITAuto-check passed

Install Bio Rna Quantification Tximport Workflow

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-rna-quantification-tximport-workflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-rna-quantification-tximport-workflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/rna-quantification/tximport-workflow .claude/skills/bio-rna-quantification-tximport-workflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-rna-quantification-tximport-workflow
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.8k tokens
SKILL.md length
965 words
Files
4
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Import transcript-level quantifications from Salmon/kallisto/RSEM into R for gene-level analysis with DESeq2/edgeR using tximport or tximeta.

  • Summarizing transcript abundances to gene counts with the correct length offset
  • SKILL.md covers Version Compatibility, Why this is not just a sum, Basic tximport and Decision: countsFromAbundance…, plus 9 more sections
  • Runs R scripts from its folder
  • Choosing a countsFromAbundance mode (full-length vs 3-tag vs DTU)

What it does

Bio Rna Quantification Tximport Workflow is an agent skill from GPTomics/bioSkills. Import transcript-level quantifications from Salmon/kallisto/RSEM into R for gene-level analysis with DESeq2/edgeR using tximport or tximeta. Use when summarizing transcript abundances to gene counts with the correct length offset, choosing a countsFromAbundance mode (full-length vs 3'-tag vs DTU), resolving transcript-ID version mismatches, or handing off to DESeq2/edgeR without double-applying the offset.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `usage-guide.md`).

The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Summarizing transcript abundances to gene counts with the correct length offset
  • Choosing a countsFromAbundance mode (full-length vs 3-tag vs DTU)
  • Resolving transcript-ID version mismatches
  • Handing off to DESeq2/edgeR without double-applying the offset

Example prompts

  • “/bio-rna-quantification-tximport-workflow”

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Rna Quantification Tximport Workflow loads about 2.8k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 965 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 965 words, ~2,783 tokens.

Download SKILL.mdSave it as .claude/skills/bio-rna-quantification-tximport-workflow/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-rna-quantification-tximport-workflow
description
Import transcript-level quantifications from Salmon/kallisto/RSEM into R for gene-level analysis with DESeq2/edgeR using tximport or tximeta. Use when summarizing transcript abundances to gene counts with the correct length offset, choosing a countsFromAbundance mode (full-length vs 3'-tag vs DTU), resolving transcript-ID version mismatches, or handing off to DESeq2/edgeR without double-applying the offset.
tool_type
r
primary_tool
tximport

Version Compatibility

Reference examples tested with: tximport 1.30+, tximeta 1.20+, DESeq2 1.42+, edgeR 4.0+, txdbmaker 1.0+, Salmon 1.10+, kallisto 0.50+

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

tximport Workflow

"Import Salmon/kallisto results into DESeq2" -> Summarize transcript-level abundance estimates to gene-level counts AND compute a per-gene, per-sample length offset that the DE model consumes.

  • R: tximport::tximport(files, type='salmon', tx2gene=tx2gene)

Why this is not just a sum

tximport does not merely add transcript counts to a gene total. The fragment count a gene produces depends on the average length of the isoforms expressed in that sample, because longer molecules yield more fragments (more start positions). When isoform usage shifts between conditions (differential transcript usage), the gene's average effective length changes, so a naive summed count is length-biased in a condition-correlated way and masquerades as differential expression. tximport corrects this by returning a per-gene, per-sample average-length matrix (txi$length) and passing it as a normalization offset to DESeq2/edgeR. A single per-gene length cannot capture this because the bias is sample-specific (Soneson, Love, Robinson 2015).

Basic tximport

Goal: Import transcript-level quantifications into R as gene-level counts plus the length offset for DESeq2 or edgeR.

Approach: Build a transcript-to-gene map, then run tximport over the quant files; the returned txi$counts/txi$length carry both the gene counts and the offset.

r
library(tximport)

files <- c(sample1 = 'sample1_quant/quant.sf',
           sample2 = 'sample2_quant/quant.sf',
           sample3 = 'sample3_quant/quant.sf')

tx2gene <- read.csv('tx2gene.csv')                 # column order: TXNAME, then GENEID
txi <- tximport(files, type = 'salmon', tx2gene = tx2gene)

txi is a list: $abundance (TPM), $counts (estimated counts), $length (the average-length offset source), $countsFromAbundance.

Decision: countsFromAbundance mode

This argument silently determines correctness; nothing errors when it is wrong.

ModeWhat it returnsUse when
'no' (default)Estimated counts + separate length offsetFull-length library -> DESeq2/edgeR (they consume the offset). The cleanest path.
'lengthScaledTPM'Counts with the length correction baked in, no separate offsetA tool that cannot take an offset (e.g. limma-voom)
'scaledTPM'TPM scaled to library size, no length scalingTranscript-level DTU with txOut=TRUE (DRIMSeq/DEXSeq); the established Love et al. workflow input
'dtuScaledTPM'Scaled by median isoform lengthDTU alternative (tximport >= 1.10), needs tx2gene; helps when isoform lengths within a gene differ widely

For DTU, scaledTPM is the established default; dtuScaledTPM is the newer purpose-built mode, preferable when a gene's isoforms span very different lengths.

The 3'-tag exception: for 3'-end protocols (10x, QuantSeq, Lexogen) a read count does not scale with transcript length, so there is no length bias to correct, and length-correcting injects one. Do not use the length-scaled modes (lengthScaledTPM/dtuScaledTPM) for tag-seq. Import with the default, but build the DESeqDataSet from the plain counts so the length offset is NOT auto-applied:

r
# 3'-tag: bypass the length offset that DESeqDataSetFromTximport would otherwise apply
dds <- DESeqDataSetFromMatrix(round(txi$counts), colData = coldata, design = ~ condition)
r
# Full-length, DESeq2/edgeR (default): keep the offset path
txi <- tximport(files, type = 'salmon', tx2gene = tx2gene)

# Transcript-level for DTU (hand off to alternative-splicing/isoform-switching)
txi_tx <- tximport(files, type = 'salmon', txOut = TRUE,
                   countsFromAbundance = 'scaledTPM')

Creating tx2gene

The map is a two-column data frame; column ORDER is load-bearing (TXNAME first, GENEID second), names do not matter.

Goal: Map every quantified transcript ID to its gene, with IDs that exactly match the quant files.

Approach: Derive from the annotation that built the index (GTF, ensembldb, biomaRt, or the index t2g); strip version suffixes to match.

r
# From a GTF: makeTxDbFromGFF moved to txdbmaker in Bioconductor >= 3.19
# (defunct in GenomicFeatures >= 1.61.1; on older Bioconductor use GenomicFeatures::makeTxDbFromGFF)
library(txdbmaker)
txdb <- makeTxDbFromGFF('annotation.gtf')
k <- keys(txdb, keytype = 'TXNAME')
tx2gene <- AnnotationDbi::select(txdb, keys = k, keytype = 'TXNAME',
                                 columns = c('TXNAME', 'GENEID'))

# From biomaRt (useEnsembl; useMart is deprecated)
library(biomaRt)
mart <- useEnsembl(biomart = 'genes', dataset = 'hsapiens_gene_ensembl')
tx2gene <- getBM(attributes = c('ensembl_transcript_id', 'ensembl_gene_id'), mart = mart)

The #1 silent failure: transcript-ID version mismatch

If quant.sf IDs carry version suffixes (ENST00000456328.4) but tx2gene does not (or vice versa), the IDs do not match. Total non-overlap raises an error; partial mismatch silently drops the non-matching transcripts and prints a summary, deflating affected genes toward zero. Fix by stripping versions consistently or with ignoreTxVersion:

r
txi <- tximport(files, type = 'salmon', tx2gene = tx2gene,
                ignoreTxVersion = TRUE, ignoreAfterBar = TRUE)
Show full SKILL.md (412 more words)Show less

Handoff to DESeq2 (offset applied automatically)

Goal: Build a DESeqDataSet that uses the tximport length offset without any manual step.

Approach: DESeqDataSetFromTximport stores txi$length as the avgTxLength assay and converts it to per-gene normalization factors inside DESeq().

r
library(DESeq2)
coldata <- data.frame(condition = factor(c('control', 'control', 'treated', 'treated')),
                      row.names = names(files))
dds <- DESeqDataSetFromTximport(txi, colData = coldata, design = ~ condition)
dds <- dds[rowSums(counts(dds)) >= 10, ]   # light pre-filter (speed); results() does the inferential filter
dds <- DESeq(dds)
res <- results(dds)

Passing a countsFromAbundance='no' txi prints "using counts and average transcript lengths from tximport"; a length-scaled txi prints "using just counts" and applies no offset. Both are handled correctly by the function.

Handoff to edgeR (manual offset)

Goal: Carry the length offset into an edgeR DGEList.

Approach: Geometric-mean-center the length matrix, fold in composition-corrected library sizes, log it, attach via scaleOffset.

r
library(edgeR)
cts <- txi$counts
normMat <- txi$length / exp(rowMeans(log(txi$length)))   # center each gene on its geometric mean
normCts <- cts / normMat
eff.lib <- calcNormFactors(normCts) * colSums(normCts)
normMat <- sweep(normMat, 2, eff.lib, '*')
y <- scaleOffset(DGEList(cts), log(normMat))
y <- y[filterByExpr(y, group = coldata$condition), , keep.lib.sizes = FALSE]   # group-aware filter

Do not double-apply the offset: if countsFromAbundance='lengthScaledTPM' already baked the correction into the counts, do not also attach a length offset. Use 'no' for the offset path, the scaled modes for the no-offset path, never both.

Transcript-level uncertainty (DTE/DTU)

Gene-level estimates are robust because per-isoform assignment uncertainty cancels on summation. Transcript-level testing must propagate it: edgeR catchSalmon deflates counts by per-transcript overdispersion (differential-expression/edger-basics), swish/fishpond tests across Salmon Gibbs samples (alternative-splicing/isoform-switching), and sleuth uses kallisto bootstraps (expression-matrix/counts-ingest). Generate the replicates at quantification time (rna-quantification/alignment-free-quant).

tximeta: provenance by checksum

tximeta hashes the index's reference sequences and looks the digest up against known GENCODE/Ensembl/RefSeq releases, attaching transcript ranges and release metadata automatically, so the exact reference becomes a verified property of the object rather than lab lore.

r
library(tximeta)
makeLinkedTxome(indexDir = 'salmon_index', source = 'Ensembl', organism = 'Homo sapiens',
                release = '110', genome = 'GRCh38', fasta = 'transcripts.fa', gtf = 'annotation.gtf')
coldata <- data.frame(names = names(files), files = files,
                      condition = c('control', 'control', 'treated', 'treated'))
se <- tximeta(coldata)
gse <- summarizeToGene(se)
dds <- DESeqDataSet(gse, design = ~ condition)

Common Errors

SymptomCauseFix
Many genes import as zero or deflatedTranscript-ID version mismatch (partial drop)ignoreTxVersion = TRUE; or strip \.\d+$ from both sides
Error: none of the transcripts present in tx2geneTotal ID mismatch (versions or wrong annotation)Rebuild tx2gene from the annotation that built the index
Summarized at the wrong level, no errortx2gene columns reversed (GENEID first)Order as TXNAME, then GENEID
Length bias appears in 3'-tag dataDESeqDataSetFromTximport auto-applied the length offsetBuild via DESeqDataSetFromMatrix(round(txi$counts), ...) so no offset is applied
Fold changes inflated near isoform switches with manual edgeROffset double-applied or omittedOne path only: 'no'+offset, or scaled mode without offset
  • rna-quantification/alignment-free-quant - Upstream Salmon/kallisto and inferential replicates
  • differential-expression/deseq2-basics - Gene-level DE from a DESeqDataSet
  • differential-expression/edger-basics - edgeR DE and catchSalmon transcript DTE
  • alternative-splicing/isoform-switching - DTU and swish from transcript-level import
  • expression-matrix/counts-ingest - sleuth and other quantifier ingestion paths
  • genome-intervals/gtf-gff-handling - Building tx2gene from a GTF

References

  • Soneson C, Love MI, Robinson MD. 2015. Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences. F1000Research 4:1521. doi:10.12688/f1000research.7563
  • Love MI, Soneson C, Hickey PF, et al. 2020. Tximeta: Reference sequence checksums for provenance identification in RNA-seq. PLoS Comput Biol 16(2):e1007664. doi:10.1371/journal.pcbi.1007664

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in rna-quantification/tximport-workflow of GPTomics/bioSkills.

  • SKILL.md
  • examples/create_tx2gene.R
  • examples/tximport_deseq2.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Rna Quantification Tximport Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Rna Quantification Tximport Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Rna Quantification Tximport Workflow this skillGPTomics/bioSkills1.2k1 repos~2.8kAutomated safety check: PassMIT
Salmon Rna Quantificationjaechang-hits/SciAgent-Skills3701 repos~4kAutomated safety check: PassGPL-3.0
Bio Rna Quantification Alignment Free Quantmajiayu000/claude-skill-registry6661 repos~1.3kAutomated safety check: PassMIT
Bio Splicing QuantificationFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~1.3kAutomated safety check: PassNone
Bio Rna Quantification Count Matrix Qcmajiayu000/claude-skill-registry6661 repos~1.5kAutomated safety check: PassMIT
Bio Proteomics QuantificationFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~1.2kAutomated safety check: PassNone

Similar skills

  • Salmon Rna Quantification

    jaechang-hits/SciAgent-Skills

    Ultra-fast RNA-seq transcript/gene quantification via quasi-mapping (no BAM).

    370 GitHub starsUsed in 1 repo~4k tokens
    Research & ScienceAuto-check passed
  • Bio Splicing Quantification

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quantifies alternative splicing events (PSI/percent spliced in) from RNA-seq using SUPPA2 from transcript TPM or rMATS-turbo from BAM files.

    3.1k GitHub stars~1.3k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Bio Rna Quantification Count Matrix Qc

    majiayu000/claude-skill-registry

    Quality control and exploration of RNA-seq count matrices before differential expression.

    666 GitHub starsUsed in 1 repo~1.5k tokens
    Research & ScienceAuto-check passed
  • Bio Proteomics Quantification

    FreedomIntelligence/OpenClaw-Medical-Skills

    Protein quantification from mass spectrometry data including label-free (LFQ, intensity-based), isobaric labeling (TMT, iTRAQ), and metabolic labeling (SILAC) approaches.

    3.1k GitHub starsUsed in 1 repo~1.2k tokens
    Research & ScienceAuto-check passed
  • Baoyu Youtube Transcript

    JimLiu/baoyu-skills

    Downloads YouTube video transcripts/subtitles and cover images by URL or video ID.

    26k GitHub starsUsed in 1 repo~2.4k tokens
    Media & CreativeAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Rna Quantification Tximport Workflow

What does Bio Rna Quantification Tximport Workflow do?

Import transcript-level quantifications from Salmon/kallisto/RSEM into R for gene-level analysis with DESeq2/edgeR using tximport or tximeta. Bio Rna Quantification Tximport Workflow is an agent skill from GPTomics/bioSkills. Import transcript-level quantifications from Salmon/kallisto/RSEM into R for gene-level analysis with DESeq2/edgeR using tximport or tximeta.

When should I use Bio Rna Quantification Tximport Workflow?

Bio Rna Quantification Tximport Workflow fits situations like: summarizing transcript abundances to gene counts with the correct length offset; choosing a countsFromAbundance mode (full-length vs 3-tag vs DTU); resolving transcript-ID version mismatches; handing off to DESeq2/edgeR without double-applying the offset.

How do I install Bio Rna Quantification Tximport Workflow in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-rna-quantification-tximport-workflow -a claude-code`. Or copy the skill folder (rna-quantification/tximport-workflow in GPTomics/bioSkills) into .claude/skills/bio-rna-quantification-tximport-workflow in your project. Claude Code loads it when a task matches its description.

How do I install Bio Rna Quantification Tximport Workflow in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-rna-quantification-tximport-workflow -a codex`. Or copy the skill folder (rna-quantification/tximport-workflow in GPTomics/bioSkills) into .agents/skills/bio-rna-quantification-tximport-workflow in your project. Codex loads it when a task matches its description.

Can I use Bio Rna Quantification Tximport Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-rna-quantification-tximport-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-rna-quantification-tximport-workflow, .gemini/skills/bio-rna-quantification-tximport-workflow, .github/skills/bio-rna-quantification-tximport-workflow and .opencode/skills/bio-rna-quantification-tximport-workflow in your project.

What does Bio Rna Quantification Tximport Workflow need to run?

Going by SKILL.md and its folder, Bio Rna Quantification Tximport Workflow needs R for the scripts in its folder.

Does Bio Rna Quantification Tximport Workflow access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Rna Quantification Tximport Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Rna Quantification Tximport Workflow use?

Bio Rna Quantification Tximport Workflow is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Rna Quantification Tximport Workflow use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Rna Quantification Tximport Workflow?

Skills that share tags, products or a category with Bio Rna Quantification Tximport Workflow: Salmon Rna Quantification (jaechang-hits/SciAgent-Skills, 370 stars), Bio Rna Quantification Alignment Free Quant (majiayu000/claude-skill-registry, 666 stars), Bio Splicing Quantification (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars) and Bio Rna Quantification Count Matrix Qc (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Rna Quantification Tximport Workflow?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.