Agent skill

Bio Rna Quantification Featurecounts Counting

by GPTomics in GPTomics/bioSkills

Count reads per gene from aligned BAM files using Subread featureCounts.

MITAuto-check passed

Install Bio Rna Quantification Featurecounts Counting

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-rna-quantification-featurecounts-counting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-rna-quantification-featurecounts-counting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/rna-quantification/featurecounts-counting .claude/skills/bio-rna-quantification-featurecounts-counting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-rna-quantification-featurecounts-counting
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.3k tokens
SKILL.md length
820 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Count reads per gene from aligned BAM files using Subread featureCounts.

  • Turning STAR/HISAT2 BAMs into a gene-level count matrix for DESeq2/edgeR
  • SKILL.md covers Version Compatibility, Basic Usage, Decision 1: strandedness (-s)… and Decision 2: paired-end…, plus 8 more sections
  • Runs Shell and Python scripts from its folder; calls pip
  • Deciding library strandedness

What it does

Bio Rna Quantification Featurecounts Counting is an agent skill from GPTomics/bioSkills. Count reads per gene from aligned BAM files using Subread featureCounts. Use when turning STAR/HISAT2 BAMs into a gene-level count matrix for DESeq2/edgeR, deciding library strandedness, handling paired-end fragment counting, choosing how to treat multi-mapping and multi-overlapping reads, or diagnosing a low assignment rate from the summary file.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/count_genes.sh`, `examples/process_counts.py` and `usage-guide.md`).

The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Turning STAR/HISAT2 BAMs into a gene-level count matrix for DESeq2/edgeR
  • Deciding library strandedness
  • Handling paired-end fragment counting
  • Choosing how to treat multi-mapping and multi-overlapping reads

Example prompts

  • “/bio-rna-quantification-featurecounts-counting”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Rna Quantification Featurecounts Counting loads about 2.3k tokens when it runs. Until then it costs about 99 tokens; SKILL.md has 820 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 820 words, ~2,282 tokens.

Download SKILL.mdSave it as .claude/skills/bio-rna-quantification-featurecounts-counting/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-rna-quantification-featurecounts-counting
description
Count reads per gene from aligned BAM files using Subread featureCounts. Use when turning STAR/HISAT2 BAMs into a gene-level count matrix for DESeq2/edgeR, deciding library strandedness, handling paired-end fragment counting, choosing how to treat multi-mapping and multi-overlapping reads, or diagnosing a low assignment rate from the summary file.
tool_type
cli
primary_tool
featureCounts

Version Compatibility

Reference examples tested with: Subread 2.0+, STAR 2.7.11+, HISAT2 2.2.1+, DESeq2 1.42+, edgeR 4.0+, pandas 2.2+

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags
  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

featureCounts Counting

"Count reads per gene from my BAM files" -> Assign each aligned read to at most one gene by overlap with a GTF, discarding ambiguous reads, to produce an integer gene-by-sample matrix for differential expression.

  • CLI: featureCounts -a genes.gtf -o counts.txt sample1.bam sample2.bam

featureCounts is bookkeeping, not inference: it tallies reads to genes and discards anything ambiguous. That is correct for gene-level DE, where almost every read's gene of origin is unambiguous even when its isoform is not. The two settings that silently corrupt the matrix if wrong are strandedness (-s) and, for paired-end data, fragment counting (--countReadPairs).

Basic Usage

bash
# Multiple samples in one run -> a single aligned matrix (recommended)
featureCounts -a annotation.gtf -o counts.txt sample1.bam sample2.bam sample3.bam

# Defaults: -t exon -g gene_id (count reads over exons, aggregate by gene)

Decision 1: strandedness (-s) is the load-bearing setting

-sMeaningRead 1
0Unstrandedstrand ignored
1Forward strandedread 1 is sense
2Reverse strandedread 1 is antisense; read 2 is sense

The dominant chemistry, dUTP / Illumina TruSeq Stranded / NEBNext Directional, is reverse, -s 2. Setting the wrong strand does not error; it silently destroys the matrix. For a truly stranded library, the correct -s assigns ~80-90% of reads while the opposite setting collapses to ~5-20% (counting only antisense background). Do not trust the kit name; determine it empirically:

bash
# Method A: RSeQC reports the strand pattern fractions
infer_experiment.py -r genes.bed -i sample.bam

# Method B: run all three and pick the one that maximizes Assigned in the .summary
for s in 0 1 2; do featureCounts -s $s -a annotation.gtf -o counts_s$s.txt sample.bam; done

If -s 1 and -s 2 give wildly different Assigned fractions, the data are stranded (use the higher); if both are roughly equal and about half of -s 0, the data are unstranded. STAR --quantMode GeneCounts provides a free cross-check (see below).

Decision 2: paired-end fragment counting

bash
# Subread >= 2.0.2: -p only declares paired input; --countReadPairs is REQUIRED to count fragments
featureCounts -p --countReadPairs -a annotation.gtf -o counts.txt *.bam

# Stricter: require both ends mapped, exclude chimeric/discordant pairs
featureCounts -p --countReadPairs -B -C -a annotation.gtf -o counts.txt *.bam

Omitting --countReadPairs on paired-end data counts each mate separately, roughly doubling counts and breaking the count model. -B requires both ends aligned; -C excludes pairs mapping across chromosomes or in the wrong orientation.

Decision 3: multi-mapping and multi-overlap reads

bash
# Default (recommended for gene-level DE): discard both -> uniquely, unambiguously assigned reads only
featureCounts -a annotation.gtf -o counts.txt *.bam

# Count multimappers fractionally (1/N) or fully (1 each) -- NOT recommended for DE
featureCounts -M --fraction -a annotation.gtf -o counts.txt *.bam
featureCounts -M -a annotation.gtf -o counts.txt *.bam

# Count reads overlapping >1 gene in all of them
featureCounts -O -a annotation.gtf -o counts.txt *.bam

Discarding multimappers is the right default for gene-level DE. -M --fraction looks principled but biases exactly the genes where resolution matters: a read truly from gene A that also maps to paralog A' is split 0.5/0.5, diluting both. This is the regime where alignment-free EM quantifiers (rna-quantification/alignment-free-quant) outperform featureCounts, because they reassign by full likelihood rather than a flat split.

Quality and feature options

bash
featureCounts -Q 10 -a annotation.gtf -o counts.txt *.bam      # min MAPQ (aligner-specific scale)
featureCounts --primary -a annotation.gtf -o counts.txt *.bam  # primary alignments only
featureCounts -t CDS -g gene_id -a annotation.gtf -o counts.txt *.bam  # count CDS instead of exon

-Q thresholds mapping quality, but MAPQ conventions are aligner-specific (STAR assigns 255 to unique reads, low values to multimappers), so confirm the scheme before choosing a cutoff. Do NOT add --ignoreDup for standard RNA-seq: high duplication is expected from highly expressed genes, and position-based deduplication discards real signal. Deduplicate only with UMIs. For exon-level usage testing (DEXSeq), use a flattened annotation rather than gene-level counting (alternative-splicing/isoform-switching).

Show full SKILL.md (335 more words)Show less

Output

counts.txt:           Geneid Chr Start End Strand Length sample1.bam sample2.bam ...
counts.txt.summary:   Status              sample1.bam  sample2.bam
                      Assigned            1523456      1678234
                      Unassigned_NoFeatures 234567     245678

Reading the .summary is the primary QC step. A good poly-A library assigns ~70-90% of mapped reads (rRNA-depletion libraries run lower).

Dominant unassigned categoryLikely causeAction
Unassigned_NoFeatures highGTF/genome mismatch (chr naming 1 vs chr1, wrong release), DNA contaminationMatch GTF release and chromosome naming to the BAM
Unassigned_MultiMapping highrRNA carryover or repetitive contentCheck rRNA depletion; inspect with FastQ Screen
Unassigned_Ambiguity highOverlapping/nested gene models or wrong feature levelExpected in gene-dense regions; reconsider -O/feature type
Assigned low, others spread thinWrong strandednessRe-test -s (see Decision 1)

STAR cross-check

If aligned with STAR --quantMode GeneCounts, ReadsPerGene.out.tab gives a free independent count: column 2 = unstranded (≈ -s 0), column 3 = forward (≈ -s 1), column 4 = reverse (≈ -s 2). The larger of columns 3 vs 4 reveals the strand directly, and the per-gene counts should track featureCounts at the matching -s.

Extract the matrix

bash
cut -f1,7- counts.txt | tail -n +2 > count_matrix.txt   # drop the 6 annotation columns
python
import pandas as pd
counts = pd.read_csv('counts.txt', sep='\t', comment='#')
mat = counts.set_index('Geneid').iloc[:, 5:]
mat.columns = [c.replace('.bam', '') for c in mat.columns]
mat.to_csv('count_matrix.csv')

Common Errors

SymptomCauseFix
Assigned ~half of expected, no errorWrong -s, or paired-end without --countReadPairs (double-counting)Determine strand empirically; add --countReadPairs for paired-end
Near-zero counts for known genesgene_id attribute or feature type mismatch with the GTFConfirm -t/-g match the annotation; check the GTF attribute names
Counts much higher than read countPaired-end mates counted separatelyAdd -p --countReadPairs
Inflated correlated paralog counts-M/-O fractional counting enabledDrop -M/-O for DE; use alignment-free EM for paralog-heavy genes
Low Assigned across all -s valuesGTF does not match the aligned genomeUse the GTF release and contig names matching the alignment reference
  • alignment-files/sam-bam-basics - Input BAM handling and filtering
  • read-alignment/star-alignment - Producing BAMs and ReadsPerGene.out.tab strand cross-check
  • genome-intervals/gtf-gff-handling - GTF/GFF annotation files
  • rna-quantification/alignment-free-quant - EM-based alternative; better for multimappers/paralogs
  • rna-quantification/count-matrix-qc - QC the resulting matrix before DE
  • differential-expression/deseq2-basics - Gene-level DE from these counts

References

  • Liao Y, Smyth GK, Shi W. 2014. featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30(7):923-930. doi:10.1093/bioinformatics/btt656
  • Liao Y, Smyth GK, Shi W. 2013. The Subread aligner: fast, accurate and scalable read mapping by seed-and-vote. Nucleic Acids Res 41(10):e108. doi:10.1093/nar/gkt214

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in rna-quantification/featurecounts-counting of GPTomics/bioSkills.

  • SKILL.md
  • examples/count_genes.sh
  • examples/process_counts.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Rna Quantification Featurecounts Counting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Rna Quantification Featurecounts Counting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Rna Quantification Featurecounts Counting this skillGPTomics/bioSkills1.2k1 repos~2.3kAutomated safety check: PassMIT
Featurecounts Rna Countingjaechang-hits/SciAgent-Skills3711 repos~3.7kAutomated safety check: NotesGPL-3.0
Salmon Rna Quantificationjaechang-hits/SciAgent-Skills3711 repos~4kAutomated safety check: PassGPL-3.0
Ccs Alignthedotmack/claude-mem99k—~6.1kAutomated safety check: PassApache-2.0
Star Rna Seq Alignerjaechang-hits/SciAgent-Skills3711 repos~4.1kAutomated safety check: PassMIT
Gene Databasedavila7/claude-code-templates32k10 repos~1.6kAutomated safety check: PassMIT

Similar skills

  • Featurecounts Rna Counting

    jaechang-hits/SciAgent-Skills

    Counts RNA-seq reads overlapping GTF gene features. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub starsUsed in 1 repo~3.7k tokens
    Research & ScienceAuto-check: notes
  • Salmon Rna Quantification

    jaechang-hits/SciAgent-Skills

    Ultra-fast RNA-seq transcript/gene quantification via quasi-mapping (no BAM).

    371 GitHub starsUsed in 1 repo~4k tokens
    Research & ScienceAuto-check passed
  • Ccs Align

    thedotmack/claude-mem

    Run the CCS Align seat's hourly breathing cycle — prove the local claude-mem worker is healthy, pull needle observations through search → timeline → getobservations, land them in a seat-owned middle…

    99k GitHub stars~6.1k tokensUpdated today
    Auto-check passed
  • Star Rna Seq Aligner

    jaechang-hits/SciAgent-Skills

    Splice-aware RNA-seq aligner producing sorted BAM and splice junction tables.

    371 GitHub starsUsed in 1 repo~4.1k tokens
    Research & ScienceAuto-check passed
  • Gene Database

    davila7/claude-code-templates

    Query NCBI Gene via E-utilities/Datasets API. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~1.6k tokens
    Research & ScienceAuto-check passed
  • Word Count

    thedaviddias/Front-End-Checklist

    A skill your agent uses when applies to key landing pages, blog posts, product pages, and any page targeting competitive queries.

    74k GitHub stars~577 tokensUpdated 3 days ago
    Frontend & DesignAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Rna Quantification Featurecounts Counting

What does Bio Rna Quantification Featurecounts Counting do?

Count reads per gene from aligned BAM files using Subread featureCounts. Bio Rna Quantification Featurecounts Counting is an agent skill from GPTomics/bioSkills. Count reads per gene from aligned BAM files using Subread featureCounts.

When should I use Bio Rna Quantification Featurecounts Counting?

Bio Rna Quantification Featurecounts Counting fits situations like: turning STAR/HISAT2 BAMs into a gene-level count matrix for DESeq2/edgeR; deciding library strandedness; handling paired-end fragment counting; choosing how to treat multi-mapping and multi-overlapping reads.

How do I install Bio Rna Quantification Featurecounts Counting in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-rna-quantification-featurecounts-counting -a claude-code`. Or copy the skill folder (rna-quantification/featurecounts-counting in GPTomics/bioSkills) into .claude/skills/bio-rna-quantification-featurecounts-counting in your project. Claude Code loads it when a task matches its description.

How do I install Bio Rna Quantification Featurecounts Counting in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-rna-quantification-featurecounts-counting -a codex`. Or copy the skill folder (rna-quantification/featurecounts-counting in GPTomics/bioSkills) into .agents/skills/bio-rna-quantification-featurecounts-counting in your project. Codex loads it when a task matches its description.

Can I use Bio Rna Quantification Featurecounts Counting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-rna-quantification-featurecounts-counting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-rna-quantification-featurecounts-counting, .gemini/skills/bio-rna-quantification-featurecounts-counting, .github/skills/bio-rna-quantification-featurecounts-counting and .opencode/skills/bio-rna-quantification-featurecounts-counting in your project.

What does Bio Rna Quantification Featurecounts Counting need to run?

Going by SKILL.md and its folder, Bio Rna Quantification Featurecounts Counting needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Bio Rna Quantification Featurecounts Counting access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Rna Quantification Featurecounts Counting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Rna Quantification Featurecounts Counting use?

Bio Rna Quantification Featurecounts Counting is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Rna Quantification Featurecounts Counting use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Rna Quantification Featurecounts Counting?

Skills that share tags, products or a category with Bio Rna Quantification Featurecounts Counting: Featurecounts Rna Counting (jaechang-hits/SciAgent-Skills, 371 stars), Salmon Rna Quantification (jaechang-hits/SciAgent-Skills, 371 stars), Ccs Align (thedotmack/claude-mem, 99k stars) and Star Rna Seq Aligner (jaechang-hits/SciAgent-Skills, 371 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Rna Quantification Featurecounts Counting?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.