Agent skill

Bio Read Alignment Star Alignment

by GPTomics in GPTomics/bioSkills

Aligns RNA-seq reads to a genome with STAR, the fast splice-aware aligner whose splice-junction database (built from a GTF at sjdbOverhang = readlength-1) and two-pass mode set junction sensitivity…

MITAuto-check passedResearch & Science

Install Bio Read Alignment Star Alignment

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-read-alignment-star-alignment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-read-alignment-star-alignment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/read-alignment/star-alignment .claude/skills/bio-read-alignment-star-alignment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-read-alignment-star-alignment
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.7k tokens
SKILL.md length
1,870 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Aligns RNA-seq reads to a genome with STAR, the fast splice-aware aligner whose splice-junction database (built from a GTF at sjdbOverhang = readlength-1) and two-pass mode set junction sensitivity…

  • Works in 3 steps: The splice-junction database is part of… → STAR's MAPQ is 255-for-unique and a… → STAR reports library strandedness for…
  • RNA reads must be placed on the genome for novel-isoform discovery
  • SKILL.md covers Version Compatibility, The Single Most Important…, How STAR Places a Splice (the… and Tool Taxonomy, plus 12 more sections
  • Runs Shell scripts from its folder

What it does

Bio Read Alignment Star Alignment is an agent skill from GPTomics/bioSkills. Aligns RNA-seq reads to a genome with STAR, the fast splice-aware aligner whose splice-junction database (built from a GTF at sjdbOverhang = readlength-1) and two-pass mode set junction sensitivity, whose 255-for-unique MAPQ breaks GATK, and whose GeneCounts output reveals library strandedness. Use when RNA reads must be placed on the genome for novel-isoform discovery, fusion detection, RNA variant calling, coverage tracks, splicing QC, or single-cell (STARsolo). Memory-constrained RNA alignment is…

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/align_star.sh` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • RNA reads must be placed on the genome for novel-isoform discovery
  • Fusion detection
  • RNA variant calling
  • Coverage tracks

Example prompts

  • “Use the bio-read-alignment-star-alignment skill to align RNA-seq reads to a genome with STAR, the fast splice-aware aligner whose splice-junction…”
  • “/bio-read-alignment-star-alignment”

Requirements

  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The splice-junction database is part of the index/run config, and a wrong sjdbOverhang or per-sample two-pass silently changes the answer…
  2. STAR's MAPQ is 255-for-unique and a multiplicity code, so it breaks GATK and a copied MAPQ filter deletes every multimapper. Unique reads…
  3. STAR reports library strandedness for free in GeneCounts, and getting strand wrong roughly halves counts. --quantMode GeneCounts emits a…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Read Alignment Star Alignment loads about 4.7k tokens when it runs. Until then it costs about 198 tokens; SKILL.md has 1,870 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~198
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,870 words, ~4,693 tokens.

Download SKILL.mdSave it as .claude/skills/bio-read-alignment-star-alignment/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-read-alignment-star-alignment
description
Aligns RNA-seq reads to a genome with STAR, the fast splice-aware aligner whose splice-junction database (built from a GTF at sjdbOverhang = readlength-1) and two-pass mode set junction sensitivity, whose 255-for-unique MAPQ breaks GATK, and whose GeneCounts output reveals library strandedness. Use when RNA reads must be placed on the genome for novel-isoform discovery, fusion detection, RNA variant calling, coverage tracks, splicing QC, or single-cell (STARsolo). Memory-constrained RNA alignment is hisat2-alignment; DE on known transcripts only should skip alignment for rna-quantification/alignment-free-quant; the QC gate and contig-naming reconciliation are alignment-files; counting is rna-quantification; DNA is bwa-alignment/bowtie2-alignment.
tool_type
cli
primary_tool
STAR

Version Compatibility

Reference examples tested with: STAR 2.7.11+, samtools 1.19+

Before using code patterns, verify installed versions match. If versions differ:

  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

STAR RNA-seq Alignment -- The Junction Database, the 255 MAPQ, and the Strand Column Decide the Result

"Align my RNA-seq reads" -> Map reads across exon-exon junctions to the genome, building or reusing a splice-aware index from a GTF -- because an RNA read spans introns the genome does not contain, so the junction database (and its sjdbOverhang), the two-pass choice, the 255-unique MAPQ, and the strand column STAR reports are what actually determine the downstream counts and variant calls.

  • CLI: STAR --runMode alignReads --genomeDir idx/ --readFilesIn R1.fq.gz R2.fq.gz --readFilesCommand zcat --outSAMtype BAM SortedByCoordinate

Scope: RNA splice-aware mapping with STAR -- index generation, two-pass, GeneCounts, chimeric/fusion output, and the MAPQ/strandedness traps. Contig-naming reconciliation and the QC gate -> alignment-files. Low-memory RNA alignment -> hisat2-alignment. Counting reads over genes/transcripts -> rna-quantification. Single-cell droplet/plate counting (STARsolo) -> single-cell. DE on known transcripts without a BAM -> rna-quantification/alignment-free-quant. OUT OF SCOPE: DNA (bwa-alignment/bowtie2-alignment), long reads (long-read-sequencing/long-read-alignment).

The Single Most Important Modern Insight

  1. The splice-junction database is part of the index/run config, and a wrong sjdbOverhang or per-sample two-pass silently changes the answer. STAR builds short artificial junction-flank sequences from the GTF so reads with a short overhang on one exon can still be placed; --sjdbOverhang sets the flank length and should equal max(readlength) - 1 (default 100 is fine near 100 bp reads but degrades junction sensitivity for very short reads). Two-pass (--twopassMode Basic) discovers novel junctions and re-aligns, raising novel-junction sensitivity -- but it is PER-SAMPLE: each sample is re-aligned against its own augmented index, so a junction found only in the deep/disease sample is rescued asymmetrically, a batch effect that confounds junction/splicing/sQTL comparisons. The cohort-correct recipe is to pool every sample's pass-1 SJ.out.tab, filter, and feed one common --sjdbFileChrStartEnd to a uniform second pass. Per-sample two-pass is fine for plain gene-level DE; it bites splicing analyses.
  2. STAR's MAPQ is 255-for-unique and a multiplicity code, so it breaks GATK and a copied MAPQ filter deletes every multimapper. Unique reads get MAPQ 255 -- the SAM "mapping quality unavailable" value -- which GATK treats as missing and drops, the classic silently-empty RNA VCF; fix it at align time with --outSAMmapqUnique 60. Multimappers get 3 / 1 / 0 for 2 / 3-4 / >=5 loci (pure locus count, no score information), so a generic samtools view -q 10 or featureCounts -Q 10 copied from a DNA pipeline DELETES every multimapper while keeping every unique -- a directional "discard paralog/gene-family/rRNA/pseudogene reads" filter that biases against recently-duplicated gene families. A hard MAPQ filter is almost never what an RNA analysis wants.
  3. STAR reports library strandedness for free in GeneCounts, and getting strand wrong roughly halves counts. --quantMode GeneCounts emits a 4-column ReadsPerGene.out.tab: gene, unstranded, forward, reverse. Summing columns 3 and 4 reveals the protocol -- roughly equal is unstranded (use col 2), col 3 dominant is forward, col 4 dominant is reverse (the common dUTP/TruSeq case, use col 4). Never assume the strand: feeding the wrong column (or the wrong -s to a counter) sends sense reads to "no feature" and roughly halves the counts, distorting DE.

How STAR Places a Splice (the mechanism in brief)

STAR seeds with Maximal Mappable Prefix search on an uncompressed suffix array: it finds the longest exact prefix, and when that prefix ends (at a junction, mismatch, or read end) it restarts from the next base -- so a junction-spanning read naturally decomposes into seeds on two exons that STAR then stitches across the intron with an N (skipped-region) CIGAR. The stitch is scored with splice priors that penalize non-canonical motifs (--scoreGapNoncan -8), long introns (--scoreGenomicLengthLog2scale -0.25), and reward annotated junctions (--sjdbScore 2), plus a hard anchor floor (--alignSJoverhangMin): a novel junction on a tiny anchor carries almost no information, so it must clear both the soft score prior and the hard overhang minimum. This is why annotation (the sjdb) and two-pass matter for short-overhang and novel junctions.

Tool Taxonomy

Mode / outputCitationMechanism / roleWhen
STAR --runMode alignReadsDobin 2013 Bioinformatics 29:15MMP seed + stitch; spliced genomic BAMthe default RNA-to-genome alignment
--twopassMode BasicDobin 2013 (STAR manual)discover novel junctions, re-alignnovel-isoform / variant / splicing work (per-sample); pool for cohorts
--quantMode GeneCountsSTAR manualper-gene counts + the strand-detection 3 columnsquick counts and strandedness inference
--quantMode TranscriptomeSAMSTAR manualtranscriptome-coord BAM for RSEM / Salmon-alnisoform-level EM quantification -> rna-quantification
--chimSegmentMin -> Chimeric.out.junctionSTAR manualsplit-across-loci reads for fusion callingSTAR-Fusion / Arriba fusion detection
STARsolo (--soloType)STAR manualcell-barcode + UMI single-cell countingscRNA-seq (route OUT) -> single-cell
HISAT2Kim 2019 Nat Biotechnol 37:907graph FM-index, ~1/4 the RAMmemory-constrained RNA (route OUT) -> hisat2-alignment
Salmon / kallistoPatro 2017 Nat Methods 14:417alignment-free transcript quantificationDE on known transcripts only (route OUT) -> rna-quantification/alignment-free-quant

Decision Tree by Scenario

ScenarioRecommendedWhy
RNA-seq, ample RAM (>=32 GB), need a genomic BAMSTARfastest splice-aware aligner; native counts, fusions, 2-pass
RNA variant callingSTAR 2-pass + --outSAMmapqUnique 60 then GATK SplitNCigarReads2-pass splices novel junctions; 60 avoids the 255 drop
Novel-isoform / splicing / sQTL across a cohortcohort 2-pass (pool SJ.out.tab -> common --sjdbFileChrStartEnd)per-sample 2-pass is a junction batch effect
Fusion detection--chimSegmentMin 12 --chimOutType ... -> STAR-Fusion / Arribachimeric junctions are the fusion signal
Single-cell RNASTARsolo (route OUT)barcode+UMI counting -> single-cell
Memory-constrained (<32 GB)route OUT to hisat2-alignmentSTAR needs ~30 GB for human
DE on known transcripts onlyroute OUT to rna-quantification/alignment-free-quantSalmon/kallisto are faster and model multimapping better
Small genome (bacterial/viral/plasmid)STAR with reduced --genomeSAindexNbasesthe default 14 silently builds a bad index / segfaults

Default when uncertain: STAR with --outSAMtype BAM SortedByCoordinate, --quantMode GeneCounts (to also read strandedness), and --twopassMode Basic for per-sample novel-junction work; set --outSAMmapqUnique 60 for any STAR -> GATK path.

Generate the Genome Index

bash
# sjdbOverhang = max(readlength) - 1 (149 for 2x150 reads). Default 100 degrades junctions for short reads.
STAR --runMode genomeGenerate --runThreadN 8 \
    --genomeDir star_index/ \
    --genomeFastaFiles genome.fa \
    --sjdbGTFfile annotation.gtf \
    --sjdbOverhang 149
# Small genome (e.g. 5 Mb): add --genomeSAindexNbases <= min(14, log2(GenomeLength)/2 - 1), or STAR segfaults.

Basic Alignment

bash
STAR --runThreadN 8 \
    --genomeDir star_index/ \
    --readFilesIn reads_1.fq.gz reads_2.fq.gz \
    --readFilesCommand zcat \
    --outFileNamePrefix sample_ \
    --outSAMtype BAM SortedByCoordinate
samtools index sample_Aligned.sortedByCoord.out.bam
# STAR already coordinate-sorts -- a subsequent `samtools sort` is redundant. Single-end: one file in --readFilesIn.

Two-Pass + Gene Counts + Strandedness

bash
STAR --runThreadN 8 --genomeDir star_index/ \
    --readFilesIn r1.fq.gz r2.fq.gz --readFilesCommand zcat \
    --outFileNamePrefix sample_ \
    --outSAMtype BAM SortedByCoordinate \
    --twopassMode Basic \
    --quantMode GeneCounts \
    --outSAMattrRGline ID:sample1 SM:sample1 PL:ILLUMINA LB:lib1 \
    --outSAMmapqUnique 60        # so a downstream GATK RNA-variant step does not drop the 255 uniques
# STAR read groups use --outSAMattrRGline (SPACE-separated tags), NOT bwa's tab-delimited -R '@RG\t...'.
# GATK requires read groups; comma-with-spaces separates groups for multiple --readFilesIn files.

# Detect strandedness from ReadsPerGene.out.tab (skip the 4 N_* summary rows, sum cols 3 vs 4):
awk 'NR>4 {f+=$3; r+=$4} END {printf "fwd(col3)=%d  rev(col4)=%d -> use col %s\n", f, r, (f>2*r?"3 fwd": r>2*f?"4 rev":"2 unstranded")}' sample_ReadsPerGene.out.tab

ENCODE Long-RNA-seq Parameter Set

bash
STAR --runThreadN 8 --genomeDir star_index/ --readFilesIn r1.fq.gz r2.fq.gz --readFilesCommand zcat \
    --outFileNamePrefix sample_ --outSAMtype BAM SortedByCoordinate \
    --outFilterType BySJout \
    --outFilterMultimapNmax 20 \
    --outFilterMismatchNmax 999 --outFilterMismatchNoverReadLmax 0.04 \
    --alignIntronMin 20 --alignIntronMax 1000000 --alignMatesGapMax 1000000 \
    --alignSJoverhangMin 8 --alignSJDBoverhangMin 1 \
    --sjdbScore 1 --outSAMattributes NH HI AS NM MD
# BySJout keeps only reads whose junctions passed the dataset-wide collapse; multimapNmax 20 retains real
# multi-locus genes; the 0.04 mismatch ratio scales with read length; intron caps at ~1 Mb cover human genes.

Fusion Detection

bash
# Chimeric junctions for STAR-Fusion (params per the STAR-Fusion wiki).
STAR --runThreadN 8 --genomeDir star_index/ --readFilesIn r1.fq.gz r2.fq.gz --readFilesCommand zcat \
    --outFileNamePrefix sample_ --outSAMtype BAM SortedByCoordinate \
    --chimSegmentMin 12 --chimJunctionOverhangMin 12 --chimOutJunctionFormat 1 \
    --chimOutType Junctions
# Arriba instead reads chimeric alignments from the BAM: --chimSegmentMin 10 --chimOutType WithinBAM SoftClip.

Key Parameters (STAR defaults unless noted)

ParameterDefaultDescription
--sjdbOverhang100junction-flank length at index build; set to readlength-1
--twopassModeNoneBasic for per-sample novel-junction discovery
--outSAMmapqUnique255MAPQ for unique reads; set 60 for GATK
--outFilterMultimapNmax10max loci to report (ENCODE uses 20 for RNA)
--alignIntronMin / Max21 / 0 (auto)gap < min is a deletion; cap Max ~1 Mb for human
--alignSJoverhangMin / SJDBoverhangMin5 / 3novel / annotated junction anchor floor (ENCODE 8 / 1)
--quantMode--GeneCounts and/or TranscriptomeSAM
--chimSegmentMin0 (off)turn on chimeric/fusion detection
--genomeSAindexNbases14reduce to min(14, log2(L)/2-1) for small genomes
--genomeLoad / --limitBAMsortRAMNoSharedMemory / --shared-memory reuse; explicit sort RAM

Per-Method Failure Modes

STAR 255 MAPQ into GATK

Trigger: STAR BAM (uniques at 255) fed to GATK RNA variant calling. Mechanism: GATK reads 255 as "MAPQ unavailable" and drops the read. Symptom: a silently empty or near-empty RNA VCF. Fix: --outSAMmapqUnique 60 at align time (then SplitNCigarReads in the GATK RNA workflow).

Show full SKILL.md (727 more words)Show less
MAPQ filter deletes every multimapper

Trigger: a -q 10 / -Q 10 MAPQ filter (copied from DNA) on STAR output. Mechanism: STAR multimappers are MAPQ <= 3, uniques 255, so the filter keeps only uniques. **Symptom:** systematic under-counting of paralog/gene-family/rRNA/pseudogene loci; DE driven by multimapper fraction. **Fix:** do not MAPQ-filter RNA for counting; handle multimappers in the counter (NH-aware) or via EM -> rna-quantification/featurecounts-counting.

Per-sample two-pass as a batch effect

Trigger: --twopassMode Basic per sample for a splicing/junction comparison. Mechanism: each sample is aligned against its own novel-junction-augmented index. Symptom: junction recovery correlates with depth/condition, confounding sQTL/differential splicing. Fix: pool pass-1 SJ.out.tab across the cohort, filter, feed one --sjdbFileChrStartEnd to a uniform second pass.

sjdbOverhang mismatched to read length

Trigger: index built with default 100 for 36-50 bp reads, or a mismatch between index build and align values. Mechanism: junction flanks far longer than the reads degrade short-overhang sensitivity; a mismatch errors at align time. Symptom: reduced novel-junction-spanning sensitivity, or "present sjdbOverhang not equal to genome generation step." Fix: rebuild with sjdbOverhang = max(readlength)-1.

Small genome with default SAindexNbases

Trigger: indexing a bacterial/viral/plasmid genome with --genomeSAindexNbases 14. Mechanism: the SA pre-index string is too long for the genome. Symptom: STAR silently builds a bad index or segfaults at align time. Fix: set --genomeSAindexNbases min(14, log2(GenomeLength)/2 - 1) (1 Mb -> 9, 100 kb -> 7).

Index built with a different STAR version

Trigger: an index built months ago loaded by a bumped STAR module. Mechanism: STAR refuses an index whose versionGenome differs. Symptom: "Genome version is INCOMPATIBLE with running STAR version," or a subtly different version that loads and behaves differently across a cohort. Fix: rebuild the index with the exact aligning STAR version; pin the version for the whole cohort.

Quantitative Thresholds

ThresholdSourceRationale
sjdbOverhang = max(readlength) - 1STAR manualsets the junction-flank length for short-overhang reads
--outSAMmapqUnique 60 for GATKSTAR manual / GATK RNA best practice255 is "unavailable" and is dropped by GATK
--outFilterMultimapNmax 20, --outFilterMismatchNoverReadLmax 0.04ENCODE long-RNA-seq pipelineretains real multi-locus genes; mismatch budget scales with read length
--genomeSAindexNbases <= min(14, log2(L)/2 - 1)STAR manualthe default 14 corrupts/segfaults small-genome indexes
STAR human-genome index RAM ~30 GBSTAR docs (approximate)the reason to route memory-constrained jobs to HISAT2
--chimSegmentMin 12 (STAR-Fusion)STAR-Fusion wikiminimum chimeric-segment length for fusion calling

Common Errors

Error / symptomCauseSolution
Empty / tiny RNA VCF after GATKSTAR 255 MAPQ dropped as "unavailable"--outSAMmapqUnique 60 (and SplitNCigarReads)
GATK "no read group" on a STAR BAMSTAR omits @RG unless askedadd --outSAMattrRGline ID:.. SM:.. PL:ILLUMINA LB:.. (space-separated, not bwa's -R)
htseq-count silently miscounts STAR outputhtseq wants name-sorted (or -r pos) input; STAR coordinate-sortsemit --outSAMtype BAM Unsorted for htseq, or samtools sort -n; featureCounts accepts either
Counts ~halved, antisense artifactswrong strandedness column / -sinfer from GeneCounts cols 3 vs 4 (or RSeQC); TruSeq/dUTP = reverse
Paralog/gene-family genes under-counteda -q 10 MAPQ filter deleted multimappersdrop the MAPQ filter; handle NH>1 in the counter -> rna-quantification
"not enough memory for BAM sorting"--limitBAMsortRAM too smallset it explicitly (e.g. 10000000000)
Segfault / bad index on a small genome--genomeSAindexNbases 14reduce per min(14, log2(L)/2 - 1)
"Genome version INCOMPATIBLE"index built with another STAR versionrebuild with the running version; pin it
0 counts despite high mapping rategenome/GTF contig-naming mismatchreconcile chr1 vs 1, chrM vs MT (same source/release) -> alignment-files
STAR reads .gz FASTQ as garbage / failsSTAR does not auto-detect gzipadd --readFilesCommand zcat (bwa/bowtie2/HISAT2 auto-detect gzip)

References

  • Dobin A, Davis CA, Schlesinger F, et al. 2013. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29:15-21.
  • Kim D, Paggi JM, Park C, Bennett C, Salzberg SL. 2019. Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype. Nat Biotechnol 37:907-915.
  • Patro R, Duggal G, Love MI, Irizarry RA, Kingsford C. 2017. Salmon provides fast and bias-aware quantification of transcript expression. Nat Methods 14:417-419.
  • Burset M, Seledtsov IA, Solovyev VV. 2000. Analysis of canonical and non-canonical splice sites in mammalian genomes. Nucleic Acids Res 28:4364-4375.
  • hisat2-alignment - Low-memory splice-aware alternative to STAR
  • bwa-alignment - DNA short-read mapping (when reads do not cross junctions)
  • read-qc/rnaseq-qc - RNA destination metrics: rRNA, gene-body coverage, strandedness
  • read-qc/fastp-workflow - Trim adapters/poly-A before alignment
  • alignment-files/bam-statistics - flagstat/idxstats QC gate; what a high mapping rate hides; contig naming
  • rna-quantification/featurecounts-counting - Count aligned reads over genes (NH-aware multimapper handling)
  • rna-quantification/alignment-free-quant - Salmon/kallisto when only known-transcript DE is needed
  • differential-expression/deseq2-basics - Downstream DE from the count matrix
  • single-cell/data-io - STARsolo single-cell counts into a single-cell workflow

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in read-alignment/star-alignment of GPTomics/bioSkills.

  • SKILL.md
  • examples/align_star.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Read Alignment Star Alignment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Read Alignment Star Alignment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Read Alignment Star Alignment this skillGPTomics/bioSkills1.2k1 repos~4.7kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Read Alignment Star Alignment

What does Bio Read Alignment Star Alignment do?

Aligns RNA-seq reads to a genome with STAR, the fast splice-aware aligner whose splice-junction database (built from a GTF at sjdbOverhang = readlength-1) and two-pass mode set junction sensitivity…. Bio Read Alignment Star Alignment is an agent skill from GPTomics/bioSkills. Aligns RNA-seq reads to a genome with STAR, the fast splice-aware aligner whose splice-junction database (built from a GTF at sjdbOverhang = readlength-1) and two-pass mode set junction sensitivity, whose 255-for-unique MAPQ breaks GATK, and whose GeneCounts output reveals library strandedness.

When should I use Bio Read Alignment Star Alignment?

Bio Read Alignment Star Alignment fits situations like: RNA reads must be placed on the genome for novel-isoform discovery; fusion detection; RNA variant calling; coverage tracks.

How do I install Bio Read Alignment Star Alignment in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-read-alignment-star-alignment -a claude-code`. Or copy the skill folder (read-alignment/star-alignment in GPTomics/bioSkills) into .claude/skills/bio-read-alignment-star-alignment in your project. Claude Code loads it when a task matches its description.

How do I install Bio Read Alignment Star Alignment in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-read-alignment-star-alignment -a codex`. Or copy the skill folder (read-alignment/star-alignment in GPTomics/bioSkills) into .agents/skills/bio-read-alignment-star-alignment in your project. Codex loads it when a task matches its description.

Can I use Bio Read Alignment Star Alignment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-read-alignment-star-alignment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-read-alignment-star-alignment, .gemini/skills/bio-read-alignment-star-alignment, .github/skills/bio-read-alignment-star-alignment and .opencode/skills/bio-read-alignment-star-alignment in your project.

What does Bio Read Alignment Star Alignment need to run?

Going by SKILL.md and its folder, Bio Read Alignment Star Alignment needs a shell for the scripts in its folder. Our summary lists: A Bash shell.

Does Bio Read Alignment Star Alignment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bio Read Alignment Star Alignment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Read Alignment Star Alignment use?

Bio Read Alignment Star Alignment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Read Alignment Star Alignment use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Read Alignment Star Alignment?

Skills that share tags, products or a category with Bio Read Alignment Star Alignment: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Read Alignment Star Alignment?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.