Agent skill

Bio Alignment Validation

by GPTomics in GPTomics/bioSkills

Validate alignment quality with insert size distribution, proper pairing rates, GC bias, strand balance, and other post-alignment metrics.

MITAuto-check passedData & Analytics

Install Bio Alignment Validation

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-alignment-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-alignment-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/alignment-files/alignment-validation .claude/skills/bio-alignment-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-alignment-validation
GitHub stars
1.2k
Used in
2 other repos
Token cost
~3.7k tokens
SKILL.md length
1,004 words
Files
4
Skills in repo
552
Repo updated
First seen
Licence
MIT

At a glance

Validate alignment quality with insert size distribution, proper pairing rates, GC bias, strand balance, and other post-alignment metrics.

  • Verifying alignment data quality before variant calling
  • SKILL.md covers Version Compatibility, Two Different Validations, Insert Size Distribution and Proper Pairing Rate, plus 5 more sections
  • Runs Python and Shell scripts from its folder; calls java and pip
  • Tasks that involve Bioinformatics

What it does

Bio Alignment Validation is an agent skill from GPTomics/bioSkills. Validate alignment quality with insert size distribution, proper pairing rates, GC bias, strand balance, and other post-alignment metrics. Use when verifying alignment data quality before variant calling or quantification.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/validate_alignment.py`, `examples/validate_alignment.sh` and `usage-guide.md`).

It sits in Data & Analytics, covering Bioinformatics and Data cleaning. It works with Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Verifying alignment data quality before variant calling
  • Tasks that involve Bioinformatics
  • Tasks that involve Data cleaning

Example prompts

  • “/bio-alignment-validation”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • java
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Alignment Validation loads about 3.7k tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 1,004 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,004 words, ~3,715 tokens.

Download SKILL.mdSave it as .claude/skills/bio-alignment-validation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-alignment-validation
description
Validate alignment quality with insert size distribution, proper pairing rates, GC bias, strand balance, and other post-alignment metrics. Use when verifying alignment data quality before variant calling or quantification.
tool_type
mixed
primary_tool
samtools

Version Compatibility

Reference examples tested with: matplotlib 3.8+, numpy 1.26+, picard 3.1+, pysam 0.22+, samtools 1.19+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Alignment Validation

Post-alignment quality control to verify alignment quality and identify issues.

"Check alignment quality" -> Compute post-alignment QC metrics (mapping rate, pairing, insert size, strand balance) to identify issues before downstream analysis.

  • CLI: samtools flagstat, samtools stats, Picard CollectAlignmentSummaryMetrics
  • Python: pysam.AlignmentFile iteration with metric calculations

Two Different Validations

ConcernToolsWhat it catches
File integritysamtools quickcheck, picard ValidateSamFileTruncation, missing EOF, malformed records, wrong CIGAR, MAPQ out of range
Sequence dictionary identitysamtools dict + M5 diffBAM aligned to wrong reference flavor / different decoy / chr vs no-chr
QC metricssamtools stats, flagstat, mosdepth, Picard CollectMultipleMetrics / CollectHsMetrics / CollectWgsMetricsAre the data biologically reasonable for the assay?
Contamination / sample swapverifybamid2, somalier, Picard CrosscheckFingerprintsCross-sample contamination, tumor-normal swap, mislabeled sample

A file can pass quickcheck and still be malformed in ways that crash GATK three hours into HaplotypeCaller. Conversely, a QC-poor BAM can be structurally valid.

File Integrity
bash
# Fast: header + EOF block check (misses mid-file truncation, invalid CIGAR)
samtools quickcheck -v in.bam || echo "QUICKCHECK FAILED"
samtools quickcheck -v *.bam > bad_bams.fofn   # one fail-line per bad file

# Slow but thorough: structural validation
picard ValidateSamFile I=in.bam MODE=SUMMARY R=ref.fa

# Production: ignore expected-but-noisy
picard ValidateSamFile I=in.bam MODE=SUMMARY R=ref.fa \
    IGNORE=INVALID_MAPPING_QUALITY \
    IGNORE=MISMATCH_FLAG_MATE_NEG_STRAND

CI-safe one-liner:

bash
test -s in.bam \
  && samtools quickcheck -v in.bam \
  && [ $(samtools view -c -F 2304 in.bam) -gt 1000 ] \
  || { echo "BAM failed integrity"; exit 1; }
Sequence Dictionary Cross-Validation (M5)
bash
# Compare per-contig MD5 between BAM and reference
diff \
    <(samtools view -H in.bam | grep '^@SQ' | tr '\t' '\n' | grep '^M5:' | sort) \
    <(samtools dict ref.fa | grep '^@SQ' | tr '\t' '\n' | grep '^M5:' | sort)

If M5s differ, the BAM was aligned to a different sequence than the current reference (even if contig names match). Concrete failure modes: GRCh38 vs GRCh38.p13 vs GRCh38_no_alt (alt contigs differ); UCSC chr1 vs Ensembl 1 (names differ, M5s match -- pure renaming); soft-masked vs unmasked (M5 matches -- case is normalized to uppercase before hashing, so lowercase soft-masking is invisible to M5). Note hard-masking is different: it replaces bases with N, changing the sequence, so its M5 does NOT match the unmasked reference. The M5 tag is the only definitive identity check.

Contamination and Sample Swap

No alignment QC is complete without these in production:

bash
# Cross-sample contamination
verifybamid2 --SVDPrefix /resources/1000g.b38.vcf.gz.SVD \
    --Reference ref.fa --BamFile sample.bam --Output sample.contam
# FREEMIX > 0.03 is the commonly used contamination-concern threshold; values escalate from there.

# Relatedness, sex check, sample swap detection
somalier extract -d extracted/ -s /resources/sites.GRCh38.vcf.gz \
    -f ref.fa sample.bam
somalier relate --infer extracted/*.somalier

# Tumor/normal pairing verification
picard CrosscheckFingerprints I=tumor.bam I=normal.bam \
    HAPLOTYPE_MAP=Homo_sapiens_assembly38.haplotype_database.txt
# LOD > 5 = same individual; < -5 = different

Sample-swap rates of 0.5-1% in production cohorts are typical. Without somalier or CrosscheckFingerprints, swaps are detected only when a downstream finding contradicts clinical expectation.

Insert Size Distribution

Goal: Verify that the fragment length distribution matches the library preparation protocol.

Approach: Extract template_length from properly paired reads and compare the distribution to expected values for the library type.

samtools stats
bash
samtools stats input.bam > stats.txt
grep "^IS" stats.txt | cut -f2,3 > insert_sizes.txt
Picard CollectInsertSizeMetrics
bash
java -jar picard.jar CollectInsertSizeMetrics \
    I=input.bam \
    O=insert_metrics.txt \
    H=insert_histogram.pdf
Expected Insert Sizes by Library
LibraryMean insertDistribution shapeDiagnostic
TruSeq DNA PCR-free WGS400-500 bpRoughly Gaussian, tight/sharpSharpest, most symmetric peak (no PCR bias); bimodality = degraded sample
TruSeq DNA Nano (PCR) WGS300-400 bpGaussian, broadened by PCR amplification bias
Twist / IDT exome capture250-350 bpGaussian
TruSeq Stranded mRNA200-300 bpRight-skewed (transcript distribution)Long tail = poor size selection
Ribo-Zero rRNA-depleted250-400 bpRight-skewed
Smart-seq2 / Smart-seq3200-700 bpBroad
10x Chromium (3')n/a -- not informativen/a
TruSeq ChIP200-400 bpSharp
ATAC-seq (Buenrostro / Omni-ATAC)MultimodalPeaks at ~50, ~180, ~370 bpMissing multimodal pattern = bad library; missing ~180 bp mononucleosome peak = over-transposition (Tn5 over-titrated) or degraded DNA
Hi-C / Micro-CMultimodalPeak at ligation-junction size
cfDNA / ctDNA160-180 bpMultimodal; ~167 bp mononucleosomal + ~340 dinucTumor-derived shorter (~145 bp); shape itself is a biomarker
FFPE100-250 bpRight-skewed, broad
aDNA30-80 bpSharp left-skewed
ONT (native)1-30 kbn/a
PacBio HiFi10-25 kbSharp peak

For ATAC, the multimodal pattern is the QC. If the mononucleosomal peak (~180 bp) is absent, Tn5 was over-titrated, under-titrated, or DNA was degraded. Use ATACseqQC fragSizeDist() for the standard ATAC fragment-size diagnostic.

Python Insert Size Analysis
python
import pysam
import numpy as np
import matplotlib.pyplot as plt

def get_insert_sizes(bam_file, max_reads=100000):
    sizes = []
    bam = pysam.AlignmentFile(bam_file, 'rb')
    for i, read in enumerate(bam.fetch()):
        if i >= max_reads:
            break
        if read.is_proper_pair and not read.is_secondary and read.template_length > 0:
            sizes.append(read.template_length)
    bam.close()
    return sizes

sizes = get_insert_sizes('sample.bam')
print(f'Median insert size: {np.median(sizes):.0f}')
print(f'Mean insert size: {np.mean(sizes):.0f}')
print(f'Std dev: {np.std(sizes):.0f}')

plt.hist(sizes, bins=100, range=(0, 1000))
plt.xlabel('Insert Size')
plt.ylabel('Count')
plt.savefig('insert_size_dist.pdf')

Proper Pairing Rate

Percentage of reads correctly paired.

samtools flagstat
bash
samtools flagstat input.bam

samtools flagstat input.bam | grep "properly paired"
Calculate Pairing Rate
bash
proper=$(samtools view -c -f 2 input.bam)
mapped=$(samtools view -c -F 4 input.bam)
rate=$(echo "scale=4; $proper / $mapped * 100" | bc)
echo "Proper pairing rate: ${rate}%"
Show full SKILL.md (410 more words)Show less
Expected Rates
MetricGoodMarginalPoor
Proper pair> 90%80-90%< 80%
Mapped> 95%90-95%< 90%
Singletons< 5%5-10%> 10%

GC Bias

GC content correlation with coverage.

Picard CollectGcBiasMetrics
bash
java -jar picard.jar CollectGcBiasMetrics \
    I=input.bam \
    O=gc_bias_metrics.txt \
    CHART=gc_bias_chart.pdf \
    S=gc_summary.txt \
    R=reference.fa
deepTools computeGCBias
bash
computeGCBias \
    -b input.bam \
    --effectiveGenomeSize 2913022398 \
    -g hg38.2bit \
    -o gc_bias.txt \
    --biasPlot gc_bias.pdf
Interpret GC Bias
IssueSymptom
Under-representationLow GC coverage drops
Over-representationHigh GC coverage elevated
PCR biasStrong correlation

Strand Balance

A balanced 0.48-0.52 forward/reverse ratio applies to WGS / WES / generic DNA-seq on autosomes. Expected to deviate for: stranded RNA-seq (deliberately strand-asymmetric -- verify with RSeQC infer_experiment.py), bisulfite (CT vs GA), small-RNA / strand-specific RNA-seq, and chrY/chrM regions. Per-chromosome strand imbalance >5% on autosomes is a field-convention rule of thumb (no single primary citation) — it picks up aligner artifacts; on chrX/chrY it suggests sex-mismatch.

Calculate Strand Ratio
bash
forward=$(samtools view -c -F 16 input.bam)
reverse=$(samtools view -c -f 16 input.bam)
echo "Forward: $forward"
echo "Reverse: $reverse"
ratio=$(echo "scale=4; $forward / $reverse" | bc)
echo "F/R ratio: $ratio"
Check Strand Bias per Chromosome
bash
for chr in chr1 chr2 chr3; do
    fwd=$(samtools view -c -F 16 input.bam $chr)
    rev=$(samtools view -c -f 16 input.bam $chr)
    echo "$chr: F=$fwd R=$rev ratio=$(echo "scale=2; $fwd/$rev" | bc)"
done

Mapping Quality Distribution

Extract MAPQ Distribution
bash
samtools view input.bam | cut -f5 | sort -n | uniq -c | sort -k2 -n
Calculate Mean MAPQ
bash
samtools view input.bam | awk '{sum+=$5; count++} END {print "Mean MAPQ:", sum/count}'
MAPQ Distribution Is Bimodal and Aligner-Specific

Mean MAPQ is misleading; distributions are bimodal (0 and aligner-max). For aligner-specific scales and "unique mapping" sentinels, see sam-bam-basics. The fraction of primary mapped reads at MAPQ >= 30 is a more informative summary than the mean.

Chromosome Coverage Balance

Calculate Per-Chromosome Coverage
bash
samtools idxstats input.bam | awk '{print $1, $3/$2}' | head -25
Check for Aneuploidy / Sex Chromosome Imbalance

Median-normalized per-autosome coverage (1.0 = expected diploid; 0.5 = monosomy/sex; 1.5 = trisomy):

bash
samtools idxstats in.bam | awk '$2>0 && $1!~/^chr[XYM]|^GL|^KI|^chrUn|^chrEBV/ {
    cov[$1] = $3 / $2
}
END {
    n = asort(cov, sorted)
    med = sorted[int(n/2)+1]
    for (c in cov) printf "%s\t%.3f\n", c, cov[c]/med
}'

For full ancestry / contamination / relatedness checking, use verifybamid2, somalier, or peddy -- they account for population AFs, not just per-contig depth.

Mismatch Rate

Picard CollectAlignmentSummaryMetrics
bash
java -jar picard.jar CollectAlignmentSummaryMetrics \
    I=input.bam \
    R=reference.fa \
    O=alignment_summary.txt
Key Metrics
MetricDescriptionGood Value
PCT_PF_READS_ALIGNEDMapped %> 95%
PF_MISMATCH_RATEMismatches< 1%
PF_INDEL_RATEIndels< 0.1%
STRAND_BALANCEStrand ratio~0.5

Comprehensive Validation Script

Goal: Run all key alignment QC checks in a single pass and generate a summary report.

Approach: Combine samtools flagstat, stats, idxstats, and strand counts into one script that outputs pass/warn/fail calls.

bash
#!/bin/bash
BAM=$1
REF=$2
NAME=$(basename $BAM .bam)
OUTDIR=${3:-qc}

mkdir -p $OUTDIR

echo "=== Alignment Validation: $NAME ===" | tee $OUTDIR/report.txt

echo -e "\n--- Flagstat ---" | tee -a $OUTDIR/report.txt
samtools flagstat $BAM | tee -a $OUTDIR/report.txt

echo -e "\n--- Mapping Rate ---" | tee -a $OUTDIR/report.txt
mapped=$(samtools view -c -F 4 $BAM)
total=$(samtools view -c $BAM)
rate=$(echo "scale=2; $mapped / $total * 100" | bc)
echo "Mapping rate: ${rate}%" | tee -a $OUTDIR/report.txt

echo -e "\n--- Proper Pairing ---" | tee -a $OUTDIR/report.txt
proper=$(samtools view -c -f 2 $BAM)
pair_rate=$(echo "scale=2; $proper / $mapped * 100" | bc)
echo "Proper pairing: ${pair_rate}%" | tee -a $OUTDIR/report.txt

echo -e "\n--- Insert Size ---" | tee -a $OUTDIR/report.txt
samtools stats $BAM | grep "insert size average" | tee -a $OUTDIR/report.txt

echo -e "\n--- Strand Balance ---" | tee -a $OUTDIR/report.txt
fwd=$(samtools view -c -F 16 $BAM)
rev=$(samtools view -c -f 16 $BAM)
strand_ratio=$(echo "scale=3; $fwd / $rev" | bc)
echo "Forward: $fwd, Reverse: $rev, Ratio: $strand_ratio" | tee -a $OUTDIR/report.txt

echo -e "\n--- Chromosome Coverage ---" | tee -a $OUTDIR/report.txt
samtools idxstats $BAM | head -25 | tee -a $OUTDIR/report.txt

echo -e "\nReport: $OUTDIR/report.txt"

Python Validation Module

A skeleton; full implementation is in examples/validate_alignment.py:

python
import pysam

class AlignmentValidator:
    def __init__(self, bam_file):
        self.bam = pysam.AlignmentFile(bam_file, 'rb')

    def report(self, sample_size=100000):
        # Sample reads, compute mapping rate, proper-pair rate, MAPQ dist, strand balance
        # See examples/validate_alignment.py for full implementation
        ...

The first-N-reads sampling pattern is biased toward chr1 (different GC content and complexity than chrM/chrX/chrY/alt contigs). For unbiased per-chromosome statistics, use samtools view -s 42.01 input.bam (the INT.FRAC form uses INT as the seed; specify an explicit nonzero seed so the subsample is documented and consistent across paired runs) instead of head-of-file iteration.

Quality Thresholds Summary

MetricGoodWarningFail
Mapping rate> 95%90-95%< 90%
Proper pairing> 90%80-90%< 80%
Duplicate rate (assay-specific)see bam-statistics decision table----
Strand balance0.48-0.520.45-0.55Outside
Mean MAPQ> 4030-40< 30
GC bias< 1.2x1.2-1.5x> 1.5x (field-convention bands; Picard CollectGcBiasMetrics does not prescribe specific cutoffs)
  • bam-statistics - Per-assay metric thresholds, depth/coverage tools, mosdepth
  • alignment-filtering - Aligner-specific MAPQ thresholds (canonical home)
  • duplicate-handling - Library-aware dedup decisions
  • sam-bam-basics - MAPQ-by-aligner table
  • chip-seq/chipseq-qc - ChIP-specific QC (FRiP, NSC, RSC)

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in alignment-files/alignment-validation of GPTomics/bioSkills.

  • SKILL.md
  • examples/validate_alignment.py
  • examples/validate_alignment.sh
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Alignment Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Alignment Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Alignment Validation this skillGPTomics/bioSkills1.2k2 repos~3.7kAutomated safety check: PassMIT
Bio Splicing QcFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~1.6kAutomated safety check: PassNone
Bio Proteomics Data ImportFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~1.2kAutomated safety check: PassNone
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Openbb Data Fetchermonarchjuno/vibe-investing299—~2.9kAutomated safety check: NotesMIT
Credit Risk Data Cleaninggithub/awesome-copilot40k1 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Bio Splicing Qc

    FreedomIntelligence/OpenClaw-Medical-Skills

    Assesses RNA-seq data quality for splicing analysis including junction saturation curves, splice site strength scoring, and junction coverage metrics using RSeQC.

    3.1k GitHub stars~1.6k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Bio Proteomics Data Import

    FreedomIntelligence/OpenClaw-Medical-Skills

    Load and parse mass spectrometry data formats including mzML, mzXML, and quantification tool outputs like MaxQuant proteinGroups.txt.

    3.1k GitHub starsUsed in 1 repo~1.2k tokens
    Research & ScienceAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Aic Collector Op Development

    ai-dynamo/aiconfigurator

    Design, add, review, or modify AIC Collector operations and their case population.

    455 GitHub stars~3k tokensUpdated 19 days ago
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 552 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Works with

Questions about Bio Alignment Validation

What does Bio Alignment Validation do?

Validate alignment quality with insert size distribution, proper pairing rates, GC bias, strand balance, and other post-alignment metrics. Bio Alignment Validation is an agent skill from GPTomics/bioSkills. Validate alignment quality with insert size distribution, proper pairing rates, GC bias, strand balance, and other post-alignment metrics.

When should I use Bio Alignment Validation?

Bio Alignment Validation fits situations like: verifying alignment data quality before variant calling; tasks that involve Bioinformatics; tasks that involve Data cleaning.

How do I install Bio Alignment Validation in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-alignment-validation -a claude-code`. Or copy the skill folder (alignment-files/alignment-validation in GPTomics/bioSkills) into .claude/skills/bio-alignment-validation in your project. Claude Code loads it when a task matches its description.

How do I install Bio Alignment Validation in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-alignment-validation -a codex`. Or copy the skill folder (alignment-files/alignment-validation in GPTomics/bioSkills) into .agents/skills/bio-alignment-validation in your project. Codex loads it when a task matches its description.

Can I use Bio Alignment Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-alignment-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-alignment-validation, .gemini/skills/bio-alignment-validation, .github/skills/bio-alignment-validation and .opencode/skills/bio-alignment-validation in your project.

What does Bio Alignment Validation need to run?

Going by SKILL.md and its folder, Bio Alignment Validation needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (java and pip). Our summary lists: Python 3; A Bash shell.

Does Bio Alignment Validation access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Alignment Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Alignment Validation use?

Bio Alignment Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Alignment Validation use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Alignment Validation?

Skills that share tags, products or a category with Bio Alignment Validation: Bio Splicing Qc (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Bio Proteomics Data Import (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Pandas Pro (Jeffallan/claude-skills, 12k stars) and Openbb Data Fetcher (monarchjuno/vibe-investing, 299 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Alignment Validation?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 552 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.