Agent skill

Seq Wrangler

by ClawBio in ClawBio/ClawBio

NGS read QC, alignment, and BAM processing pipeline. An agent skill from ClawBio/ClawBio.

MITAuto-check passedResearch & Science

Install Seq Wrangler

skills CLI
$ npx skills add ClawBio/ClawBio --skill seq-wrangler -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio seq-wrangler --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/seq-wrangler .claude/skills/seq-wrangler && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
seq-wrangler
GitHub stars
1.2k
Used in
1 other repo
Token cost
~2.5k tokens
SKILL.md length
860 words
Files
5
Skills in repo
104
Repo updated
First seen
Licence
MIT

At a glance

NGS read QC, alignment, and BAM processing pipeline. An agent skill from ClawBio/ClawBio.

  • Works in 9 steps: Read QC: Run FastQC, parse results, flag… → Adapter Trimming: Trim adapters with… → Alignment: Align reads to reference… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Trigger, Why This Exists, Core Capabilities and Input Formats, plus 11 more sections
  • Runs Python scripts from its folder; calls python and conda

What it does

Seq Wrangler is an agent skill from ClawBio/ClawBio. NGS read QC, alignment, and BAM processing pipeline. Wraps FastQC, BWA/Bowtie2/Minimap2, SAMtools, and MultiQC for automated read-to-BAM workflows.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `README.md`, `examples/demo-results/report.md` and `seq_wrangler.py`).

It sits in Research & Science, covering Bioinformatics. It works with Cloudflare Workers. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/seq-wrangler”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Read QC: Run FastQC, parse results, flag quality issues
  2. Adapter Trimming: Trim adapters with fastp (optional)
  3. Alignment: Align reads to reference genomes (BWA-MEM2, Bowtie2, Minimap2)
  4. BAM Processing: MAPQ filter → name sort → fixmate → coordinate sort → markdup → index
  5. Statistics: flagstat, per-chromosome coverage, insert size (paired-end)
  6. MultiQC Report: Aggregate QC metrics across samples (optional)
  7. Pipeline Generation: Export the full workflow as a shell script or Nextflow pipeline
  8. Reproducibility Bundle: commands.sh, environment.yml, checksums.sha256, run_metadata.json
  9. Demo Mode: Synthetic data run, no external tools required

What it can do on your machine

Read from SKILL.md and the folder at commit 5e045e3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Seq Wrangler loads about 2.5k tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 860 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit 5e045e3, republished under its MIT licence (© ClawBio). 860 words, ~2,498 tokens.

Download SKILL.mdSave it as .claude/skills/seq-wrangler/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
seq-wrangler
description
NGS read QC, alignment, and BAM processing pipeline. Wraps FastQC, BWA/Bowtie2/Minimap2, SAMtools, and MultiQC for automated read-to-BAM workflows.
license
MIT
metadata.author
Daniel Garbozo
metadata.domain
genomics
metadata.emoji
🦖
metadata.os
darwin, linux
metadata.tags
fastq, alignment, bam, qc, bowtie2, bwa, minimap2, samtools, ngs, genomics
metadata.version
0.1.0

🦖 Seq Wrangler

You are the Seq Wrangler, a specialised agent for sequence data QC, alignment, and BAM processing.

Trigger

Fire this skill when the user says any of:

  • "align reads", "align fastq", "align paired-end"
  • "run QC on my reads"
  • "map reads to reference"
  • "process my fastq files"
  • "sort and index this BAM"
  • "what is the coverage of this BAM"
  • "trim adapters and align"
  • "bowtie2", "bwa mem", "minimap2"

Do NOT fire when:

  • User wants variant annotation from a BAM/VCF (route to vcf-annotator)
  • User wants differential expression from a BAM (route to rnaseq-de)
  • User wants methylation analysis (route to methylation-clock)

Why This Exists

Without this skill, aligning FASTQ reads to a reference genome requires manually coordinating 6+ tools (FastQC, fastp, BWA/Bowtie2/Minimap2, samtools sort/fixmate/markdup/index), managing intermediate files, and producing no reproducibility record. Seq Wrangler automates the full read-to-BAM pipeline, enforces MAPQ filtering, marks duplicates, computes per-sample statistics, and generates a reproducibility bundle in a single command.

Core Capabilities

  1. Read QC: Run FastQC, parse results, flag quality issues
  2. Adapter Trimming: Trim adapters with fastp (optional)
  3. Alignment: Align reads to reference genomes (BWA-MEM2, Bowtie2, Minimap2)
  4. BAM Processing: MAPQ filter → name sort → fixmate → coordinate sort → markdup → index
  5. Statistics: flagstat, per-chromosome coverage, insert size (paired-end)
  6. MultiQC Report: Aggregate QC metrics across samples (optional)
  7. Pipeline Generation: Export the full workflow as a shell script or Nextflow pipeline
  8. Reproducibility Bundle: commands.sh, environment.yml, checksums.sha256, run_metadata.json
  9. Demo Mode: Synthetic data run, no external tools required

Input Formats

FormatExtensionRequired fields
FASTQ (SE).fastq.gz, .fq.gzSingle-end reads
FASTQ (PE).fastq.gz, .fq.gzR1 + R2 paired reads
Samplesheet.csvsample, fastq1, fastq2 (optional)
Aligner indexprefixPre-built BWA/Bowtie2/Minimap2 index

Workflow

  1. Validate input files and tools
  2. Run FastQC on all FASTQs (if --run-fastqc)
  3. Trim adapters with fastp (if --trim)
  4. Align reads with selected aligner → SAM
  5. Filter by MAPQ threshold with samtools view
  6. Sort by read name with samtools sort -n
  7. Fix mate-pair information with samtools fixmate
  8. Coordinate sort with samtools sort
  9. Mark (or remove) duplicates with samtools markdup
  10. Index final BAM with samtools index
  11. Compute flagstat, coverage, insert size
  12. Aggregate with MultiQC (if --run-multiqc)
  13. Generate Markdown report and reproducibility bundle

CLI Reference

bash
# Demo (no external tools needed)
python skills/seq-wrangler/seq_wrangler.py --demo --output /tmp/demo

# Single sample paired-end
python skills/seq-wrangler/seq_wrangler.py \
  --r1 sample_R1.fastq.gz \
  --r2 sample_R2.fastq.gz \
  --index ref/hg38 \
  --aligner bowtie2 \
  --output results/

# Single sample single-end
python skills/seq-wrangler/seq_wrangler.py \
  --r1 sample.fastq.gz \
  --index ref/hg38 \
  --aligner bwa \
  --output results/

# Batch mode via samplesheet
python skills/seq-wrangler/seq_wrangler.py \
  --samplesheet samples.csv \
  --index ref/hg38 \
  --output results/

# With trimming and duplicate removal
python skills/seq-wrangler/seq_wrangler.py \
  --r1 sample_R1.fastq.gz --r2 sample_R2.fastq.gz \
  --index ref/hg38 --aligner bowtie2 \
  --trim --remove-duplicates --keep-sam \
  --output results/

Demo

bash
python skills/seq-wrangler/seq_wrangler.py --demo --output /tmp/demo

Expected output: Markdown report with synthetic flagstat (97.5% mapped, 8.7% duplicates) and coverage statistics for two demo samples (CTRL_REP1 paired-end, TREAT_REP1 single-end). No external tools required.

Output Structure

bash
output/
├── report.md # Full alignment and QC report
├── summary.json # Per-sample statistics as JSON
├── bam/
│ └── sample_sorted.bam # Final sorted, markdup BAM
│ └── sample_sorted.bam.bai # BAM index
├── alignment/
│ └── sample.sam # Intermediate SAM (only with --keep-sam)
├── fastqc/ # FastQC reports (if --run-fastqc)
├── trimmed/ # Trimmed FASTQs (if --trim)
├── multiqc/ # MultiQC report (if --run-multiqc)
└── reproducibility/
│ └── commands.sh # Exact command to reproduce this run
│ └── environment.yml # Conda environment spec 
│ └── checksums.sha256 # SHA-256 of all input files
│ └── run_metadata.json # Full run parameters and timestamp

Dependencies

Required:

  • samtools (BAM manipulation)
  • One of: bwa, bowtie2, or minimap2 (alignment)

Optional:

  • fastqc: per-sample read QC
  • fastp: adapter trimming
  • multiqc: aggregated QC report

Install via conda:

bash
conda install -c bioconda samtools bowtie2 bwa minimap2 fastqc fastp multiqc

Gotchas

  • Memory for samtools sort: Uses 2G RAM per thread by default. On machines with <8G RAM, use --threads 2 or --threads 3 to avoid OOM errors.

  • python3 vs python on Windows: Tests use sys.executable instead of python3 for cross-platform compatibility. On Windows, python3 may not exist in PATH.

  • Index prefix vs file: --index expects the aligner index prefix (e.g. hg38_chr22), not a .fa or .bt2 file path. Build with bowtie2-build genome.fa prefix first.

  • SAM files are deleted by default: Use --keep-sam to retain intermediate SAM files. They can be 10x larger than the final BAM.

  • MAPQ filter removes unaligned reads: Default --mapq 20 filters out reads that did not align or aligned poorly. Lower this value if you expect low-quality data.

  • GRCh37 vs GRCh38: The --genome-build flag is for metadata and reporting only. It does not affect alignment — always build your index from the correct reference genome.

Show full SKILL.md (286 more words)Show less

Agent Boundary

The agent (LLM) dispatches the FASTQ files and explains results. The skill (Python) executes all tool calls and generates files. The agent must NOT invent flagstat percentages, coverage values, or insert size statistics.

Safety

  • Local-first: no data is uploaded to external servers
  • Network calls: none
  • Disclaimer: Seq Wrangler is a research and educational tool. Results must be validated before use in clinical or production settings
  • No hardcoded credentials or absolute paths
  • MAPQ filtering applied by default (≥20) to reduce spurious alignments

Integration with Bio Orchestrator

Trigger conditions:

  • User provides FASTQ files and asks for alignment or QC
  • Keywords: align, fastq, bam, coverage, paired-end, bowtie2, bwa

Chaining partners:

  • → rnaseq-de: pass final BAM for differential expression
  • → methylation-clock: pass BAM for methylation analysis
  • → equity-scorer: pass BAM for population equity metrics
  • → acmg: pass aligned BAM for variant calling upstream

Example Queries

  • "Run QC on these FASTQ files and show me the quality summary"
  • "Align paired-end reads to GRCh38 and sort the output BAM"
  • "What is the mean coverage of this BAM file?"
  • "Trim adapters and re-align these reads"
  • "Process this samplesheet of 10 samples with bowtie2 and remove duplicates"
  • "Run the seq-wrangler demo so I can see what the output looks like"
  • "Align these single-end reads with minimap2 and keep the SAM file"

Citations

  • Li H. et al. (2009) The Sequence Alignment/Map format and SAMtools. Bioinformatics
  • Langmead B. & Salzberg S. (2012) Fast gapped-read alignment with Bowtie 2. Nature Methods
  • Li H. & Durbin R. (2009) Fast and accurate short read alignment with BWA. Bioinformatics
  • Li H. (2018) Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics
  • Chen S. et al. (2018) fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics
  • Ewels P. et al. (2016) MultiQC: summarize analysis results for multiple tools. Bioinformatics

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/seq-wrangler of ClawBio/ClawBio.

  • SKILL.md
  • README.md
  • examples/demo-results/report.md
  • seq_wrangler.py
  • tests/test_seq_wrangler.py

Open the folder on GitHubat commit 5e045e3

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in ClawBio/ClawBio, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Seq Wrangler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Seq Wrangler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Seq Wrangler this skillClawBio/ClawBio1.2k1 repos~2.5kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k3 repos~3.4kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
MFA Pipeline Orchestratoraiming-lab/AutoResearchClaw15k—~923Automated safety check: PassMIT

Similar skills

  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 3 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Singlecell Qc

    xuzhougeng/wisp-science

    A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.

    1k GitHub stars~1.6k tokensUpdated today
    Research & ScienceAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.3k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated today
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Questions about Seq Wrangler

What does Seq Wrangler do?

NGS read QC, alignment, and BAM processing pipeline. An agent skill from ClawBio/ClawBio. Seq Wrangler is an agent skill from ClawBio/ClawBio. NGS read QC, alignment, and BAM processing pipeline.

When should I use Seq Wrangler?

Seq Wrangler fits situations like: tasks that involve Bioinformatics.

How do I install Seq Wrangler in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill seq-wrangler -a claude-code`. Or copy the skill folder (skills/seq-wrangler in ClawBio/ClawBio) into .claude/skills/seq-wrangler in your project. Claude Code loads it when a task matches its description.

How do I install Seq Wrangler in Codex?

Run `npx skills add ClawBio/ClawBio --skill seq-wrangler -a codex`. Or copy the skill folder (skills/seq-wrangler in ClawBio/ClawBio) into .agents/skills/seq-wrangler in your project. Codex loads it when a task matches its description.

Can I use Seq Wrangler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill seq-wrangler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/seq-wrangler, .gemini/skills/seq-wrangler, .github/skills/seq-wrangler and .opencode/skills/seq-wrangler in your project.

What does Seq Wrangler need to run?

Going by SKILL.md and its folder, Seq Wrangler needs Python for the scripts in its folder and the command-line tools its instructions call (python and conda). Our summary lists: Python 3.

Does Seq Wrangler access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Seq Wrangler safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Seq Wrangler use?

Seq Wrangler is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Seq Wrangler use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Seq Wrangler?

Skills that share tags, products or a category with Seq Wrangler: Dbsnp Database (google-deepmind/science-skills, 3.2k stars), Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars) and Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Seq Wrangler?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 7, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.