Agent skill

Bio Atac Seq Atac Qc

by GPTomics in GPTomics/bioSkills

ATAC-seq library quality control -- TSS enrichment, FRiP, fragment-size periodicity, library complexity (NRF/PBC1/PBC2), mitochondrial fraction, and ENCODE 4 thresholds.

MITAuto-check passedResearch & Science

Install Bio Atac Seq Atac Qc

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-atac-seq-atac-qc -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-atac-seq-atac-qc --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/atac-seq/atac-qc .claude/skills/bio-atac-seq-atac-qc && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-atac-seq-atac-qc
GitHub stars
1.2k
Used in
2 other repos
Token cost
~5k tokens
SKILL.md length
1,929 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

ATAC-seq library quality control -- TSS enrichment, FRiP, fragment-size periodicity, library complexity (NRF/PBC1/PBC2), mitochondrial fraction, and ENCODE 4 thresholds.

  • Assessing whether an ATAC-seq library passes ENCODE acceptance criteria
  • SKILL.md covers Version Compatibility, ENCODE 4 ATAC-seq Acceptance…, TSS Enrichment: ENCODE Method… and Fragment-Size Periodicity…, plus 12 more sections
  • Runs R scripts from its folder; calls pip
  • Diagnosing transposition artefacts

What it does

Bio Atac Seq Atac Qc is an agent skill from GPTomics/bioSkills. ATAC-seq library quality control -- TSS enrichment, FRiP, fragment-size periodicity, library complexity (NRF/PBC1/PBC2), mitochondrial fraction, and ENCODE 4 thresholds. Use when assessing whether an ATAC-seq library passes ENCODE acceptance criteria, diagnosing transposition artefacts, comparing Omni-ATAC vs standard prep quality, or selecting which replicates to drop before peak calling.

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Assessing whether an ATAC-seq library passes ENCODE acceptance criteria
  • Diagnosing transposition artefacts
  • Comparing Omni-ATAC vs standard prep quality
  • Selecting which replicates to drop before peak calling

Example prompts

  • “/bio-atac-seq-atac-qc”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Atac Seq Atac Qc loads about 5k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 1,929 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,929 words, ~5,006 tokens.

Download SKILL.mdSave it as .claude/skills/bio-atac-seq-atac-qc/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-atac-seq-atac-qc
description
ATAC-seq library quality control -- TSS enrichment, FRiP, fragment-size periodicity, library complexity (NRF/PBC1/PBC2), mitochondrial fraction, and ENCODE 4 thresholds. Use when assessing whether an ATAC-seq library passes ENCODE acceptance criteria, diagnosing transposition artefacts, comparing Omni-ATAC vs standard prep quality, or selecting which replicates to drop before peak calling.
tool_type
mixed
primary_tool
deeptools

Version Compatibility

Reference examples tested with: deepTools 3.5+, Picard 3.1+, samtools 1.19+, bedtools 2.31+, ATACseqQC 1.26+, pysam 0.22+, pyBigWig 0.3+, numpy 1.26+, pandas 2.2+, MultiQC 1.21+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt.

ATAC-seq Quality Control

"Does my ATAC library pass ENCODE quality criteria?" -> Compute the seven canonical metrics (depth, alignment rate, mitochondrial fraction, library complexity, fragment-size periodicity, TSS enrichment, FRiP) and compare against ENCODE 4 thresholds, then diagnose failures.

  • CLI: picard CollectInsertSizeMetrics, samtools flagstat, samtools idxstats
  • CLI: deeptools plotFingerprint, computeMatrix reference-point + plotProfile
  • R: ATACseqQC::TSSEscore, ATACseqQC::fragSizeDist, ATACseqQC::PTscore
  • Python: custom NRF/PBC from coordinate hash; pyBigWig for TSS enrichment

ENCODE 4 ATAC-seq Acceptance Thresholds

MetricDefinitionIdealAcceptableRejectSource
Nuclear reads (after dedup, no chrM)Mapped, MAPQ >= 30, non-chrM, deduped>= 50M25-50M< 25MENCODE 4 ATAC-seq Standards
Alignment rateMapped / total reads>= 95%80-95%< 80%ENCODE 4
Mitochondrial fractionchrM / total mapped< 5% (Omni-ATAC), < 20% (standard)20-50%> 50%Corces 2017 (Omni-ATAC)
NRF (Non-Redundant Fraction)Distinct positions / total reads>= 0.90.7-0.9< 0.7Landt 2012
PBC1 (PCR Bottlenecking Coefficient 1)Positions w/ 1 read / Positions w/ >= 1 read>= 0.90.7-0.9< 0.7Landt 2012
PBC2Positions w/ 1 read / Positions w/ 2 reads>= 3.01.0-3.0< 1.0Landt 2012
TSS enrichment (hg38, GENCODE v29)Avg signal at TSS / avg flanking>= 75-7< 5ENCODE 4
FRiP (Fraction Reads in Peaks)Reads in MACS peaks / total>= 0.30.2-0.3< 0.2ENCODE 4, Landt 2012
Insert-size periodicityNFR + mono-nuc + di-nuc peaks visibleClear 3+ peaksNFR + mono onlyFlat / single peakBuenrostro 2013

ENCODE thresholds are organism-specific. Mouse (mm10, GENCODE M21) TSS enrichment >= 5 is acceptable; non-model organisms have no published threshold (use cohort percentile rank instead). Methodology evolves; verify against the current ENCODE ATAC-seq Standards before reporting.

TSS Enrichment: ENCODE Method vs ATACseqQC Method

The two most common implementations DO NOT produce identical scores.

MethodNumeratorDenominatorScaling
ENCODE pyTSSe / Kundaje gtsseMean signal in 100 bp window centered at TSSMean signal in 100 bp window at +/- 1900 to +/- 2000 bp (flanks)Per-base normalization to flanks; reported as fold-enrichment
ATACseqQC TSSEscoreSum signal in TSS +/- 100 bpSum signal at +/- 1000 bp flanking windowsDifferent window sizes; ratios are larger
deeptools plotProfileVisual; numeric ratio not standardizedReference-point matrixNo standard score; for visualization only

Trigger: Comparing a TSS score across studies.

Mechanism: Different normalization windows shift the absolute number; ATACseqQC's TSSEscore is typically 2-3x ENCODE's because of the wider flank.

Symptom: Reported score 21 vs ENCODE-ideal 7 mismatch. Likely the calculator was ATACseqQC; the equivalent ENCODE score might be 8.

Fix: State which implementation was used. For ENCODE comparisons, use pyTSSe (Kundaje lab) or implement the ENCODE recipe directly.

python
import numpy as np
import pyBigWig

def encode_tss_enrichment(bw_path, tss_bed, flank=2000):
    """ENCODE-style TSS enrichment: signal at TSS center / signal at flanks."""
    bw = pyBigWig.open(bw_path)
    profiles = []
    for line in open(tss_bed):
        chrom, start, end, *rest = line.strip().split('\t')
        tss = int(start)
        strand = rest[2] if len(rest) > 2 else '+'
        try:
            vals = bw.values(chrom, tss - flank, tss + flank)
            if vals is None or len(vals) != 2 * flank: continue
            if strand == '-': vals = vals[::-1]
            profiles.append(np.nan_to_num(vals))
        except RuntimeError:
            continue
    avg = np.nanmean(profiles, axis=0)
    flank_signal = np.mean(np.concatenate([avg[:100], avg[-100:]]))
    center_signal = np.mean(avg[flank - 50: flank + 50])
    return center_signal / flank_signal if flank_signal > 0 else 0.0

Fragment-Size Periodicity Patterns

PatternVisual signatureInterpretationAction
Strong tri-modalNFR (~50bp) >> mono (~200bp) > di (~400bp) > tri (~600bp) peaksExcellent transposition; well-positioned chromatinPass
Clear bi-modalNFR + mono only, di and tri faintAcceptable; common in Omni-ATACPass
Single broad peakFlat after NFR or no NFROver-transposition (too much Tn5) OR degraded chromatinReject; cannot distinguish nucleosomes
Inverted (mono >> NFR)Mono peak dominant, NFR weakUnder-transposition OR chromatin condensationCaution; peak counts will be low
Sharp 147 bp spike with no flanksTight peak at 147 bpChIP-seq input contamination (MNase-like)Reject; not ATAC-grade
10.4 bp helical periodicity overlaySub-peaks at 50, 60, 70, 80 bp on NFRExcellent chromatin structure resolution; helical phasing visiblePass; high-quality

The 10.4 bp helical periodicity is a Buenrostro 2013 hallmark: it reflects the helical pitch of B-form DNA, with Tn5 preferring outward-facing minor grooves on nucleosomal DNA. Its presence is a positive QC indicator but not required.

Per-Metric Failure Modes

Mitochondrial fraction > 50%

Trigger: Standard ATAC-seq protocol on intact cells (no nuclear isolation), or insufficient detergent in lysis.

Mechanism: Mitochondrial DNA is naked (no histones), so Tn5 hyperactively cuts it. Without nuclear-isolation steps (Omni-ATAC pre-spin, OR digitonin lysis with mt removal), chrM dominates the library.

Symptom: samtools idxstats sample.bam | awk '$1=="chrM"' shows >50% of mapped reads on chrM.

Fix: Re-prep with Omni-ATAC (Corces 2017) or fast-ATAC. Re-running QC on chrM-stripped BAM hides the underlying problem; the wasted sequencing remains. If chrM fraction is 30-50%, the library may still be salvageable via chrM removal but yield is reduced.

NRF / PBC1 / PBC2 below threshold

Trigger: Over-amplified library; low input cell count combined with high PCR cycles.

Mechanism: Each PCR cycle doubles starting fragments. With low complexity input (<5000 cells) and >12 cycles, distinct fragments saturate and reads pile up at identical positions. NRF measures unique fragments / total; PBC2 specifically detects multi-copy duplication.

Symptom: NRF < 0.7; PBC2 < 1.0; massive duplicate-removal loss in samtools markdup.

Fix: No fix post-hoc. Re-prep with more starting cells and fewer PCR cycles. Note: ATAC has legitimate duplicates at hyperaccessible sites (Tn5 cuts identically there), so NRF < 0.9 is not by itself fatal. The combined PBC1 < 0.7 + PBC2 < 1.0 + visual coverage pile-ups confirm true bottlenecking.

TSS enrichment < 5

Trigger: Generic chromatin opening throughout the genome (over-transposition), OR genome build mismatch between TSS BED and BAM, OR strand-flip in TSS file.

Mechanism: TSS enrichment requires that signal at TSSs is >> signal in genomic flanks. Over-transposition flattens the signal landscape. Strand-flipped TSSs subtract real signal because TSSs on - strand are calculated from the wrong direction.

Symptom: TSS profile is flat or shows a slight dip at TSS center. Genome browser shows accessibility everywhere, not concentrated at promoters.

Fix: Verify genome build (mm10 vs mm39 differ in TSS positions); verify GTF strand column; confirm signal track was generated post-deduplication. If TSS profile is genuinely flat, library is over-transposed and not recoverable; lower transposition time / Tn5 concentration in next prep.

FRiP < 0.2

Trigger: Signal too diffuse to call peaks (over-transposition), low TSS enrichment, OR peak set is too narrow / restrictive.

Mechanism: FRiP correlates with TSS enrichment because both measure how concentrated the signal is. A diffuse library will have low FRiP regardless of peak count.

Symptom: Peak count looks normal but FRiP < 0.15.

Fix: Check TSS enrichment first. If TSS is also low, the library is over-transposed. If TSS is OK but FRiP is low, the peak caller may be undercalling -- try -p 0.01 (looser) and recalculate FRiP.

Replicate correlation < 0.85

Trigger: Batch effect, technical artefact, or cell-state drift between replicate biological collections.

Mechanism: Pearson correlation on log-scaled binned counts (deepTools multiBamSummary bins -bs 10000) tracks coverage similarity. Below 0.85 indicates non-trivial divergence; ENCODE wants >= 0.9 for biological reps.

Fix: Check PCA; if reps cluster apart from condition, drop the outlier or rerun. If the divergence aligns with batch, add batch as a covariate downstream (DiffBind ~Batch + Condition). Do not silently merge with bad correlation.

Library Complexity (NRF, PBC1, PBC2)

Goal: Detect over-amplification or low-input bottlenecks.

Approach: Hash mapped read positions (or fragment 5' coordinates), tally how many positions have 1, 2, or more reads, and compute the three metrics.

python
import pysam
from collections import Counter

def library_complexity(bam):
    pos_counts = Counter()
    total = 0
    with pysam.AlignmentFile(bam, 'rb') as bf:
        for r in bf.fetch():
            if r.is_unmapped or r.is_secondary or r.is_supplementary:
                continue
            if r.is_duplicate:                          # Mark, not skip; PBC counts pre-dedup
                pass
            total += 1
            key = (r.reference_name, r.reference_start, r.is_reverse)
            pos_counts[key] += 1
    distinct = len(pos_counts)
    histogram = Counter(pos_counts.values())            # {1: N1, 2: N2, ...}
    n1 = histogram.get(1, 0)
    n2 = histogram.get(2, 0)
    nrf = distinct / total if total else 0.0
    pbc1 = n1 / distinct if distinct else 0.0
    pbc2 = n1 / n2 if n2 else float('inf')
    return {'NRF': nrf, 'PBC1': pbc1, 'PBC2': pbc2, 'total': total, 'distinct': distinct}

r.is_duplicate is informational only here; ENCODE NRF/PBC are computed pre-deduplication on the raw mapped BAM.

Show full SKILL.md (753 more words)Show less

Cross-Replicate QC

bash
# Spearman correlation (more robust than Pearson for ATAC)
multiBamSummary bins -bs 10000 -p 8 \
    --bamfiles rep1.bam rep2.bam rep3.bam \
    -o multi.npz

plotCorrelation -in multi.npz \
    --corMethod spearman --whatToPlot heatmap --skipZeros \
    -o spearman_heatmap.png

# Fingerprint (per-bin signal cumulative -- diagonal = no enrichment, sharp curve = good)
plotFingerprint -p 8 -b rep1.bam rep2.bam rep3.bam \
    --labels rep1 rep2 rep3 \
    --skipZeros --numberOfSamples 50000 \
    -o fingerprint.png \
    --outQualityMetrics fingerprint_metrics.txt

deepTools fingerprint quality metrics report a synthetic JS distance without a reference; the (non-synthetic) Jensen-Shannon distance column is only computed when a reference sample is supplied via --JSDsample. Larger values indicate stronger enrichment.

Library Complexity Extrapolation (preseq)

Goal: Predict whether re-sequencing would rescue a low-NRF library, separating "library is bottlenecked" from "we just sequenced too shallow."

Approach: Fit preseq's rational-function (Pade) approximation of the Good-Toulmin power-series estimator on observed BAM read positions; extrapolate distinct-fragment yield as a function of additional sequencing depth.

bash
# c_curve: observed complexity at current depth
preseq c_curve -B sample.bam -o sample.ccurve.tsv -s 1e6

# lc_extrap: predicted complexity at higher depth (extrapolation -e here 200M; preseq default -e is 1e10, step -s default 1M)
preseq lc_extrap -B sample.bam -o sample.lcextrap.tsv -e 200000000 -s 5000000

Interpretation: if lc_extrap shows distinct-fragment count flattening before 100M reads, the library is bottlenecked (re-sequencing won't help; re-prep needed). If it continues to climb, re-sequencing will recover more unique reads. Use alongside NRF/PBC1/PBC2 to decide library re-prep vs deeper sequencing.

Sex-Chromosome QC

Trigger: Clinical-grade ATAC; biobank-scale studies; sample-mix-up detection.

Mechanism: chrY has minimal coverage in female samples; XIST locus (chrX) is highly accessible only in female cells (X-inactivation). Sample-swap or sex-misassignment detectable from these two loci.

bash
# chrY read fraction
samtools idxstats sample.bam | awk '$1=="chrY"{print $3 / $2}'   # reads per bp

# XIST locus accessibility (chrX:73820651-73852753 in hg38)
samtools view -c sample.bam chrX:73820651-73852753

Female: chrY reads/bp ~0; XIST count high. Male: chrY reads/bp ~male coverage; XIST count low. Discrepancy with sample metadata flags swap.

Cell-Cycle Effect on Accessibility

Trigger: Proliferating cell lines (K562, HEK293, HeLa); samples with high S/G2M signature.

Mechanism: Replication-associated chromatin opening adds 5-15% global accessibility shift in proliferating cells; without correction, condition-specific cell-cycle differences confound differential analysis.

Detection: Score cells/samples for S-phase signature (Macosko 2015 cell cycle gene set adapted for chromatin: regulated origin loci, replication-stress-response genes); for bulk ATAC, compute per-sample peak intersection with replication-origin atlas (Repli-seq peaks).

Fix for differential: Add S-phase score as covariate in DESeq2 design (~Sphase + Condition); for scATAC, regress on TF-IDF residuals analogous to Seurat CellCycleScoring.

Spike-in QC (Drosophila or E. coli Chromatin)

Trigger: Studies where global accessibility shift is biological (HDAC inhibitor, DNMT inhibitor, differentiation).

Mechanism: Per-library normalization (RPM, CPM) erases global accessibility shifts because total reads are nominally constant. Exogenous chromatin spike-in (Drosophila S2 or E. coli Tn5-naive chromatin added pre-Tn5) provides an external scaling reference.

Pipeline: Align reads to a concatenated human + Drosophila reference; count spike-in reads per sample; normalize by spike-in (not by total reads). Reske 2020 Epigenetics Chromatin shows that normalization-method choice materially changes differential-accessibility results when a global accessibility shift is expected (ARID1A/PIK3CA endometrial-epithelium case study), motivating an external reference such as a chromatin spike-in.

QC threshold: spike-in fraction 0.5-5% of total reads is the workable range. Below 0.1% spike-in is unreliable; above 10% suggests too much spike-in (loss of cellular reads).

Comprehensive QC Aggregation

Goal: Produce a per-sample report card with PASS/FAIL flags against ENCODE thresholds.

Approach: Compute each metric independently, compare to thresholds, write a tab-delimited report consumable by MultiQC.

python
import json, subprocess, sys
from pathlib import Path

ENCODE_THRESHOLDS = {
    'nuclear_reads_M': (25, 50),                      # (min acceptable, ideal)
    'mt_fraction': (0.5, 0.05),                       # (max acceptable, ideal); inverted
    'NRF': (0.7, 0.9), 'PBC1': (0.7, 0.9), 'PBC2': (1.0, 3.0),
    'TSS_enrichment': (5.0, 7.0), 'FRiP': (0.2, 0.3),
}

def grade(value, thr_acceptable, thr_ideal, inverted=False):
    if inverted:
        return 'FAIL' if value > thr_acceptable else ('PASS' if value <= thr_ideal else 'WARN')
    return 'FAIL' if value < thr_acceptable else ('PASS' if value >= thr_ideal else 'WARN')

def report(metrics, out_tsv):
    rows = []
    for k, (acc, ideal) in ENCODE_THRESHOLDS.items():
        if k not in metrics: continue
        inverted = (k == 'mt_fraction')
        flag = grade(metrics[k], acc, ideal, inverted=inverted)
        rows.append((k, metrics[k], acc, ideal, flag))
    with open(out_tsv, 'w') as f:
        f.write('metric\tvalue\tacceptable\tideal\tflag\n')
        for r in rows: f.write('\t'.join(map(str, r)) + '\n')

MultiQC Aggregation

bash
# Run after generating per-sample QC outputs
multiqc \
    fastqc/ \
    picard/ \
    samtools_stats/ \
    macs2/ \
    deeptools/ \
    -o multiqc_report

MultiQC ingests Picard CollectInsertSizeMetrics, samtools flagstat, deepTools plotFingerprint output, and MACS peaks tables. It does NOT compute TSS enrichment or NRF; pipe a custom _mqc.tsv for those.

Common Errors

Error / symptomCauseSolution
TSS enrichment off by 3x from expectedWrong implementation (ENCODE vs ATACseqQC)State the formula; convert by recomputing
NRF = 1.0 exactlyBAM was already deduplicated -> all positions distinctCompute NRF on raw mapped BAM (pre-dedup)
PBC2 = infNo positions with 2 readsLibrary is too sparse; PBC2 unreliable below ~5M reads
Mt fraction reported but BAM has no chrMMitochondrial chromosome named MT, Mt, or chromosome:MTMatch samtools idxstats chromosome name to the filter
Insert size distribution flat after PicardSample is single-endInsert size only valid for paired-end; switch to deeptools fragmentSize
Replicates correlate poorly but PCA looks fineHigh background dominates correlationUse --skipZeros; or compute correlation on peak counts only
FRiP differs by 2x between identical pipeline runsPeak set differs (q-value cutoff drift)Pin caller version + cutoff; FRiP is peak-set-dependent
TSS enrichment lower than expected on Omni-ATACUsed standard TSS BED on FFPE-prepped sampleFFPE TSSs are degraded; use peak-based metric instead

References

  • Buenrostro JD et al 2013 Nat Methods 10:1213 (ATAC-seq protocol; fragment-size periodicity)
  • Corces MR et al 2017 Nat Methods 14:959 (Omni-ATAC; mt fraction reduction protocol)
  • Landt SG et al 2012 Genome Res 22:1813 (ENCODE/modENCODE QC framework, NRF/PBC definitions; the PBC1/PBC2 split is a later ENCODE-pipeline refinement)
  • ENCODE 4 ATAC-seq Data Standards (encodeproject.org/atac-seq) -- canonical thresholds
  • Ou J et al 2018 BMC Genomics 19:169 (ATACseqQC R package; TSSEscore implementation)
  • Ramirez F et al 2016 Nucleic Acids Res 44:W160 (deepTools, plotFingerprint JSD)
  • Daley T & Smith AD 2013 Nat Methods 10:325 (preseq library-complexity extrapolation model; the lc_extrap re-sequencing decision)
  • atac-seq/atac-peak-calling - FRiP requires peaks; QC drives accept/reject before calling
  • atac-seq/nucleosome-positioning - Fragment-size analysis
  • atac-seq/single-cell-atac - per-cell QC has different thresholds
  • read-qc/quality-reports - upstream FastQC
  • alignment-files/bam-statistics - samtools flagstat / idxstats
  • alignment-files/duplicate-handling - dedup before NRF/PBC computation

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in atac-seq/atac-qc of GPTomics/bioSkills.

  • SKILL.md
  • examples/atac_qc_metrics.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Atac Seq Atac Qc next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Atac Seq Atac Qc compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Atac Seq Atac Qc this skillGPTomics/bioSkills1.2k2 repos~5kAutomated safety check: PassMIT
Uipath GenomeUiPath/skills167—~7.7kAutomated safety check: PassMIT
Pride FetchClawBio/ClawBio1.2k—~4.2kAutomated safety check: PassMIT
Regulomedb Databasejaechang-hits/SciAgent-Skills3741 repos~5.3kAutomated safety check: PassCC-BY-4.0
Conventional Non Oncology Hub Gene Research Planneraipoch/medical-research-skills1.9k—~4.6kAutomated safety check: PassMIT
Conventional Oncology Hub Gene Research Planneraipoch/medical-research-skills1.9k—~4.6kAutomated safety check: PassMIT

Similar skills

  • Uipath Genome

    UiPath/skills

    UiPath automation genome (-genome.md): always invoke to build, execute or edit a genome file, before any skill it names.

    167 GitHub stars~7.7k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Pride Fetch

    ClawBio/ClawBio

    Query metadata and download data from the PRIDE Archive, EMBL-EBI's proteomics identifications database, via the PRIDE Archive REST API v3.

    1.2k GitHub stars~4.2k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • Regulomedb Database

    jaechang-hits/SciAgent-Skills

    Query RegulomeDB v2 GET REST API to score variants for regulatory function and retrieve overlapping evidence (TF binding, histone marks, DNase peaks, footprints, motifs, eQTLs, chromatin state).

    374 GitHub starsUsed in 1 repo~5.3k tokens
    Research & ScienceAuto-check passed
  • Generates complete conventional non-oncology bioinformatics research designs from a user-provided disease context, process-related gene family or biological theme, and validation direction.

    1.9k GitHub stars~4.6k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Generates complete conventional oncology bulk-transcriptome biomarker and hub-gene research designs from a user-provided cancer type and study direction.

    1.9k GitHub stars~4.6k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Generates complete cross-disease shared-biomarker bioinformatics research designs from a user-provided disease pair and validation direction.

    1.9k GitHub stars~4.5k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Atac Seq Atac Qc

What does Bio Atac Seq Atac Qc do?

ATAC-seq library quality control -- TSS enrichment, FRiP, fragment-size periodicity, library complexity (NRF/PBC1/PBC2), mitochondrial fraction, and ENCODE 4 thresholds. Bio Atac Seq Atac Qc is an agent skill from GPTomics/bioSkills. ATAC-seq library quality control -- TSS enrichment, FRiP, fragment-size periodicity, library complexity (NRF/PBC1/PBC2), mitochondrial fraction, and ENCODE 4 thresholds.

When should I use Bio Atac Seq Atac Qc?

Bio Atac Seq Atac Qc fits situations like: assessing whether an ATAC-seq library passes ENCODE acceptance criteria; diagnosing transposition artefacts; comparing Omni-ATAC vs standard prep quality; selecting which replicates to drop before peak calling.

How do I install Bio Atac Seq Atac Qc in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-atac-seq-atac-qc -a claude-code`. Or copy the skill folder (atac-seq/atac-qc in GPTomics/bioSkills) into .claude/skills/bio-atac-seq-atac-qc in your project. Claude Code loads it when a task matches its description.

How do I install Bio Atac Seq Atac Qc in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-atac-seq-atac-qc -a codex`. Or copy the skill folder (atac-seq/atac-qc in GPTomics/bioSkills) into .agents/skills/bio-atac-seq-atac-qc in your project. Codex loads it when a task matches its description.

Can I use Bio Atac Seq Atac Qc in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-atac-seq-atac-qc -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-atac-seq-atac-qc, .gemini/skills/bio-atac-seq-atac-qc, .github/skills/bio-atac-seq-atac-qc and .opencode/skills/bio-atac-seq-atac-qc in your project.

What does Bio Atac Seq Atac Qc need to run?

Going by SKILL.md and its folder, Bio Atac Seq Atac Qc needs R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Atac Seq Atac Qc access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Atac Seq Atac Qc safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Atac Seq Atac Qc use?

Bio Atac Seq Atac Qc is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Atac Seq Atac Qc use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Atac Seq Atac Qc?

Skills that share tags, products or a category with Bio Atac Seq Atac Qc: Uipath Genome (UiPath/skills, 167 stars), Pride Fetch (ClawBio/ClawBio, 1.2k stars), Regulomedb Database (jaechang-hits/SciAgent-Skills, 374 stars) and Conventional Non Oncology Hub Gene Research Planner (aipoch/medical-research-skills, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Atac Seq Atac Qc?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.