Agent skill

Bio Clinical Databases Gnomad Frequencies

by GPTomics in GPTomics/bioSkills

Queries gnomAD v4 (807k samples), v3, v2.1.1, and constraint metrics with grpmax FAF95, bottleneck-group exclusion, LOEUF interpretation, SV/CNV/mtDNA catalogs, and Whiffin max-credible-AF framework.

MITAuto-check passedResearch & Science

Install Bio Clinical Databases Gnomad Frequencies

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-clinical-databases-gnomad-frequencies -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-clinical-databases-gnomad-frequencies --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clinical-databases/gnomad-frequencies .claude/skills/bio-clinical-databases-gnomad-frequencies && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-clinical-databases-gnomad-frequencies
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6.2k tokens
SKILL.md length
2,455 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Queries gnomAD v4 (807k samples), v3, v2.1.1, and constraint metrics with grpmax FAF95, bottleneck-group exclusion, LOEUF interpretation, SV/CNV/mtDNA catalogs, and Whiffin max-credible-AF framework.

  • Filtering rare variants
  • SKILL.md covers Version Compatibility, v2.1.1 / v3.1.2 / v4.x: When…, v4 Ancestry Groups: popmax ->… and Filtering Allele Frequency…, plus 17 more sections
  • Runs Python scripts from its folder; calls pip; reaches gnomad.broadinstitute.org and clinicalgenome.org
  • Applying ACMG BS1/BA1

What it does

Bio Clinical Databases Gnomad Frequencies is an agent skill from GPTomics/bioSkills. Queries gnomAD v4 (807k samples), v3, v2.1.1, and constraint metrics with grpmax FAF95, bottleneck-group exclusion, LOEUF interpretation, SV/CNV/mtDNA catalogs, and Whiffin max-credible-AF framework. Use when filtering rare variants, applying ACMG BS1/BA1, ranking genes by LoF intolerance, or selecting between v2 (GRCh37 + chrX/Y constraint) and v4 (GRCh38 + 807k samples).

Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/gnomad_query.py` and `usage-guide.md`).

It sits in Research & Science. It works with Google Cloud and Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Filtering rare variants
  • Applying ACMG BS1/BA1
  • Ranking genes by LoF intolerance
  • Selecting between v2 (GRCh37 + chrX/Y constraint) and v4 (GRCh38 + 807k samples)

Example prompts

  • “Use the bio-clinical-databases-gnomad-frequencies skill to query gnomAD v4 (807k samples), v3, v2.1.1, and constraint metrics with grpmax FAF95…”
  • “/bio-clinical-databases-gnomad-frequencies”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • gnomad.broadinstitute.org
    • clinicalgenome.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Clinical Databases Gnomad Frequencies loads about 6.2k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 2,455 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,455 words, ~6,203 tokens.

Download SKILL.mdSave it as .claude/skills/bio-clinical-databases-gnomad-frequencies/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-clinical-databases-gnomad-frequencies
description
Queries gnomAD v4 (807k samples), v3, v2.1.1, and constraint metrics with grpmax FAF95, bottleneck-group exclusion, LOEUF interpretation, SV/CNV/mtDNA catalogs, and Whiffin max-credible-AF framework. Use when filtering rare variants, applying ACMG BS1/BA1, ranking genes by LoF intolerance, or selecting between v2 (GRCh37 + chrX/Y constraint) and v4 (GRCh38 + 807k samples).
tool_type
python
primary_tool
requests

Version Compatibility

Reference examples tested with: requests 2.31+, hail 0.2.130+, pandas 2.2+, myvariant 1.0+. Current gnomAD release is v4.1 (May 2024); v4.1 fixed the v4.0 AN under-counting issue that inflated rare-variant AF estimates by 5-10%.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • Hail: hl.version(); pin to >=0.2.130 for v4 schema

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying. The gnomAD browser GraphQL API at https://gnomad.broadinstitute.org/api is the supported public endpoint; Hail Tables on Google Cloud Storage at gs://gcp-public-data--gnomad/ are the supported bulk access.

gnomAD Frequency Queries and Constraint

'How rare is this variant in the general population?' -> Pull allele frequency, grpmax FAF95 (the ACMG-grade frequency), LOEUF gene-level constraint, structural variant catalog, mtDNA frequencies, and the appropriate dataset version per use case.

  • Python (single variant): GraphQL via requests.post('https://gnomad.broadinstitute.org/api', json={'query': ..., 'variables': ...})
  • Python (aggregator): myvariant.MyVariantInfo().getvariant(hgvs, fields=['gnomad_exome', 'gnomad_genome'])
  • Python (bulk): hl.read_table('gs://gcp-public-data--gnomad/release/4.1/ht/exomes/gnomad.exomes.v4.1.sites.ht')

v2.1.1 / v3.1.2 / v4.x: When to Use Which

This is the most consequential decision in any gnomAD query. The releases are not interchangeable; choice determines what can and cannot be said about a variant.

ReleaseBuildSamplesUse whenFails when
v2.1.1GRCh37125,748 exomes + 15,708 genomesConstraint metrics needed (LOEUF v2 most-validated); chrX/Y constraint required; GRCh37 native non-negotiableGRCh38 native cohort; modern rare-variant FAF95 (use v4)
v3.1.2GRCh3876,156 genomes (NO exomes)Non-coding region rare variants on GRCh38; mtDNA frequenciesExome variants needed (no exomes); 76k cohort smaller than v4
v4.0/v4.1GRCh38730,947 exomes + 76,215 genomes = 807,162 totalDefault for everything; rare-variant filtering, FAF95, gene querieschrX/Y constraint (not released); cancer-cohort analysis (no TCGA in v4)

Critical caveats:

  • v4 genomes are the SAME 76,215 v3 samples reprocessed against GRCh38 with updated pipelines; not independent.
  • ~81% of v2 genomes are also in v3; joint v2+v3 meta-analysis must dedupe at sample ID.
  • v4 includes 416,555 UK Biobank exomes under a specific collaboration agreement; check use terms.
  • v4 does NOT include TCGA, so the non_cancer subset is unnecessary; the v4 subset is non_ukb (excludes UKB exomes for ancestry rebalancing).
  • Liftover v2 (GRCh37 -> GRCh38) is NOT equivalent to v4 native; variant representation differs at ~0.5-1% of sites due to assembly fixes.

v4 Ancestry Groups: popmax -> grpmax Terminology

v4 ancestry groups: AFR, AMR, ASJ, EAS, FIN, MID, NFE, SAS, AMI, REMAINING. The MID (Middle Eastern) group was new in v4; previously absorbed into "OTH". The REMAINING group (31,256 v4 samples) is individuals who did not cluster with any reference; they contribute to overall AF but not to grpmax.

Terminology shift: gnomAD documentation and ACMG-facing narrative uses grpmax (genetic ancestry group max) -- replacing the older popmax ("population max") term -- to disambiguate genetic ancestry from self-reported race/ethnicity. The public GraphQL schema still exposes legacy field names containing popmax (e.g. faf95.popmax, faf95.popmax_population); these are the grpmax values under the modern terminology. Always check the schema version when writing queries; new browsers may rename these fields.

grpmax_faf95 is the operational ACMG field. It computes the maximum 95% lower-CI allele frequency, excluding bottleneck groups (AMI, ASJ, FIN, REMAINING) because pathogenic founder variants in those groups would otherwise falsely trigger BS1/BA1. MID is included in grpmax but is the smallest non-bottleneck group with highest per-allele variance.

Filtering Allele Frequency (FAF95): The ACMG-Grade AF

Whiffin 2017 Genet Med 19:1151 introduced FAF95 = Poisson lower bound of 95% CI for AF. By construction, AF > FAF95; FAF95 is the conservative frequency for ACMG application.

Max-credible-AF formula: (prevalence x heterogeneity x allelic-contribution) / (penetrance x 2). Plug in disease parameters to get the gene-specific BA1 / BS1 threshold; compare against grpmax_faf95.

CodeThresholdNotes
BA1AF > 5% in any non-bottleneck groupClinGen SVI default; VCEPs may override (Hearing Loss VCEP uses 0.5%)
BS1AF > gene-specific max-credible-AFComputed per gene via Whiffin formula
PM2_SupportingAbsent or ultra-rare in gnomADDowngraded from PM2_Moderate in SVI 2020

Use grpmax_faf95, not raw AF, for BS1/BA1 application; this is the ClinGen-recommended approach.

Constraint Metrics: pLI, LOEUF, missense Z

Karczewski 2020 Nature 581:434 defined LOEUF as the upper bound of the 90% CI of observed/expected pLoF count per gene. LOEUF is recommended over pLI because it is continuous and accounts for gene size more rigorously.

MetricWhatInterpretation
LOEUFUpper bound of 90% CI of LoF observed/expected ratioLower = more LoF-intolerant; first decile (LOEUF < 0.35 v2; < 0.6 v4) = strongly intolerant
pLIProbability LoF intolerantStill used; gnomAD team recommends LOEUF for ranking
Missense ZZ-score of observed-vs-expected missenseZ > 3.09 = 'constrained' set (~p < 0.001, one-tailed)
Missense O/EObserved/expected missense ratioContinuous form of missense Z

Critical version mismatch:

  • v2.1.1 constraint metrics published 2020; v4 constraint published March 2024 (4 months after v4 data release).
  • v4 constraint is autosomes only; chrX and chrY constraint metrics in v4 are NOT released. For X/Y constraint, fall back to v2.1.1.
  • LOEUF first decile shifted v2 to v4: v2 < 0.35; v4 < 0.6 (larger sample shifted the distribution). Gene rank in deciles is stable across versions but absolute thresholds are NOT interchangeable.

Subsets: non_cancer, non_neuro, controls

ReleaseSubsetRemovesUse when
v2.1.1non_cancerTCGACancer-related variant analysis (avoids circularity)
v2.1.1non_neuroPsychiatric/neuro cohortsNeuropsychiatric variant analysis
v2.1.1controlsCases with known disease (~60k samples)Disease-association calibration
v3.1.2non_v2v2 overlapping samplesIndependent of v2
v3.1.2controls_and_biobanksDisease cases retained, biobanks emphasizedPopulation-level reference
v4non_ukbUK Biobank exomesWhen EUR-skew of UKB problematic
v4non_neuroDeprecated--
v4non_cancerUnnecessary (no TCGA in v4)--

SV Catalog and CNV

ResourceReleaseSamplesCoverage
gnomAD-SV v2Collins 2020 Nature 581:44414,891 unrelated WGS433k SVs, GRCh37
gnomAD-SV v4Nov 202363,046 unrelated WGS1,199,117 high-confidence SVs, GRCh38
gnomAD-CNV v4Nov 2023464,297 individuals (exome-derived gCNV)Rare (AF < 1%) autosomal coding CNVs

gnomAD-CNV v4 is the resource that democratized exome-derived CNV background frequencies; previously only ExAC-CNV provided this at scale.

mtDNA (Laricchia 2022 Genome Res 32:569)

10,850 unique mtDNA variants across 56,434 individuals (v3.1). Frequencies reported per nuclear-ancestry AND per mitochondrial-haplogroup. Heteroplasmy >=10% threshold; ~1/250 individuals carry pathogenic mtDNA variant at heteroplasmy >=10%. mtDNA inheritance is non-Mendelian; standard ACMG criteria do not apply directly; use MITOMAP and HmtVar in parallel.

VEP Version Pinning

Each gnomAD release pins to a VEP version:

  • v4 uses VEP 105 with GENCODE 39 / Ensembl 105 transcripts
  • v2.1.1 uses VEP 85

A variant's consequence prediction can flip between v2 and v4 due to MANE Select adoption and transcript-set updates. Always pin VEP version when reproducing gnomAD annotations.

Decision Tree by Query Scenario

ScenarioRecommended pathWhy
Single variant AF lookupGraphQL API or myvariant.infoLowest latency; returns full per-ancestry breakdown
ACMG BS1/BA1 applicationgrpmax_faf95 from v4The ClinGen-recommended field
Gene-level LoF constraint (autosomes)LOEUF from v4 March 2024 releaseLarger sample, more stable
Gene-level LoF constraint (chrX/Y)LOEUF from v2.1.1v4 X/Y constraint NOT released
Bulk rare-variant filter (cohort-scale)Hail Table on GCSNo rate limits; full schema
SV frequencygnomAD-SV v4 (WGS) or gnomAD-CNV v4 (exome)Choose by data type
mtDNA frequencyv3.1 mtDNA release (Laricchia 2022)Only gnomAD release with mtDNA
Cancer-variant analysisv2.1.1 non_cancer subset OR v4 (no TCGA)Avoid TCGA circularity in v2
Comparison across buildsUse canonical SPDI or CA ID, normalize firstLiftover != native

Single Variant Query (GraphQL)

Goal: Retrieve exome + genome AF, grpmax, FAF95, and per-ancestry breakdown for one variant.

Approach: Hit gnomAD's GraphQL API with explicit dataset version; parse the nested response.

python
import requests

GNOMAD_API = 'https://gnomad.broadinstitute.org/api'

def query_variant(chrom, pos, ref, alt, dataset='gnomad_r4'):
    '''Query gnomAD GraphQL for variant frequency + grpmax FAF95.

    dataset options: gnomad_r4 (v4.1, default), gnomad_r3, gnomad_r2_1
    '''
    query = '''
    query VariantById($variantId: String!, $dataset: DatasetId!) {
      variant(variantId: $variantId, dataset: $dataset) {
        variant_id
        rsids
        exome {
          ac
          an
          af
          homozygote_count
          filters
          populations { id ac an }
          faf95 { popmax popmax_population }
        }
        genome {
          ac
          an
          af
          homozygote_count
          filters
          populations { id ac an }
          faf95 { popmax popmax_population }
        }
      }
    }
    '''
    variant_id = f'{chrom}-{pos}-{ref}-{alt}'
    r = requests.post(GNOMAD_API,
                      json={'query': query, 'variables': {'variantId': variant_id, 'dataset': dataset}},
                      timeout=30)
    r.raise_for_status()
    return r.json().get('data', {}).get('variant')


def grpmax_faf95(payload):
    '''Extract the grpmax FAF95; the ACMG-grade frequency. Excludes bottleneck groups.'''
    exome = payload.get('exome') if payload else None
    if exome and exome.get('faf95'):
        return {
            'faf95': exome['faf95'].get('popmax'),
            'grpmax_ancestry': exome['faf95'].get('popmax_population'),
            'source': 'exome'
        }
    genome = payload.get('genome') if payload else None
    if genome and genome.get('faf95'):
        return {
            'faf95': genome['faf95'].get('popmax'),
            'grpmax_ancestry': genome['faf95'].get('popmax_population'),
            'source': 'genome'
        }
    return {'faf95': 0.0, 'grpmax_ancestry': None, 'source': 'absent'}

ACMG BS1/BA1 Application

Goal: Apply Whiffin max-credible-AF framework to a candidate variant.

Approach: Compute the gene-specific BS1 threshold from disease parameters, compare to grpmax_faf95.

python
def max_credible_af(prevalence, max_allelic_contribution=1.0, max_genetic_contribution=1.0,
                    penetrance=1.0):
    '''Whiffin 2017 max-credible-AF formula.

    Args:
        prevalence: disease prevalence (e.g., 1/10000 = 1e-4)
        max_allelic_contribution: max contribution of single allele to disease in any case
        max_genetic_contribution: max contribution of this gene to disease in any case
        penetrance: probability that variant carriers develop disease

    Returns: max-credible per-allele frequency under dominant inheritance (use /2 for AR)
    '''
    return (prevalence * max_genetic_contribution * max_allelic_contribution) / (penetrance * 2)


def apply_bs1_ba1(grpmax_faf95_val, max_credible, ba1_threshold=0.05):
    '''Apply ClinGen SVI BS1/BA1 criteria.

    BA1 default 5% per ClinGen SVI; VCEP-specific overrides exist (Hearing Loss = 0.5%).
    BS1 = max-credible-AF specific to gene+disease.
    '''
    if grpmax_faf95_val is None:
        return 'PM2_Supporting'  # Absent or ultra-rare
    if grpmax_faf95_val > ba1_threshold:
        return 'BA1'
    if grpmax_faf95_val > max_credible:
        return 'BS1'
    return None  # No criterion triggered; variant is consistent with rare-disease causation

Gene-Level Constraint (LOEUF)

Goal: Retrieve gene constraint metrics with awareness of version mismatch for chrX/Y.

Approach: Use v4 LOEUF for autosomes; fall back to v2.1.1 for chrX/Y. Report LOEUF decile, not raw value, to avoid cross-version comparison errors.

python
def query_gene_constraint(gene_symbol, dataset='gnomad_r4'):
    '''Pull gene constraint metrics. Note: v4 has no chrX/Y constraint; use v2 fallback.'''
    query = '''
    query GeneById($symbol: String!) {
      gene(gene_symbol: $symbol, reference_genome: GRCh38) {
        gene_id
        symbol
        chrom
        gnomad_constraint {
          oe_lof
          oe_lof_lower
          oe_lof_upper
          oe_mis
          oe_mis_upper
          pli
          mis_z
        }
      }
    }
    '''
    r = requests.post(GNOMAD_API,
                      json={'query': query, 'variables': {'symbol': gene_symbol}},
                      timeout=30)
    r.raise_for_status()
    gene = r.json().get('data', {}).get('gene')
    if gene is None:
        return None
    if gene.get('chrom') in ('X', 'Y'):
        gene['constraint_note'] = ('v4 constraint NOT released for chrX/Y; query v2.1.1 '
                                   'via gnomad_r2_1 dataset on the v2 endpoint')
    return gene

Bulk Query via Hail (cohort-scale)

Goal: Filter millions of variants by AF, grpmax, or LOEUF without API rate limits.

Approach: Read gnomAD v4 Hail Table from Google Cloud Storage; use hl.read_table() + filter operations.

python
import hail as hl

def init_hail_for_gnomad():
    '''Initialize Hail for gnomAD v4 GCS access. Requires Hail 0.2.130+.'''
    hl.init(default_reference='GRCh38')


def filter_rare_variants_hail(input_vcf, max_grpmax_faf95=0.0001, output_path='filtered.mt'):
    '''Filter input MT to variants below grpmax FAF95 threshold using gnomAD v4 exomes.'''
    ht_v4 = hl.read_table('gs://gcp-public-data--gnomad/release/4.1/ht/exomes/'
                          'gnomad.exomes.v4.1.sites.ht')
    mt = hl.import_vcf(input_vcf, reference_genome='GRCh38')
    mt = mt.annotate_rows(gnomad=ht_v4[mt.locus, mt.alleles])
    mt = mt.filter_rows(
        (hl.is_missing(mt.gnomad.grpmax_faf95)) |
        (mt.gnomad.grpmax_faf95.faf95 < max_grpmax_faf95)
    )
    mt.write(output_path, overwrite=True)
    return mt
Show full SKILL.md (1,145 more words)Show less

Per-Operation Failure Modes

1. Using popmax/AF where grpmax_faf95 belongs

  • Trigger: Apply BS1 with raw AF instead of FAF95.
  • Mechanism: Raw AF inflates for low-N populations; FAF95 is the lower-bound CI; conservative.
  • Symptom: Pathogenic variants falsely categorized BS1 in small-N ancestry groups (especially MID with v4's smallest sample size).
  • Fix: Use grpmax_faf95.popmax field; not populations[i].af.

2. Failing to exclude bottleneck groups

  • Trigger: Compute grpmax including AMI, ASJ, FIN, REMAINING.
  • Mechanism: Founder variants in bottleneck groups can reach AF > 5% but are not population-general; would falsely trigger BA1.
  • Symptom: Founder-population pathogenic variants reported benign.
  • Fix: Use gnomAD's pre-computed grpmax_faf95 which excludes bottleneck groups by design.

3. Querying v4 constraint for chrX/Y

  • Trigger: Pull LOEUF for DMD or USP9Y from v4 release.
  • Mechanism: v4 March 2024 constraint release excluded sex chromosomes.
  • Symptom: Missing or stale constraint metrics for X/Y genes.
  • Fix: Query v2.1.1 LOEUF for chrX/Y; use v4 for autosomes; report LOEUF decile rather than raw value.

4. Comparing LOEUF absolute values across v2/v4

  • Trigger: "v4 LOEUF for GENE-X is 0.45; v2 was 0.30; has it become more tolerant?"
  • Mechanism: Larger v4 sample shifts the LOEUF distribution upward; first-decile threshold shifted v2 < 0.35 -> v4 < 0.6.
  • Symptom: Genes appear to lose constraint between versions when they have not.
  • Fix: Compare deciles, not absolute values; or stay within one version.

5. v2 -> v4 liftover assumed equivalent

  • Trigger: Project v2 GRCh37 variants onto GRCh38 with CrossMap, treat as v4 native.
  • Mechanism: ~0.5-1% of sites have different representations after liftover due to assembly fixes (e.g., gaps closed, contigs joined).
  • Symptom: Inconsistent AFs at low rate; failed cross-version reproducibility.
  • Fix: Query v4 native by GRCh38 coordinates directly; do not use liftover output as v4-equivalent.

6. UKB sample contamination of grpmax

  • Trigger: Compute grpmax across v4 default subset; observe inflated NFE/SAS.
  • Mechanism: 416,555 UK Biobank exomes dominate the v4 NFE+SAS subsets.
  • Symptom: Variants common in UKB but rare globally falsely look common.
  • Fix: Use non_ukb subset for grpmax when ancestry composition matters.

7. v3 vs v4 confusion; "I want WGS"

  • Trigger: User says "I want WGS AFs" and pipeline pulls v4 genomes.
  • Mechanism: v4 genomes are the SAME 76,215 v3 samples reprocessed against GRCh38; not independent.
  • Symptom: WGS AFs appear identical to v3.1.2; not a bug, but worth flagging.
  • Fix: Document that v4 genomes = v3 genomes reprocessed; for true independent WGS, no such resource yet exists at scale.

8. Constraint applied to multi-isoform gene without transcript awareness

  • Trigger: Apply LOEUF "for the gene" when LoF is isoform-specific.
  • Mechanism: gnomAD constraint is computed on the canonical transcript; tissue-specific or alternative isoforms may have different LoF tolerance.
  • Symptom: Mis-prioritization of variants on minor transcripts.
  • Fix: Cross-check with MANE Select; for isoform-specific LoF, use isoform-level constraint where available (rare).

Reconciliation: When Sources Disagree

PatternLikely causeAction
ClinVar P vs gnomAD grpmax_faf95 > 1%Founder-population pathogenic; or ClinVar is stale low-starApply Whiffin max-credible-AF for the gene; check ClinVar star/freshness
v2 LOEUF < 0.35 vs v4 LOEUF = 0.5Distribution shifted with v4 sample size, not biologyUse deciles; v4 first decile = < 0.6
v2 AF != v4 AF for same variantSample overlap (v3 in v4) + new exomes; expectedTrust v4 default; non-overlapping subsets via non_v2 or non_ukb
Variant present in v3 genomes, absent v4 exomesVariant outside exome capture region (intronic, intergenic)Use v3.1.2 or v4 genomes for non-coding
gnomAD-SV v2 vs v4 different breakpointsv2 GRCh37, v4 GRCh38; assembly fixes shift coordsUse v4 native; document build
Browser shows lower AF than Hail TableBrowser pre-filters with filters=PASS; Hail Table includes allApply filters filter in Hail explicitly

Quantitative Thresholds and Conventions

ThresholdConventionSource
BA1 defaultgrpmax_faf95 > 5% in non-bottleneck groupRichards 2015 + ClinGen SVI
BS1grpmax_faf95 > gene-specific max-credible-AFWhiffin 2017
PM2_SupportingAbsent or ultra-rare in gnomADSVI 2020 downgrade
LOEUF first decile v2< 0.35Karczewski 2020
LOEUF first decile v4< 0.6gnomAD constraint release March 2024
Missense Z constrainedZ > 3.09 (~p < 0.001)Samocha 2014
mtDNA heteroplasmy carrier threshold>=10% heteroplasmyLaricchia 2022
v4 sample size730,947 exomes + 76,215 genomes = 807,162gnomAD v4.0 release Nov 2023
Bottleneck groups (excluded from grpmax)AMI, ASJ, FIN, REMAININGgnomAD v4 documentation
API rate limitNone published; ~10 req/s practicalgnomAD browser GraphQL

Common Errors

SymptomCauseSolution
Cannot read property 'af' of undefinedVariant not in dataset; variant returned nullCheck if payload is None; absence is biologically informative
FAF95 = 0 for a known common variantgrpmax_faf95 only computed when AN sufficientCheck AC and AN directly; FAF95 is 0 when N too low to estimate
Variant filter status AC0 or RFFailed gnomAD QCVariants with non-PASS should usually be excluded from analysis
Different AFs between gnomAD browser and Hail TableBrowser auto-applies PASS filter; Hail does notFilter filters.size() == 0 (i.e., PASS) in Hail
LOEUF appears worse in v4 vs v2Distribution shifted with larger sampleCompare deciles, not absolute values
SV not found in v4-SVv2-SV is GRCh37, v4-SV is GRCh38; or variant not called in WGSTry v2-SV with liftover; or check gnomAD-CNV for exome-derived
mtDNA variant missingOnly v3.1 has mtDNA; not in v4Query v3.1 directly

Anticipated Reviewer Pushback

PushbackStandard response
"Why FAF95 instead of AF?"Raw AF is point estimate; FAF95 is Poisson lower-bound 95% CI; ClinGen SVI recommendation for BS1/BA1.
"Why exclude FIN and ASJ from grpmax?"Founder-population pathogenic variants reach high AF locally; including them would trigger false BA1.
"This LOEUF differs from the 2020 paper"We use v4 March 2024 constraint (807k samples); 2020 paper used v2 (141k samples). Decile rank is stable; absolute shifted.
"Why not v4 for chrX constraint?"v4 March 2024 constraint release is autosomes only; chrX/Y not yet released as of 2025. Fall back to v2.
"Why v3 if v4 exists?"v4 genomes = v3 genomes reprocessed; for genome-only analysis they are equivalent.
"Variant exists in liftover v2 but not v4"~0.5-1% of sites differ post-assembly fixes; use v4 native, not liftover, as ground truth.
"Browser AF higher than this value"Browser includes flagged variants by default; we filter on PASS.

References

  • Chen S et al. 2024. A genomic mutational constraint map using variation in 76,156 human genomes. Nature 625:92.
  • Karczewski KJ et al. 2020. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581:434.
  • Samocha KE et al. 2014. A framework for the interpretation of de novo mutation in human disease. Nat Genet 46:944.
  • Collins RL et al. 2020. A structural variation reference for medical and population genetics. Nature 581:444.
  • Laricchia KM et al. 2022. Mitochondrial DNA variation across 56,434 individuals in gnomAD. Genome Res 32:569.
  • Whiffin N et al. 2017. Using high-resolution variant frequencies to empower clinical genome interpretation. Genet Med 19:1151.
  • ClinGen guidance on gnomAD v4 (March 2024): https://clinicalgenome.org/site/assets/files/9445/clingen_guidance_to_vceps_regarding_the_use_of_gnomad_v4_march_2024.pdf
  • gnomAD v4 release notes: https://gnomad.broadinstitute.org/news/2023-11-gnomad-v4-0/
  • gnomAD v4.1 updates: https://gnomad.broadinstitute.org/news/2024-05-gnomad-v4-1-updates/
  • clinical-databases/clinvar-lookup - Pathogenicity classification (gnomAD AF used for BS1/BA1)
  • clinical-databases/acmg-classification - Whiffin FAF95 framework applied to ACMG criteria
  • clinical-databases/variant-prioritization - Rare-disease pipeline using grpmax_faf95
  • clinical-databases/myvariant-queries - Aggregated queries including gnomAD overlay
  • population-genetics/population-structure - Population stratification background

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in clinical-databases/gnomad-frequencies of GPTomics/bioSkills.

  • SKILL.md
  • examples/gnomad_query.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Clinical Databases Gnomad Frequencies next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Clinical Databases Gnomad Frequencies compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Clinical Databases Gnomad Frequencies this skillGPTomics/bioSkills1.2k2 repos~6.2kAutomated safety check: PassMIT
GitHub Deep Researchbytedance/deer-flow84k4 repos~1.3kAutomated safety check: PassMIT
Last30daysmvanhorn/last30days-skill64k—~7.9kAutomated safety check: NotesMIT
NetworkxzLanqing/codex-claude-academic-skills4.7k15 repos~3.2kAutomated safety check: PassBSD-3-Clause
Nature-Style Scientific FiguresYuan1z0825/nature-skills47k—~3.1kAutomated safety check: PassApache-2.0
Citation ManagementK-Dense-AI/claude-scientific-writer2.4k2 repos~3.9kAutomated safety check: NotesMIT

Similar skills

  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    84k GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Last30days

    mvanhorn/last30days-skill

    Research what people actually say about any topic in the last 30 days.

    64k GitHub stars~7.9k tokensUpdated yesterday
    Research & ScienceAuto-check: notes
  • Networkx

    zLanqing/codex-claude-academic-skills

    Comprehensive toolkit for creating, analyzing, and visualizing complex networks and graphs in Python.

    4.7k GitHub starsUsed in 15 repos~3.2k tokens
    Research & ScienceAuto-check passed
  • Nature-Style Scientific Figures

    Yuan1z0825/nature-skills

    Creates, revises, audits and exports manuscript-ready scientific figures in Python or R, and routes AI-generated graphical abstracts to a separate workflow.

    47k GitHub stars~3.1k tokensUpdated today
    Research & ScienceAuto-check passed
  • Citation Management

    K-Dense-AI/claude-scientific-writer

    Finds papers in OpenAlex, PubMed and Google Scholar, turns DOIs, PMIDs and arXiv IDs into clean BibTeX, and validates citations for a manuscript or thesis.

    2.4k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Preprint Search on bioRxiv

    LigphiDonk/Oh-my--paper

    Searches bioRxiv life sciences preprints by keyword, author, date range or category with a Python script, returning JSON metadata and optional PDF downloads.

    739 GitHub starsUsed in 12 repos~3.7k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Clinical Databases Gnomad Frequencies

What does Bio Clinical Databases Gnomad Frequencies do?

Queries gnomAD v4 (807k samples), v3, v2.1.1, and constraint metrics with grpmax FAF95, bottleneck-group exclusion, LOEUF interpretation, SV/CNV/mtDNA catalogs, and Whiffin max-credible-AF framework. Bio Clinical Databases Gnomad Frequencies is an agent skill from GPTomics/bioSkills.1, and constraint metrics with grpmax FAF95, bottleneck-group exclusion, LOEUF interpretation, SV/CNV/mtDNA catalogs, and Whiffin max-credible-AF framework.

When should I use Bio Clinical Databases Gnomad Frequencies?

Bio Clinical Databases Gnomad Frequencies fits situations like: filtering rare variants; applying ACMG BS1/BA1; ranking genes by LoF intolerance; selecting between v2 (GRCh37 + chrX/Y constraint) and v4 (GRCh38 + 807k samples).

How do I install Bio Clinical Databases Gnomad Frequencies in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-gnomad-frequencies -a claude-code`. Or copy the skill folder (clinical-databases/gnomad-frequencies in GPTomics/bioSkills) into .claude/skills/bio-clinical-databases-gnomad-frequencies in your project. Claude Code loads it when a task matches its description.

How do I install Bio Clinical Databases Gnomad Frequencies in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-gnomad-frequencies -a codex`. Or copy the skill folder (clinical-databases/gnomad-frequencies in GPTomics/bioSkills) into .agents/skills/bio-clinical-databases-gnomad-frequencies in your project. Codex loads it when a task matches its description.

Can I use Bio Clinical Databases Gnomad Frequencies in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-gnomad-frequencies -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clinical-databases-gnomad-frequencies, .gemini/skills/bio-clinical-databases-gnomad-frequencies, .github/skills/bio-clinical-databases-gnomad-frequencies and .opencode/skills/bio-clinical-databases-gnomad-frequencies in your project.

What does Bio Clinical Databases Gnomad Frequencies need to run?

Going by SKILL.md and its folder, Bio Clinical Databases Gnomad Frequencies needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Clinical Databases Gnomad Frequencies access the network?

SKILL.md names 2 domains. In commands or code: gnomad.broadinstitute.org and clinicalgenome.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Clinical Databases Gnomad Frequencies safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Clinical Databases Gnomad Frequencies use?

Bio Clinical Databases Gnomad Frequencies is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Clinical Databases Gnomad Frequencies use?

About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Clinical Databases Gnomad Frequencies?

Skills that share tags, products or a category with Bio Clinical Databases Gnomad Frequencies: GitHub Deep Research (bytedance/deer-flow, 84k stars), Last30days (mvanhorn/last30days-skill, 64k stars), Networkx (zLanqing/codex-claude-academic-skills, 4.7k stars) and Nature-Style Scientific Figures (Yuan1z0825/nature-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Clinical Databases Gnomad Frequencies?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.