Agent skill

Bio Clinical Databases Variant Prioritization

by GPTomics in GPTomics/bioSkills

Prioritizes rare-disease variants from trio/quad WES/WGS with de novo (DeNovoGear, Triodenovo), compound-heterozygous phasing (WhatsHap), mosaic VAF tiering, phenotype-driven ranking (Exomiser…

MITAuto-check passedProduct & Project Management

Install Bio Clinical Databases Variant Prioritization

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-clinical-databases-variant-prioritization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-clinical-databases-variant-prioritization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clinical-databases/variant-prioritization .claude/skills/bio-clinical-databases-variant-prioritization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-clinical-databases-variant-prioritization
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6.3k tokens
SKILL.md length
2,172 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Prioritizes rare-disease variants from trio/quad WES/WGS with de novo (DeNovoGear, Triodenovo), compound-heterozygous phasing (WhatsHap), mosaic VAF tiering, phenotype-driven ranking (Exomiser…

  • Running diagnostic exome / genome pipelines
  • SKILL.md covers Version Compatibility, Pipeline Architecture: The…, Inheritance-Based Filtering and De Novo Calling: Trio Analysis, plus 12 more sections
  • Runs Python scripts from its folder; calls pip; reaches search.clinicalgenome.org and cspec.genome.network
  • Identifying candidate Mendelian disease genes

What it does

Bio Clinical Databases Variant Prioritization is an agent skill from GPTomics/bioSkills. Prioritizes rare-disease variants from trio/quad WES/WGS with de novo (DeNovoGear, Triodenovo), compound-heterozygous phasing (WhatsHap), mosaic VAF tiering, phenotype-driven ranking (Exomiser, Phen2Gene, AMELIE), ClinGen gene-disease validity gating, and ACMG SF v3.2 secondary findings reporting. Use when running diagnostic exome / genome pipelines, identifying candidate Mendelian disease genes, screening for incidental findings, or auditing VUS reclassification cycles. The ACMG/AMP classification framework…

Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/prioritize_variants.py` and `usage-guide.md`).

It sits in Product & Project Management, covering Prioritization frameworks, Bioinformatics and Performance reviews. It works with Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Running diagnostic exome / genome pipelines
  • Identifying candidate Mendelian disease genes
  • Screening for incidental findings
  • Auditing VUS reclassification cycles

Example prompts

  • “Use the bio-clinical-databases-variant-prioritization skill to prioritiz rare-disease variants from trio/quad WES/WGS with de novo (DeNovoGear…”
  • “/bio-clinical-databases-variant-prioritization”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • search.clinicalgenome.org
    • cspec.genome.network
    • hpo.jax.org
    • gimjournal.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Clinical Databases Variant Prioritization loads about 6.3k tokens when it runs. Until then it costs about 170 tokens; SKILL.md has 2,172 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~170
When it runs · the whole SKILL.md, loaded when a task matches
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,172 words, ~6,259 tokens.

Download SKILL.mdSave it as .claude/skills/bio-clinical-databases-variant-prioritization/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-clinical-databases-variant-prioritization
description
Prioritizes rare-disease variants from trio/quad WES/WGS with de novo (DeNovoGear, Triodenovo), compound-heterozygous phasing (WhatsHap), mosaic VAF tiering, phenotype-driven ranking (Exomiser, Phen2Gene, AMELIE), ClinGen gene-disease validity gating, and ACMG SF v3.2 secondary findings reporting. Use when running diagnostic exome / genome pipelines, identifying candidate Mendelian disease genes, screening for incidental findings, or auditing VUS reclassification cycles. The ACMG/AMP classification framework (PVS1 decision tree, Pejaver PP3/BP4 calibration, Tavtigian point system) is in clinical-databases/acmg-classification.
tool_type
python
primary_tool
pandas

Version Compatibility

Reference examples tested with: pandas 2.2+, cyvcf2 0.30+, pyhgvs 0.12+, Exomiser 14.0+ (Smedley 2015), Phen2Gene 1.2+ (Zhao 2020), DeNovoGear 1.1.1+ (Ramu 2013), WhatsHap 2.0+ (Patterson 2015), HPO 2024+ (Human Phenotype Ontology). ACMG Secondary Findings list is v3.2 (Miller 2023): 81 genes.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying. Phenotype-driven prioritization REQUIRES high-quality HPO terms; without rich phenotypic input Exomiser/AMELIE degrade significantly.

Rare-Disease Variant Prioritization Pipeline

'Prioritize candidate disease-causing variants from this trio exome' -> Filter to rare + functional + inheritance-consistent variants; rank by phenotype concordance; flag ACMG SF v3.2 incidental findings; report tiers with classification logic deferred to acmg-classification.

  • Python (filtering pipeline): pandas + cyvcf2 + myvariant.info aggregation
  • CLI (phenotype-driven ranking): exomiser --analysis hiPHIVE-prioritised.yml
  • Python (de novo calling): DeNovoGear / Triodenovo / PossibleDeNovo
  • CLI (compound het phasing): whatshap phase --indels for singletons; trio-based for families
  • Python (HPO concordance): Phen2Gene / AMELIE / Phenolyzer
  • VCEP curations: https://cspec.genome.network/cspec/ui/svi/all

Pipeline Architecture: The Standard Rare-Disease Funnel

Typical trio exome enters as 40,000-100,000 variants per individual; reaches diagnostic candidate list of 1-10 variants through cascading filters:

StageFilterVariant count (typical trio)
Raw joint-called--100k-150k
QC filter (PASS, depth, GQ, missingness)GATK best practices + Hail QC80k-120k
Population frequencygnomAD grpmax_faf95 < 0.0001 (or disease-specific Whiffin max-credible-AF)5k-15k
Functional consequenceCoding / splice / regulatory1k-3k
Inheritance patternde novo / AR-hom / AR-compoundhet / X-linked / mosaic50-500
Phenotype concordanceExomiser hiPHIVE / Phen2Gene / AMELIE score5-50
ACMG classificationDefer to acmg-classification1-10
ACMG SF v3.2 cross-checkMiller 2023 (81 genes)Separate output

Inheritance-Based Filtering

PatternFilter
De novo (DNV)Variant in proband, absent in both parents; needs trio
Autosomal recessive; homozygousHom-alt in proband; het in both parents
Autosomal recessive; compound hetTwo het variants in same gene on opposite alleles
X-linked recessiveMale proband hemizygous; carrier mother het
X-linked dominantHet in affected; consider XCI skewing in females
Mitochondrial heteroplasmymtDNA variant present at varying heteroplasmy across tissues
MosaicSub-clonal VAF in proband; absent in inherited transmissions

De Novo Calling: Trio Analysis

Goal: Identify variants present in proband but absent in both parents with high specificity.

Approach: Use specialized DNV callers; supplement with manual IGV inspection.

ToolApproachUse case
DeNovoGear (Ramu 2013 Nat Methods)Bayesian, considers parent-of-originStandard for trio WES
Triodenovo (Wei 2015)Bayesian + family-awareAlternative
GATK PossibleDeNovo annotationHard filterQuick prefilter; not standalone
DeNovoCNN (2022)Deep learning trio callerMost accurate as of 2022-2026

False-DNV rate: ~10-30% without manual IGV inspection; concentrated in:

  • Tandem repeat regions (DNM rate inflated)
  • Heterozygous parent with low coverage
  • Mosaic parents (parental mosaicism transmitted to >1 offspring)
  • Mapping errors in segmental duplications

Phenotype-Driven Prioritization

ToolApproachPerformance (typical benchmark)Fails when
Exomiser (Smedley 2015 Nat Protoc)hiPHIVE: phenotype + interactome + sequence damage74% top-1; 94% top-5 (Cipriani 2020)Sparse HPO (< 5 specific terms); novel-disease gene
Phen2Gene (Zhao 2020 NARGAB)HPO-to-gene mapping; faster than ExomiserSimilar top-5Phenotype-only filtering insufficient
AMELIE (Birgmeier 2020 Sci Transl Med)Literature-mining + phenotypeBest when literature is richNew / rare disease without literature; specific patient HPO unmatched
Phenolyzer (Yang 2015 Nat Methods)Phenotype-based gene scoringLegacyModern multi-feature tools (Exomiser, AMELIE) preferred
GADO (Deelen 2019 Nat Commun)Gene Network-based; HPO-free optionWhen HPO is sparsePhenotype-rich cases where Exomiser hiPHIVE wins
CADA (Peng 2021)Cross-species gene prioritizationAnimal model integrationGenes without orthologs; rare-disease without animal model

Critical requirement: all phenotype-driven tools degrade significantly with sparse HPO terms. Capture 5-10 specific HPO terms; avoid generic "intellectual disability" alone.

ClinGen Gene-Disease Validity: Mandatory Gating

Strande et al. 2017 AJHG + ClinGen ongoing curation: Limited / Moderate / Strong / Definitive evidence per gene-disease pair.

CategoryWhen to apply
DefinitiveStrong literature evidence + functional / population genetic evidence
Strong--
Moderate--
LimitedSingle case report or weak segregation
DisputedContradicting evidence
No Known Disease RelationshipGene not associated with the queried disease

Many commercial panels include genes with only Limited validity. ClinGen-curated https://search.clinicalgenome.org/kb/gene-validity is the authoritative directory.

ACMG Secondary Findings v3.2 (Miller 2023 Genet Med 25:100866)

81 genes for opt-in/opt-out reporting on clinical exome/genome. Growth: 56 -> 59 -> 73 -> 78 -> 81. v3.2 additions: CALM1, CALM2, CALM3 (calmodulinopathy; long QT / CPVT; high actionability via beta-blockade + ICD).

Inclusion criteria: ClinGen Strong or Definitive gene-disease validity + ClinGen ADWG actionability scoring.

python
ACMG_SF_V3_2_GENES = [
    # Cardiomyopathies
    'ACTA2', 'ACTC1', 'BAG3', 'COL3A1', 'DES', 'FBN1', 'FLNC', 'GLA', 'LMNA', 'MYBPC3',
    'MYH11', 'MYH7', 'MYL2', 'MYL3', 'PRKAG2', 'PKP2', 'RBM20', 'SCN5A', 'SMAD3',
    'TGFBR1', 'TGFBR2', 'TMEM43', 'TNNC1', 'TNNI3', 'TNNT2', 'TPM1', 'TTN',
    # CALM v3.2 additions (calmodulinopathies)
    'CALM1', 'CALM2', 'CALM3',
    # Arrhythmias and channelopathies
    'CACNA1S', 'KCNH2', 'KCNQ1', 'RYR1', 'RYR2',
    # Vascular
    'ACVRL1', 'ENG',
    # Cancer predisposition
    'APC', 'ATM', 'BAP1', 'BMPR1A', 'BRCA1', 'BRCA2', 'BRIP1', 'CDH1', 'CDKN2A',
    'CHEK2', 'GREM1', 'HOXB13', 'MAX', 'MEN1', 'MLH1', 'MSH2', 'MSH6', 'MUTYH',
    'NF2', 'PALB2', 'PMS2', 'PTEN', 'RAD51C', 'RAD51D', 'RB1', 'RET', 'SDHAF2',
    'SDHB', 'SDHC', 'SDHD', 'SMAD4', 'STK11', 'TMEM127', 'TP53', 'TSC1', 'TSC2',
    'VHL', 'WT1',
    # Other
    'FH', 'GAA', 'HFE', 'HNF1A', 'LDLR', 'OTC', 'PCSK9', 'TTR'
]
# Note: above list is illustrative; pin to Miller 2023 supplement for exact set.

Decision Tree by Scenario

ScenarioRecommended pathWhy
Trio WES, suspected MendelianFull pipeline with DeNovoGear + Exomiser + HPOStandard rare-disease workflow
Singleton WESWhatsHap read-based phasing + AR-hom + AR-compoundhet candidatesCompound het hard without trio
Suspected mosaicLower VAF threshold (2-30%); deep coverage (>200x)Standard tools miss mosaic
Long-read genomeAdd SV calling + STR repeat expansionSVs miss in short-read
Newborn screening (BabyScreen+)605-gene Mendelian panel with current ACMG SF v3.2Lunke 2025 Nat Med 31:4236
Cancer predispositionClinGen Hereditary Cancer VCEPs + ACMG SF cancer subsetUse VCEP CSpec
Cardiomyopathy / arrhythmiaClinGen HCM / DCM / LQT VCEPsStrict gene-disease validity
Population screeningACMG SF v3.2 (81 genes) opt-in/opt-outMiller 2023

Standard Pipeline Workflow

Goal: From a trio joint-called VCF, output ranked candidate variants with inheritance pattern, phenotype concordance, and ACMG SF flags.

Approach: Cascading filters with QC, population frequency, functional consequence, inheritance, phenotype.

python
from cyvcf2 import VCF
import pandas as pd
from pathlib import Path

# Quality + population frequency filter (apply first)
def filter_qc_and_frequency(vcf_path, max_grpmax_faf95=0.0001, min_dp=10, min_gq=20):
    '''Stage 1: QC + frequency filter. Reduces 100k -> ~5-15k variants.'''
    vcf = VCF(vcf_path)
    samples = vcf.samples  # e.g., [proband, mother, father]
    rows = []
    for v in vcf:
        if v.FILTER is not None:
            continue
        if min(v.gt_depths) < min_dp:
            continue
        if v.QUAL is not None and v.QUAL < min_gq:
            continue
        gnomad = (v.INFO.get('grpmax_faf95') or v.INFO.get('AF_grpmax') or
                  v.INFO.get('AF_popmax') or 0)
        if gnomad > max_grpmax_faf95:
            continue
        rows.append({
            'chrom': v.CHROM, 'pos': v.POS, 'ref': v.REF, 'alt': v.ALT[0],
            'genotypes': dict(zip(samples, v.gt_types.tolist())),
            'depth': dict(zip(samples, v.gt_depths.tolist())),
            'gnomad_faf95': gnomad,
            'consequence': v.INFO.get('CSQ', '').split('|')[1] if v.INFO.get('CSQ') else None
        })
    return pd.DataFrame(rows)


def call_de_novo(df, proband, mother, father):
    '''Stage 2: identify DNV candidates: hom-ref both parents, het/hom-alt proband.

    Implements Mendelian-violation logic; supplement with DeNovoGear or DeNovoCNN
    for production (this implementation has 10-30% false-positive rate without IGV).
    '''
    is_dnv = []
    for _, row in df.iterrows():
        gts = row['genotypes']
        if gts[mother] == 0 and gts[father] == 0 and gts[proband] in (1, 3):
            # Mother hom-ref AND father hom-ref AND proband het OR hom-alt
            # Confidence boost: depth at parent sites should be >= 10 to trust hom-ref
            if row['depth'][mother] >= 10 and row['depth'][father] >= 10:
                is_dnv.append(True)
                continue
        is_dnv.append(False)
    df['is_de_novo_candidate'] = is_dnv
    return df


def call_compound_het(df, proband, mother, father, gene_col='gene'):
    '''Stage 3: identify compound het: two het variants in same gene, one from each parent.

    Trio phasing is gold standard; singletons require WhatsHap read-based phasing.
    '''
    het_in_proband = df[df['genotypes'].apply(lambda gts: gts[proband] == 1)]
    candidate_genes = []
    for gene in het_in_proband[gene_col].unique():
        if pd.isna(gene):
            continue
        gene_variants = het_in_proband[het_in_proband[gene_col] == gene]
        # Need >= 2 variants; one inherited from each parent
        maternal_het = gene_variants[gene_variants['genotypes'].apply(
            lambda gts: gts[mother] == 1 and gts[father] == 0)]
        paternal_het = gene_variants[gene_variants['genotypes'].apply(
            lambda gts: gts[father] == 1 and gts[mother] == 0)]
        if len(maternal_het) >= 1 and len(paternal_het) >= 1:
            candidate_genes.append(gene)
    df['is_compound_het_candidate'] = df[gene_col].isin(candidate_genes)
    return df


def flag_acmg_sf(df, acmg_sf_genes, gene_col='gene', clnsig_col='clinvar_sig'):
    '''Stage: flag ACMG Secondary Findings (Miller 2023 v3.2; 81 genes).

    Only P/LP variants in SF genes are reportable as secondary findings.
    '''
    df['is_acmg_sf_candidate'] = (
        df[gene_col].isin(acmg_sf_genes) &
        df[clnsig_col].astype(str).str.contains('athogenic', na=False)
    )
    return df


def filter_by_clingen_validity(df, validity_table, gene_col='gene',
                                min_validity='Moderate'):
    '''Gate on ClinGen gene-disease validity. Limited or Disputed -> low confidence.

    validity_table: DataFrame from `https://search.clinicalgenome.org/kb/gene-validity`
    '''
    rank = {'No Known Disease Relationship': 0, 'Disputed': 0, 'Limited': 1,
            'Moderate': 2, 'Strong': 3, 'Definitive': 4}
    min_rank = rank[min_validity]
    df_merged = df.merge(validity_table, on=gene_col, how='left')
    df_merged['validity_rank'] = df_merged['gene_validity'].map(rank).fillna(0)
    df_merged['pass_validity'] = df_merged['validity_rank'] >= min_rank
    return df_merged


def phenotype_score_with_exomiser_yml(yml_path, vcf_path, hpo_terms, output_dir):
    '''Emit Exomiser command for phenotype-driven ranking.

    HPO terms (e.g., HP:0001250 for seizures) must be SPECIFIC.
    Sparse generic HPO degrades Exomiser hiPHIVE accuracy significantly.
    '''
    return (f'java -jar exomiser-cli-14.0.0.jar --analysis {yml_path} '
            f'--vcf {vcf_path} --hpo {",".join(hpo_terms)} '
            f'--output-dir {output_dir}')

Per-Operation Failure Modes

1. De novo with false-positive rate 10-30%

  • Trigger: Report DNV candidates from Mendelian-violation analysis without IGV inspection.
  • Mechanism: Tandem-repeat regions, low-coverage parents, parental mosaicism, mapping errors in segmental duplications all produce false DNVs.
  • Symptom: 10-30% of reported DNVs are artifacts.
  • Fix: Use DeNovoGear / DeNovoCNN (Bayesian frameworks); manually inspect candidates in IGV; check parental coverage at site.

2. Compound het without phasing

  • Trigger: Report two hets in same gene as compound het without confirming phase.
  • Mechanism: Trans (compound het) vs cis (same chromosome) is critical for AR mechanism.
  • Symptom: False-positive compound het when both variants are in cis.
  • Fix: Trio phasing if available; WhatsHap read-based phasing for variants within ~500 bp; consider long-read for broader phasing.

3. Limited-validity gene reported as diagnostic

  • Trigger: Gene appears on commercial panel; variant labeled disease-causing.
  • Mechanism: Commercial panels often include Limited or Disputed validity genes.
  • Symptom: False-positive diagnostic report.
  • Fix: Cross-check ClinGen gene-disease validity; reject Limited / Disputed without VCEP curation.

4. Sparse HPO terms degrading Exomiser

  • Trigger: Submit Exomiser with single generic HPO (e.g., HP:0001250 "Seizure" only).
  • Mechanism: Phenotype-driven prioritization relies on HPO-to-gene network; sparse terms reduce discriminative power.
  • Symptom: Top-5 rank includes implausible genes; correct diagnosis sub-rank.
  • Fix: Capture 5-10 specific HPO terms (e.g., "infantile spasms with hypsarrhythmia", "facial dysmorphism with hypertelorism").

5. ACMG SF v3.1 used instead of v3.2

  • Trigger: Pipeline reports SF based on 78-gene v3.1 list; misses CALM1/2/3 calmodulinopathies.
  • Mechanism: v3.2 (Miller 2023) added CALM1, CALM2, CALM3.
  • Symptom: Misses calmodulinopathy SF; high-actionability long-QT/CPVT not flagged.
  • Fix: Use Miller 2023 v3.2 list (81 genes); re-run prior cohorts.

6. Mosaic variants below standard VAF threshold

  • Trigger: Filter at VAF >= 30% on standard pipeline.
  • Mechanism: Mosaic variants frequently 2-30% VAF; below threshold filters them out.
  • Symptom: Mosaic disease missed (e.g., Proteus syndrome PIK3CA, McCune-Albright GNAS).
  • Fix: For suspected mosaic disorders, deep coverage (>= 200x); VAF threshold 2-5%; sample affected tissue when possible.

7. ClinVar P variant in Limited-validity gene

  • Trigger: Variant labeled P in ClinVar; gene-disease validity is Limited.
  • Mechanism: ClinVar P is variant-level assertion; gene-disease validity is the upstream question.
  • Symptom: Reported P variant in non-disease-associated gene.
  • Fix: Apply ClinGen gene-disease validity gate BEFORE variant-level interpretation.

8. VUS reclassification gaps

  • Trigger: VUS labeled 2017 still in active diagnostic report 2025.
  • Mechanism: VUS are reclassified as evidence accrues in actively-curated genes; a one-time classification has an expiry date.
  • Symptom: Stale classifications drive incorrect clinical decisions.
  • Fix: Annual VUS re-review for active diagnostic variants; tools like Genome Alert! (Yauy 2022) automate detection of monthly ClinVar changes.

9. Inheritance pattern assumed wrong

  • Trigger: Assume AD inheritance for a gene with variable expressivity / incomplete penetrance.
  • Mechanism: AD genes can have AR variants in functionally significant compound het pattern.
  • Symptom: Miss AR mechanism in mostly-AD gene.
  • Fix: Allow multi-inheritance candidate generation; cross-check ClinGen gene-disease inheritance.
Show full SKILL.md (766 more words)Show less

Reconciliation: When Sources Disagree

PatternLikely causeAction
Exomiser ranks low; ClinVar says PSparse or wrong HPO terms; rare disease in atypical geneRe-run with full HPO; manual review
ClinVar P + ClinGen Limited validityVariant-level vs gene-disease tensionTreat as candidate; require VCEP curation or functional evidence
DeNovoGear high posterior; trio coverage unevenParental mosaicism or mapping errorIGV review; consider parent-of-origin testing
Compound het in phasing-ambiguous geneDistance > 500 bp; can't phase from readsTrio phasing; long-read confirmation
SF gene with V3.1 list; missing CALMMiller 2023 v3.2 updateRe-run with v3.2 (81 genes)
Phenotype tool disagrees with clinicalTool-specific phenotype model; literature gapCross-check with AMELIE for literature-mining alternative
Mosaic suspected but standard pipeline negativeVAF below 30% thresholdDeep targeted sequencing or affected tissue

Quantitative Thresholds and Conventions

ThresholdConventionSource
Rare-disease frequency filtergrpmax_faf95 < 0.0001ClinGen SVI
Recessive disease filtergrpmax_faf95 < 0.005ClinGen SVI
Whiffin gene-specific max-credible-AFComputed per gene + diseaseWhiffin 2017
DNV minimum parental coverage>= 10x both parentsStandard
DNV manual IGV reviewRequired for all reportable DNVsStandard
Compound het phasing<= 500 bp read-based; trio gold standardWhatsHap
Exomiser top-1 diagnostic rank74%; top-5 94% (with rich HPO)Cipriani 2020
ACMG SF v3.2 genes81 (Miller 2023)Miller 2023 Genet Med
VUS reclassification cycleReassess as evidence accrues; ClinGen recommends periodic re-reviewconvention
Mosaic VAF threshold2-30%Convention
ClinGen gene-disease validity gateModerate or Strong minimum for diagnostic reportingClinGen SVI

Common Errors

SymptomCauseSolution
Too many candidate variants (>50)Frequency filter too looseTighten to grpmax_faf95 < 0.0001 (dominant) or 0.005 (recessive)
No DNV candidates in obvious DNV phenotypeFalse-negative DNV callingDeNovoGear / DeNovoCNN; check parental sample swap
Compound het in gene known AD onlyPhasing not validatedConfirm phase via trio or long-read
Exomiser top hit unrelated to phenotypeHPO too generic or wrongAdd specific HPO; check ontology version
Mosaic disease missedVAF threshold too highDeep coverage; affected tissue sampling; VAF 2-5%
SF gene match flagged but variant benignWrong variant classificationApply ACMG framework via acmg-classification skill
Genotype-phenotype discordanceLocus heterogeneity OR multi-gene contributionRun digenic / oligogenic analysis tools

Anticipated Reviewer Pushback

PushbackStandard response
"Why grpmax_faf95 instead of AF?"grpmax_faf95 is the Whiffin 2017 ClinGen-recommended frequency; excludes bottleneck groups; per ACMG SVI specifications.
"Compound het without phase confirmation"Trio phased; if singleton, WhatsHap read-based for variants within 500 bp; long-read otherwise.
"DNV call without IGV review?"All reportable DNVs underwent IGV inspection; we report posterior probability + parental coverage.
"ClinGen Limited validity gene"Excluded per gate; we require Moderate or higher for reportable diagnostic candidates.
"Why ACMG SF v3.2 not v3.1?"v3.2 (Miller 2023) added CALM1/2/3 calmodulinopathies (high actionability). We use current.
"Phenotype-driven prioritization with single HPO term?"We submit 5-10 specific HPO terms; sparse input degrades Exomiser.
"ACMG classification logic?"Variant prioritization (this skill) outputs candidates; ACMG classification (PVS1 / PP3 / BS1 / etc.) is in acmg-classification skill.
"Why not VarSome / Franklin automated ACMG?"We report aggregated annotations via myvariant.info; ACMG classification per acmg-classification skill using Tavtigian point system + Pejaver 2022 calibration.

References

  • Richards S et al. 2015. Standards and guidelines for the interpretation of sequence variants. Genet Med 17:405. (ACMG/AMP)
  • Miller DT et al. 2023. ACMG SF v3.2 list for reporting of secondary findings in clinical exome and genome sequencing. Genet Med 25:100866.
  • Smedley D et al. 2015. Next-generation diagnostics and disease-gene discovery with the Exomiser. Nat Protoc 10:2004.
  • Zhao M et al. 2020. Phen2Gene: rapid phenotype-driven gene prioritization for rare diseases. NARGAB 2:lqaa032.
  • Birgmeier J et al. 2020. AMELIE speeds Mendelian diagnosis by matching patient phenotype and genotype to primary literature. Sci Transl Med 12:eaau9113.
  • Cipriani V et al. 2020. An improved phenotype-driven tool for rare Mendelian variant prioritization. Genes 11:460.
  • Ramu A et al. 2013. DeNovoGear: de novo indel and point mutation discovery and phasing. Nat Methods 10:985.
  • Patterson M et al. 2015. WhatsHap: weighted haplotype assembly for future-generation sequencing reads. J Comput Biol 22:498.
  • Strande NT et al. 2017. Evaluating the clinical validity of gene-disease associations: an evidence-based framework developed by ClinGen. AJHG 100:895.
  • Whiffin N et al. 2017. Using high-resolution variant frequencies to empower clinical genome interpretation. Genet Med 19:1151.
  • Lunke S et al. 2025. Feasibility, acceptability and clinical outcomes of the BabyScreen+ genomic newborn screening study. Nat Med 31:4236.
  • Yauy K et al. 2022. Genome Alert! Genet Med 24:1316. (VUS reclassification monitoring)
  • ClinGen gene-disease validity: https://search.clinicalgenome.org/kb/gene-validity
  • HPO: https://hpo.jax.org/
  • ACMG SF v3.2 supplement: https://www.gimjournal.org/article/S1098-3600(23)00879-1/fulltext
  • clinical-databases/acmg-classification - PVS1 / PP3 / BS1 / PM2 calibration and Tavtigian point system
  • clinical-databases/clinvar-lookup - Variant pathogenicity database query
  • clinical-databases/gnomad-frequencies - Population frequency filtering
  • clinical-databases/myvariant-queries - Aggregated annotation
  • clinical-databases/pharmacogenomics - PGx variant handling
  • variant-calling/clinical-interpretation - Clinical reporting workflow
  • variant-calling/filtering-best-practices - Upstream QC

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in clinical-databases/variant-prioritization of GPTomics/bioSkills.

  • SKILL.md
  • examples/prioritize_variants.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Clinical Databases Variant Prioritization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Clinical Databases Variant Prioritization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Clinical Databases Variant Prioritization this skillGPTomics/bioSkills1.2k2 repos~6.3kAutomated safety check: PassMIT
Product Manager Toolkitdavila7/claude-code-templates33k7 repos~2.2kAutomated safety check: PassMIT
Product Manager Toolkitmajiayu000/spellbook287—~2.2kAutomated safety check: PassMIT
Bio Clinical Databases Variant PrioritizationFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~2kAutomated safety check: PassNone
Outcome Trackermohitagw15856/pm-claude-skills1.4k—~1.6kAutomated safety check: PassMIT
Mapping Mitre Attack Techniquesmukul975/Anthropic-Cybersecurity-Skills34k—~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • Product Manager Toolkit

    davila7/claude-code-templates

    Scores feature requests with RICE, mines customer interview transcripts for pain points, and offers PRD templates, with two Python scripts behind it.

    33k GitHub starsUsed in 7 repos~2.2k tokens
    Product & Project ManagementAuto-check passed
  • Product Manager Toolkit

    majiayu000/spellbook

    Product management helpers: a RICE scoring script, an interview transcript analyzer and PRD templates for prioritizing features, synthesizing research and writing requirements.

    287 GitHub stars~2.2k tokensUpdated yesterday
    Product & Project ManagementAuto-check passed
  • Bio Clinical Databases Variant Prioritization

    FreedomIntelligence/OpenClaw-Medical-Skills

    Filter and prioritize variants by pathogenicity, population frequency, and clinical evidence for rare disease analysis.

    3.1k GitHub stars~2k tokensUpdated 2 mo ago
    Product & Project ManagementAuto-check passed
  • Outcome Tracker

    mohitagw15856/pm-claude-skills

    Record the testable predictions inside a decision, then score them against reality later — so frameworks earn trust from outcomes, not vibes.

    1.4k GitHub stars~1.6k tokensUpdated today
    Product & Project ManagementAuto-check passed
  • Mapping Mitre Attack Techniques

    mukul975/Anthropic-Cybersecurity-Skills

    Maps observed adversary behaviors, security alerts, and detection rules to MITRE ATT&CK techniques and sub-techniques to quantify detection coverage and guide control prioritization.

    34k GitHub stars~1.8k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Regulomedb Database

    jaechang-hits/SciAgent-Skills

    Query RegulomeDB v2 GET REST API to score variants for regulatory function and retrieve overlapping evidence (TF binding, histone marks, DNase peaks, footprints, motifs, eQTLs, chromatin state).

    374 GitHub starsUsed in 1 repo~5.3k tokens
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Clinical Databases Variant Prioritization

What does Bio Clinical Databases Variant Prioritization do?

Prioritizes rare-disease variants from trio/quad WES/WGS with de novo (DeNovoGear, Triodenovo), compound-heterozygous phasing (WhatsHap), mosaic VAF tiering, phenotype-driven ranking (Exomiser…. Bio Clinical Databases Variant Prioritization is an agent skill from GPTomics/bioSkills.2 secondary findings reporting.

When should I use Bio Clinical Databases Variant Prioritization?

Bio Clinical Databases Variant Prioritization fits situations like: running diagnostic exome / genome pipelines; identifying candidate Mendelian disease genes; screening for incidental findings; auditing VUS reclassification cycles.

How do I install Bio Clinical Databases Variant Prioritization in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-variant-prioritization -a claude-code`. Or copy the skill folder (clinical-databases/variant-prioritization in GPTomics/bioSkills) into .claude/skills/bio-clinical-databases-variant-prioritization in your project. Claude Code loads it when a task matches its description.

How do I install Bio Clinical Databases Variant Prioritization in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-variant-prioritization -a codex`. Or copy the skill folder (clinical-databases/variant-prioritization in GPTomics/bioSkills) into .agents/skills/bio-clinical-databases-variant-prioritization in your project. Codex loads it when a task matches its description.

Can I use Bio Clinical Databases Variant Prioritization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-variant-prioritization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clinical-databases-variant-prioritization, .gemini/skills/bio-clinical-databases-variant-prioritization, .github/skills/bio-clinical-databases-variant-prioritization and .opencode/skills/bio-clinical-databases-variant-prioritization in your project.

What does Bio Clinical Databases Variant Prioritization need to run?

Going by SKILL.md and its folder, Bio Clinical Databases Variant Prioritization needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Clinical Databases Variant Prioritization access the network?

SKILL.md names 4 domains. In commands or code: search.clinicalgenome.org, cspec.genome.network, hpo.jax.org and gimjournal.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Clinical Databases Variant Prioritization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Clinical Databases Variant Prioritization use?

Bio Clinical Databases Variant Prioritization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Clinical Databases Variant Prioritization use?

About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Clinical Databases Variant Prioritization?

Skills that share tags, products or a category with Bio Clinical Databases Variant Prioritization: Product Manager Toolkit (davila7/claude-code-templates, 33k stars), Product Manager Toolkit (majiayu000/spellbook, 287 stars), Bio Clinical Databases Variant Prioritization (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars) and Outcome Tracker (mohitagw15856/pm-claude-skills, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Clinical Databases Variant Prioritization?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.