Agent skill

Bio Clinical Databases Dbsnp Queries

by GPTomics in GPTomics/bioSkills

Resolves rsIDs, navigates RsMergeArch/SNPHistory merge chains, and converts between rsID, SPDI, HGVS, and VCF representations using the dbSNP Build 156 JSON architecture.

MITAuto-check passedResearch & Science

Install Bio Clinical Databases Dbsnp Queries

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-clinical-databases-dbsnp-queries -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-clinical-databases-dbsnp-queries --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clinical-databases/dbsnp-queries .claude/skills/bio-clinical-databases-dbsnp-queries && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-clinical-databases-dbsnp-queries
GitHub stars
1.2k
Used in
2 other repos
Token cost
~5.2k tokens
SKILL.md length
1,769 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Resolves rsIDs, navigates RsMergeArch/SNPHistory merge chains, and converts between rsID, SPDI, HGVS, and VCF representations using the dbSNP Build 156 JSON architecture.

  • Normalizing variant identifiers
  • SKILL.md covers Version Compatibility, rsID Is a Cluster Identifier,…, Build 156 Schema Overhaul:… and RsMergeArch: The Multi-Hop…, plus 13 more sections
  • Runs Python scripts from its folder; calls pip; reaches api.ncbi.nlm.nih.gov and ncbi.nlm.nih.gov
  • Joining variant databases by cluster ID

What it does

Bio Clinical Databases Dbsnp Queries is an agent skill from GPTomics/bioSkills. Resolves rsIDs, navigates RsMergeArch/SNPHistory merge chains, and converts between rsID, SPDI, HGVS, and VCF representations using the dbSNP Build 156 JSON architecture. Use when normalizing variant identifiers, joining variant databases by cluster ID, or tracking deprecated rsIDs through historical merges.

Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/dbsnp_lookup.py` and `usage-guide.md`).

It sits in Research & Science. It works with NCBI and Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Normalizing variant identifiers
  • Joining variant databases by cluster ID
  • Tracking deprecated rsIDs through historical merges

Example prompts

  • “Use the bio-clinical-databases-dbsnp-queries skill to resolve rsIDs, navigates RsMergeArch/SNPHistory merge chains, and converts between rsID, SPDI…”
  • “/bio-clinical-databases-dbsnp-queries”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.ncbi.nlm.nih.gov
    • ncbi.nlm.nih.gov
    • reg.clinicalgenome.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Clinical Databases Dbsnp Queries loads about 5.2k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 1,769 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,769 words, ~5,162 tokens.

Download SKILL.mdSave it as .claude/skills/bio-clinical-databases-dbsnp-queries/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-clinical-databases-dbsnp-queries
description
Resolves rsIDs, navigates RsMergeArch/SNPHistory merge chains, and converts between rsID, SPDI, HGVS, and VCF representations using the dbSNP Build 156 JSON architecture. Use when normalizing variant identifiers, joining variant databases by cluster ID, or tracking deprecated rsIDs through historical merges.
tool_type
python
primary_tool
myvariant

Version Compatibility

Reference examples tested with: myvariant 1.0+, requests 2.31+, biopython 1.83+, Entrez Direct 21.0+. dbSNP Build 156 (September 2022) is the current schema; Build 151 (2017) was the last with relational SQL dumps. Builds 152-155 dual-released JSON+SQL; 156+ is JSON-only.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • CLI: <tool> --version then <tool> --help to confirm flags

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying. The Variation Services REST API uses path-based versioning (/variation/v0/); E-utilities db=snp returns thin legacy summaries missing build-156 schema fields.

dbSNP Queries and rsID Normalization

'Look up this rsID / normalize variant representations' -> Resolve rsIDs through merge chains, compute canonical SPDI, and convert between rsID, HGVS-g, HGVS-c, and VCF allele representations.

  • Python (aggregator): myvariant.MyVariantInfo().getvariant(rsid, fields=['dbsnp', 'clinvar', 'gnomad_exome'])
  • Python (direct): requests.get(f'https://api.ncbi.nlm.nih.gov/variation/v0/refsnp/{rsid_int}')
  • Python (E-utilities, legacy): Bio.Entrez.esearch(db='snp', term=rsid); returns thin summary
  • Bulk: ftp.ncbi.nlm.nih.gov/snp/latest_release/JSON/refsnp-chr{N}.json.bz2

rsID Is a Cluster Identifier, Not a Variant Identifier

This is the load-bearing concept. dbSNP cluster definition: ss records (submitted SNPs) are mapped to the genome and clustered into RefSNPs by position + variant type, not by allele. A single rsID can point to a locus with multiple alleles:

  • rs12345 may resolve to {A>G, A>T, A>C} at one position; the RefSNP JSON primary_snapshot_data.placements_with_allele[*].alleles enumerates them.
  • ~6-8% of dbSNP rsIDs are multi-allelic.
  • PLINK and many older tools historically misuse rsIDs as if they were variant identifiers, which fails for multi-allelic sites and yields wrong genotype assignments.

Rule: Use rsID as a human-facing label only; use SPDI or ClinGen Allele Registry CA ID for joins.

Build 156 Schema Overhaul: What Changed

AspectBuild 151 (2017)Build 156 (2022) and current
DistributionRelational SQL dumps + XMLJSON per RefSNP, partitioned by chromosome
FTP pathftp/snp/organisms/human_9606/ftp.ncbi.nlm.nih.gov/snp/latest_release/JSON/
Primary keysnp_id, ss_idrefsnp_id, with primary_snapshot_data block
Frequency dataEmbedded sparseALFA aggregated populations
Merge trackingRsMergeArch.bcp.gzrefsnp-merged.json.bz2 (also RsMergeArch.bcp.gz retained for legacy)
WithdrawnSNPHistory.bcp.gzrefsnp-withdrawn.json.bz2
API accessLegacy E-utilities db=snp onlyVariation Services REST /v0/refsnp/{id} returns the full JSON

E-utilities still works for db=snp but returns a thin pre-156 summary missing key fields like primary_snapshot_data.placements_with_allele; pipelines reliant on Entrez get out-of-date data.

RsMergeArch: The Multi-Hop Merge Footgun

When two rsIDs are found to refer to the same allele cluster, the higher (later-assigned) rsID is merged into the lower. RsMergeArch.bcp.gz stores (rsHigh, rsLow, rsCurrent) tuples.

The trap: rsCurrent in any given row is the merge target at the time of that merge event, NOT the current dbSNP rsID. A multi-merge chain (rs3 -> rs2 -> rs1, then later rs1 -> rs0) appears as multiple rows. Naive one-hop lookup resolves to a stale ID.

Withdrawn rsIDs (submitter-withdrawn or QC-failed) live in SNPHistory.bcp.gz, not RsMergeArch. Both tables must be consulted to resolve any historical rsID.

SPDI: The Canonical Variant Representation

SPDI (Sequence:Position:Deletion:Insertion) format: NC_000017.11:43044294:G:A. Position is 0-based, half-open (differs from HGVS's 1-based, fully-closed).

The Contextual Allele transformation (Variant Overprecision Correction Algorithm) returns the right-aligned, normalized canonical form across left-aligned VCFs and right-aligned HGVS conventions. This is the basis for ClinGen Allele Registry CA ID computation.

RepresentationBuild dependencyTranscript dependencyBijective?Best for
VCF (chrom-pos-ref-alt)YesNoYes (same build)Pipelines, bulk
SPDIYes (via RefSeq accession)NoYes for SNV/small indelCanonical normalization
HGVS-gYes (NC_xxxxx.N)NoYes for SNV/small indelHuman-readable genomic
HGVS-cIndirect (via transcript)YesNo (one HGVS-c -> many HGVS-g)Clinical reporting
HGVS-pIndirectYesDegenerate (one HGVS-p -> many HGVS-c)Protein-level annotation
rsIDNone (cluster identifier)NoneNO (multi-allelic)Human label only
CA IDNone (canonical)NoneYesCross-database join

Decision Tree by Query Scenario

ScenarioRecommended pathWhy
Resolve single rsID to coordinates + allelesVariation Services /v0/refsnp/{id}Returns full Build 156 JSON, including merge history
Resolve historical/deprecated rsIDVariation Services /v0/refsnp/{id} -> follow merged_snapshot_data chainSingle-hop RsMergeArch lookup misses multi-hop chains
Batch query 100-10k rsIDsmyvariant.info getvariants(rsids)Aggregated with ClinVar/gnomAD overlay; rate-limit safe
Convert coords <-> rsIDmyvariant.info HGVS query or Variation Services /spdi/{spdi}/rsidSPDI is the canonical bridge
Normalize variant representationsVariation Services /hgvs/{hgvs}/contextualsReturns canonical SPDI, right-aligned
Bulk genomic-wide rsID -> coordsLocal download of refsnp-chr{N}.json.bz2 + parserNo rate limits; weekly snapshots
Joining dbSNP with gnomAD by IDUse SPDI or CA ID, never rsID alonersID is a cluster; alleles may not match
Get population AF for common variantALFA (via Variation Services) for array-genotyped variants; gnomAD for sequencing-derivedDifferent sample compositions

Single rsID Resolution

Goal: Resolve an rsID to full Build 156 RefSNP JSON, including coordinates, alleles, gene context, and merge history.

Approach: Hit Variation Services /v0/refsnp/{id_without_rs}; the response includes primary_snapshot_data (current) and merged_snapshot_data (if this rsID is itself a merge target).

python
import requests

VARSVC = 'https://api.ncbi.nlm.nih.gov/variation/v0'

def refsnp(rsid):
    '''Fetch full Build 156 RefSNP JSON. rsid can be 'rs121913529' or 121913529.'''
    rs_int = str(rsid).lstrip('rs')
    r = requests.get(f'{VARSVC}/refsnp/{rs_int}', timeout=30)
    if r.status_code == 404:
        return None
    r.raise_for_status()
    return r.json()

def summarize_refsnp(payload):
    '''Extract minimal fields. Handles multi-allelic cluster correctly.

    The placement JSON nests assembly metadata; the precise path varies by
    Build / API version. Common variants seen in the wild:
        placement['seq_id_traits_by_assembly'][0]['assembly_name']
        placement['placement_annot']['seq_id_traits_by_assembly'][0]['assembly_name']
    Inspect the actual JSON returned for the current dbSNP Build before
    relying on either path in production.
    '''
    if payload is None or payload.get('is_withdrawn'):
        return None
    primary = payload.get('primary_snapshot_data', {})
    placements = primary.get('placements_with_allele', [])
    def assembly_name(p):
        traits = (p.get('placement_annot') or p).get('seq_id_traits_by_assembly') or []
        return traits[0].get('assembly_name') if traits else ''
    grch38 = next((p for p in placements if 'GRCh38' in (assembly_name(p) or '')), None)
    if grch38 is None:
        return None
    alleles = []
    for allele in grch38.get('alleles', []):
        spdi = allele.get('allele', {}).get('spdi', {})
        alleles.append({
            'ref': spdi.get('deleted_sequence'),
            'alt': spdi.get('inserted_sequence'),
            'seq_id': spdi.get('seq_id'),
            'pos_0based': spdi.get('position')
        })
    return {
        'rsid': payload.get('refsnp_id'),
        'gene': primary.get('allele_annotations', [{}])[0].get('assembly_annotation', [{}])[0].get('genes', [{}])[0].get('locus'),
        'placements_grch38': alleles,
        'is_multiallelic': len(alleles) > 2,
        'merge_history': payload.get('merged_snapshot_data', [])
    }

Multi-Hop Merge Resolution

Goal: Resolve a possibly-deprecated rsID to the current canonical rsID, following the full merge chain.

Approach: Recursively follow merged_snapshot_data until the response has no further merge entries, with cycle detection.

python
def resolve_merge_chain(rsid, max_hops=10):
    '''Follow multi-hop merge chain. Cycle-safe with max_hops cap.'''
    seen = set()
    current = str(rsid).lstrip('rs')
    for _ in range(max_hops):
        if current in seen:
            return {'error': 'merge cycle detected', 'chain': list(seen)}
        seen.add(current)
        payload = refsnp(current)
        if payload is None:
            return {'error': 'not found', 'final_rsid': current, 'chain': list(seen)}
        if payload.get('is_withdrawn'):
            return {'status': 'withdrawn', 'final_rsid': current, 'chain': list(seen)}
        primary = payload.get('primary_snapshot_data')
        if primary is not None:
            return {'status': 'resolved', 'final_rsid': payload.get('refsnp_id'), 'chain': list(seen)}
        merged = payload.get('merged_snapshot_data', [])
        if not merged:
            return {'status': 'orphan', 'final_rsid': current, 'chain': list(seen)}
        current = str(merged[0].get('merged_into', ''))
    return {'error': 'hop limit', 'chain': list(seen)}

SPDI <-> HGVS <-> VCF Conversion

Goal: Move between variant representations using Variation Services as the canonical bridge.

Approach: SPDI endpoints handle build resolution and right-alignment; HGVS contextuals applies the Variant Overprecision Correction Algorithm.

python
def hgvs_to_spdi_canonical(hgvs):
    '''Resolve HGVS to canonical SPDI via the Variant Overprecision Correction Algorithm.'''
    r = requests.get(f'{VARSVC}/hgvs/{hgvs}/contextuals', timeout=30)
    if not r.ok:
        return None
    contextuals = r.json().get('data', {}).get('spdis', [])
    return contextuals[0] if contextuals else None

def spdi_to_rsid(spdi_str):
    '''SPDI 'NC_000017.11:43044294:G:A' -> rsID if a cluster exists.'''
    r = requests.get(f'{VARSVC}/spdi/{spdi_str}/rsids', timeout=30)
    if not r.ok:
        return None
    rsids = r.json().get('data', {}).get('rsids', [])
    return rsids[0] if rsids else None

def vcf_to_canonical_spdi(chrom, pos, ref, alt, assembly='GRCh38'):
    '''VCF (1-based) -> SPDI (0-based, right-aligned).'''
    refseq_map = {('1', 'GRCh38'): 'NC_000001.11', ('17', 'GRCh38'): 'NC_000017.11'}
    refseq = refseq_map.get((str(chrom).lstrip('chr'), assembly))
    if refseq is None:
        return None
    raw_spdi = f'{refseq}:{pos - 1}:{ref}:{alt}'
    r = requests.get(f'{VARSVC}/spdi/{raw_spdi}/canonical_representative', timeout=30)
    return r.json().get('data', {}).get('spdi') if r.ok else None

ALFA Frequencies vs gnomAD

SourceSample basisVariants coveredWhen to use
ALFA~1M dbGaP subjects (array + WGS) across 12 ancestry groups447M+ sites; broader (includes array-only)Common variants, dbGaP-deposited cohorts
gnomAD v4807k WGS+WES individuals across 9 ancestry groupsSequencing-derived (deeper at rare variants)Rare-variant FAF95, ACMG BS1/BA1

ALFA does NOT provide FAF95-style upper-bound CIs; raw AF only. ALFA captures consent-tier metadata enabling consent-respecting lookups for variants gnomAD doesn't carry.

python
def alfa_frequency(rsid, ancestry='Total'):
    '''Pull ALFA per-population AF via Variation Services.'''
    payload = refsnp(rsid)
    if payload is None:
        return None
    freq_records = payload.get('primary_snapshot_data', {}).get('allele_annotations', [{}])[0].get('frequency', [])
    alfa_records = [f for f in freq_records if 'ALFA' in f.get('study_name', '')]
    for record in alfa_records:
        if record.get('common_name') == ancestry:
            return {
                'allele': record.get('observation', {}).get('inserted_sequence'),
                'count': record.get('allele_count'),
                'total': record.get('total_count'),
                'freq': record.get('allele_count') / record.get('total_count') if record.get('total_count') else None
            }
    return None

Per-Operation Failure Modes

1. Treating rsID as a unique variant identifier

  • Trigger: Join two databases by rsID expecting a single variant.
  • Mechanism: rsID is a cluster identifier; multi-allelic clusters have 2-4 alleles at one position.
  • Symptom: Allele mismatches at low rate (~6-8% of sites); silent merger of unrelated variants.
  • Fix: Normalize both sides to SPDI or CA ID before joining.

2. Single-hop RsMergeArch lookup

  • Trigger: Read one row of RsMergeArch.bcp.gz and treat rsCurrent as the final answer.
  • Mechanism: Multi-hop merges (rs3 -> rs2 -> rs1 -> rs0) span multiple rows; each row records one hop only.
  • Symptom: Resolved rsID is itself stale; subsequent queries return outdated annotation.
  • Fix: Follow merge chains recursively via Variation Services merged_snapshot_data (handles multi-hop in one call).

3. Confusing withdrawn vs merged

  • Trigger: Query a withdrawn rsID and find no merge target.
  • Mechanism: Withdrawn rsIDs live in SNPHistory.bcp.gz, not RsMergeArch. They are NOT merged into anything; the cluster was QC-failed or submitter-retracted.
  • Symptom: 404 from naive lookup; pipelines proceed with stale rsID.
  • Fix: Check is_withdrawn field in RefSNP JSON; flag the variant for manual review.

4. E-utilities thin summary

  • Trigger: Use Entrez.esummary(db='snp') and treat output as authoritative.
  • Mechanism: E-utilities db=snp returns a pre-Build-156 summary missing primary_snapshot_data.placements_with_allele, frequency data, and merge history.
  • Symptom: Annotations look incomplete or out-of-date.
  • Fix: Use Variation Services REST /v0/refsnp/{id} for full JSON.

5. SPDI 0-based vs HGVS 1-based mismatch

  • Trigger: Build SPDI string from VCF position without converting to 0-based.
  • Mechanism: SPDI position is 0-based half-open; VCF is 1-based.
  • Symptom: Coordinate off-by-one; SPDI does not resolve to expected rsID.
  • Fix: SPDI position = (VCF position - 1). For indels, also normalize ref/alt.

6. Strand/orientation ambiguity for A/T C/G variants

  • Trigger: Merge variants across builds or platforms relying on rsID alone.
  • Mechanism: rsID is locus-level; opposite-strand alleles get the same rsID with different ref/alt representation.
  • Symptom: Strand-flipped genotypes after merge.
  • Fix: Use SPDI (which encodes strand via deleted/inserted sequence) or MAF-match for ambiguous variants.
Show full SKILL.md (550 more words)Show less

Reconciliation: When Sources Disagree

PatternLikely causeAction
dbSNP rsID returns 404 in current buildWithdrawn (in refsnp-withdrawn.json.bz2)Check withdrawal reason; consider manual curation
rsID resolves to different coords across buildsGenome assembly change (GRCh37 -> GRCh38)Use SPDI with explicit RefSeq accession; lift over via pyliftover
ALFA AF and gnomAD AF disagree by >2xDifferent sample compositions; ALFA includes array-only sites under-represented in gnomADTrust gnomAD for sequencing data; ALFA for array-derived; use the more relevant source
Multiple rsIDs map to one SPDITrue duplicates from independent submissions; rare since Build 152 enforced cluster mergingPick the lowest rsID per RsMergeArch convention
One rsID has different ref allele in dbSNP vs gnomADdbSNP uses NCBI's ref; gnomAD aligns to its build assemblyNormalize to SPDI before joining

Quantitative Thresholds and Conventions

ThresholdConventionSource
Build 156Current as of Sep 2022; JSON-only distributiondbSNP NCBI
Multi-allelic rate~6-8% of dbSNP rsIDs are multi-allelicoperational estimate
Variation Services rate limit10 req/s with API key; 3 req/s withoutNCBI E-utilities policy
SPDI position0-based, half-openNCBI SPDI specification
HGVS position1-based, fully-closedHGVS nomenclature
ALFA samples~1M individuals across 12 ancestries (2024 release)NCBI dbGaP aggregation
Bulk download chunksOne file per chromosome (refsnp-chr{N}.json.bz2)NCBI FTP

Common Errors

SymptomCauseSolution
404 from Variation Services on valid rsIDrsID is withdrawn or never assignedCheck refsnp-withdrawn.json.bz2; consider strand-flipped equivalent
Merge chain resolves but final rsID has different allelesMulti-allelic cluster; pick allele matching the variantFilter placements_with_allele.alleles[*] by allele match
ALFA frequency missing for common variantVariant not in dbGaP-deposited studiesFall back to gnomAD (sequencing-derived)
Entrez.esummary returns old dataLegacy E-utilities, not Build 156 schemaSwitch to Variation Services REST /v0/refsnp/{id}
HGVS-c conversion fails for synonymous variantsSome HGVS-c rely on non-MANE transcripts not in NCBI defaultSpecify transcript explicitly; use VEP --mane_select
SPDI for indel does not round-tripLeft/right alignment mismatchUse /spdi/{spdi}/canonical_representative for normalization
Bulk JSON parse OOMrefsnp-chr1.json.bz2 is ~20GB uncompressedStream parse with bz2.BZ2File + line-by-line JSON; do not load whole file

Anticipated Reviewer Pushback

PushbackStandard response
"Why not just use rsID for the join?"rsID is a cluster identifier; ~6-8% of clusters are multi-allelic, causing silent mismatches. We use SPDI / CA ID.
"This annotation says rs12345 but the literature says rs67890"rsIDs are merged; we resolved through RsMergeArch / merged_snapshot_data to the current canonical rsID.
"dbSNP frequency != gnomAD frequency"ALFA (dbSNP-embedded) and gnomAD use different sample sets and different ascertainment (array vs sequencing); reconciled per use case.
"Why wasn't Entrez used?"Entrez db=snp returns the pre-Build-156 summary missing key fields; we use Variation Services REST for the full JSON.
"Coordinate off-by-one in the SPDI"SPDI is 0-based half-open; VCF is 1-based; intentional conversion applied.

References

  • Phan L et al. 2025. The evolution of dbSNP: 25 years of impact in genomic research. Nucleic Acids Res 53:D925.
  • Sayers EW et al. 2024. Database resources of the National Center for Biotechnology Information. Nucleic Acids Res 52:D33.
  • Holmes JB et al. 2020. SPDI: data model for variants and applications at NCBI. Bioinformatics 36:1902.
  • NCBI Variation Services API: https://api.ncbi.nlm.nih.gov/variation/v0/
  • dbSNP FTP layout: ftp.ncbi.nlm.nih.gov/snp/latest_release/JSON/
  • ALFA release notes: https://www.ncbi.nlm.nih.gov/snp/docs/gsr/alfa/
  • ClinGen Allele Registry: https://reg.clinicalgenome.org/docs/cg-car/
  • clinical-databases/myvariant-queries - Aggregated rsID + annotation queries
  • clinical-databases/clinvar-lookup - ClinVar VariationID vs rsID linkage
  • clinical-databases/gnomad-frequencies - Frequency lookups by canonical SPDI
  • clinical-databases/variant-prioritization - Pipeline using normalized variant IDs
  • database-access/entrez-search - General Entrez query patterns

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in clinical-databases/dbsnp-queries of GPTomics/bioSkills.

  • SKILL.md
  • examples/dbsnp_lookup.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Clinical Databases Dbsnp Queries next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Clinical Databases Dbsnp Queries compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Clinical Databases Dbsnp Queries this skillGPTomics/bioSkills1.2k2 repos~5.2kAutomated safety check: PassMIT
ETE Toolkit for Phylogenetic Treesdavila7/claude-code-templates33k11 repos~4.5kAutomated safety check: NotesMIT
Biopythondavila7/claude-code-templates33k12 repos~3.4kAutomated safety check: PassMIT
BiopythonK-Dense-AI/scientific-agent-skills48k1 repos~4.3kAutomated safety check: NotesMIT
Biopythonlamm-mit/scienceclaw246—~3.9kAutomated safety check: PassApache-2.0
Clinvar Databaseaipoch/medical-research-skills1.9k—~856Automated safety check: PassMIT

Similar skills

  • ETE Toolkit for Phylogenetic Trees

    davila7/claude-code-templates

    Guides your agent through building, editing, comparing and drawing phylogenetic trees with the ETE Python toolkit, including orthology calls and NCBI taxonomy lookups.

    33k GitHub starsUsed in 11 repos~4.5k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    davila7/claude-code-templates

    Primary Python toolkit for molecular biology. An agent skill from davila7/claude-code-templates.

    33k GitHub starsUsed in 12 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • Biopython

    K-Dense-AI/scientific-agent-skills

    Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez).

    48k GitHub starsUsed in 1 repo~4.3k tokens
    Research & ScienceAuto-check: notes
  • Biopython

    lamm-mit/scienceclaw

    Computational molecular biology library (sequence I/O, alignment, phylogenetics).

    246 GitHub stars~3.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Clinvar Database

    aipoch/medical-research-skills

    Utilities for querying the NCBI ClinVar database to retrieve variant records, clinical significance, and phenotype relationships; use when searching variants by gene/condition/significance…

    1.9k GitHub stars~856 tokensUpdated 23 days ago
    Research & ScienceAuto-check passed
  • Gget

    aipoch/medical-research-skills

    Unified CLI/Python interface for querying genomic, proteomic, structure, and expression data across 20+ bioinformatics databases; use when you need fast, scriptable retrieval by gene/protein IDs or…

    1.9k GitHub stars~816 tokensUpdated 23 days ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Clinical Databases Dbsnp Queries

What does Bio Clinical Databases Dbsnp Queries do?

Resolves rsIDs, navigates RsMergeArch/SNPHistory merge chains, and converts between rsID, SPDI, HGVS, and VCF representations using the dbSNP Build 156 JSON architecture. Bio Clinical Databases Dbsnp Queries is an agent skill from GPTomics/bioSkills. Resolves rsIDs, navigates RsMergeArch/SNPHistory merge chains, and converts between rsID, SPDI, HGVS, and VCF representations using the dbSNP Build 156 JSON architecture.

When should I use Bio Clinical Databases Dbsnp Queries?

Bio Clinical Databases Dbsnp Queries fits situations like: normalizing variant identifiers; joining variant databases by cluster ID; tracking deprecated rsIDs through historical merges.

How do I install Bio Clinical Databases Dbsnp Queries in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-dbsnp-queries -a claude-code`. Or copy the skill folder (clinical-databases/dbsnp-queries in GPTomics/bioSkills) into .claude/skills/bio-clinical-databases-dbsnp-queries in your project. Claude Code loads it when a task matches its description.

How do I install Bio Clinical Databases Dbsnp Queries in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-dbsnp-queries -a codex`. Or copy the skill folder (clinical-databases/dbsnp-queries in GPTomics/bioSkills) into .agents/skills/bio-clinical-databases-dbsnp-queries in your project. Codex loads it when a task matches its description.

Can I use Bio Clinical Databases Dbsnp Queries in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-dbsnp-queries -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clinical-databases-dbsnp-queries, .gemini/skills/bio-clinical-databases-dbsnp-queries, .github/skills/bio-clinical-databases-dbsnp-queries and .opencode/skills/bio-clinical-databases-dbsnp-queries in your project.

What does Bio Clinical Databases Dbsnp Queries need to run?

Going by SKILL.md and its folder, Bio Clinical Databases Dbsnp Queries needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Clinical Databases Dbsnp Queries access the network?

SKILL.md names 3 domains. In commands or code: api.ncbi.nlm.nih.gov, ncbi.nlm.nih.gov and reg.clinicalgenome.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Bio Clinical Databases Dbsnp Queries safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Clinical Databases Dbsnp Queries use?

Bio Clinical Databases Dbsnp Queries is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Clinical Databases Dbsnp Queries use?

About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Clinical Databases Dbsnp Queries?

Skills that share tags, products or a category with Bio Clinical Databases Dbsnp Queries: ETE Toolkit for Phylogenetic Trees (davila7/claude-code-templates, 33k stars), Biopython (davila7/claude-code-templates, 33k stars), Biopython (K-Dense-AI/scientific-agent-skills, 48k stars) and Biopython (lamm-mit/scienceclaw, 246 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Clinical Databases Dbsnp Queries?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.