Pubchem Database Skill
aipoch/medical-research-skills
Programmatic access to the PubChem database (via PUG-REST API and PubChemPy) for searching chemical compounds, retrieving physicochemical properties, performing structure similarity/substructure…
Queries ClinVar for variant pathogenicity classifications, ClinGen VCEP curations, and somatic-vs-germline interpretations via REST API, weekly VCF, or bulk XML.
$ npx skills add GPTomics/bioSkills --skill bio-clinical-databases-clinvar-lookup -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-clinical-databases-clinvar-lookup --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clinical-databases/clinvar-lookup .claude/skills/bio-clinical-databases-clinvar-lookup && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-clinical-databases-clinvar-lookup" agent skill from https://github.com/GPTomics/bioSkills/tree/main/clinical-databases/clinvar-lookup into .claude/skills/bio-clinical-databases-clinvar-lookup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-clinical-databases-clinvar-lookup", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/clinical-databases/clinvar-lookupType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-clinical-databases-clinvar-lookup -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-clinical-databases-clinvar-lookup --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/clinical-databases/clinvar-lookup .agents/skills/bio-clinical-databases-clinvar-lookup && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-clinical-databases-clinvar-lookup" agent skill from https://github.com/GPTomics/bioSkills/tree/main/clinical-databases/clinvar-lookup into .agents/skills/bio-clinical-databases-clinvar-lookup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-clinical-databases-clinvar-lookup", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-clinical-databases-clinvar-lookup -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-clinical-databases-clinvar-lookup --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/clinical-databases/clinvar-lookup .cursor/skills/bio-clinical-databases-clinvar-lookup && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-clinical-databases-clinvar-lookup" agent skill from https://github.com/GPTomics/bioSkills/tree/main/clinical-databases/clinvar-lookup into .cursor/skills/bio-clinical-databases-clinvar-lookup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-clinical-databases-clinvar-lookup", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path clinical-databases/clinvar-lookup--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-clinical-databases-clinvar-lookup -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-clinical-databases-clinvar-lookup --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/clinical-databases/clinvar-lookup .gemini/skills/bio-clinical-databases-clinvar-lookup && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-clinical-databases-clinvar-lookup" agent skill from https://github.com/GPTomics/bioSkills/tree/main/clinical-databases/clinvar-lookup into .gemini/skills/bio-clinical-databases-clinvar-lookup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-clinical-databases-clinvar-lookup", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-clinical-databases-clinvar-lookupInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-clinical-databases-clinvar-lookup -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/clinical-databases/clinvar-lookup .github/skills/bio-clinical-databases-clinvar-lookup && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-clinical-databases-clinvar-lookup" agent skill from https://github.com/GPTomics/bioSkills/tree/main/clinical-databases/clinvar-lookup into .github/skills/bio-clinical-databases-clinvar-lookup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-clinical-databases-clinvar-lookup", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-clinical-databases-clinvar-lookup -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-clinical-databases-clinvar-lookup --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/clinical-databases/clinvar-lookup .opencode/skills/bio-clinical-databases-clinvar-lookup && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-clinical-databases-clinvar-lookup" agent skill from https://github.com/GPTomics/bioSkills/tree/main/clinical-databases/clinvar-lookup into .opencode/skills/bio-clinical-databases-clinvar-lookup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-clinical-databases-clinvar-lookup", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-clinical-databases-clinvar-lookupQueries ClinVar for variant pathogenicity classifications, ClinGen VCEP curations, and somatic-vs-germline interpretations via REST API, weekly VCF, or bulk XML.
Bio Clinical Databases Clinvar Lookup is an agent skill from GPTomics/bioSkills. Queries ClinVar for variant pathogenicity classifications, ClinGen VCEP curations, and somatic-vs-germline interpretations via REST API, weekly VCF, or bulk XML. Use when determining clinical significance, triangulating conflicting interpretations, or aggregating evidence against the ACMG/AMP framework with ClinGen SVI specifications.
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/clinvar_query.py` and `usage-guide.md`).
It sits in Backend & APIs, covering REST APIs. It works with Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
wgetpipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
reg.clinicalgenome.orgcspec.genome.networkftp.ncbi.nlm.nih.goveutils.ncbi.nlm.nih.govFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Clinical Databases Clinvar Lookup loads about 5.7k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 2,288 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,288 words, ~5,651 tokens.
.claude/skills/bio-clinical-databases-clinvar-lookup/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: requests 2.31+, cyvcf2 0.30+, pandas 2.2+, bcftools 1.19+, Entrez Direct 21.0+, lxml 5.0+ (for v2 XML schema).
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signatures<tool> --version then <tool> --help to confirm flagsIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying. ClinVar XML schema v2 (rolled out in 2024) replaces <ClinVarSet> with <VariationArchive> as the top-level anchor; XSLT or parsers targeting the legacy element silently emit zero records.
'Look up the clinical significance of this variant' -> Retrieve ClinVar VCV-level aggregate, SCV-level submissions, ClinGen Variant Curation Expert Panel (VCEP) overrides, and conflict-resolution status.
requests.get() against the E-utilities clinvar databasecyvcf2.VCF('clinvar.vcf.gz') for batch queries against the weekly snapshotbcftools annotate -a clinvar.vcf.gz -c INFO/CLNSIG,INFO/CLNREVSTAT,INFO/CLNDNhttps://reg.clinicalgenome.org/| Level | Format | What it aggregates | When to use | Fails when |
|---|---|---|---|---|
| SCV | SCVxxxxxxxxx.N | One submitter, one variant, one condition (atomic submission unit) | Auditing who said what; conflict triangulation | Aggregated reporting (use VCV); cross-condition analysis |
| RCV | RCVxxxxxxxxx.N | All SCVs for a single (variant, condition) pair | Condition-stratified analysis; legacy aggregation | Variant-level reporting across all conditions (use VCV) |
| VCV | VCVxxxxxxxxx.N | All RCVs for one variant across all conditions | Canonical anchor since 2017; default API entrypoint | Condition-specific clinical action (use RCV); CLNSIG collapses multi-condition |
Operational footgun: the clinvar.vcf.gz CLNSIG field is the variant-level (VCV) aggregate. A variant Pathogenic for disease A but VUS for disease B collapses to "Pathogenic/Conflicting". For condition-stratified analysis, parse RCV-level XML, never CLNSIG alone.
2024 XML schema overhaul: ClinVar v2 XML separates GermlineClassification, SomaticClinicalImpact, and OncogenicityClassification under one <VariationArchive> anchor. The legacy <ClinicalSignificance> element is gone. Pipelines built before September 2024 against <ClinVarSet> silently emit zero records on new XML. The dual-release period ended December 2024.
| Stars | Review status | What it means operationally |
|---|---|---|
| 4 | Practice guideline | ACMG/CAP CFTR-level (vanishingly rare) |
| 3 | Expert panel reviewed (ClinGen VCEP) | FDA-recognized tier; overrides lower-star records for clinical action |
| 2 | Multiple submitters, criteria provided, no conflicts | Reliable aggregate |
| 1 | Single submitter OR conflicting interpretations (often mis-reported as 2-star) | Use with scrutiny |
| 0 | No assertion criteria provided | Literature-only or legacy submissions |
ClinVar does NOT retract or hide lower-star records when a VCEP publishes; a variant can simultaneously display "Pathogenic (3-star VCEP)" and "Conflicting interpretations (1-star)". Tools handle this differently (VarSeq, Franklin, GenoOx each pick a winner via different rules); this is a major source of inter-tool disagreement.
As of 2025, ~80-90 VCEPs are approved or in progress across RASopathies, hereditary cancer (ENIGMA BRCA1/2, InSiGHT MMR), cardiomyopathy (sarcomere genes), hearing loss, RPE65/IRD, inborn errors of metabolism, and FH. The current count is moving; the authoritative directory is the Criteria Specification Registry at https://cspec.genome.network/cspec/ui/svi/all.
Each VCEP publishes a gene-disease-specific CSpec that re-weights ACMG/AMP criteria. The Hearing Loss VCEP downgrades PM2 to supporting by default and upgrades PS3 thresholds for OTOF. Treating "ACMG/AMP" as a single rubric across all genes is the most common error in non-specialist tooling.
The Richards 2015 28-criterion framework is the foundation, but every modern automated classifier (InterVar, GeneBe, Franklin, VarSome) implements the Tavtigian 2018/2020 Bayesian point system, not the original combining rules. Strengths map to points: Supporting=1, Moderate=2, Strong=4, Very Strong=8 (benign codes negative). Final categories: P >=10, LP 6-9, VUS 0-5, LB -1 to -6, B <=-7.
For variant interpretation framework details, calibrated in-silico thresholds, and PVS1 decision-tree logic, defer to clinical-databases/acmg-classification. This skill focuses on querying ClinVar; it intentionally does not re-implement classification.
Harrison 2017 Genet Med 19:1096 (PMID 28301460) showed 87% of inter-lab conflicts were resolvable by reassessment plus data sharing. As of 2024, only 3.8% of conflicting BRCA1 missense VUS reached consensus despite years of effort; conflict resolution is slow even in best-curated genes.
Submission staleness is non-trivial: ClinVar does not push reclassifications to submitters; a 2017 SCV can persist on an active label in 2026 if the lab has not re-submitted. Genome Alert! (Yauy 2022 Genet Med) was built specifically to detect classification drift between weekly releases. The median delta is ~1,247 classification changes per month with potential clinical impact.
| Scenario | Recommended path | Why |
|---|---|---|
| Single variant, known gene/condition | E-utilities esummary against clinvar DB | Lowest latency, returns VCV-level summary |
| Batch (10-1000 variants) by HGVS or rsID | myvariant.info with fields=clinvar | Aggregated, includes ClinVar review status |
| Batch (>1000) or coordinate-based | Local clinvar.vcf.gz with bcftools annotate or cyvcf2 | No rate limits; weekly snapshot |
| Condition-stratified (variant in disease A vs B) | Bulk XML VariationArchive parsing | RCV is the only level that preserves per-condition classification |
| Cross-database join with gnomAD / dbSNP / COSMIC | ClinGen Allele Registry CA ID | Build-agnostic, transcript-agnostic canonical identifier |
| Reproducible analysis with citable date | First-Thursday-of-month archive on FTP | Only monthly snapshots are archived; weekly releases disappear |
Goal: Retrieve VCV-level ClinVar summary for a single variant by ID, gene, or HGVS.
Approach: Hit esummary.fcgi or esearch.fcgi against db=clinvar, parse JSON, then optionally hydrate to full record with efetch.
import requests
EUTILS = 'https://eutils.ncbi.nlm.nih.gov/entrez/eutils'
def clinvar_summary(variation_id):
'''Retrieve VCV-level summary by ClinVar VariationID (do not confuse with CA ID).
The germline / somatic / oncogenicity classification nesting shown below
follows the ClinVar 2024 eSummary v2 schema described in the data-access
documentation. Field names have changed between API versions -- inspect
the actual JSON returned by eSummary for the live ClinVar version before
pinning these key paths in production.
'''
r = requests.get(f'{EUTILS}/esummary.fcgi',
params={'db': 'clinvar', 'id': variation_id, 'retmode': 'json'},
timeout=30)
r.raise_for_status()
record = r.json()['result'][str(variation_id)]
return {
'vcv': record.get('accession'),
'name': record.get('title'),
'germline_class': record.get('germline_classification', {}).get('description'),
'germline_review_status': record.get('germline_classification', {}).get('review_status'),
'somatic_clinical': record.get('clinical_impact_classification', {}).get('description'),
'oncogenicity': record.get('oncogenicity_classification', {}).get('description'),
'last_evaluated': record.get('germline_classification', {}).get('last_evaluated')
}
def clinvar_search_gene(gene, pathogenic_only=False, retmax=500):
term = f'{gene}[gene]'
if pathogenic_only:
term += ' AND (clinsig_pathogenic[Properties] OR clinsig_likely_pathogenic[Properties])'
r = requests.get(f'{EUTILS}/esearch.fcgi',
params={'db': 'clinvar', 'term': term, 'retmax': retmax, 'retmode': 'json'},
timeout=30)
return r.json()['esearchresult']['idlist']Goal: Annotate or look up thousands of variants without rate limits.
Approach: Download the weekly clinvar.vcf.gz (note: only first-Thursday-of-month is archived; for longitudinal stability pin to monthly archives), query by genomic coordinates with cyvcf2 or annotate VCFs with bcftools.
mkdir -p clinvar/$(date +%Y%m); cd clinvar/$(date +%Y%m)
wget https://ftp.ncbi.nlm.nih.gov/pub/clinvar/vcf_GRCh38/clinvar.vcf.gz
wget https://ftp.ncbi.nlm.nih.gov/pub/clinvar/vcf_GRCh38/clinvar.vcf.gz.tbifrom cyvcf2 import VCF
clinvar = VCF('clinvar.vcf.gz')
def lookup(chrom, pos, ref, alt):
'''Look up by GRCh38 coords. Returns variant-level (VCV) aggregate; not RCV.'''
for v in clinvar(f'{chrom}:{pos}-{pos}'):
if v.REF == ref and alt in v.ALT:
info = v.INFO
return {
'vcv_id': info.get('ALLELEID'),
'clnsig': info.get('CLNSIG'),
'clnsig_conf': info.get('CLNSIGCONF'),
'clnrevstat': info.get('CLNREVSTAT'),
'clndn': info.get('CLNDN'),
'clnvc': info.get('CLNVC'),
'clnhgvs': info.get('CLNHGVS'),
'clndisdb': info.get('CLNDISDB'),
'oncdn': info.get('ONCDN'),
'scidn': info.get('SCIDN')
}
return Nonebcftools annotate \
-a clinvar.vcf.gz \
-c INFO/CLNSIG,INFO/CLNREVSTAT,INFO/CLNDN,INFO/CLNVC,INFO/CLNHGVS,INFO/CLNSIGCONF \
input.vcf.gz -O z -o annotated.vcf.gz
bcftools index -t annotated.vcf.gzClinGen Allele Registry (https://reg.clinicalgenome.org/) computes a build-agnostic, transcript-agnostic CA ID (format CA######) for any allele projectable onto NCBI references (GRCh37, GRCh38, T2T-CHM13, any RefSeq transcript). The Registry covers ~700M+ alleles, vastly more than ClinVar. CA ID and ClinVar VariationID are one-to-one when a variant exists in ClinVar.
def car_id(hgvs_g):
'''Resolve HGVS-g to canonical ClinGen Allele Registry CA ID.'''
r = requests.put(f'https://reg.clinicalgenome.org/allele',
headers={'Content-Type': 'text/plain'},
data=hgvs_g, timeout=30)
return r.json().get('@id', '').split('/')[-1] if r.ok else NoneUse CA ID for any join touching non-ClinVar resources (gnomAD, dbSNP, COSMIC, MAVEdb). VariationID was renumbered during the 2017 ClinVar schema redesign; treating it as a stable cross-build identifier is unsafe.
1. Treating CLNSIG as gospel for condition-specific work
CLNSIG=Pathogenic from clinvar.vcf.gz and report variant as pathogenic for the patient's specific phenotype.<RCVAccession> per condition); cross-check CLNDN and report per-condition classifications.2. Parsing legacy XML against 2024 schema
<ClinVarSet> or <ClinicalSignificance>.<VariationArchive> + germline/somatic/oncogenicity tripartite classifications.<VariationArchive> and read GermlineClassification, SomaticClinicalImpact, OncogenicityClassification separately.3. Counting variant_summary.txt rows naively
wc -l variant_summary.txt to estimate variant count.awk -F'\t' '$17=="GRCh38"' variant_summary.txt | wc -l.4. Trusting VariationID as a stable cross-build identifier
5. Ignoring star-rating override hierarchy
review_status rank (4>3>2>1>0); use the highest-star record. For ties, sort by date.6. Aggregating "Conflicting" without inspecting the conflict
CLNSIG=Conflicting interpretations as VUS.CLNSIGCONF to see exact conflict; weight by submitter star.7. Missing somatic interpretations
CLNSIG.ONCDN, SCIDN, CLNSIGSOMATIC) since 2024.ONCDN (oncogenicity disease name), SCIDN (somatic clinical impact disease name), and the somatic-specific significance fields.| Pattern | Likely cause | Action |
|---|---|---|
| ClinVar P vs gnomAD AF > 1% | Variant is true founder allele in unstratified gnomAD subset, OR ClinVar P is a stale low-star assertion | Check grpmax_faf95 excluding bottleneck groups; check ClinVar star rating |
| ClinVar P vs AlphaMissense < 0.1 | Variant in NMD-escape region, alternative isoform, or ClinVar P is mis-curated | Check Pejaver 2022 calibration in acmg-classification skill; cross-check VCEP |
| VCEP 3-star P vs commercial-lab 1-star B | VCEP supersedes for clinical action | Use VCEP; flag submitter for resubmission |
| ClinVar VCV-level P vs RCV-level VUS for actual condition | VCV averages across conditions | Always report at RCV level for clinical action |
| ClinVar P vs LOVD/HGMD discordant | LOVD/HGMD use different classification systems; HGMD "DM" != ACMG P | Triangulate against published evidence; do not auto-translate labels |
| ClinVar P missing for a known disease variant | Submission lag (~6-12 months typical for new findings) | Check published literature; flag for ClinVar submission |
| Threshold | Convention | Source |
|---|---|---|
| Star >= 2 | Acceptable confidence for clinical action without further review | ClinGen SVI operational guidance |
| Star = 3 | VCEP-curated; supersedes lower-star records | ClinGen FDA Recognition 2018 |
CLNSIG includes 'Pathogenic' OR 'Likely_pathogenic' | Treat as actionable for ACMG | ClinVar field schema |
CLNSIGCONF present | Multiple SCVs disagree; do NOT auto-action | ClinVar field schema |
| Monthly archive | Use first-Thursday-of-month FTP snapshot for reproducible analyses | NCBI FTP retention policy |
| Submission staleness | Re-check classification annually for active diagnostic variants | Yauy 2022 Genet Med (Genome Alert!) |
| AF > 5% in gnomAD | BA1 standalone benign per ClinGen SVI default (VCEP overrides exist) | Richards 2015; SVI specs |
| 1247 changes/month | Median variants with classification change per release | Yauy 2022 |
The 2024 schema separates three orthogonal classifications, each with its own ReviewStatus and DateLastEvaluated:
A single VCV can carry all three with distinct evaluations; the legacy "Pathogenic" label is now ambiguous if not qualified by classification type.
| Symptom | Cause | Solution |
|---|---|---|
Empty result from efetch db=clinvar | rsID passed where VariationID expected | Use esearch first to resolve rsID to VariationID |
CLNSIG is _None or comma-separated mess | Variant has multi-condition RCVs; VCF collapses them | Parse RCV XML for per-condition values |
| Variant present in ClinVar XML but absent from VCF | Variant lacks GRCh38 coordinates (legacy GRCh37-only submission) | Check <VariationArchive><SequenceLocation> per assembly |
| 2024-format XML parser silently emits zero records | XML schema v2 incompatibility | Re-target to <VariationArchive> |
| Conflicting interpretations with same star rating across two submitters | True scientific disagreement; sometimes resolved by VCEP later | Apply Tavtigian point system to manually reconcile; flag for VCEP review |
| Variant has CA ID but no VariationID | Variant in Allele Registry but never submitted to ClinVar | Use AlleleRegistry as canonical; submit to ClinVar if novel pathogenic |
CLNSIG says Pathogenic but no associated condition CLNDN | Orphan classification (older submissions) | Treat as low confidence; cross-check publication |
| Variants pulled by gene return only some isoforms | RefSeq transcript priority differences | Use MANE Select transcript explicitly; cross-check with VEP --mane_select |
| Pushback | Standard response |
|---|---|
| "Why is this pathogenic variant 1-star?" | We report star rating per submission; clinical action requires star >=2 OR VCEP curation per ClinGen SVI 2018. |
| "ClinVar says P but gnomAD AF = 2%" | Reconciled via Whiffin FAF95 max-credible-AF framework; bottleneck-group rule applied. |
| "This VCV count differs from ClinVar.gov" | We pulled from the monthly archive (first-Thursday-of-month) for reproducibility; the live web is post-most-recent-weekly. |
| "Why wasn't the somatic variant flagged?" | Pre-2024 XML schema had no separate somatic field; we now read ONCDN/SCIDN/SomaticClinicalImpact per v2 schema. |
| "VarSome says LP but this says VUS" | Tool-specific aggregation rule differences; VarSome auto-applies PP3+PM2 by default per Tavtigian point system; we apply VCEP-specific PP3 calibration per CSpec. |
| "rsID match returned wrong variant" | rsID is a cluster identifier; multi-allelic rsIDs require allele-level resolution; we use SPDI or CA ID. |
| "Why retest a 2022-curated variant?" | Classifications drift as evidence accrues; ClinGen recommends annual re-review for active diagnostic variants. |
https://reg.clinicalgenome.org/docs/cg-car/https://cspec.genome.network/cspec/ui/svi/all© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in clinical-databases/clinvar-lookup of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Clinical Databases Clinvar Lookup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Clinical Databases Clinvar Lookup this skillGPTomics/bioSkills | 1.2k | 2 repos | ~5.7k | Automated safety check: Pass | MIT | |
| Pubchem Database Skillaipoch/medical-research-skills | 1.9k | — | ~954 | Automated safety check: Pass | MIT | |
| Zhihu Searchitwanger/toBeBetterJavaer | 18k | — | ~1.5k | Automated safety check: Pass | None | |
| Fastcrudbenavlabs/fastcrud | 1.6k | — | ~5k | Automated safety check: Pass | MIT | |
| Cloudflare Email Servicehodgef/apiker | 127 | 2 repos | ~2k | Automated safety check: Pass | MIT | |
| Build X402 Clientcoinbase/cdp-sdk | 203 | — | ~3k | Automated safety check: Pass | MIT |
aipoch/medical-research-skills
Programmatic access to the PubChem database (via PUG-REST API and PubChemPy) for searching chemical compounds, retrieving physicochemical properties, performing structure similarity/substructure…
itwanger/toBeBetterJavaer
Search Zhihu for content using the searchv3 API. An agent skill from itwanger/toBeBetterJavaer.
benavlabs/fastcrud
A skill your agent uses when building or modifying CRUD endpoints with FastCRUD (the fastcrud PyPI package) in a FastAPI project — covers FastCRUD, crudrouter, EndpointCreator, FilterConfig…
hodgef/apiker
Send and receive transactional emails with Cloudflare Email Service (Email Sending + Email Routing).
coinbase/cdp-sdk
Write code that pays for an HTTP API returning 402 Payment Required, using the x402 protocol and a CDP-managed wallet.
kappa90/dinobase
Writes a new Dinobase YAML connector for a REST API that has no verified dlt source, covering auth, pagination, read and write endpoints and incremental loading.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Works with
Categories
Queries ClinVar for variant pathogenicity classifications, ClinGen VCEP curations, and somatic-vs-germline interpretations via REST API, weekly VCF, or bulk XML. Bio Clinical Databases Clinvar Lookup is an agent skill from GPTomics/bioSkills. Queries ClinVar for variant pathogenicity classifications, ClinGen VCEP curations, and somatic-vs-germline interpretations via REST API, weekly VCF, or bulk XML.
Bio Clinical Databases Clinvar Lookup fits situations like: determining clinical significance; triangulating conflicting interpretations; aggregating evidence against the ACMG/AMP framework with ClinGen SVI specifications.
Run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-clinvar-lookup -a claude-code`. Or copy the skill folder (clinical-databases/clinvar-lookup in GPTomics/bioSkills) into .claude/skills/bio-clinical-databases-clinvar-lookup in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-clinvar-lookup -a codex`. Or copy the skill folder (clinical-databases/clinvar-lookup in GPTomics/bioSkills) into .agents/skills/bio-clinical-databases-clinvar-lookup in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clinical-databases-clinvar-lookup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clinical-databases-clinvar-lookup, .gemini/skills/bio-clinical-databases-clinvar-lookup, .github/skills/bio-clinical-databases-clinvar-lookup and .opencode/skills/bio-clinical-databases-clinvar-lookup in your project.
Going by SKILL.md and its folder, Bio Clinical Databases Clinvar Lookup needs Python for the scripts in its folder and the command-line tools its instructions call (wget and pip). Our summary lists: Python 3.
SKILL.md names 4 domains. In commands or code: reg.clinicalgenome.org, cspec.genome.network, ftp.ncbi.nlm.nih.gov and eutils.ncbi.nlm.nih.gov; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Clinical Databases Clinvar Lookup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Clinical Databases Clinvar Lookup: Pubchem Database Skill (aipoch/medical-research-skills, 1.9k stars), Zhihu Search (itwanger/toBeBetterJavaer, 18k stars), Fastcrud (benavlabs/fastcrud, 1.6k stars) and Cloudflare Email Service (hodgef/apiker, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.