Queries the 1000 Genomes Project dataset (3,202 whole-genome-sequenced individuals, GRCh38) at the level of individual participants.

MITAuto-check: notesResearch & Science

Install Onekgpd

skills CLI
$ npx skills add K-Dense-AI/scientific-agent-skills --skill onekgpd -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install K-Dense-AI/scientific-agent-skills onekgpd --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/onekgpd .claude/skills/onekgpd && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
onekgpd
GitHub stars
48k
Used in
1 other repo
Token cost
~5.6k tokens
SKILL.md length
2,162 words
Files
6 (incl. scripts, references, assets)
Skills in repo
153
Repo updated
First seen
Licence
MIT

At a glance

Queries the 1000 Genomes Project dataset (3,202 whole-genome-sequenced individuals, GRCh38) at the level of individual participants.

  • Works in 4 steps: uv: This skill's script is run with uv… → Data use terms: The 1000 Genomes Project… → Access constraints: There is no API key,… → …
  • A question is about individuals
  • SKILL.md covers Scope, When to Use, Prerequisites and Core Rules, plus 10 more sections
  • Runs Python scripts from its folder; calls uv; reaches ncbi.nlm.nih.gov

What it does

Onekgpd is an agent skill from K-Dense-AI/scientific-agent-skills. Queries the 1000 Genomes Project dataset (3,202 whole-genome-sequenced individuals, GRCh38) at the level of individual participants. Use when a question is about individuals or variants in the 1000 Genomes Project cohort: which individuals carry variants matching specific criteria in a gene or region, which individuals are homozygous-reference at a position, which variants exist in the dataset or carried by specified individuals in a gene or region, the relatedness between two specified individuals. Variants are…

Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts, reference files and assets (for example `assets/kgpe.json`, `references/annotation_vocabularies.md` and `references/onekgpd_commands.md`). Compatibility notes: Requires Python =3.11. Variant and sample queries require outbound network access to the public 1000 Genomes query endpoint over TLS; the sample/population…

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: Turn any AI agent into an AI Scientist. The 1 Agent Skills library for science, used by 250,000+ scientists worldwide. 177 ready-to-use validated skills plus 100+ scientific… The licence is MIT.

When your agent uses it

  • A question is about individuals
  • Variants in the 1000 Genomes Project cohort: which individuals carry variants matching specific criteria in a gene
  • Which individuals are homozygous-reference at a position
  • Which variants exist in the dataset

Example prompts

  • “Use the onekgpd skill to query the 1000 Genomes Project dataset (3,202 whole-genome-sequenced individuals, GRCh38) at the level of individual…”
  • “/onekgpd”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python >=3.11. Variant and sample queries require outbound network access to the public 1000 Genomes query endpoint over TLS; the sample/population metadata commands run fully offline over a data file bundled in the skill. No credentials, API keys, or environment variables are used.
  • Pre-approved tools (allowed-tools): Write, Bash

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. uv: This skill's script is run with uv run, which reads the script's
  2. Data use terms: The 1000 Genomes Project data is open; users should be
  3. Access constraints: There is no API key, no .env file, and no
  4. Timeouts: network commands use --timeout 30 seconds per RPC by default.

What it can do on your machine

Read from SKILL.md and the folder at commit 92ace75. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Write
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ncbi.nlm.nih.gov

    Also links to:

    • internationalgenome.org
    • docs.astral.sh
    • dnaerys.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python >=3.11. Variant and sample queries require outbound network access to the public 1000 Genomes query endpoint over TLS; the sample/population metadata commands run fully offline over a data file bundled in the skill. No credentials, API keys, or environment variables are used.

    From compatibility in the SKILL.md frontmatter.

Context cost

Onekgpd loads about 5.6k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 164 tokens; SKILL.md has 2,162 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~164
When it runs · the whole SKILL.md, loaded when a task matches
~5.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:91
    constraints**: There is no API key, no `.env` file, and no
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Write, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from K-Dense-AI/scientific-agent-skills at commit 92ace75, republished under its MIT licence (© K-Dense-AI). 2,162 words, ~5,623 tokens.

Download SKILL.mdSave it as .claude/skills/onekgpd/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
onekgpd
description
Queries the 1000 Genomes Project dataset (3,202 whole-genome-sequenced individuals, GRCh38) at the level of individual participants. Use when a question is about individuals or variants in the 1000 Genomes Project cohort: which individuals carry variants matching specific criteria in a gene or region, which individuals are homozygous-reference at a position, which variants exist in the dataset or carried by specified individuals in a gene or region, the relatedness between two specified individuals. Variants are returned with 1000 Genomes allele frequencies (AF), gnomAD v4.1 exome and genome AF, AlphaMissense score, and HGVSp annotations.
allowed-tools
Write, Bash
compatibility
Requires Python >=3.11. Variant and sample queries require outbound network access to the public 1000 Genomes query endpoint over TLS; the sample/population metadata commands run fully offline over a data file bundled in the skill. No credentials, API keys, or environment variables are used.
license
MIT
metadata.version
1.4
metadata.last-reviewed
2026-09-30
metadata.upstream-version
dnaerys 0.2.1; proto R1.20.0
metadata.skill-author
Dnaerys

OneKGPd: Individual-Level Queries over the 1000 Genomes Project

Scope

This skill queries the 1000 Genomes Project dataset — the extended high-coverage cohort of 3,202 whole-genome-sequenced individuals, on the GRCh38 assembly. All results are drawn from this cohort, and sample names returned by the skill (for example HG00096 or NA21130) identify its participants.

Queries resolve against the cohort's per-individual genotype data. This supports two complementary classes of question: selecting variants carried within a region (across the whole cohort or within a specified set of individuals), and selecting the individuals who carry variants matching given criteria. Variant selection can be filtered by allele frequency, predicted consequence, clinical significance, AlphaMissense classification, and the other annotation axes listed below. Relatedness between two named individuals is also available.

The genotype state in which a variant is carried — heterozygous or homozygous — is a criterion that queries may specify; results are returned as variants or as sample names, not as raw genotypes.

The public service is TLS gRPC at db.dnaerys.org:443, accessed with dnaerys 0.2.1 (Python 3.11+); it is not a REST base URL. The maintained service snapshot advertises VEP 115 / GENCODE 49, ClinVar 202502, and gnomAD 4.1. These are the service's annotation releases, not the latest release of each upstream resource. Record them when interpreting results.

When to Use

Use this skill when you need to:

  • Find variants carried in a region or set of regions matching some criteria across the whole cohort (select-variants).
  • Find variants carried in a region or set of regions matching some criteria in specific set of individuals (select-variants-in-samples).
  • Find which 1000 Genomes individuals carry variants matching some criteria in a region or set of regions (select-samples).
  • Count how many individuals carry specific variants (count-samples).
  • Restrict any variant query to heterozygous-only or homozygous-only carriage, or query both together (default).
  • Identify which individuals are homozygous reference at a single position (select-samples-hom-ref).
  • Determine the relatedness between two named 1000 Genomes individuals — both the degree (twin / 1st / 2nd / 3rd / unrelated) and the KING kinship coefficient (kinship).
  • Get dataset totals — sample count, sex split, variant count, assembly (dataset-info).
  • Variant selection can be specified by KGP allele frequency, gnomAD 4.1 exome and gnomAD 4.1 genome allele frequency, AlphaMissense Score and AlphaMissense Class, ClinVar significance (202502), and VEP annotations (impact, biotype, feature type, variant class, consequences).

Do NOT use this skill for:

  • Resolving a gene symbol, rsID, or transcript to coordinates, or fetching reference sequence. Resolve coordinates first (see Coordinate Provenance below), then query this skill with the resolved GRCh38 region.
  • Any cohort other than the 1000 Genomes Project — this skill serves only that dataset.

Prerequisites

  1. uv: This skill's script is run with uv run, which reads the script's inline dependency metadata and provisions an ephemeral environment. Ensure uv is installed and on PATH (https://docs.astral.sh/uv/).
  2. Data use terms: The 1000 Genomes Project data is open; users should be aware of the 1000 Genomes Project / IGSR data-use terms (https://www.internationalgenome.org/data).
  3. Access constraints: There is no API key, no .env file, and no rate-limit token to configure.
  4. Timeouts: network commands use --timeout 30 seconds per RPC by default. A positive finite override is allowed. Pagination makes several RPCs and retryable failures retry the whole fetch up to three times, so this is not a deadline for the whole command.

Core Rules

  • Use the Wrappers: ALWAYS execute the provided helper scripts rather than constructing your own client calls or network requests. Use scripts/onekgpd_api.py for variant/sample/kinship queries (it handles the connection, streaming, pagination, and JSON serialization), and scripts/onekgpd_meta.py for sample/population metadata (offline, see Sample & population metadata).
  • Coordinates MUST be resolved against an authoritative source first — see Coordinate Provenance. This is mandatory, not advisory.
  • Count before you select: every variant and sample selection has a paired counting command. Call the count command FIRST to size the result set, then select only if the count is manageable.
  • Zygosity defaults to both: selection and counting commands include both heterozygous and homozygous carriage by default. Narrow with --het-only or --hom-only when the question is specifically about one state. (You do not need to pass anything to get both.)
  • Output: scripts write full JSON to a file (--output, default under /tmp/) and print a concise summary to stdout. Do not read large JSON files into context — use jq or a small disposable uv run python snippet to extract fields. --page-size retrieves every page but accumulates all variants in RAM; it is not a bounded-memory export. Size the query first.
  • Check completeness: result_incomplete=true means results cannot support a definitive zero/absence claim. Re-run after service recovery. For capped variant selections, truncated=true means the limit was reached and more records may exist, even if the cluster result itself was complete.

Coordinate Provenance (MANDATORY FIRST STEP)

Before any region-based query, resolve the gene or feature to GRCh38 coordinates against an authoritative source (for example Ensembl or NCBI), and query with those resolved coordinates. Inputs are 1-based, inclusive: a BED interval [start0, end0) becomes start=start0+1, end=end0. Record the source accession, annotation release, and retrieval date; gene boundaries can differ by annotation release even on the same assembly. The assembly must be explicit, and a gene-range must be resolved to precise positions before use. This is structural, not advisory: there is no source-side guardrail that would catch a misplaced region, so an unverified coordinate produces results for an unintended location with no error.

bash
# Resolve gene symbol -> GRCh38 region with an authoritative source FIRST,
# then pass the verified coordinates to the OneKGPd query below.

[!CAUTION] The dataset is GRCh38. A GRCh37 coordinate, or any region that does not correctly correspond to the intended feature on GRCh38, will return results for an unintended location without raising an error. Verify the assembly and the resolved coordinates before querying.

Command Selection Guide

Match the question to the command. Counting commands are cheap and should precede their selection counterpart.

  • Which individuals carry matching variants in a region → count-samples then select-samples
  • Which variants are carried in a region, cohort-wide → count-variants then select-variants
  • Which variants are carried in a region, within a named set of individuals → count-variants-in-samples then select-variants-in-samples
  • Who is homozygous-reference at a single position → count-samples-hom-ref then select-samples-hom-ref
  • Relatedness (degree + coefficient) between two named individuals → kinship
  • Dataset totals (sample count, sex split, variant total, assembly) → dataset-info

Cohort interpretation

The 3,202-sample cohort includes relatives: the additional 698 high-coverage samples extend the original 2,504-sample panel. Carrier counts therefore are not counts of independent observations, and cohort AF is not a population prevalence estimate. For association or frequency comparisons, document the selected populations and relatedness policy; use the bundled pedigree metadata and kinship when choosing or auditing the analysis set. See the IGSR cohort announcement.

Annotation filters (shared across variant and sample selection/counting)

All variant- and sample-selection commands (count-variants, select-variants, their -in-samples forms, count-samples, select-samples) accept the same annotation filters. Different filter fields are combined with AND; multiple values within one field are combined with OR. Enum values are case-insensitive (e.g. missense_variant or MISSENSE_VARIANT).

These are selection criteria applied on the server. The fields returned on a selected variant are listed under Variant-returning commands; a criterion used for filtering is not necessarily echoed back on the returned variant.

  • --af-lt / --af-gt: 1000 Genomes dataset allele frequency bounds
  • --gnomad-exomes-af-lt / --gnomad-exomes-af-gt: gnomAD v4.1 exome AF bounds
  • --gnomad-genomes-af-lt / --gnomad-genomes-af-gt: gnomAD v4.1 genome AF bounds
  • --clin-significance: ClinVar significance terms, CSV (e.g. PATHOGENIC,LIKELY_PATHOGENIC)
  • --consequence: Sequence Ontology consequence terms, CSV (e.g. MISSENSE_VARIANT,STOP_GAINED)
  • --impact: VEP impact, CSV (HIGH,MODERATE,LOW,MODIFIER)
  • --variant-type, --feature-type, --bio-type: SO variant class / VEP feature / VEP biotype, CSV
  • --alpha-missense-class: AM_LIKELY_BENIGN,AM_LIKELY_PATHOGENIC,AM_AMBIGUOUS (CSV)
  • --alpha-missense-score-lt / --alpha-missense-score-gt: AlphaMissense score bounds
  • --biallelic-only / --multiallelic-only
  • --exclude-males / --exclude-females
  • --min-len-bp / --max-len-bp: alternate-allele length bounds (bp)

[!NOTE] --alpha-missense-class and --alpha-missense-score-* are mutually exclusive (the engine ignores the class when a score bound is set). --biallelic-only and --multiallelic-only are mutually exclusive. --exclude-males and --exclude-females are mutually exclusive. Setting a *-gt bound greater than or equal to its matching *-lt bound defines an empty range and will return nothing.

[!NOTE] gnomad_exomes_af, gnomad_genomes_af, and am_score use 0.0 for not annotated in this service snapshot. This does not establish absence from the current gnomAD release, biological rarity, or a benign prediction. The dataset's own af field is a different statistic, not this sentinel.

[!CAUTION] A zero numeric filter is unset on the server, so --gnomad-exomes-af-gt 0 does not exclude missing annotations. The wrapper rejects zero, nonfinite, out-of-range, and float32-underflowing bounds. Choose an explicit positive threshold (for example --gnomad-exomes-af-gt 0.000001 means AF > 1e-6, not merely annotation presence). For exact > 0, retrieve a complete variant set and post-filter the returned AF locally. A < X filter alone includes unannotated zero values. Apply the same missing-score caution to AlphaMissense.

Categorical annotations are retained across transcripts. Combining consequence and impact filters does not establish that they describe the same transcript. amino_acids may contain multiple HGVSp entries; the generic gRPC service places canonical annotations first, whereas the separate MCP layer trims its output. Preserve transcript identifiers, and do not treat a model's likely-pathogenic class as a clinical diagnosis or a participant phenotype.

Show full SKILL.md (708 more words)Show less

Quick Start

bash
# Step 1. NCBI Gene 672, GRCh38.p14 / NC_000017.11, RS_2025_08:
# BRCA1 spans chr17:43044295-43170327 (1-based inclusive).
# Source: https://www.ncbi.nlm.nih.gov/gene/672 ; re-resolve for your analysis.
# Step 2. Size the result set: how many individuals carry predicted likely-pathogenic
#    missense variants in this region?
uv run scripts/onekgpd_api.py count-samples \
  --chrom chr17 --start 43044295 --end 43170327 \
  --consequence MISSENSE_VARIANT \
  --alpha-missense-class AM_LIKELY_PATHOGENIC \
  --output /tmp/count.json
# Step 3. If the count is manageable, list those individuals.
uv run scripts/onekgpd_api.py select-samples \
  --chrom chr17 --start 43044295 --end 43170327 \
  --consequence MISSENSE_VARIANT \
  --alpha-missense-class AM_LIKELY_PATHOGENIC \
  --output /tmp/samples.json
# Step 4. Count then select variants for actual returned sample IDs.
# HG03169,NA20506 below are illustrative IDs; substitute the Step 3 results.
uv run scripts/onekgpd_api.py count-variants-in-samples \
  --chrom chr17 --start 43044295 --end 43170327 \
  --samples HG03169,NA20506 \
  --consequence MISSENSE_VARIANT --alpha-missense-class AM_LIKELY_PATHOGENIC \
  --output /tmp/variant_count.json
uv run scripts/onekgpd_api.py select-variants-in-samples \
  --chrom chr17 --start 43044295 --end 43170327 \
  --samples HG03169,NA20506 \
  --consequence MISSENSE_VARIANT --alpha-missense-class AM_LIKELY_PATHOGENIC \
  --output /tmp/variants.json

Commands

Each command writes full JSON to a file (--output PATH, default a temp file) and prints a concise stdout summary. All region/sample commands share: the region input (--chrom/--start/--end with optional --ref/--alt, or one or more repeated --region CHR:START-END), the zygosity flags (--het-only/--hom-only, default both), and the annotation filters above. The full per-flag tables live in references/onekgpd_commands.md.

Variant-returning commands

select-* return matching variants; count-* return an integer count.

  • count-variants — count variants in a region, cohort-wide.
  • select-variants — select variants in a region, cohort-wide. Use --limit N (hard cap, default 200) or --page-size N (retrieve the full set in pages); the two are mutually exclusive. The summary flags truncated when the cap is reached.
  • count-variants-in-samples — as count-variants, restricted to --samples NAME1,NAME2,... (required).
  • select-variants-in-samples — as select-variants, restricted to --samples NAME1,NAME2,... (required).

Each returned variant carries these 22 keys: chr, start, end, ref, alt, af, ac, an, hom_samples, het_samples, mis_samples, hom_samples_fx, het_samples_fx, mis_samples_fx, hom_samples_mxy, het_samples_mxy, mis_samples_mxy, gnomad_exomes_af, gnomad_genomes_af, am_score, amino_acids, biallelic. ClinVar significance and VEP consequence are filter criteria only and are not returned. Full schema: references/onekgpd_commands.md.

Sample-returning commands
  • count-samples — count individuals carrying a matching variant in a region.
  • select-samples — list the names of individuals carrying a matching variant. Supports --skip N and --limit N. Returns names only; to see which variants qualified an individual, feed the names into select-variants-in-samples.
Homozygous-reference commands

Single position via --chrom + --position (not a region).

  • count-samples-hom-ref — count individuals with a 0/0 call at the position. The count uses a sentinel: -1 = no variant exists at that position at all; 0 = a variant exists but no individual is homozygous reference; >0 = the number of homozygous-reference individuals. These interpretations require result_incomplete=false; otherwise variant_present is null. No variant record is not evidence that all 3,202 individuals have callable 0/0 genotypes.
  • select-samples-hom-ref — list the individuals with a 0/0 call at the position.
Relatedness command
  • kinship --sample1 NAME --sample2 NAME — relatedness between two named individuals: the degree (TWINS_MONOZYGOTIC / FIRST_DEGREE / SECOND_DEGREE / THIRD_DEGREE / UNRELATED) and the KING kinship coefficient (phi_bwf).
Dataset metadata command
  • dataset-info — dataset totals: samples_total (3,202), female/male split, variants_total, assembly (GRCh38), and the cohort breakdown. No region required; doubles as a connectivity check.

Sample & population metadata (offline)

Population, sex, pedigree, and superpopulation questions are answered by a second script, scripts/onekgpd_meta.py, from a data file bundled in the skill — no network, no credentials, no coordinates. The sample IDs are the same names the variant commands use, so the two layers compose (e.g. pick a cohort by population, then query its variants). Run uv run scripts/onekgpd_meta.py <command>.

The cohort has 5 superpopulations (AFR, AMR, EAS, EUR, SAS) and 26 populations. Population/superpopulation values match case-insensitively by short code or full name; sample IDs are case-sensitive.

  • sample-metadata --samples NA19240,HG00096 — family, gender, parents, children, population, superpopulation, and phase3 status for the given samples.
  • list-populations — all 26 populations with superpopulation and sample count (use to discover valid values).
  • list-superpopulations — the 5 superpopulations with sample count and constituent populations.
  • population-stats --populations YRI [--populations CHS …] — per-population sex split, phase3 count, and trio membership. Repeat --populations for multiple values (full names contain commas, so they are not comma-separated).
  • superpopulation-summary --superpopulations EAS [--superpopulations EUR …] — per-superpopulation totals with a per-population breakdown.
  • select-samples-by-population --population YRI and/or --superpopulation AFR, with optional --skip/--limit (default 0 / 50, max 3202) — the sample IDs in a population and/or superpopulation; both given intersects. Feed the names into select-variants-in-samples to see their variants.

See references/onekgpd_commands.md for full argument tables and JSON output schemas.

Typical Workflows

Which individuals, then which variants they carry

The following is an illustrative template; replace all angle-bracket placeholders.

bash
# Step 1: resolve gene -> verified GRCh38 region (authoritative source).
# Step 2: count individuals carrying a qualifying variant in the region.
uv run scripts/onekgpd_api.py count-samples \
  --chrom <chr> --start <start> --end <end> \
  --consequence MISSENSE_VARIANT --alpha-missense-class AM_LIKELY_PATHOGENIC \
  --output /tmp/n.json
# Step 3: list those individuals.
uv run scripts/onekgpd_api.py select-samples \
  --chrom <chr> --start <start> --end <end> \
  --consequence MISSENSE_VARIANT --alpha-missense-class AM_LIKELY_PATHOGENIC \
  --output /tmp/who.json
# Step 4: count variants for those individuals before selecting.
uv run scripts/onekgpd_api.py count-variants-in-samples \
  --chrom <chr> --start <start> --end <end> \
  --samples <name1,name2,...> \
  --consequence MISSENSE_VARIANT --alpha-missense-class AM_LIKELY_PATHOGENIC \
  --output /tmp/variant_count.json
uv run scripts/onekgpd_api.py select-variants-in-samples \
  --chrom <chr> --start <start> --end <end> \
  --samples <name1,name2,...> \
  --consequence MISSENSE_VARIANT --alpha-missense-class AM_LIKELY_PATHOGENIC \
  --output /tmp/variants.json
Homozygous-reference carriers at a position of interest

Illustrative template; replace the placeholders with verified coordinates.

bash
# After identifying a position of interest (verified coordinate):
uv run scripts/onekgpd_api.py count-samples-hom-ref \
  --chrom <chr> --position <pos> --output /tmp/homref_n.json
uv run scripts/onekgpd_api.py select-samples-hom-ref \
  --chrom <chr> --position <pos> --output /tmp/homref.json

Common Mistakes

  • Mistake: Querying with an unverified coordinate. Fix: Always resolve gene/feature → GRCh38 against an authoritative source first. A misplaced region returns results for an unintended location without error.
  • Mistake: Calling a selection command before its counting command. Fix: Count first; selection result sets can be large.
  • Mistake: Assuming a GRCh37 coordinate will work. Fix: The dataset is GRCh38 only.

References

© K-Dense-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references, assets) in skills/onekgpd of K-Dense-AI/scientific-agent-skills.

  • SKILL.md
  • assets/kgpe.json
  • references/annotation_vocabularies.md
  • references/onekgpd_commands.md
  • scripts/onekgpd_api.py
  • scripts/onekgpd_meta.py

Open the folder on GitHubat commit 92ace75

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in K-Dense-AI/scientific-agent-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Onekgpd next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Onekgpd compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Onekgpd this skillK-Dense-AI/scientific-agent-skills48k1 repos~5.6kAutomated safety check: NotesMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0
MFA Pipeline Orchestratoraiming-lab/AutoResearchClaw15k—~923Automated safety check: PassMIT

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Singlecell Qc

    xuzhougeng/wisp-science

    A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.

    1k GitHub stars~1.6k tokensUpdated today
    Research & ScienceAuto-check passed

More from K-Dense-AI/scientific-agent-skills

All 153 skills in this repo
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Analytical Method Validation Planner

    K-Dense-AI/scientific-agent-skills

    Plans, runs, and documents analytical method validation, verification, or transfer studies under ICH Q2(R2)/Q14, USP, ICH M10, CLSI EP, or ISO/IEC 17025.

    48k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check: notes
  • Cantera Ignition Delay

    K-Dense-AI/scientific-agent-skills

    Runs Cantera constant-volume or constant-pressure ignition simulations and reports temperature-based ignition delay with mechanism provenance and checks.

    48k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • DiffDock Molecular Docking

    K-Dense-AI/scientific-agent-skills

    Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

    48k GitHub starsUsed in 1 repo~3k tokens
    Auto-check: notes
  • HypoGeniC Hypothesis Generation

    K-Dense-AI/scientific-agent-skills

    Plans and audits runs of the HypoGeniC and HypoRefine packages, which propose hypotheses from labeled text datasets, with local checks before any model call.

    48k GitHub starsUsed in 1 repo~3.6k tokens
    Auto-check: notes
  • ISO Standards Readiness Evidence

    K-Dense-AI/scientific-agent-skills

    Organizes scope, controlled documents, risk files and traceability into draft evidence for human review against ISO 13485, 14971, 17025 and 15189.

    48k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check: notes

Questions about Onekgpd

What does Onekgpd do?

Queries the 1000 Genomes Project dataset (3,202 whole-genome-sequenced individuals, GRCh38) at the level of individual participants. Onekgpd is an agent skill from K-Dense-AI/scientific-agent-skills. Queries the 1000 Genomes Project dataset (3,202 whole-genome-sequenced individuals, GRCh38) at the level of individual participants.

When should I use Onekgpd?

Onekgpd fits situations like: A question is about individuals; variants in the 1000 Genomes Project cohort: which individuals carry variants matching specific criteria in a gene; which individuals are homozygous-reference at a position; which variants exist in the dataset.

How do I install Onekgpd in Claude Code?

Run `npx skills add K-Dense-AI/scientific-agent-skills --skill onekgpd -a claude-code`. Or copy the skill folder (skills/onekgpd in K-Dense-AI/scientific-agent-skills) into .claude/skills/onekgpd in your project. Claude Code loads it when a task matches its description.

How do I install Onekgpd in Codex?

Run `npx skills add K-Dense-AI/scientific-agent-skills --skill onekgpd -a codex`. Or copy the skill folder (skills/onekgpd in K-Dense-AI/scientific-agent-skills) into .agents/skills/onekgpd in your project. Codex loads it when a task matches its description.

Can I use Onekgpd in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add K-Dense-AI/scientific-agent-skills --skill onekgpd -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/onekgpd, .gemini/skills/onekgpd, .github/skills/onekgpd and .opencode/skills/onekgpd in your project.

What does Onekgpd need to run?

Going by SKILL.md and its folder, Onekgpd needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Write, Bash. Compatibility (from SKILL.md): Requires Python >=3.11. Variant and sample queries require outbound network access to the public 1000 Genomes query endpoint over TLS; the sample/population metadata commands run fully offline over a data file bundled in the skill. No credentials, API keys, or environment variables are used..

Does Onekgpd access the network?

SKILL.md names 4 domains. In commands or code: ncbi.nlm.nih.gov; the agent is likely to contact it when it follows the instructions. As links in the text: internationalgenome.org, docs.astral.sh and dnaerys.org. This is read from the text; nothing was executed.

Is Onekgpd safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Onekgpd use?

Onekgpd is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Onekgpd use?

About 5.6k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.7k tokens, read only when the agent opens those files.

What are the alternatives to Onekgpd?

Skills that share tags, products or a category with Onekgpd: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars), Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars) and Dbsnp Database (google-deepmind/science-skills, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Onekgpd?

K-Dense-AI (a GitHub organization) maintains it in K-Dense-AI/scientific-agent-skills, which has 48,215 GitHub stars. The repository holds 153 skills in this directory. The repository was last updated on October 5, 2026.

Source: K-Dense-AI/scientific-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.