Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
View, query, and interpret VCF/BCF variant files with bcftools and cyvcf2.
$ npx skills add GPTomics/bioSkills --skill bio-vcf-basics -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-vcf-basics --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/variant-calling/vcf-basics .claude/skills/bio-vcf-basics && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-vcf-basics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/vcf-basics into .claude/skills/bio-vcf-basics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-vcf-basics", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/variant-calling/vcf-basicsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-vcf-basics -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-vcf-basics --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/variant-calling/vcf-basics .agents/skills/bio-vcf-basics && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-vcf-basics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/vcf-basics into .agents/skills/bio-vcf-basics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-vcf-basics", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-vcf-basics -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-vcf-basics --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/variant-calling/vcf-basics .cursor/skills/bio-vcf-basics && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-vcf-basics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/vcf-basics into .cursor/skills/bio-vcf-basics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-vcf-basics", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path variant-calling/vcf-basics--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-vcf-basics -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-vcf-basics --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/variant-calling/vcf-basics .gemini/skills/bio-vcf-basics && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-vcf-basics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/vcf-basics into .gemini/skills/bio-vcf-basics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-vcf-basics", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-vcf-basicsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-vcf-basics -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/variant-calling/vcf-basics .github/skills/bio-vcf-basics && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-vcf-basics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/vcf-basics into .github/skills/bio-vcf-basics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-vcf-basics", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-vcf-basics -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-vcf-basics --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/variant-calling/vcf-basics .opencode/skills/bio-vcf-basics && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-vcf-basics" agent skill from https://github.com/GPTomics/bioSkills/tree/main/variant-calling/vcf-basics into .opencode/skills/bio-vcf-basics/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-vcf-basics", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-vcf-basicsView, query, and interpret VCF/BCF variant files with bcftools and cyvcf2.
Bio Vcf Basics is an agent skill from GPTomics/bioSkills. View, query, and interpret VCF/BCF variant files with bcftools and cyvcf2. Use when inspecting variants, extracting fields with query format strings, converting VCF/BCF, or correctly reading a field -- QUAL (site) vs GQ (genotype) vs PL/GL likelihoods, AD vs DP and allele balance, GT phasing/ploidy/PS and missing-vs-hom-ref, INFO/FORMAT Number A/R/G semantics, symbolic alleles (<DEL, <NONREF, spanning ) and END, or telling a raw gVCF apart from a filtered callset.
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/view_vcf.py` and `usage-guide.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
samtools.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Vcf Basics loads about 5.3k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 2,316 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,316 words, ~5,265 tokens.
.claude/skills/bio-vcf-basics/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: bcftools 1.19+, cyvcf2 0.30+, numpy 1.26+
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signatures<tool> --version then <tool> --help to confirm flagsIf code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Show me and extract fields from this VCF" -> Parse the VCF/BCF format, then view, subset, or pull specific columns into a flat table.
bcftools view / bcftools query -fcyvcf2.VCF (iterate records with attribute access)A VCF field is only meaningful once its LEVEL and its Number are known. QUAL is a site-level property; GQ, PL, AD, DP, GT are per-sample. QUAL and GQ answer different questions and are NOT interchangeable. A field's header Number (A/R/G/.) dictates how many values it carries and how it must be re-subset after a multiallelic split. And several encodings are load-bearing traps: . (missing) is never 0/0 (hom-ref); a bare * ALT is a spanning-deletion placeholder, not an allele; a gVCF <NON_REF> record is a reference-confidence intermediate, not a filtered call. Read the header, read the Number, read the level -- a structurally valid VCF read at the wrong level silently produces wrong numbers with no error.
| Format | Description | Use Case |
|---|---|---|
| VCF | Text format, human-readable | Debugging, small files |
| VCF.gz | Compressed VCF (bgzip) | Standard distribution |
| BCF | Binary VCF | Fast processing, large files |
##fileformat=VCFv4.2
##INFO=<ID=DP,Number=1,Type=Integer,Description="Total Depth">
##FORMAT=<ID=GT,Number=1,Type=String,Description="Genotype">
##FORMAT=<ID=DP,Number=1,Type=Integer,Description="Read Depth">
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE1
chr1 1000 rs123 A G 30 PASS DP=50 GT:DP 0/1:25##fileformat - VCF version##INFO / ##FORMAT - INFO / FORMAT field definitions (ID, Number, Type)##FILTER - Filter definitions##contig - Reference contigs (required for region indexing and contig order)##reference - Reference genomeEvery INFO/FORMAT tag used in the body MUST be declared in a ##INFO/##FORMAT line giving its ID, Number, and Type; parsers (bcftools, cyvcf2, pysam) read these declarations to know how many values a field holds and how to type it. An out-of-sync header -- a tag used but not declared, or declared with the wrong Number/Type -- silently breaks parsing: a Number=1 declaration over data that holds a vector, or a missing ##contig, makes tools drop, mistype, or mis-subset values with NO error thrown. After any hand-edit or annotation that adds a field, update the header to match (bcftools +fill-tags and bcftools annotate manage this automatically).
| Column | Description |
|---|---|
| CHROM | Chromosome |
| POS | 1-based position of the first base in REF (contrast BED's 0-based half-open) |
| ID | Variant identifier (e.g., rs number) or . if novel |
| REF | Reference allele (matches the reference exactly) |
| ALT | Alternate allele(s), comma-separated. * = allele missing due to an overlapping deletion at this site |
| QUAL | Phred-scaled quality of the ALT assertion, -10*log10 P(no variant); higher = more confident a variant exists (site-level, NOT per-sample) |
| FILTER | PASS or semicolon-separated filter names. . means filters were not applied |
| INFO | Semicolon-separated key=value pairs (site-level annotations) |
| FORMAT | Colon-separated format keys defining per-sample field order |
| SAMPLE | Colon-separated values matching FORMAT order |
What each field actually measures -- and what it does not -- drives every filtering and interpretation decision. QUAL, GQ, and PL answer three DIFFERENT questions.
| Field | Level | Scale | Question answered |
|---|---|---|---|
| QUAL (col 6) | Site | Phred: -10*log10 P(no variant) | "Is there ANY variant at this site?" |
| GQ (FORMAT) | Genotype | Phred, capped at 99 | "Is THIS sample's assigned genotype correct?" |
| PL (FORMAT) | Genotype | Phred, rebased to min=0 | Relative likelihood of every possible genotype |
| GL (FORMAT) | Genotype | log10, <=0, raw | Same info as PL, unscaled (PL = -10*GL, rebased) |
QUAL is computed once across all samples and SCALES with total depth, so a high-coverage artifact can carry a large QUAL -- hence QD (QUAL normalized by depth) is preferred for filtering. GQ is per-sample and does not scale with cohort size. They are NOT interchangeable: QUAL can be high while an individual genotype is uncertain (low GQ), and a sample can have a confident genotype (high GQ) at a site with only moderate QUAL. Filter site-level junk on QUAL/QD; no-call untrustworthy genotypes on GQ.
PL holds phred-scaled genotype likelihoods, rebased so the CALLED (most likely) genotype is exactly 0 and every other value is its phred penalty relative to that call. For a biallelic diploid site PL is ordered [PL(0/0), PL(0/1), PL(1/1)] -- the index of the 0 IS the genotype the caller assigned. GL is the same information as raw log10 likelihoods (<=0, larger is better). GQ = the difference between the two SMALLEST PL values, i.e. phred confidence in the call versus the next-best genotype; GQ=0 means the top two genotypes are tied (uninformative), GQ is capped at 99 by convention.
For a site with n alleles, diploid genotype j/k (j<=k) sits at PL index k*(k+1)/2 + j (this is the Number=G ordering). Getting this index formula wrong is the classic bug when re-parsing PL after a multiallelic split -- the vector must be re-subset by the formula, never sliced positionally.
AD (FORMAT, Number=R) is per-allele read depth [ref_depth, alt1_depth, ...], REF first. DP is total depth. sum(AD) is often LESS than DP -- expected, not an error:
Allele balance for a het is DERIVED from AD (GATK does not emit it directly): AB = alt_AD / (ref_AD + alt_AD). A true het sits near 0.5; hets far from 0.5 (e.g. <0.2 or >0.8) suggest a mapping artifact, CNV, or contamination.
Every ##INFO/##FORMAT header declares a Number telling a parser how many values a field holds AND how to re-subset it when a multiallelic record is split:
| Number | One value per | Examples | On multiallelic split |
|---|---|---|---|
A | ALT allele | AF, AC | take the k-th element for the k-th ALT |
R | allele incl. REF | AD | REF value first, then per-ALT (off-by-one vs A) |
G | genotype | PL, GL | re-subset via the k*(k+1)/2+j index formula |
. | variable/unknown | -- | parser CANNOT auto-subset; carried whole onto every record |
0 | flag (presence only) | -- | -- |
Load-bearing for correctness: bcftools norm -m- uses these codes to reapportion fields on split. A field mis-declared Number=. when it is really A keeps its full multiallelic vector on every split record, so downstream tools read the WRONG allele's value with no error. See variant-calling/variant-normalization for split/join reapportionment.
| Annotation | Meaning | What It Detects |
|---|---|---|
| QD | QUAL / allele depth | Low values suggest variant quality not supported by reads |
| FS | Fisher strand bias (phred-scaled) | Variant reads predominantly on one strand (artifact) |
| SOR | Strand odds ratio | Same as FS but handles high-depth sites better |
| MQ | Root mean square mapping quality | Low values indicate reads map ambiguously (paralogous regions) |
| MQRankSum | MQ difference: ref vs alt reads | Very negative = alt reads map much worse than ref (suspicious) |
| ReadPosRankSum | Read position: ref vs alt reads | Very negative = variant only at read ends (misalignment artifact) |
| Genotype | Meaning |
|---|---|
0/0 | Homozygous reference (confidently called ref) |
0/1 | Heterozygous |
1/1 | Homozygous alternate |
1/2 | Heterozygous for two different ALT alleles (compound het at multiallelic site) |
./. | Missing genotype (no confident call) |
0|1 | Phased heterozygous (allele before | is on haplotype 1) |
/ separates unphased alleles -- the two chromosomal copies are known, but which came from which parent is not| separates phased alleles -- haplotype assignment is known (read-backed phasing, trio analysis, or long-read sequencing)A | is only meaningful WITHIN a phase set. The FORMAT/PS tag (an integer, usually the POS of the block's first variant) groups variants phased relative to EACH OTHER; 0|1 in two different PS blocks are not guaranteed to lie on the same physical haplotype. Read-backed phasers (WhatsHap) and trio phasing emit PS. A | with no consistent PS across records carries no global phase -- a subtle trap when merging phased VCFs.
Ploidy is read from the NUMBER of alleles in GT: 0/1 diploid, 0 haploid (chrY, chrM, male chrX outside the PAR), 0/1/1 triploid. Per-region ploidy (PAR, chrX in males, mito) must match the sample karyotype.
. (missing) is NOT reference. ./. = no-call (genotype could not be determined, usually low depth); 0/0 = confidently called homozygous reference. Treating ./. as 0/0 inflates the reference-allele count and biases allele frequencies, missingness, and burden tests. This is load-bearing: never impute ./. as reference. In a gVCF, the ABSENCE of a record also does not mean reference -- see the reference-confidence model below.
At multiallelic sites (e.g. ALT = G,T), allele indices reference the comma-separated ALT list: 0=REF, 1=first ALT, 2=second ALT. 1/2 means one copy of each ALT. Splitting multiallelics into biallelic records with bcftools norm -m- converts 1/2 into two 0/1 records, losing compound-heterozygosity information -- see variant-calling/variant-normalization for caveats.
Not every ALT spells out a sequence. Symbolic alleles are angle-bracketed placeholders for events whose sequence is not given inline:
| ALT | Meaning |
|---|---|
<DEL> <DUP> <INS> <INV> <CNV> | Structural-variant classes (sequence not spelled out) |
<NON_REF> | gVCF: "any allele not yet observed" (reference-confidence model) |
<*> | Same role as <NON_REF> in some callers' gVCF/mpileup output |
* (bare) | Spanning deletion: allele MISSING because an upstream deletion on ANOTHER line overlaps this position |
Two parsing traps:
INFO/END gives the end coordinate of a symbolic/large event. A tool that infers a record's span from len(REF) is WRONG for symbolic alleles -- it must read END. gVCF reference blocks also use END to mark the last position of the band.* ALT is interpretable only relative to the overlapping deletion on another record; it is not a real alternate allele here. Splitting/subsetting can strand a * from the deletion it references (see variant-calling/variant-normalization).<NON_REF> Reference-Confidence ModelA gVCF (GATK HaplotypeCaller -ERC GVCF) is fundamentally different from a filtered callset: it emits a record for EVERY position or block, not just variant sites. Non-variant stretches are compressed into END-delimited blocks (bands) grouped by GQ, so a gVCF is not one line per base.
<NON_REF> ALT with PL/AD computed against "any unseen allele." This lets joint genotyping evaluate a site in THIS sample even when the variant was only discovered in ANOTHER cohort sample -- the <NON_REF> likelihood supplies the evidence.GenomicsDBImport/CombineGVCFs -> GenotypeGVCFs) to yield a normal VCF. Do NOT filter, annotate, or count variants on a raw gVCF, and never build a multi-sample callset by bcftools merge-ing single-sample project VCFs when gVCF joint-genotyping is available -- merging fabricates hom-ref genotypes. See variant-calling/joint-calling.Goal: View, subset, and convert VCF/BCF files from the command line.
Approach: Use bcftools view with flags for header control, region selection, sample extraction, and format conversion.
bcftools view input.vcf.gz | head # full records
bcftools view -h input.vcf.gz # header only
bcftools view -H input.vcf.gz | head # skip header
bcftools view input.vcf.gz chr1:1000000-2000000 # region (needs index)
bcftools view -s sample1,sample2 input.vcf.gz # keep samples
bcftools view -s ^sample3 input.vcf.gz # exclude samplesGoal: Extract specific fields from a VCF in a custom tabular format.
Approach: Use bcftools query -f with format specifiers for CHROM, POS, INFO, and FORMAT fields. Square brackets [...] loop over samples.
bcftools query -f '%CHROM\t%POS\t%REF\t%ALT\n' input.vcf.gz
bcftools query -f '%CHROM\t%POS\t%INFO/DP\t%INFO/AF\n' input.vcf.gz
bcftools query -f '%CHROM\t%POS[\t%GT]\n' input.vcf.gz # per-sample GT
bcftools query -f '%CHROM\t%POS[\t%SAMPLE=%GT]\n' -s sample1 input.vcf.gz
bcftools query -H -f '%CHROM\t%POS\t%REF\t%ALT\n' input.vcf.gz # column header| Specifier | Description |
|---|---|
%CHROM %POS %ID | Position fields |
%REF %ALT | Alleles |
%QUAL %FILTER | Site quality / filter status |
%INFO/TAG | INFO field value |
%TYPE | Variant type (snp, indel, etc.) |
[%GT] [%DP] [%AD] [%GQ] | Per-sample FORMAT fields (loop in [...]) |
[%SAMPLE] | Sample name |
\n \t | Newline / tab |
Goal: Convert between VCF, compressed VCF, and BCF, and index for region queries.
Approach: Use bcftools view output flags (-Ov/-Oz/-Ou/-Ob), then bgzip + index.
bcftools view -Ob -o output.bcf input.vcf.gz # VCF -> BCF
bcftools view -Ov -o output.vcf input.bcf # BCF -> VCF
bgzip input.vcf # -> input.vcf.gz (bgzip, NOT gzip)
bcftools index input.vcf.gz # -> .csi index
bcftools index -t input.vcf.gz # -> .tbi (tabix) index| Flag | Format |
|---|---|
-Ov | Uncompressed VCF |
-Oz | Compressed VCF (bgzip) |
-Ou | Uncompressed BCF (fast piping) |
-Ob | Compressed BCF |
BCF is the binary encoding of VCF: faster to parse and smaller for large callsets. Region queries (chr1:1-1000) require a bgzipped+indexed VCF or a BCF -- plain .gz (gzip) is not seekable and fails.
Goal: Read, query, and write VCF files programmatically in Python.
Approach: Use cyvcf2's VCF reader to iterate variants with attribute access to fields, and Writer to emit filtered output.
"Parse this VCF in Python" -> Open with cyvcf2 and iterate variant records.
from cyvcf2 import VCF
vcf = VCF('input.vcf.gz')
for variant in vcf:
# ALT is a list; QUAL is site-level and may be None
print(variant.CHROM, variant.POS, variant.REF, variant.ALT)
print(variant.ID, variant.QUAL, variant.FILTER, variant.var_type)
dp = variant.INFO.get('DP') # INFO field, None if absent
af = variant.INFO.get('AF')
break
vcf.close()from cyvcf2 import VCF
vcf = VCF('input.vcf.gz')
samples = vcf.samples
for variant in vcf:
# gt_types: 0=HOM_REF, 1=HET, 2=UNKNOWN(missing), 3=HOM_ALT
gts = variant.gt_types
depths = variant.format('DP') # numpy array, one row per sample
gqs = variant.format('GQ') # per-sample genotype quality
ad = variant.format('AD') # per-allele depth, Number=R
print(dict(zip(samples, gts)))
break
vcf.close()Note: cyvcf2 codes missing genotypes as gt_types == 2 (UNKNOWN) -- treat that as no-call, never as HOM_REF.
from cyvcf2 import VCF
vcf = VCF('input.vcf.gz')
print(vcf.samples, vcf.seqnames) # sample names, contig names
for info in vcf.header_iter():
if info['HeaderType'] == 'INFO':
print(info['ID'], info['Description'])
for variant in vcf('chr1:1000000-2000000'): # requires an index
print(variant.CHROM, variant.POS)from cyvcf2 import VCF, Writer
vcf = VCF('input.vcf.gz')
writer = Writer('output.vcf', vcf) # inherit the input header
for variant in vcf:
if variant.QUAL is not None and variant.QUAL > 30: # QUAL is site-level
writer.write_record(variant)
writer.close()
vcf.close()| Task | bcftools | cyvcf2 |
|---|---|---|
| View VCF | bcftools view file.vcf.gz | VCF('file.vcf.gz') |
| View header | bcftools view -h file.vcf.gz | vcf.header_iter() |
| Get region | bcftools view file.vcf.gz chr1:1-1000 | vcf('chr1:1-1000') |
| Query fields | bcftools query -f '%CHROM\t%POS\n' | Loop with properties |
| Count variants | bcftools view -H file.vcf.gz | wc -l | sum(1 for _ in vcf) |
| VCF to BCF | bcftools view -Ob -o out.bcf in.vcf.gz | Use Writer |
| Error | Cause | Solution |
|---|---|---|
no BGZF EOF marker | Not bgzipped (plain gzip) | Recompress with bgzip, not gzip |
index required / region query fails | Missing index | Run bcftools index (-t for tabix) |
sample not found | Wrong sample name | Check with bcftools query -l |
| INFO/FORMAT field missing or mistyped | Header out of sync with body | Fix ##INFO/##FORMAT Number/Type; use bcftools +fill-tags |
| Every hom-alt or missing site vanishes on filter | Treated ././. as failing or as ref | Missing != hom-ref; make missing pass, never impute 0/0 |
| Wrong allele's AF/AD after split | Number=. field not re-subset | Declare the true Number (A/R/G) so bcftools reapportions |
* overlapping-deletion allele, Number=A/R/G, PL/GL/GQ, gVCF <NON_REF>)© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in variant-calling/vcf-basics of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Vcf Basics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Vcf Basics this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5.3k | Automated safety check: Pass | MIT | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
Categories
View, query, and interpret VCF/BCF variant files with bcftools and cyvcf2. Bio Vcf Basics is an agent skill from GPTomics/bioSkills. View, query, and interpret VCF/BCF variant files with bcftools and cyvcf2.
Bio Vcf Basics fits situations like: inspecting variants; extracting fields with query format strings; converting VCF/BCF; correctly reading a field -- QUAL (site) vs GQ (genotype) vs PL/GL likelihoods.
Run `npx skills add GPTomics/bioSkills --skill bio-vcf-basics -a claude-code`. Or copy the skill folder (variant-calling/vcf-basics in GPTomics/bioSkills) into .claude/skills/bio-vcf-basics in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-vcf-basics -a codex`. Or copy the skill folder (variant-calling/vcf-basics in GPTomics/bioSkills) into .agents/skills/bio-vcf-basics in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-vcf-basics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-vcf-basics, .gemini/skills/bio-vcf-basics, .github/skills/bio-vcf-basics and .opencode/skills/bio-vcf-basics in your project.
Going by SKILL.md and its folder, Bio Vcf Basics needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: samtools.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Vcf Basics is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Vcf Basics: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.