Agent skill

Vcf Variant Filtering

by jaechang-hits in jaechang-hits/SciAgent-Skills

Guide to quality filtering raw VCF files before computing summary stats (Ts/Tv ratio, variant counts, AF distributions).

CC-BY-4.0Auto-check passedResearch & Science

Install Vcf Variant Filtering

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill vcf-variant-filtering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills vcf-variant-filtering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/genomics-bioinformatics/variant/vcf-variant-filtering .claude/skills/vcf-variant-filtering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vcf-variant-filtering
GitHub stars
374
Used in
1 other repo
Token cost
~4.8k tokens
SKILL.md length
2,022 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
CC-BY-4.0

At a glance

Guide to quality filtering raw VCF files before computing summary stats (Ts/Tv ratio, variant counts, AF distributions).

  • Works in 7 steps: Always inspect the VCF before computing… → Use established CLI tools instead of… → Report what filtering was applied (or… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, Key Concepts, Pre-flight Interview and Decision Framework, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Vcf Variant Filtering is an agent skill from jaechang-hits/SciAgent-Skills. Guide to quality filtering raw VCF files before computing summary stats (Ts/Tv ratio, variant counts, AF distributions). Covers detecting raw VCFs via FILTER column and QUAL inspection, QUAL-based filtering with bcftools, Ts/Tv interpretation, and when NOT to filter. Read before any variant-level QC task. See bcftools-variant-manipulation for advanced filters, gatk-variant-calling for caller config, samtools-bam-processing for upstream alignment QC.

Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-4.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/vcf-variant-filtering”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Always inspect the VCF before computing statistics. Check the FILTER column distribution and QUAL score distribution before running any…
  2. Use established CLI tools instead of custom parsers. bcftools, vcftools, and similar domain-specific tools handle edge cases that custom…
  3. Report what filtering was applied (or not). Every time you present VCF summary statistics, state clearly whether the data was filtered…
  4. Use QUAL>=30 as a sensible default, not a universal rule. QUAL>=30 (1-in-1000 error rate) is a widely used default for first-pass…
  5. Combine filtering and statistics in a single pipeline. Piping the filtered output directly into the statistics command avoids creating…
  6. Validate filtering by checking the Ts/Tv ratio. After filtering, the Ts/Tv ratio should be in the expected range for the assay type. If it…
  7. Preserve the original VCF. Never overwrite the raw VCF with filtered output. Keep the raw file for auditability and in case you need to…

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • samtools.github.io
    • gatk.broadinstitute.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vcf Variant Filtering loads about 4.8k tokens when it runs. Until then it costs about 119 tokens; SKILL.md has 2,022 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-4.0 licence (© jaechang-hits). 2,022 words, ~4,830 tokens.

Download SKILL.mdSave it as .claude/skills/vcf-variant-filtering/SKILL.md (or your agent's skills folder).
name
vcf-variant-filtering
description
Guide to quality filtering raw VCF files before computing summary stats (Ts/Tv ratio, variant counts, AF distributions). Covers detecting raw VCFs via FILTER column and QUAL inspection, QUAL-based filtering with bcftools, Ts/Tv interpretation, and when NOT to filter. Read before any variant-level QC task. See bcftools-variant-manipulation for advanced filters, gatk-variant-calling for caller config, samtools-bam-processing for upstream alignment QC.
license
CC-BY-4.0

VCF Variant Filtering Guide

Overview

Raw VCF files produced by variant callers (GATK HaplotypeCaller, bcftools mpileup, DeepVariant, etc.) contain a mixture of true variants and artifacts from sequencing errors, alignment issues, and low-coverage regions. Computing summary statistics -- Ts/Tv ratio, variant counts, allele frequency distributions -- on unfiltered data yields unreliable results because false-positive calls disproportionately inflate transversion counts and depress the Ts/Tv ratio. This guide covers how to detect whether a VCF is raw, how to apply appropriate quality filters, when filtering is not appropriate, and how to interpret the resulting statistics correctly.

Key Concepts

VCF Quality Scores (QUAL Field)

The QUAL column in a VCF file represents the Phred-scaled probability that the variant site is polymorphic. A QUAL score of 30 means a 1-in-1000 chance the call is wrong; a QUAL of 20 means 1-in-100. Variant callers assign QUAL scores based on read evidence, base qualities, and mapping qualities. Low-QUAL variants (below 20-30) are enriched for sequencing errors and alignment artifacts. Filtering on QUAL is the simplest and most widely used first-pass quality control step.

Common QUAL thresholds and their interpretation:

QUAL ScoreError ProbabilityTypical Use
101 in 10Very permissive; rarely appropriate for final calls
201 in 100Lenient filtering; useful for somatic calling with supporting evidence
301 in 1,000Standard default for germline variant filtering
501 in 100,000Stringent; used in clinical or high-confidence applications
100+Extremely lowHighly supported variants; may over-filter low-coverage regions

For GATK-based workflows, VQSR or hard filtering on INFO annotations (QD, FS, MQ, etc.) may supplement or replace QUAL filtering. GATK's recommended hard filters for SNPs include QD < 2.0, FS > 60.0, MQ < 40.0, MQRankSum < -12.5, and ReadPosRankSum < -8.0.

Ts/Tv Ratio Significance

The transition-to-transversion (Ts/Tv) ratio is a key quality metric for variant call sets. Transitions (A<->G, C<->T) are chemically favored over transversions (all other substitutions) due to the molecular structure of nucleotide bases. Expected Ts/Tv values serve as benchmarks:

  • Whole-genome sequencing (WGS): approximately 2.0-2.1
  • Whole-exome sequencing (WES): approximately 2.8-3.3 (higher due to CpG enrichment in coding regions)
  • Raw/unfiltered call sets: often 1.5-1.8 or lower

A Ts/Tv ratio significantly below the expected range indicates contamination by false-positive transversion calls, which are the hallmark of sequencing errors. After proper quality filtering, the Ts/Tv ratio should rise to the expected range for the assay type.

Raw vs Filtered VCFs

A "raw" VCF is the direct output of a variant caller before any quality filtering has been applied. A "filtered" VCF has had quality thresholds applied, either by hard filtering (QUAL, DP, QD, etc.) or by model-based filtering (GATK VQSR, CNN). Distinguishing between the two is critical because computing statistics on raw data without disclosure leads to incorrect conclusions.

Indicators that a VCF is raw or unfiltered:

  • The filename contains "raw" (e.g., sample_raw_variants.vcf)
  • The FILTER column contains only . (missing) for all records
  • A large fraction of variants have QUAL scores below 30
  • The Ts/Tv ratio is well below the expected range for the assay type

Indicators that a VCF has already been filtered:

  • The FILTER column contains meaningful values (PASS, LowQual, VQSRTrancheSNP99.90to100.00)
  • A filter command is recorded in the VCF header (##FILTER= and ##bcftools_viewCommand= lines)
FILTER Column Semantics

The FILTER column in VCF format has specific semantics defined by the VCF specification:

  • . (dot) -- filter status has not been applied; the variant is unassessed
  • PASS -- the variant passed all filters
  • Any other value -- the variant failed the named filter(s); multiple filters are semicolon-separated

A common misconception is that . means the variant passed. In reality, . means no filter has been evaluated, so the variant's quality is unknown. Another misconception is that PASS in every row means the file is unfiltered -- some callers (e.g., DeepVariant) mark all emitted variants as PASS because they only output high-confidence calls.

To inspect the FILTER column programmatically:

bash
# Count occurrences of each FILTER value
bcftools query -f '%FILTER\n' input.vcf | sort | uniq -c | sort -rn | head

# Check if any non-'.' FILTER values exist
bcftools query -f '%FILTER\n' input.vcf | grep -v '^\.$' | head

Understanding the FILTER column is the first step in any VCF quality assessment. Always inspect it before deciding whether additional filtering is needed.

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: analysisGoal
    kind: required
    source: user
    ask: "What is the filtered call set for - a rare-disease search, a somatic variant list, a population study, or a genotyping panel?"
    default: null

  - id: D2
    param: filteringApproach
    kind: required
    source: user
    depends_on: [D1]
    ask: "Fixed hard thresholds, or a recalibration model trained on the cohort?"
    default: "hard filters; recalibration needs a cohort large enough to train on"

  - id: D3
    param: qualityThresholds
    kind: required
    source: user
    depends_on: [D2]
    ask: "Where should depth, genotype quality, and strand-bias cuts sit?"
    default: null

  - id: D4
    param: frequencyFilter
    kind: required
    source: user
    depends_on: [D1]
    ask: "Should variants common in reference populations be removed, and at what frequency?"
    default: "no frequency filter"

  - id: D5
    param: consequenceFilter
    kind: required
    source: user
    depends_on: [D1]
    ask: "Should the set be restricted by predicted consequence - coding only, high-impact only, or everything?"
    default: "everything"

  - id: D6
    param: inheritanceModel
    kind: optional_conditional
    source: user
    depends_on: [D1]
    ask: "Is there a family structure to filter on - de novo, recessive, compound heterozygous?"
    default: "none"
    skip_if: "single unrelated sample"

Every decision here hangs on D1, and that is the point: there is no generally correct filter. A rare-disease search removes anything common in the population; a population-genetics study needs exactly those variants kept. Applying a remembered threshold set without asking what the call set is for silently answers a different question.

Decision Framework

Is the VCF raw or unfiltered?
├── Yes (FILTER='.', many low-QUAL variants)
│   ├── Does the user ask for RAW statistics specifically?
│   │   ├── Yes → Do NOT filter; compute on data as-is, note it is raw
│   │   └── No → Apply QUAL>=30 filter before computing statistics
│   └── Does the user specify a custom threshold?
│       ├── Yes → Use their threshold
│       └── No → Default to QUAL>=30
├── No (FILTER has PASS/other values, filtered header present)
│   ├── Does the user ask for additional filtering?
│   │   ├── Yes → Apply requested filter
│   │   └── No → Compute statistics on existing filtered set
│   └── Does the Ts/Tv ratio look suspicious despite filtering?
│       └── Yes → Investigate; may need stricter thresholds
└── Uncertain
    └── Inspect FILTER column and QUAL distribution to determine status
ScenarioActionRationale
Raw VCF, general statistics requestedApply QUAL>=30 before computingLow-QUAL variants are enriched for errors and bias all statistics
Raw VCF, user explicitly asks for raw statsCompute as-is, report that data is unfilteredRespect the user's intent; they may be evaluating caller performance
Raw VCF, user specifies threshold (e.g., QUAL>=50)Use the user-specified thresholdUser has domain-specific requirements
Pre-filtered VCF (FILTER=PASS present)Compute without additional QUAL filterDouble-filtering may remove true variants unnecessarily
VCF with mixed FILTER valuesFilter to PASS-only with bcftools view -f PASSNon-PASS variants were flagged for a reason
VQSR-filtered VCFRespect existing tranche filteringVQSR is a more sophisticated filter than simple QUAL thresholds
Small panel or targeted sequencingConsider lower QUAL threshold (e.g., 20)Small panels have different error profiles and fewer variants

Best Practices

  1. Always inspect the VCF before computing statistics. Check the FILTER column distribution and QUAL score distribution before running any summary. A 30-second inspection prevents publishing misleading numbers. Use bcftools query -f '%FILTER\n' input.vcf | sort | uniq -c | sort -rn | head to get a quick overview.

  2. Use established CLI tools instead of custom parsers. bcftools, vcftools, and similar domain-specific tools handle edge cases that custom parsers typically miss: multi-allelic sites, complex variants, missing genotypes, and proper transition/transversion classification. Writing a Python parser to count Ts/Tv is error-prone and unnecessary.

    bash
    # Preferred: use bcftools for Ts/Tv computation
    bcftools stats input.vcf | grep ^TSTV
    
    # Preferred: use bcftools for variant counts
    bcftools stats input.vcf | grep ^SN
  3. Report what filtering was applied (or not). Every time you present VCF summary statistics, state clearly whether the data was filtered, what threshold was used, and how many variants passed. This is essential for reproducibility and for the reader to assess the validity of the results.

  4. Use QUAL>=30 as a sensible default, not a universal rule. QUAL>=30 (1-in-1000 error rate) is a widely used default for first-pass filtering of germline variant calls. However, different applications may warrant different thresholds: somatic variant calling may use lower thresholds with additional evidence, while clinical applications may demand stricter cutoffs. When in doubt, QUAL>=30 is a defensible starting point.

  5. Combine filtering and statistics in a single pipeline. Piping the filtered output directly into the statistics command avoids creating intermediate files and ensures you never accidentally compute statistics on the wrong file.

    bash
    # Filter and compute statistics in one pipeline
    bcftools view -i 'QUAL>=30' input.vcf | bcftools stats - > filtered_stats.txt
  6. Validate filtering by checking the Ts/Tv ratio. After filtering, the Ts/Tv ratio should be in the expected range for the assay type. If it is still low, consider that the variant caller may have systematic issues, the sample may have quality problems, or a stricter filter is needed.

  7. Preserve the original VCF. Never overwrite the raw VCF with filtered output. Keep the raw file for auditability and in case you need to re-filter with different parameters.

Show full SKILL.md (1,016 more words)Show less

Common Pitfalls

  1. Computing statistics on raw VCF files without any filtering. This is the most common and most consequential mistake. Raw Ts/Tv ratios can be 1.5 or lower, suggesting much worse data quality than actually exists, and variant counts will be inflated by thousands of false positives.

    • How to avoid: Always check the FILTER column before computing statistics. If FILTER is . for all records, apply QUAL>=30 filtering first.
  2. Assuming FILTER='.' means the variant passed. The dot in the FILTER column means "not assessed," not "passed." Treating unassessed variants as high-quality leads to inclusion of many false positives.

    • How to avoid: Distinguish between . (unassessed), PASS (passed all filters), and named filters (failed). Use bcftools query -f '%FILTER\n' | sort | uniq -c to inspect.
  3. Writing custom Python parsers for Ts/Tv calculation. Custom parsers frequently mishandle multi-allelic sites, complex variants, or edge cases in the VCF format. They also tend to be slower than optimized C-based tools.

    • How to avoid: Use bcftools stats for Ts/Tv computation. It is well-tested, fast, and handles all VCF edge cases correctly.
  4. Double-filtering a VCF that was already filtered. Applying QUAL>=30 to a VCF that already went through VQSR or hard filtering can remove true variants that were assigned lower QUAL scores by the caller but passed model-based assessment.

    • How to avoid: Inspect the VCF header for filter command records (##bcftools_viewCommand, ##GATKCommandLine) and the FILTER column for non-dot values before adding more filters.
  5. Using the wrong Ts/Tv expectation for the assay type. Comparing a WES Ts/Tv of 3.0 against a WGS expectation of 2.1 would incorrectly suggest the data is too good, while comparing a WGS Ts/Tv of 2.0 against WES expectations would incorrectly suggest poor quality.

    • How to avoid: Know the expected Ts/Tv range for your assay type. WGS: 2.0-2.1; WES: 2.8-3.3; targeted panels: variable depending on target regions.
  6. Forgetting to report the filtering status in results. Presenting a Ts/Tv ratio without stating whether or how the data was filtered makes the number uninterpretable and non-reproducible.

    • How to avoid: Always include a statement like "Ts/Tv computed after QUAL>=30 filtering (N variants passed)" or "Ts/Tv computed on unfiltered raw calls (N total variants)."
  7. Filtering on QUAL alone when richer annotations are available. GATK and other callers provide INFO-level annotations (QD, FS, MQ, SOR, MQRankSum, ReadPosRankSum) that are more informative than QUAL alone. Relying solely on QUAL when these are available leaves quality on the table.

    • How to avoid: Check whether INFO annotations are present. If so, consider GATK hard filtering recommendations or VQSR as a more powerful alternative to simple QUAL filtering.

Workflow

  1. Inspect the VCF metadata and FILTER column

    • Check the VCF header for caller information and existing filter records
    • Examine the FILTER column distribution
    bash
    # View header for caller and filter information
    bcftools view -h input.vcf | grep -E '##(FILTER|source|GATKCommandLine|bcftools)'
    
    # Check FILTER column values
    bcftools query -f '%FILTER\n' input.vcf | sort | uniq -c | sort -rn | head
  2. Assess QUAL score distribution

    • Determine the proportion of low-QUAL variants
    bash
    # Check QUAL distribution
    bcftools query -f '%QUAL\n' input.vcf | awk '{if($1<30) low++; else high++} END {print "QUAL<30:", low+0, "QUAL>=30:", high+0}'
    • Decision point: If most variants are QUAL>=30 and FILTER=PASS, skip to Step 4. If FILTER=. and many low-QUAL variants exist, proceed to Step 3.
  3. Apply quality filtering

    • Use the appropriate threshold (default QUAL>=30 unless otherwise specified)
    bash
    # Apply QUAL>=30 filter
    bcftools view -i 'QUAL>=30' input.vcf -Oz -o filtered.vcf.gz
    bcftools index filtered.vcf.gz
  4. Compute summary statistics

    • Extract Ts/Tv ratio, variant counts, and other metrics
    bash
    # Ts/Tv ratio from filtered VCF
    bcftools view -i 'QUAL>=30' input.vcf | bcftools stats - | grep ^TSTV | cut -f5
    
    # Full summary numbers
    bcftools view -i 'QUAL>=30' input.vcf | bcftools stats - | grep ^SN
    
    # Per-sample statistics
    bcftools stats -s - input.vcf | grep ^PSC
  5. Compare before and after filtering

    • Count variants before and after to quantify the impact of filtering
    bash
    # Count variants before filtering
    bcftools view -H input.vcf | wc -l
    
    # Count variants after filtering
    bcftools view -i 'QUAL>=30' input.vcf | bcftools view -H | wc -l
    
    # Compare Ts/Tv before and after
    echo "=== Before filtering ==="
    bcftools stats input.vcf | grep ^TSTV
    echo "=== After QUAL>=30 filtering ==="
    bcftools view -i 'QUAL>=30' input.vcf | bcftools stats - | grep ^TSTV
  6. Validate and report

    • Check that Ts/Tv is within the expected range for the assay type (WGS: 2.0-2.1, WES: 2.8-3.3)
    • If Ts/Tv is still below expected range after filtering, investigate sample quality or consider stricter thresholds
    • Report the filtering applied, number of variants before and after, and the resulting statistics
    • Include the bcftools command used so the analysis is fully reproducible

Protocol Guidelines

  1. Pre-analysis inspection is non-negotiable. Before any VCF analysis, run bcftools query -f '%FILTER\n' and examine the QUAL distribution. This takes seconds and prevents hours of wasted analysis on unreliable data.

  2. Default to QUAL>=30 for unlabeled VCFs. When the filtering status is ambiguous and the user has not specified a threshold, QUAL>=30 is the standard community default. Document this choice explicitly in the output.

  3. Never silently filter. If you apply filtering, always report it. If you choose not to filter, state why. Transparency in filtering decisions is fundamental to reproducible genomics.

  4. Use piped commands for efficiency. Chain bcftools view and bcftools stats in a pipe rather than writing intermediate files. This is faster, uses less disk, and reduces the chance of analyzing the wrong file.

  5. Verify with Ts/Tv as a sanity check. After any filtering operation, compute Ts/Tv as a quality control metric. If the ratio does not improve or remains outside the expected range, reassess the filtering strategy or investigate upstream data quality issues.

Further Reading

  • bcftools-variant-manipulation -- Detailed bcftools usage for subsetting, merging, annotating, and manipulating VCF files beyond simple filtering
  • gatk-variant-calling -- Upstream variant calling with GATK HaplotypeCaller, including VQSR and hard filtering configuration
  • samtools-bam-processing -- BAM file quality control and processing steps that precede variant calling and affect downstream VCF quality

© jaechang-hits, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/genomics-bioinformatics/variant/vcf-variant-filtering of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Vcf Variant Filtering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vcf Variant Filtering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vcf Variant Filtering this skilljaechang-hits/SciAgent-Skills3741 repos~4.8kAutomated safety check: PassCC-BY-4.0
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 11 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 11 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 11 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 11 days ago
    Auto-check passed

Questions about Vcf Variant Filtering

What does Vcf Variant Filtering do?

Guide to quality filtering raw VCF files before computing summary stats (Ts/Tv ratio, variant counts, AF distributions). Vcf Variant Filtering is an agent skill from jaechang-hits/SciAgent-Skills. Guide to quality filtering raw VCF files before computing summary stats (Ts/Tv ratio, variant counts, AF distributions).

When should I use Vcf Variant Filtering?

Vcf Variant Filtering fits situations like: tasks that involve Bioinformatics.

How do I install Vcf Variant Filtering in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill vcf-variant-filtering -a claude-code`. Or copy the skill folder (skills/genomics-bioinformatics/variant/vcf-variant-filtering in jaechang-hits/SciAgent-Skills) into .claude/skills/vcf-variant-filtering in your project. Claude Code loads it when a task matches its description.

How do I install Vcf Variant Filtering in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill vcf-variant-filtering -a codex`. Or copy the skill folder (skills/genomics-bioinformatics/variant/vcf-variant-filtering in jaechang-hits/SciAgent-Skills) into .agents/skills/vcf-variant-filtering in your project. Codex loads it when a task matches its description.

Can I use Vcf Variant Filtering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill vcf-variant-filtering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vcf-variant-filtering, .gemini/skills/vcf-variant-filtering, .github/skills/vcf-variant-filtering and .opencode/skills/vcf-variant-filtering in your project.

What does Vcf Variant Filtering need to run?

SKILL.md names no scripts, command-line tools or credentials: Vcf Variant Filtering is instructions for the agent only. Our summary lists: Python 3.

Does Vcf Variant Filtering access the network?

SKILL.md names 2 domains. As links in the text: samtools.github.io and gatk.broadinstitute.org. This is read from the text; nothing was executed.

Is Vcf Variant Filtering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vcf Variant Filtering use?

Vcf Variant Filtering is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vcf Variant Filtering use?

About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vcf Variant Filtering?

Skills that share tags, products or a category with Vcf Variant Filtering: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vcf Variant Filtering?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.