Agent skill

Equity Scorer

by ClawBio in ClawBio/ClawBio

Compute HEIM diversity and equity metrics from VCF or ancestry data.

MITAuto-check passedDocuments & Office

Install Equity Scorer

skills CLI
$ npx skills add ClawBio/ClawBio --skill equity-scorer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ClawBio/ClawBio equity-scorer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ClawBio/ClawBio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/equity-scorer .claude/skills/equity-scorer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
equity-scorer
GitHub stars
1.2k
Used in
3 other repos
Token cost
~1.7k tokens
SKILL.md length
523 words
Files
5
Skills in repo
104
Repo updated
First seen
Licence
MIT

At a glance

Compute HEIM diversity and equity metrics from VCF or ancestry data.

  • Works in 6 steps: Heterozygosity Analysis: Compute… → FST Calculation: Pairwise fixation index… → PCA Visualisation: Principal Component… → …
  • Documents & Office work in your project
  • SKILL.md covers Core Capabilities, Input Formats, HEIM Equity Score Methodology and Workflow, plus 5 more sections
  • Runs Python scripts from its folder

What it does

Equity Scorer is an agent skill from ClawBio/ClawBio. Compute HEIM diversity and equity metrics from VCF or ancestry data. Generates heterozygosity, FST, PCA plots, and a composite HEIM Equity Score with markdown reports.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `api.py`, `equity_scorer.py` and `tests/test_equity_scorer.py`).

It sits in Documents & Office. The repository describes itself as: 🦖 ClawBio - The first bioinformatics-native AI agent skill library. Local-first. Reproducible. Open. Free. The licence is MIT.

When your agent uses it

  • Documents & Office work in your project

Example prompts

  • “/equity-scorer”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Heterozygosity Analysis: Compute observed and expected heterozygosity per population.
  2. FST Calculation: Pairwise fixation index between population groups.
  3. PCA Visualisation: Principal Component Analysis of genotype data, coloured by ancestry/population.
  4. HEIM Equity Score: A composite 0-100 score measuring representation equity across populations.
  5. Ancestry Distribution: Summarise and visualise the ancestry composition of a dataset.
  6. Markdown Report: Full analysis report with tables, figures, methods, and reproducibility block.

What it can do on your machine

Read from SKILL.md and the folder at commit dece754. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Equity Scorer loads about 1.7k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 523 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ClawBio/ClawBio at commit dece754, republished under its MIT licence (© ClawBio). 523 words, ~1,721 tokens.

Download SKILL.mdSave it as .claude/skills/equity-scorer/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
equity-scorer
description
Compute HEIM diversity and equity metrics from VCF or ancestry data. Generates heterozygosity, FST, PCA plots, and a composite HEIM Equity Score with markdown reports.
license
MIT
metadata.version
0.1.0

🦖 Equity Scorer

You are the Equity Scorer, a specialised bioinformatics agent for computing diversity and health equity metrics from genomic data. You implement the HEIM (Health Equity Index for Minorities) framework to quantify how well a dataset, biobank, or study represents global population diversity.

Core Capabilities

  1. Heterozygosity Analysis: Compute observed and expected heterozygosity per population.
  2. FST Calculation: Pairwise fixation index between population groups.
  3. PCA Visualisation: Principal Component Analysis of genotype data, coloured by ancestry/population.
  4. HEIM Equity Score: A composite 0-100 score measuring representation equity across populations.
  5. Ancestry Distribution: Summarise and visualise the ancestry composition of a dataset.
  6. Markdown Report: Full analysis report with tables, figures, methods, and reproducibility block.

Input Formats

VCF File

Standard Variant Call Format (.vcf or .vcf.gz) with:

  • Genotype fields (GT) for multiple samples
  • Optional: population/ancestry annotations in sample metadata
Ancestry CSV

Tabular file with columns:

  • sample_id: Unique identifier
  • population or ancestry: Population label (e.g., "EUR", "AFR", "EAS", "AMR", "SAS")
  • Optional: superpopulation, country, ethnicity
  • Optional: genotype columns for variant-level analysis

HEIM Equity Score Methodology

The HEIM Equity Score (0-100) is a composite metric:

HEIM_Score = w1 * Representation_Index
           + w2 * Heterozygosity_Balance
           + w3 * FST_Coverage
           + w4 * Geographic_Spread

where:
  Representation_Index = 1 - max_deviation_from_global_proportions
  Heterozygosity_Balance = mean_het / max_possible_het
  FST_Coverage = proportion_of_pairwise_FST_computed
  Geographic_Spread = n_continents_represented / 7

Default weights: w1=0.35, w2=0.25, w3=0.20, w4=0.20
Score Interpretation
ScoreRatingMeaning
80-100ExcellentStrong representation across global populations
60-79GoodReasonable diversity with some gaps
40-59FairNotable underrepresentation of some populations
20-39PoorSignificant diversity gaps
0-19CriticalSeverely limited population representation

Workflow

When the user asks for diversity/equity analysis:

  1. Detect input: Check if the input is VCF or CSV. Inspect headers and sample count.
  2. Extract populations: Parse population labels from metadata or ancestry columns.
  3. Compute metrics:
    • If VCF: parse genotypes, compute per-site and per-population heterozygosity, pairwise FST, run PCA
    • If CSV: compute representation statistics, ancestry distribution, geographic spread
  4. Calculate HEIM Score: Apply the composite formula above.
  5. Generate visualisations:
    • PCA scatter plot (PC1 vs PC2, coloured by population)
    • Ancestry bar chart (proportion per population)
    • Heterozygosity comparison (observed vs expected per population)
    • FST heatmap (pairwise between populations)
  6. Write report: Markdown with embedded figure paths, methods, and reproducibility block.
Show full SKILL.md (193 more words)Show less

Example Queries

  • "Score the diversity of my VCF file at data/samples.vcf"
  • "What is the HEIM Equity Score for the UK Biobank ancestry data?"
  • "Compare population representation between two cohorts"
  • "Generate a PCA plot coloured by ancestry for these samples"
  • "How underrepresented are African populations in this dataset?"

Output Structure

equity_report/
├── report.md                 # Full analysis report
├── figures/
│   ├── pca_plot.png         # PCA scatter (PC1 vs PC2)
│   ├── ancestry_bar.png     # Population proportions
│   ├── heterozygosity.png   # Observed vs expected Het
│   └── fst_heatmap.png      # Pairwise FST matrix
├── tables/
│   ├── population_summary.csv
│   ├── heterozygosity.csv
│   ├── fst_matrix.csv
│   └── heim_score.json
└── reproducibility/
    ├── commands.sh          # Commands to re-run
    ├── environment.yml      # Conda export
    └── checksums.sha256     # Output artefact checksums (tables/ + figures/)

Verify with cd <output_dir> && sha256sum -c reproducibility/checksums.sha256. The manifest names the run's derived tables and figures only: report.md and result.json embed wall-clock timestamps and would never re-verify, so they are deliberately outside the envelope (every number in result.json is still covered via the byte-identical tables/heim_score.json). Input files are not listed — input provenance is recorded as input_checksum in result.json.

Example Report Output

markdown
# HEIM Equity Report: UK Biobank Subset

**Date**: 2026-02-26
**Samples**: 1,247
**Populations**: 5 (EUR: 892, SAS: 156, AFR: 98, EAS: 67, AMR: 34)

## HEIM Equity Score: 42/100 (Fair)

### Breakdown
- Representation Index: 0.31 (EUR overrepresented at 71.5%)
- Heterozygosity Balance: 0.68 (AFR populations show highest diversity)
- FST Coverage: 1.00 (all pairwise computed)
- Geographic Spread: 0.71 (5/7 continental groups)

### Key Finding
African and American populations are underrepresented by 3.2x and 5.8x
respectively relative to global proportions. This limits the generalisability
of GWAS findings from this cohort to non-European populations.

### Recommendations
1. Prioritise recruitment from AMR and AFR communities
2. Apply ancestry-aware statistical methods for any association analyses
3. Report HEIM score alongside study demographics in publications

Dependencies

Required (Python packages):

  • biopython >= 1.82 (VCF parsing via Bio.SeqIO, population genetics)
  • pandas >= 2.0 (data wrangling)
  • numpy >= 1.24 (numerical computation)
  • scikit-learn >= 1.3 (PCA)
  • matplotlib >= 3.7 (visualisation)

Optional:

  • cyvcf2 (faster VCF parsing for large files)
  • seaborn (enhanced visualisations)
  • pysam (BAM/VCF indexing)

Safety

  • No data upload: All computation local. No external API calls for genomic data.
  • Large file warning: If VCF > 1GB, warn the user and suggest subsetting or using cyvcf2.
  • Ancestry sensitivity: Population labels are analytical categories, not identities. Include this disclaimer in reports.

© ClawBio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/equity-scorer of ClawBio/ClawBio.

  • SKILL.md
  • api.py
  • equity_scorer.py
  • tests/test_equity_scorer.py
  • tests/test_repro_bundle.py

Open the folder on GitHubat commit dece754

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in ClawBio/ClawBio, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Equity Scorer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Equity Scorer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Equity Scorer this skillClawBio/ClawBio1.2k3 repos~1.7kAutomated safety check: PassMIT
Research Writingalfonso0512/research-writing-skill4881 repos~818Automated safety check: PassMIT
Paper WritingMLNLP-World/Paper-Writing-Tips4.7k—~630Automated safety check: PassNone
PaperjurySpark-To-Paper-Skills/paperjury1.2k—~5.3kAutomated safety check: PassMIT
Literature Surveyai4s-research/ai4s-skills2372 repos~2kAutomated safety check: PassMIT
Venue TemplatesK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: PassMIT

Similar skills

  • Research Writing

    alfonso0512/research-writing-skill

    科研论文写作助手,提供 30 个 Prompt 模板覆盖论文写作全流程. An agent skill from alfonso0512/research-writing-skill.

    488 GitHub starsUsed in 1 repo~818 tokens
    Documents & OfficeAuto-check passed
  • Paper Writing

    MLNLP-World/Paper-Writing-Tips

    学术论文写作检查与优化助手。基于 MLNLP-World 社区整理的论文写作技巧,帮助检查和优化学术论文。Use when: (1) 检查论文 LaTeX 格式和排版, (2) 优化公式符号使用, (3) 改进图表设计, (4) 润色英文学术表达, (5) 检查参考文献格式, (6) 投稿前终稿检查, (7) 用户询问论文写作技巧或规范。

    4.7k GitHub stars~630 tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed
  • Paperjury

    Spark-To-Paper-Skills/paperjury

    Three modes for CS-conference papers (CVPR/ICCV/ECCV vision, ACL/EMNLP/NAACL NLP, ICLR/NeurIPS/ICML/AAAI ML).

    1.2k GitHub stars~5.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Literature Survey

    ai4s-research/ai4s-skills

    A skill your agent uses when the user wants a comprehensive literature survey on a specific research topic.

    237 GitHub starsUsed in 2 repos~2k tokens
    Documents & OfficeAuto-check passed
  • Venue Templates

    K-Dense-AI/claude-scientific-writer

    Prepare journal manuscripts, conference papers, research posters, and grant documents using venue-specific formatting guidance and bundled LaTeX scaffolds.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    Documents & OfficeAuto-check passed
  • Press Conference Revision Evidence Bound

    lensback940701/Evidence-Bound-Press-Conference-Revision-Skill

    Diagnose and revise defensive academic writing while preserving claim ceilings, evidence status, scope conditions, rival explanations, and conceptual hierarchy.

    259 GitHub stars~2.2k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from ClawBio/ClawBio

All 104 skills in this repo
  • Fetch a region of cis-eQTL summary statistics from EBI eQTL Catalogue v7+ via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~4.7k tokens
    Auto-check passed
  • Xena Tcga Gene Query

    ClawBio/ClawBio

    Query TCGA tumor biology through the ucscxenatoolspy API. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~4.7k tokensUpdated today
    Auto-check passed
  • Fetch a region of GWAS summary statistics from the NHGRI-EBI GWAS Catalog harmonised collection via tabix-on-FTP.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • Dnasp

    ClawBio/ClawBio

    Population genetics of pre-aligned DNA sequences or multi-sample VCFs using selected DnaSP 6 methods.

    1.2k GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Compute pairwise r² between a lead variant and every variant in a window using the 1000 Genomes Phase 3 GRCh38 reference panel, ancestry-stratified.

    1.2k GitHub stars~3.9k tokensUpdated today
    Auto-check passed
  • Ncbi Datasets

    ClawBio/ClawBio

    Download genomes, genes, virus sequences, and taxonomy data from NCBI using the datasets and dataformat CLI tools.

    1.2k GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed

Questions about Equity Scorer

What does Equity Scorer do?

Compute HEIM diversity and equity metrics from VCF or ancestry data. Equity Scorer is an agent skill from ClawBio/ClawBio. Compute HEIM diversity and equity metrics from VCF or ancestry data.

When should I use Equity Scorer?

Equity Scorer fits situations like: documents & Office work in your project.

How do I install Equity Scorer in Claude Code?

Run `npx skills add ClawBio/ClawBio --skill equity-scorer -a claude-code`. Or copy the skill folder (skills/equity-scorer in ClawBio/ClawBio) into .claude/skills/equity-scorer in your project. Claude Code loads it when a task matches its description.

How do I install Equity Scorer in Codex?

Run `npx skills add ClawBio/ClawBio --skill equity-scorer -a codex`. Or copy the skill folder (skills/equity-scorer in ClawBio/ClawBio) into .agents/skills/equity-scorer in your project. Codex loads it when a task matches its description.

Can I use Equity Scorer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ClawBio/ClawBio --skill equity-scorer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/equity-scorer, .gemini/skills/equity-scorer, .github/skills/equity-scorer and .opencode/skills/equity-scorer in your project.

What does Equity Scorer need to run?

Going by SKILL.md and its folder, Equity Scorer needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Equity Scorer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Equity Scorer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Equity Scorer use?

Equity Scorer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Equity Scorer use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Equity Scorer?

Skills that share tags, products or a category with Equity Scorer: Research Writing (alfonso0512/research-writing-skill, 488 stars), Paper Writing (MLNLP-World/Paper-Writing-Tips, 4.7k stars), Paperjury (Spark-To-Paper-Skills/paperjury, 1.2k stars) and Literature Survey (ai4s-research/ai4s-skills, 237 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Equity Scorer?

ClawBio (a GitHub organization) maintains it in ClawBio/ClawBio, which has 1,154 GitHub stars. The repository holds 104 skills in this directory. The repository was last updated on October 8, 2026.

Source: ClawBio/ClawBio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.