Agent skill

Nan Safe Correlation

by jaechang-hits in jaechang-hits/SciAgent-Skills

Per-feature NaN-safe Spearman/Pearson correlation across many features (genes, proteins, variants) with missing values.

CC-BY-4.0Auto-check passedData & Analytics

Install Nan Safe Correlation

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill nan-safe-correlation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills nan-safe-correlation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scientific-computing/nan-safe-correlation .claude/skills/nan-safe-correlation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nan-safe-correlation
GitHub stars
374
Used in
1 other repo
Token cost
~2.9k tokens
SKILL.md length
882 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
CC-BY-4.0

At a glance

Per-feature NaN-safe Spearman/Pearson correlation across many features (genes, proteins, variants) with missing values.

  • Works in 7 steps: Always print NaN summary before… → Use scipy.stats.spearmanr per feature in… → Set a minimum valid pair threshold… → …
  • Tasks that involve Machine learning
  • SKILL.md covers Overview, Key Concepts, Decision Framework and Best Practices, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Nan Safe Correlation is an agent skill from jaechang-hits/SciAgent-Skills. Per-feature NaN-safe Spearman/Pearson correlation across many features (genes, proteins, variants) with missing values. Covers why bulk matrix shortcuts fail, correct pairwise deletion, degenerate input filtering, and large-dataset performance. Use statistical-analysis for test choice; shap-model-explainability for interpretability.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Machine learning, Data cleaning and Statistics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is CC-BY-4.0.

When your agent uses it

  • Tasks that involve Machine learning
  • Tasks that involve Data cleaning
  • Tasks that involve Statistics

Example prompts

  • “/nan-safe-correlation”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Always print NaN summary before analysis: Report total NaN count, features with any NaN, and per-feature NaN distribution. This documents…
  2. Use scipy.stats.spearmanr per feature in a loop: This is the only method that guarantees correct pairwise NaN removal for each feature…
  3. Set a minimum valid pair threshold (min_valid): Default to 10. Features with fewer valid pairs after NaN removal produce unreliable…
  4. Filter degenerate inputs before computing correlations: Remove constant features, near-constant features, and features with excessive NaN…
  5. Track n_valid per feature in the output: The number of valid pairs varies per feature. Report it alongside rho and p-value so downstream…
  6. Report how many features were skipped or filtered: Silent feature loss is a common source of confusion. Always print the count of filtered…
  7. Use parallelization for large datasets: For > 10,000 features, use joblib to distribute the per-feature loop across cores. The per-feature…

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.scipy.org
    • pandas.pydata.org
    • stefvanbuuren.name

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nan Safe Correlation loads about 2.9k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 882 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its CC-BY-4.0 licence (© jaechang-hits). 882 words, ~2,895 tokens.

Download SKILL.mdSave it as .claude/skills/nan-safe-correlation/SKILL.md (or your agent's skills folder).
name
nan-safe-correlation
description
Per-feature NaN-safe Spearman/Pearson correlation across many features (genes, proteins, variants) with missing values. Covers why bulk matrix shortcuts fail, correct pairwise deletion, degenerate input filtering, and large-dataset performance. Use statistical-analysis for test choice; shap-model-explainability for interpretability.
license
CC-BY-4.0

NaN-Safe Correlation Computation

Overview

Computing correlations across many features (genes, proteins, variants) when missing values are present is error-prone. The most common mistake is using bulk matrix shortcuts that silently mishandle NaN, producing incorrect correlation values. This guide covers correct per-feature pairwise computation, degenerate input filtering, and performance optimization.

Key Concepts

Pairwise vs Listwise Deletion
  • Pairwise deletion: For each feature pair, remove only samples where either value is NaN. Each feature uses the maximum available data.
  • Listwise deletion: Remove any sample with NaN in any feature. Wastes valid data and biases results if missingness is not completely random.
  • Rule: Always use pairwise deletion for per-feature correlations.
Why Bulk Matrix Shortcuts Fail

Different features have different missing value patterns across samples. Bulk methods handle this inconsistently:

MethodProblem
DataFrame.rank() then corrwith()rank() assigns NaN ranks; corrwith() may drop globally or per-column inconsistently
DataFrame.corrwith(method='spearman')Implementation varies by pandas version; may use listwise deletion
np.corrcoef on ranked dataPropagates NaN to entire result if any value is missing
Impact of Incorrect Computation
  • Correlations can shift by 0.01-0.05 or more
  • Features near a threshold (e.g., 0.6) can be misclassified
  • Valid sample count per feature is unknown (may silently use fewer samples than expected)
Degenerate Inputs

Features that produce undefined or unstable correlations:

TypeDescriptionEffect
Constant featuresAll values identical (variance = 0)Correlation undefined (division by zero)
Near-constant featuresVery low varianceCorrelation numerically unstable
Too few valid valuesAfter NaN removal, fewer than min_valid pairsStatistically unreliable
Single-value after filteringOnly one unique value remains post-NaN removalCorrelation undefined

Decision Framework

Do you have missing values (NaN) in your feature matrix?
├── No NaN at all → Bulk methods are safe (corrwith, np.corrcoef)
└── Yes, NaN present
    ├── Same NaN pattern across all features? → Listwise deletion is acceptable
    └── Different NaN patterns per feature (typical)
        ├── < 10,000 features → Per-feature loop with scipy.stats.spearmanr
        └── > 10,000 features → Parallelized per-feature loop (joblib)
ScenarioRecommended ApproachRationale
No missing dataDataFrame.corrwith()Fast, correct when no NaN
Sparse NaN, < 10K featuresPer-feature spearmanr loopCorrect pairwise deletion, acceptable speed
Sparse NaN, > 10K featuresParallelized per-feature loopSame correctness, scales with cores
Dense NaN (> 50% missing)Per-feature loop + strict min_validMany features will be skipped; report skip count
Uniform NaN patternListwise deletion + bulk methodIf all features share same NaN rows, pairwise = listwise

Best Practices

  1. Always print NaN summary before analysis: Report total NaN count, features with any NaN, and per-feature NaN distribution. This documents data quality and alerts you to severe missingness patterns.

  2. Use scipy.stats.spearmanr per feature in a loop: This is the only method that guarantees correct pairwise NaN removal for each feature independently.

  3. Set a minimum valid pair threshold (min_valid): Default to 10. Features with fewer valid pairs after NaN removal produce unreliable correlations and should be skipped with NaN.

  4. Filter degenerate inputs before computing correlations: Remove constant features, near-constant features, and features with excessive NaN before the correlation loop. This avoids undefined results and speeds up computation.

  5. Track n_valid per feature in the output: The number of valid pairs varies per feature. Report it alongside rho and p-value so downstream analysis can assess reliability.

  6. Report how many features were skipped or filtered: Silent feature loss is a common source of confusion. Always print the count of filtered degenerate features and skipped low-data features.

  7. Use parallelization for large datasets: For > 10,000 features, use joblib to distribute the per-feature loop across cores. The per-feature computation is embarrassingly parallel.

Show full SKILL.md (361 more words)Show less

Common Pitfalls

  1. Using bulk rank-then-correlate with NaN present: df.rank() followed by corrwith() silently mishandles NaN, producing incorrect correlations.

    • How to avoid: Always use scipy.stats.spearmanr per feature when NaN is present.
  2. Assuming uniform sample count across features: Different features have different NaN patterns, so each correlation is computed on a different number of samples.

    • How to avoid: Track and report n_valid for every feature.
  3. Not filtering degenerate inputs: Constant or near-constant features produce undefined correlations or divide-by-zero warnings that can silently corrupt results.

    • How to avoid: Run filter_degenerate() before the correlation loop.
  4. Using listwise deletion when NaN patterns differ: Listwise deletion removes any row with NaN in any feature, potentially discarding most of your data.

    • How to avoid: Use pairwise deletion (per-feature NaN removal).
  5. Ignoring the NaN summary step: Skipping the data quality report means you cannot verify whether the NaN pattern is severe enough to affect results.

    • How to avoid: Always print NaN summary before correlation computation.
  6. Setting min_valid too low: With fewer than ~10 valid pairs, Spearman correlation is unreliable and p-values are meaningless.

    • How to avoid: Use min_valid >= 10; increase for high-dimensional studies.

Workflow

  1. Step 1: Print NaN Summary

    • Report dataset shape, total NaN, features with any NaN, mean/max NaN per feature
    • Decision point: If > 50% NaN overall, reconsider data quality before proceeding
  2. Step 2: Filter Degenerate Features

    • Remove constant features (nunique < 3)
    • Remove features with excessive NaN (< 50% non-NaN)
    • Report count of removed features
  3. Step 3: Compute Per-Feature Correlations

    • Loop over features with scipy.stats.spearmanr
    • Apply pairwise NaN removal per feature
    • Skip features with < min_valid valid pairs (record as NaN)
  4. Step 4: Assemble and Report Results

    • Create DataFrame with rho, p-value, n_valid per feature
    • Report count of skipped features
    • Verify no silent data loss (input features = output features + skipped + filtered)
Reference Implementation
python
from scipy.stats import spearmanr
import numpy as np
import pandas as pd

def nan_summary(df):
    """Print NaN summary before correlation analysis."""
    print(f"Dataset shape: {df.shape}")
    print(f"Total NaN: {df.isna().sum().sum()}")
    print(f"Features with any NaN: {(df.isna().any()).sum()}")
    print(f"NaN per feature (mean): {df.isna().sum().mean():.1f}")
    print(f"NaN per feature (max): {df.isna().sum().max()}")

def filter_degenerate(df, min_unique=3, min_nonnan_frac=0.5):
    """Remove degenerate features before correlation analysis.

    Args:
        df: DataFrame (samples x features)
        min_unique: Minimum number of unique non-NaN values required
        min_nonnan_frac: Minimum fraction of non-NaN values required

    Returns:
        Filtered DataFrame, count of removed features
    """
    n_samples = len(df)
    keep = []
    for col in df.columns:
        values = df[col].dropna()
        if len(values) < n_samples * min_nonnan_frac:
            continue
        if values.nunique() < min_unique:
            continue
        keep.append(col)
    removed = len(df.columns) - len(keep)
    print(f"Filtered {removed} degenerate features out of {len(df.columns)}")
    return df[keep], removed

def pairwise_spearman(df_x, df_y, min_valid=10):
    """Compute per-feature Spearman correlation with pairwise NaN removal.

    Args:
        df_x: DataFrame (samples x features), aligned with df_y
        df_y: DataFrame (samples x features), same shape as df_x
        min_valid: Minimum number of valid (non-NaN) pairs required

    Returns:
        DataFrame with columns: rho, pvalue, n_valid
    """
    nan_summary(df_x)
    nan_summary(df_y)

    results = []
    for feature in df_x.columns:
        x = df_x[feature].values
        y = df_y[feature].values
        mask = ~(np.isnan(x) | np.isnan(y))
        n_valid = mask.sum()
        if n_valid < min_valid:
            results.append({'feature': feature, 'rho': np.nan,
                          'pvalue': np.nan, 'n_valid': n_valid})
            continue
        rho, pval = spearmanr(x[mask], y[mask])
        results.append({'feature': feature, 'rho': rho,
                       'pvalue': pval, 'n_valid': n_valid})

    result_df = pd.DataFrame(results).set_index('feature')
    skipped = result_df['rho'].isna().sum()
    if skipped > 0:
        print(f"Skipped {skipped} features with < {min_valid} valid pairs")
    return result_df
Anti-Patterns
python
# WRONG: Bulk rank-then-correlate
ranked_x = df_x.rank()
ranked_y = df_y.rank()
corrs = ranked_x.corrwith(ranked_y)

# WRONG: Bulk corrwith with method parameter
corrs = df_x.corrwith(df_y, method='spearman')

# WRONG: numpy corrcoef on ranked arrays (propagates NaN)
corrs = np.corrcoef(df_x.rank().values.T, df_y.rank().values.T)
Performance Optimization (> 10,000 features)
python
from joblib import Parallel, delayed

def parallel_spearman(df_x, df_y, min_valid=10, n_jobs=4):
    """Parallelized per-feature Spearman correlation."""
    def compute_one(feature):
        x = df_x[feature].values
        y = df_y[feature].values
        mask = ~(np.isnan(x) | np.isnan(y))
        n = mask.sum()
        if n < min_valid:
            return feature, np.nan, np.nan, n
        rho, pval = spearmanr(x[mask], y[mask])
        return feature, rho, pval, n

    results = Parallel(n_jobs=n_jobs)(
        delayed(compute_one)(f) for f in df_x.columns
    )
    return pd.DataFrame(
        results, columns=['feature', 'rho', 'pvalue', 'n_valid']
    ).set_index('feature')

Further Reading

  • statistical-analysis -- General statistical test selection and assumption checking
  • degenerate-input-filtering -- Broader guide on filtering uninformative data before any statistical test
  • scikit-learn-machine-learning -- Feature selection and preprocessing pipelines

© jaechang-hits, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/scientific-computing/nan-safe-correlation of jaechang-hits/SciAgent-Skills.

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Nan Safe Correlation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nan Safe Correlation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nan Safe Correlation this skilljaechang-hits/SciAgent-Skills3741 repos~2.9kAutomated safety check: PassCC-BY-4.0
Sae Feature Annotationssoftnanolab/bagel148—~1.5kAutomated safety check: PassMIT
Code EngineeropenJiuwen-ai/sciencediscovery159—~2.8kAutomated safety check: PassApache-2.0
scikit-survival Time-to-Event Modelingdavila7/claude-code-templates33k11 repos~3.7kAutomated safety check: PassMIT
SHAP Model Explainabilitydavila7/claude-code-templates33k11 repos~4.6kAutomated safety check: PassMIT
Data Analysisxiaoyuge886/aigc198—~794Automated safety check: PassMIT

Similar skills

  • Sae Feature Annotations

    softnanolab/bagel

    Look up what a Biohub ESM-C sparse-autoencoder (SAE) feature means — its label, description, top-activating proteins, decoder neighbours, and activation statistics — by querying the Biohub…

    148 GitHub stars~1.5k tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed
  • Code Engineer

    openJiuwen-ai/sciencediscovery

    A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…

    159 GitHub stars~2.8k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • scikit-survival Time-to-Event Modeling

    davila7/claude-code-templates

    Fits and evaluates survival models with scikit-survival: Cox models, Random Survival Forests, boosting, survival SVMs, concordance index, Brier score and competing risks.

    33k GitHub starsUsed in 11 repos~3.7k tokens
    Data & AnalyticsAuto-check passed
  • SHAP Model Explainability

    davila7/claude-code-templates

    Explains machine learning predictions with SHAP: picking the right explainer, computing Shapley values and drawing waterfall, beeswarm, bar and force plots.

    33k GitHub starsUsed in 11 repos~4.6k tokens
    Data & AnalyticsAuto-check passed
  • Data Analysis

    xiaoyuge886/aigc

    Perform data analysis tasks including data cleaning, statistical analysis, visualization, and insight generation.

    198 GitHub stars~794 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Data Profiler

    aspi6246/Claude-Code-Skills-for-Academics

    Systematic dataset profiling protocol for empirical research.

    159 GitHub stars~2k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Questions about Nan Safe Correlation

What does Nan Safe Correlation do?

Per-feature NaN-safe Spearman/Pearson correlation across many features (genes, proteins, variants) with missing values. Nan Safe Correlation is an agent skill from jaechang-hits/SciAgent-Skills. Per-feature NaN-safe Spearman/Pearson correlation across many features (genes, proteins, variants) with missing values.

When should I use Nan Safe Correlation?

Nan Safe Correlation fits situations like: tasks that involve Machine learning; tasks that involve Data cleaning; tasks that involve Statistics.

How do I install Nan Safe Correlation in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill nan-safe-correlation -a claude-code`. Or copy the skill folder (skills/scientific-computing/nan-safe-correlation in jaechang-hits/SciAgent-Skills) into .claude/skills/nan-safe-correlation in your project. Claude Code loads it when a task matches its description.

How do I install Nan Safe Correlation in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill nan-safe-correlation -a codex`. Or copy the skill folder (skills/scientific-computing/nan-safe-correlation in jaechang-hits/SciAgent-Skills) into .agents/skills/nan-safe-correlation in your project. Codex loads it when a task matches its description.

Can I use Nan Safe Correlation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill nan-safe-correlation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nan-safe-correlation, .gemini/skills/nan-safe-correlation, .github/skills/nan-safe-correlation and .opencode/skills/nan-safe-correlation in your project.

What does Nan Safe Correlation need to run?

SKILL.md names no scripts, command-line tools or credentials: Nan Safe Correlation is instructions for the agent only. Our summary lists: Python 3.

Does Nan Safe Correlation access the network?

SKILL.md names 3 domains. As links in the text: docs.scipy.org, pandas.pydata.org and stefvanbuuren.name. This is read from the text; nothing was executed.

Is Nan Safe Correlation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nan Safe Correlation use?

Nan Safe Correlation is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nan Safe Correlation use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nan Safe Correlation?

Skills that share tags, products or a category with Nan Safe Correlation: Sae Feature Annotations (softnanolab/bagel, 148 stars), Code Engineer (openJiuwen-ai/sciencediscovery, 159 stars), scikit-survival Time-to-Event Modeling (davila7/claude-code-templates, 33k stars) and SHAP Model Explainability (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nan Safe Correlation?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.