Agent skill

Scientific Data Preprocessing

by foryourhealth111-pixel in foryourhealth111-pixel/Vibe-Skills

⚠️ CRITICAL USER EXPERIENCE-BASED SKILL - ALWAYS CONSULT BEFORE DATA PREPROCESSING ⚠️ Prevents catastrophic errors (88.9% error rate in V1.0 case study) through multi-level feature analysis, data…

Apache-2.0Auto-check passedData & Analytics

Install Scientific Data Preprocessing

skills CLI
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install foryourhealth111-pixel/Vibe-Skills scientific-data-preprocessing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/bundled/skills/scientific-data-preprocessing .claude/skills/scientific-data-preprocessing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scientific-data-preprocessing
GitHub stars
3.6k
Token cost
~5k tokens
SKILL.md length
824 words
Files
8 (incl. references)
Skills in repo
65
Repo updated
First seen
Licence
Apache-2.0

At a glance

⚠️ CRITICAL USER EXPERIENCE-BASED SKILL - ALWAYS CONSULT BEFORE DATA PREPROCESSING ⚠️ Prevents catastrophic errors (88.9% error rate in V1.0 case study) through multi-level feature analysis, data…

  • Works in 3 steps: Confirmation that data has groups (e.g.,… → Understanding of whether goal is… → Domain constraints on data ranges/units
  • Tasks that involve Machine learning
  • SKILL.md covers Core Mission, When to Use This Skill, Not For / Boundaries and Quick Reference, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Scientific Data Preprocessing is an agent skill from foryourhealth111-pixel/Vibe-Skills. ⚠️ CRITICAL USER EXPERIENCE-BASED SKILL - ALWAYS CONSULT BEFORE DATA PREPROCESSING ⚠️ Prevents catastrophic errors (88.9% error rate in V1.0 case study) through multi-level feature analysis, data leakage detection, and semantic validation. MANDATORY for: data preprocessing, feature engineering, standardization, normalization, interpolation, missing value handling, feature selection, or ANY data transformation task. Covers grouped time-series, cross-sectional, panel data. Detects: time travel leakage, causal…

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `README.md`, `UPDATE_LOG_V1.1.md` and `references/ai-common-pitfalls.md`).

It sits in Data & Analytics, covering Machine learning and Data cleaning. The repository describes itself as: Intelligent Skill routing and workflow orchestration for AI agents — +21.12 pp reward, −29.6% tokens on SkillsBench with DeepSeekV4Flash-VE. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Machine learning
  • Tasks that involve Data cleaning

Example prompts

  • “/scientific-data-preprocessing”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Confirmation that data has groups (e.g., match_id, patient_id, session_id)
  2. Understanding of whether goal is within-group (relative) or cross-group (absolute) comparison
  3. Domain constraints on data ranges/units

What it can do on your machine

Read from SKILL.md and the folder at commit ddcaa2a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scientific Data Preprocessing loads about 5k tokens when it runs, and up to ~21k if it reads all its reference files. Until then it costs about 167 tokens; SKILL.md has 824 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~167
When it runs · the whole SKILL.md, loaded when a task matches
~5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~21k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from foryourhealth111-pixel/Vibe-Skills at commit ddcaa2a, republished under its Apache-2.0 licence (© foryourhealth111-pixel). 824 words, ~4,956 tokens.

Download SKILL.mdSave it as .claude/skills/scientific-data-preprocessing/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
scientific-data-preprocessing
description
⚠️ CRITICAL USER EXPERIENCE-BASED SKILL - ALWAYS CONSULT BEFORE DATA PREPROCESSING ⚠️ Prevents catastrophic errors (88.9% error rate in V1.0 case study) through multi-level feature analysis, data leakage detection, and semantic validation. MANDATORY for: data preprocessing, feature engineering, standardization, normalization, interpolation, missing value handling, feature selection, or ANY data transformation task. Covers grouped time-series, cross-sectional, panel data. Detects: time travel leakage, causal inversion, ID misuse, semantic-numeric fallacies, distribution blindness. User's hard-won lessons from real project failures.

Scientific Data Preprocessing Skill

⚠️ CRITICAL: USER'S HARD-WON EXPERIENCE - MANDATORY CONSULTATION ⚠️

This skill encapsulates painful lessons learned from real preprocessing disasters (88.9% error rate documented). ALWAYS use this skill for planning, reflection, and validation when ANY data preprocessing is involved.

Why this skill is mandatory:

  • Based on actual project failures (V1.0, V2.0 case studies)
  • Prevents data leakage that causes production disasters
  • Catches semantic errors AI agents commonly make
  • Saves weeks of debugging and model retraining

When to invoke (DO NOT SKIP):

  • ✅ Before starting ANY data preprocessing task
  • ✅ During preprocessing for reflection and validation
  • ✅ After preprocessing for comprehensive audit
  • ✅ When reviewing AI-generated preprocessing code

Core Mission

Prevent catastrophic preprocessing errors in grouped time-series data by applying multi-level feature analysis and respecting data structure boundaries.

When to Use This Skill

MANDATORY consultation - trigger immediately when:

Data Preprocessing Tasks (ALWAYS)
  • Any data cleaning, transformation, or preparation work
  • Loading and preparing data for modeling
  • Creating training/test splits
  • Handling missing values (imputation, deletion)
  • Feature scaling/normalization/standardization
  • Encoding categorical variables
  • Feature engineering or construction
  • Feature selection or dimensionality reduction
Data Structure Types (ALWAYS)
  • Preprocesssing time-series data with natural groupings (matches, sessions, patients, experiments)
  • Sports analytics (tennis, basketball, etc.)
  • Medical/clinical data with patient groupings
  • Panel data or longitudinal studies
  • Any grouped/hierarchical data structure
Quality Assurance (ALWAYS)
  • Auditing existing preprocessing for data leakage or semantic errors
  • Reviewing AI-generated preprocessing code for common pitfalls
  • Validating preprocessing before model training
  • Debugging unexpected model performance
Critical Checkpoints (NEVER SKIP)
  • ✅ BEFORE: Planning preprocessing strategy
  • ✅ DURING: Reflecting on decisions and checking for errors
  • ✅ AFTER: Comprehensive validation and audit

Trigger keywords that MUST invoke this skill:

  • "preprocess", "preprocessing", "data cleaning", "data preparation"
  • "standardize", "normalize", "scale", "transform"
  • "impute", "fill missing", "handle NaN"
  • "encode", "one-hot", "categorical"
  • "feature engineering", "feature selection", "feature construction"
  • "train test split", "cross validation split"
  • "interpolate", "smooth", "aggregate"

Not For / Boundaries

This skill does NOT:

  • Handle purely cross-sectional data (ungrouped, single timepoint)
  • Make domain-specific feature engineering decisions (you decide business logic)
  • Choose ML models (focuses on preprocessing only)
  • Handle distributed/big data infrastructure (assumes data fits in memory)

Required inputs before proceeding:

  1. Confirmation that data has groups (e.g., match_id, patient_id, session_id)
  2. Understanding of whether goal is within-group (relative) or cross-group (absolute) comparison
  3. Domain constraints on data ranges/units

Quick Reference

Multi-Level Feature Analysis Framework

Level 1: Data Type

python
# Check data types
df.dtypes  # int64, float64, object, etc.

Level 2: Feature Type Classification

python
# Binary (0/1)
binary_features = [col for col in df.columns if df[col].nunique() == 2]

# Categorical (finite discrete values)
categorical_features = [col for col in df.select_dtypes(include='object').columns]

# Continuous (infinite possible values)
continuous_features = [col for col in df.select_dtypes(include=['float64', 'int64']).columns
                       if df[col].nunique() > 10]

Level 3: Data Structure

python
# Check for grouping
print(f"Number of groups: {df['group_id'].nunique()}")
print(f"Avg points per group: {df.groupby('group_id').size().mean():.1f}")

# Check for time-series
df_sorted = df.sort_values(['group_id', 'timestamp'])

Level 4: Physical Meaning

python
# Validate physical ranges
assert df['speed_mph'].max() < 200, "Speed exceeds physical limit"
assert df['distance_meters'].min() >= 0, "Negative distance impossible"
Critical Processing Decision Tree
python
# Decision: Within-group or global processing?
def choose_processing_scope(data, feature, goal):
    """
    goal = 'relative' → within-group (e.g., "this point was intense FOR THIS MATCH")
    goal = 'absolute' → global (e.g., "this was an intense point OVERALL")
    """
    if goal == 'relative':
        return 'within_group'
    elif goal == 'absolute':
        return 'global'
    else:
        raise ValueError("Goal must be 'relative' or 'absolute'")
Pattern 1: Within-Group Interpolation (CORRECT)
python
from scipy.interpolate import CubicSpline
import numpy as np

# ✅ CORRECT: Interpolate within each group
for group_id in df['match_id'].unique():
    mask = df['match_id'] == group_id
    group_data = df.loc[mask, 'speed_mph'].copy()

    # Get valid (non-NaN) indices
    valid_idx = group_data.notna()
    valid_positions = np.where(valid_idx)[0]
    valid_values = group_data[valid_idx].values

    if len(valid_positions) >= 4:
        cs = CubicSpline(valid_positions, valid_values)
        missing_positions = np.where(~valid_idx)[0]
        df.loc[mask & ~valid_idx, 'speed_mph'] = cs(missing_positions)
Pattern 2: Global Interpolation (WRONG - Don't Do This)
python
# ❌ WRONG: Cross-group interpolation
# This interpolates between match A's last point and match B's first point!
cs = CubicSpline(
    np.where(df['speed_mph'].notna())[0],  # ❌ All indices globally
    df['speed_mph'].dropna().values
)
df.loc[df['speed_mph'].isna(), 'speed_mph'] = cs(
    np.where(df['speed_mph'].isna())[0]
)
Pattern 3: Within-Group Standardization (for Relative Analysis)
python
from sklearn.preprocessing import StandardScaler

# ✅ CORRECT: Standardize within each match
for match_id in df['match_id'].unique():
    mask = df['match_id'] == match_id
    scaler = StandardScaler()

    df.loc[mask, 'distance_run_std_within'] = scaler.fit_transform(
        df.loc[mask, [['distance_run']]
    )

# Interpretation: z=+2 means "2 std above average FOR THIS MATCH"
Pattern 4: Global Standardization (for Absolute Comparison)
python
# ✅ CORRECT: Global standardization (when appropriate)
scaler = StandardScaler()
df['distance_run_std_global'] = scaler.fit_transform(df[['distance_run']])

# Interpretation: z=+2 means "2 std above average ACROSS ALL MATCHES"
Pattern 5: Feature Type Processing Rules
python
# Binary variables (0/1) - KEEP AS-IS
binary_cols = ['is_ace', 'is_winner', 'is_error']
# ❌ NEVER standardize these! They have semantic meaning as 0/1

# Categorical variables - ONE-HOT ENCODE
df_encoded = pd.get_dummies(df, columns=['server', 'serve_number'], dtype=int)

# Continuous variables - STANDARDIZE (within-group or global)
continuous_cols = ['distance_run', 'rally_count', 'speed_mph']
# ✅ Apply pattern 3 or 4 based on goal
Pattern 6: Sliding Window Features (for Momentum)
python
# ✅ CORRECT: Sliding window for momentum analysis
window = 10

df['win_rate_last10'] = df.groupby('match_id')['point_won'].transform(
    lambda x: x.rolling(window, min_periods=1).mean()
)

# ❌ WRONG: Cumulative features (loses temporal locality)
df['cumulative_points_won'] = df.groupby('match_id')['point_won'].cumsum()
# This just increases monotonically and correlates with point_number
Pattern 7: Data Quality Validation
python
def validate_data_quality(df, feature, expected_range):
    """Validate before processing"""
    # Check range
    assert df[feature].min() >= expected_range[0], f"{feature} below minimum"
    assert df[feature].max() <= expected_range[1], f"{feature} above maximum"

    # Check for anomalies
    mean = df[feature].mean()
    std = df[feature].std()

    if std > mean:
        print(f"⚠️ WARNING: {feature} has std > mean (highly skewed or errors)")

    # Check missing pattern
    missing_by_group = df.groupby('match_id')[feature].apply(lambda x: x.isna().sum())
    if missing_by_group.max() > len(df) / df['match_id'].nunique() * 0.5:
        print(f"⚠️ WARNING: {feature} has >50% missing in some groups")

# Example
validate_data_quality(df, 'speed_mph', expected_range=(50, 165))
Pattern 8: Detect Processing Scope Automatically
python
def detect_processing_scope(df, group_col, feature_col):
    """
    Recommend within-group vs global based on variance structure
    """
    # Calculate variance components
    within_group_var = df.groupby(group_col)[feature_col].var().mean()
    global_var = df[feature_col].var()

    # Intraclass correlation
    between_group_var = global_var - within_group_var
    icc = between_group_var / global_var

    if icc > 0.5:
        return 'within_group', f"High between-group variance (ICC={icc:.2f})"
    else:
        return 'global', f"Low between-group variance (ICC={icc:.2f})"

scope, reason = detect_processing_scope(df, 'match_id', 'distance_run')
print(f"Recommended: {scope} - {reason}")
Pattern 9: Data Leakage Detection
python
def detect_data_leakage(df, target_col, feature_cols, id_cols):
    """
    Critical checks for data leakage and AI common pitfalls
    """
    issues = []

    # 1. ID Leakage: High cardinality variables as features
    for col in feature_cols:
        if col in id_cols:
            issues.append(f"❌ FATAL: {col} is an ID - NEVER use as feature")
            continue

        # Check if looks like ID (>50% unique)
        uniqueness = df[col].nunique() / len(df)
        if uniqueness > 0.5:
            issues.append(f"⚠️ {col}: {uniqueness*100:.1f}% unique - possible ID leakage")

    # 2. Causal Inversion: Perfect correlation with target
    for col in feature_cols:
        if col == target_col:
            continue
        if df[col].dtype in ['int64', 'float64']:
            corr = abs(df[[col, target_col]].corr().iloc[0, 1])
            if corr > 0.95:
                issues.append(f"❌ FATAL: {col} correlation={corr:.3f} - likely consequence of target!")

    # 3. Meaningless Numeric: Codes treated as numbers
    for col in feature_cols:
        if df[col].dtype in ['int64', 'float64']:
            # Pattern: High values, many uniques, looks like code
            if df[col].min() > 1000 and df[col].nunique() > 100:
                issues.append(f"⚠️ {col}: Looks like code (zipcode/ID) - should be categorical")

    # 4. Time Travel: Check if standardization used global statistics
    # (Requires knowing if train/test split was done first)

    # Print report
    if issues:
        print("="*60)
        print("DATA LEAKAGE AUDIT")
        print("="*60)
        for issue in issues:
            print(issue)
        print("="*60)
    else:
        print("✅ No obvious leakage detected")

    return issues

# Example usage
issues = detect_data_leakage(
    df,
    target_col='point_won',
    feature_cols=['speed_mph', 'user_id', 'distance_run'],
    id_cols=['match_id', 'user_id']
)
Pattern 10: Distribution-Aware Scaling
python
from scipy.stats import skew, kurtosis
from sklearn.preprocessing import StandardScaler, RobustScaler

def smart_scaler_selection(df, col):
    """
    Choose scaler based on distribution characteristics
    """
    data = df[col].dropna()

    # Check distribution
    skewness = skew(data)
    kurt = kurtosis(data)

    print(f"{col}: skewness={skewness:.2f}, kurtosis={kurt:.2f}")

    if abs(skewness) < 0.5 and abs(kurt) < 3:
        # Roughly normal
        print("  → StandardScaler (data is roughly normal)")
        return StandardScaler(), None

    elif skewness > 1:
        # Right-skewed (long tail)
        print("  → Log transform + StandardScaler (right-skewed)")
        return StandardScaler(), 'log'

    else:
        # Heavy outliers or non-normal
        print("  → RobustScaler (heavy outliers)")
        return RobustScaler(), None

# Example usage
for col in continuous_features:
    scaler, transform = smart_scaler_selection(df, col)

    if transform == 'log':
        df[f'{col}_log'] = np.log1p(df[col])
        df[f'{col}_scaled'] = scaler.fit_transform(df[[f'{col}_log']])
    else:
        df[f'{col}_scaled'] = scaler.fit_transform(df[[col]])

Examples

Example 1: Tennis Match Preprocessing (Complete Pipeline)

Input:

  • CSV with 7,284 rows, 31 matches
  • Features: speed_mph, distance_run, rally_count, is_ace, server
  • Goal: Analyze momentum (relative intensity within each match)

Steps:

python
import pandas as pd
from sklearn.preprocessing import StandardScaler

# 1. Load and inspect
df = pd.read_csv('tennis_data.csv')
print(f"Matches: {df['match_id'].nunique()}")
print(f"Features: {df.dtypes}")

# 2. Classify features
binary_features = ['is_ace', 'is_winner', 'is_break_point']
categorical_features = ['server', 'serve_number']
continuous_features = ['distance_run', 'speed_mph', 'rally_count']

# 3. Validate data quality
for feat in continuous_features:
    print(f"\n{feat}:")
    print(df[feat].describe())
    # Check for impossible values
    if feat == 'speed_mph':
        assert df[feat].max() < 170, "Speed exceeds world record!"

# 4. Handle missing values (within-group)
for match_id in df['match_id'].unique():
    mask = df['match_id'] == match_id
    for feat in continuous_features:
        if df.loc[mask, feat].isna().any():
            # Simple linear interpolation within match
            df.loc[mask, feat] = df.loc[mask, feat].interpolate(method='linear')

# 5. One-hot encode categorical
df = pd.get_dummies(df, columns=categorical_features, dtype=int)

# 6. Standardize continuous features WITHIN each match
for feat in continuous_features:
    df[f'{feat}_std'] = np.nan
    for match_id in df['match_id'].unique():
        mask = df['match_id'] == match_id
        scaler = StandardScaler()
        df.loc[mask, f'{feat}_std'] = scaler.fit_transform(
            df.loc[mask, [[feat]]
        )

# 7. Create sliding window features
window = 10
df['win_rate_last10'] = df.groupby('match_id')['point_won'].transform(
    lambda x: x.rolling(window, min_periods=1).mean()
)

# 8. KEEP binary features as 0/1 (don't transform!)
# binary_features are already correct

print("\n✅ Preprocessing complete!")
print(f"Final shape: {df.shape}")
print(f"Standardized features: {[f for f in df.columns if f.endswith('_std')]}")

Expected output:

  • Binary features remain 0/1
  • Categorical features one-hot encoded (e.g., server_1, server_2)
  • Continuous features have both original and _std versions
  • _std features have mean≈0, std≈1 WITHIN each match
  • Sliding window features capture local momentum
  • No missing values
Show full SKILL.md (305 more words)Show less
Example 2: Detecting Cross-Group Contamination

Input:

  • Preprocessed data where you suspect cross-group standardization

Steps:

python
# Check if standardization was done correctly
def check_within_group_standardization(df, group_col, feature_std_col):
    """
    Verify that standardized feature has mean≈0, std≈1 within each group
    """
    results = df.groupby(group_col)[feature_std_col].agg(['mean', 'std'])

    # Within-group standardization: each group should have mean≈0, std≈1
    if (results['mean'].abs() < 0.1).all() and (results['std'].between(0.9, 1.1)).all():
        print("✅ CORRECT: Within-group standardization detected")
        return True

    # Global standardization: groups will have varying means and stds
    else:
        print("❌ WRONG: Global standardization detected!")
        print("Group means:", results['mean'].values[:5])
        print("Group stds:", results['std'].values[:5])
        return False

check_within_group_standardization(df, 'match_id', 'distance_run_std')

Expected output:

  • CORRECT: All group means ≈ 0, all group stds ≈ 1
  • WRONG: Group means vary widely, indicating global standardization
Example 3: Fixing Cumulative Feature Error

Input:

  • Existing pipeline using cumulative sums for momentum

Steps:

python
# ❌ WRONG approach (existing code)
df['cumulative_wins'] = df.groupby('match_id')['point_won'].cumsum()

# Problem: This just counts total wins up to this point
# Doesn't capture recent momentum!

# ✅ CORRECT approach (fix)
# Replace cumulative with sliding window
window = 10
df['recent_win_rate'] = df.groupby('match_id')['point_won'].transform(
    lambda x: x.rolling(window, min_periods=1).mean()
)

# Compare
print("Cumulative (wrong):", df['cumulative_wins'].values[50:60])
print("Sliding window (correct):", df['recent_win_rate'].values[50:60])

# Cumulative: [25, 26, 26, 27, 28, ...] - monotonic
# Sliding window: [0.6, 0.7, 0.5, 0.6, ...] - fluctuates with momentum

Expected output:

  • Cumulative features removed
  • Sliding window features show local variations
  • Momentum analysis now captures short-term trends

References

  • references/index.md: Navigation and overview
  • references/error-case-studies.md: Real-world preprocessing disasters from tennis data
  • references/decision-trees.md: Full decision trees for all preprocessing choices
  • references/validation-checklist.md: Pre-processing validation checklist
  • references/ai-common-pitfalls.md: AI-specific errors (data leakage, semantic fallacies, distribution blindness)

Maintenance

⚠️ CRITICAL NOTICE: USER'S PERSONAL EXPERIENCE-BASED SKILL ⚠️

This skill is NOT theoretical - it's based on real project failures:

  • V1.0 disaster: 88.9% error rate, weeks of wasted work
  • V2.0 issues: Cross-group contamination, unreliable results
  • V3.0 success: All errors fixed, production-ready

Why this matters to you (Claude):

  • These are the EXACT errors AI agents commonly make
  • User has already paid the price for these mistakes
  • Ignoring this skill = repeating documented failures
  • Following this skill = learning from experience without pain

Authority level: HIGHEST

  • Based on user's hard-won lessons from actual project
  • Validated through multiple iterations (V1.0 → V2.0 → V3.0)
  • Every error documented with impact metrics
  • Every fix validated with comprehensive testing

Sources:

  • Primary: User's personal project (2024 MCM Problem C - Tennis Momentum Analysis)
  • Secondary: Statistical best practices for grouped data
  • Tertiary: Common AI preprocessing errors observed across domains

Mandatory consultation:

  • ⚠️ ALWAYS consult before, during, and after any data preprocessing
  • ⚠️ NEVER skip validation steps outlined in this skill
  • ⚠️ When in doubt, err on the side of caution (use this skill)

Last updated: 2026-01-18 (V1.1)

Known limits:

  • Assumes data fits in memory (not for big data infrastructure)
  • Focused on numeric/categorical features (text/image preprocessing partially covered)
  • Does not prescribe domain-specific feature engineering (user decides business logic)
  • Requires basic understanding of statistics (mean, std, correlation)

© foryourhealth111-pixel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in bundled/skills/scientific-data-preprocessing of foryourhealth111-pixel/Vibe-Skills.

  • SKILL.md
  • README.md
  • UPDATE_LOG_V1.1.md
  • references/ai-common-pitfalls.md
  • references/decision-trees.md
  • references/error-case-studies.md
  • references/index.md
  • references/validation-checklist.md

Open the folder on GitHubat commit ddcaa2a

Compare with similar skills

Scientific Data Preprocessing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scientific Data Preprocessing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scientific Data Preprocessing this skillforyourhealth111-pixel/Vibe-Skills3.6k—~5kAutomated safety check: PassApache-2.0
Splitting Datasetsjeremylongshore/tons-of-skills-marketplace2.8k1 repos~836Automated safety check: PassMIT
Sap Hana Cloud Data Intelligencesecondsky/sap-skills462—~3.2kAutomated safety check: PassGPL-3.0
Rf Model Importance Analysisaipoch/medical-research-skills2k—~2.7kAutomated safety check: PassMIT
Data Cleanbrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~1.1kAutomated safety check: PassCustom licence
Nan Safe Correlationjaechang-hits/SciAgent-Skills3701 repos~2.9kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Splitting Datasets

    jeremylongshore/tons-of-skills-marketplace

    Process split datasets into training, validation, and testing sets for ML model development.

    2.8k GitHub starsUsed in 1 repo~836 tokens
    Data & AnalyticsAuto-check passed
  • Develops data processing pipelines, integrations, and machine learning scenarios in SAP Data Intelligence Cloud.

    462 GitHub stars~3.2k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Rf Model Importance Analysis

    aipoch/medical-research-skills

    A skill your agent uses when you need a standardized R CLI workflow to train a two-class random forest model from an expression-like feature matrix, rank variable importance, and generate…

    2k GitHub stars~2.7k tokensUpdated 21 days ago
    Data & AnalyticsAuto-check passed
  • Data Clean

    brycewang-stanford/Auto-Empirical-Research-Skills

    Produce documented data cleaning scripts that log every transformation with N before/after each step, generate a CONSORT-style exclusion flow diagram, create decision log entries for every…

    4.5k GitHub stars~1.1k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Nan Safe Correlation

    jaechang-hits/SciAgent-Skills

    Per-feature NaN-safe Spearman/Pearson correlation across many features (genes, proteins, variants) with missing values.

    370 GitHub starsUsed in 1 repo~2.9k tokens
    Data & AnalyticsAuto-check passed
  • Empirical Analysis Skill Python

    Drchronx/ai-agent-research-starter-kit

    Parameterized Python empirical-analysis and machine-learning workflow for applied economics, public health epidemiology, supervised ML, and ML causal inference.

    135 GitHub stars~3k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed

More from foryourhealth111-pixel/Vibe-Skills

All 65 skills in this repo
  • Market Research Reports

    foryourhealth111-pixel/Vibe-Skills

    Produces long consulting-style market research and industry reports covering market sizing, competitive landscape, market entry and investment theses.

    3.6k GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check: notes
  • Academic Venue Templates

    foryourhealth111-pixel/Vibe-Skills

    Supplies venue-specific LaTeX templates and formatting rules for journals, conferences and posters, and checks a manuscript against page limits and submission requirements.

    3.6k GitHub stars~3.9k tokensUpdated 1 mo ago
    Auto-check: notes
  • Digital Brain

    foryourhealth111-pixel/Vibe-Skills

    This skill should be used when the user asks to "write a post", "check my voice", "look up contact", "prepare for meeting", "weekly review", "track goals", or mentions personal brand, content…

    3.6k GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Smart File Writer

    foryourhealth111-pixel/Vibe-Skills

    Diagnoses why a file write failed (permissions, disk space, path length, locks, read-only mounts) before retrying, instead of repeating the same call blindly.

    3.6k GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Automated Video Studio

    foryourhealth111-pixel/Vibe-Skills

    Turns footage, audio and a storyboard plan into a finished short video with FFmpeg jump-cuts, subtitle burn-in and a final polish pass.

    3.6k GitHub stars~838 tokensUpdated 1 mo ago
    Auto-check passed
  • Citation Management

    foryourhealth111-pixel/Vibe-Skills

    Turns DOIs, PMIDs and arXiv IDs into clean BibTeX, searches Google Scholar and PubMed, and checks and deduplicates a reference list.

    3.6k GitHub stars~7.6k tokensUpdated 1 mo ago
    Auto-check: notes

Questions about Scientific Data Preprocessing

What does Scientific Data Preprocessing do?

⚠️ CRITICAL USER EXPERIENCE-BASED SKILL - ALWAYS CONSULT BEFORE DATA PREPROCESSING ⚠️ Prevents catastrophic errors (88.9% error rate in V1.0 case study) through multi-level feature analysis, data…. Scientific Data Preprocessing is an agent skill from foryourhealth111-pixel/Vibe-Skills.0 case study) through multi-level feature analysis, data leakage detection, and semantic validation.

When should I use Scientific Data Preprocessing?

Scientific Data Preprocessing fits situations like: tasks that involve Machine learning; tasks that involve Data cleaning.

How do I install Scientific Data Preprocessing in Claude Code?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a claude-code`. Or copy the skill folder (bundled/skills/scientific-data-preprocessing in foryourhealth111-pixel/Vibe-Skills) into .claude/skills/scientific-data-preprocessing in your project. Claude Code loads it when a task matches its description.

How do I install Scientific Data Preprocessing in Codex?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a codex`. Or copy the skill folder (bundled/skills/scientific-data-preprocessing in foryourhealth111-pixel/Vibe-Skills) into .agents/skills/scientific-data-preprocessing in your project. Codex loads it when a task matches its description.

Can I use Scientific Data Preprocessing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scientific-data-preprocessing, .gemini/skills/scientific-data-preprocessing, .github/skills/scientific-data-preprocessing and .opencode/skills/scientific-data-preprocessing in your project.

What does Scientific Data Preprocessing need to run?

SKILL.md names no scripts, command-line tools or credentials: Scientific Data Preprocessing is instructions for the agent only. Our summary lists: Python 3.

Does Scientific Data Preprocessing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scientific Data Preprocessing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scientific Data Preprocessing use?

Scientific Data Preprocessing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scientific Data Preprocessing use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.

What are the alternatives to Scientific Data Preprocessing?

Skills that share tags, products or a category with Scientific Data Preprocessing: Splitting Datasets (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Sap Hana Cloud Data Intelligence (secondsky/sap-skills, 462 stars), Rf Model Importance Analysis (aipoch/medical-research-skills, 2k stars) and Data Clean (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scientific Data Preprocessing?

foryourhealth111-pixel (a GitHub user) maintains it in foryourhealth111-pixel/Vibe-Skills, which has 3,604 GitHub stars. The repository holds 65 skills in this directory. The repository was last updated on August 31, 2026.

Source: foryourhealth111-pixel/Vibe-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.