Agent skill

ML Data Leakage Guard

by foryourhealth111-pixel in foryourhealth111-pixel/Vibe-Skills

Detects and prevents data leakage in machine learning and mathematical modeling.

Apache-2.0Auto-check passedData & Analytics

Install ML Data Leakage Guard

skills CLI
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill ml-data-leakage-guard -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install foryourhealth111-pixel/Vibe-Skills ml-data-leakage-guard --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/bundled/skills/ml-data-leakage-guard .claude/skills/ml-data-leakage-guard && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-data-leakage-guard
GitHub stars
3.6k
Token cost
~3.4k tokens
SKILL.md length
742 words
Files
5 (incl. references)
Skills in repo
81
Repo updated
First seen
Licence
Apache-2.0

At a glance

Detects and prevents data leakage in machine learning and mathematical modeling.

  • Works in 8 steps: Fit-Transform Order: Did I fit any… → Global Statistics: Did I compute any… → Feature Selection: Did I select features… → …
  • Tasks that involve Machine learning
  • SKILL.md covers When to Use This Skill, Not For / Boundaries, Core Principle and Quick Reference, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

ML Data Leakage Guard is an agent skill from foryourhealth111-pixel/Vibe-Skills. Detects and prevents data leakage in machine learning and mathematical modeling. Use after ML tasks involving data cleaning, feature engineering, data augmentation, algorithm development, normalization, missing value imputation, dimensionality reduction, feature selection, or time series modeling. Checks if features/statistics would be available at prediction time.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/detection-strategies.md`, `references/index.md` and `references/leakage-patterns.md`).

It sits in Data & Analytics, covering Machine learning, Data cleaning and Database schema design. The repository describes itself as: Intelligent Skill routing and workflow orchestration for AI agents — +21.12 pp reward, −29.6% tokens on SkillsBench with DeepSeekV4Flash-VE. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Machine learning
  • Tasks that involve Data cleaning
  • Tasks that involve Database schema design

Example prompts

  • “Use the ml-data-leakage-guard skill to detect and prevents data leakage in machine learning and mathematical modeling”
  • “/ml-data-leakage-guard”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Fit-Transform Order: Did I fit any transformer (scaler, encoder, imputer, PCA) on the full dataset before splitting?
  2. Global Statistics: Did I compute any statistics (mean, std, median, mode) using the entire dataset?
  3. Feature Selection: Did I select features using information from the test set?
  4. Target Information: Did any feature encoding use target values from the test set?
  5. Temporal Order: For time series, did I use random splits or include future information?
  6. Post-Event Features: Are any features only available after the outcome occurs?
  7. Cross-Validation: Did I preprocess before setting up CV folds?
  8. Production Reality: At prediction time, can I actually compute this feature with available data?

What it can do on your machine

Read from SKILL.md and the folder at commit ddcaa2a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Data Leakage Guard loads about 3.4k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 97 tokens; SKILL.md has 742 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from foryourhealth111-pixel/Vibe-Skills at commit ddcaa2a, republished under its Apache-2.0 licence (© foryourhealth111-pixel). 742 words, ~3,358 tokens.

Download SKILL.mdSave it as .claude/skills/ml-data-leakage-guard/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
ml-data-leakage-guard
description
Detects and prevents data leakage in machine learning and mathematical modeling. Use after ML tasks involving data cleaning, feature engineering, data augmentation, algorithm development, normalization, missing value imputation, dimensionality reduction, feature selection, or time series modeling. Checks if features/statistics would be available at prediction time.

ML Data Leakage Guard Skill

Automatically detects and prevents data leakage in machine learning workflows by verifying that all preprocessing steps, feature engineering, and statistical computations would be available at prediction time.

When to Use This Skill

Use this skill after work involving:

  • Data preprocessing (normalization, standardization, scaling)
  • Missing value imputation
  • Feature engineering and feature selection
  • Dimensionality reduction (PCA, SVD, t-SNE)
  • Target encoding or label encoding
  • Time series feature construction
  • Data augmentation strategies
  • Algorithm development and optimization
  • Train-test split procedures
  • Cross-validation setup

Not For / Boundaries

  • Pure theoretical ML discussions without implementation
  • Model architecture design (without data preprocessing)
  • Hyperparameter tuning (unless it involves data-dependent operations)

Core Principle

The Golden Rule: At the exact moment of prediction in production, can I access this value from the database or compute it using only information available up to that point?

If the answer is "no" or "not completely", then data leakage exists.

Quick Reference

Critical Leakage Patterns

Pattern 1: Preprocessing Before Split

python
# ❌ WRONG: Leakage - fit on entire dataset
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)  # Uses test set statistics
X_train, X_test = train_test_split(X_scaled, y)

# ✅ CORRECT: Fit only on training data
X_train, X_test, y_train, y_test = train_test_split(X, y)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)  # Fit on train only
X_test_scaled = scaler.transform(X_test)  # Transform test using train statistics

Pattern 2: Global Missing Value Imputation

python
# ❌ WRONG: Uses global statistics including test set
df['age'].fillna(df['age'].mean(), inplace=True)  # Global mean includes test data
X_train, X_test = train_test_split(df, y)

# ✅ CORRECT: Compute statistics on training set only
X_train, X_test, y_train, y_test = train_test_split(df, y)
train_mean = X_train['age'].mean()  # Only from training data
X_train['age'].fillna(train_mean, inplace=True)
X_test['age'].fillna(train_mean, inplace=True)  # Use train mean for test

Pattern 3: PCA/Dimensionality Reduction on Full Dataset

python
# ❌ WRONG: PCA learns variance structure from test set
pca = PCA(n_components=10)
X_reduced = pca.fit_transform(X)  # Includes test set variance
X_train, X_test = train_test_split(X_reduced, y)

# ✅ CORRECT: Fit PCA only on training data
X_train, X_test, y_train, y_test = train_test_split(X, y)
pca = PCA(n_components=10)
X_train_reduced = pca.fit_transform(X_train)  # Learn from train only
X_test_reduced = pca.transform(X_test)  # Apply train-learned transformation

Pattern 4: Target Encoding with Full Dataset

python
# ❌ WRONG: Uses target values from test set
category_means = df.groupby('category')['target'].mean()  # Includes test targets
df['category_encoded'] = df['category'].map(category_means)
X_train, X_test = train_test_split(df, y)

# ✅ CORRECT: Compute encoding only from training targets
X_train, X_test, y_train, y_test = train_test_split(df, y)
category_means = X_train.groupby('category')['target'].mean()  # Train only
X_train['category_encoded'] = X_train['category'].map(category_means)
X_test['category_encoded'] = X_test['category'].map(category_means)

Pattern 5: Feature Selection on Full Dataset

python
# ❌ WRONG: Feature selection sees test set
from sklearn.feature_selection import SelectKBest
selector = SelectKBest(k=10)
X_selected = selector.fit_transform(X, y)  # Uses test set for selection
X_train, X_test = train_test_split(X_selected, y)

# ✅ CORRECT: Select features using training data only
X_train, X_test, y_train, y_test = train_test_split(X, y)
selector = SelectKBest(k=10)
X_train_selected = selector.fit_transform(X_train, y_train)  # Train only
X_test_selected = selector.transform(X_test)  # Apply train-learned selection

Pattern 6: Random Split on Temporal Data

python
# ❌ WRONG: Random split on time series (uses future to predict past)
X_train, X_test = train_test_split(df, test_size=0.2, random_state=42)

# ✅ CORRECT: Time-based split for temporal data
split_date = '2024-01-01'
X_train = df[df['date'] < split_date]
X_test = df[df['date'] >= split_date]

Pattern 7: Future Function in Time Series Features

python
# ❌ WRONG: Uses future data to compute current features
df['daily_avg'] = df.groupby('date')['value'].transform('mean')  # Includes all day's data

# ✅ CORRECT: Use only past data (expanding window)
df = df.sort_values('timestamp')
df['cumulative_avg'] = df.groupby('user_id')['value'].expanding().mean().reset_index(0, drop=True)

Pattern 8: Post-Event Features

python
# ❌ WRONG: Feature only exists after the outcome
# Predicting loan default using "number of collection calls" as feature
# Collection calls only happen AFTER default occurs

# ✅ CORRECT: Use only pre-event features
# Use features available BEFORE the outcome: credit score, income, debt ratio, etc.

Pattern 9: Leakage in Cross-Validation

python
# ❌ WRONG: Preprocessing before CV split
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
scores = cross_val_score(model, X_scaled, y, cv=5)  # Each fold sees other folds' statistics

# ✅ CORRECT: Preprocessing inside CV pipeline
from sklearn.pipeline import Pipeline
pipeline = Pipeline([
    ('scaler', StandardScaler()),
    ('model', LogisticRegression())
])
scores = cross_val_score(pipeline, X, y, cv=5)  # Scaling happens per fold

Pattern 10: Data Augmentation Leakage

python
# ❌ WRONG: Augment before split (test set influenced by augmented train data)
X_augmented = augment_data(X)  # Augmentation sees all data
X_train, X_test = train_test_split(X_augmented, y)

# ✅ CORRECT: Augment only training data after split
X_train, X_test, y_train, y_test = train_test_split(X, y)
X_train_augmented = augment_data(X_train)  # Augment train only
# X_test remains unchanged
Leakage Detection Checklist

After any ML preprocessing or feature engineering, verify:

  1. Fit-Transform Order: Did I fit any transformer (scaler, encoder, imputer, PCA) on the full dataset before splitting?
  2. Global Statistics: Did I compute any statistics (mean, std, median, mode) using the entire dataset?
  3. Feature Selection: Did I select features using information from the test set?
  4. Target Information: Did any feature encoding use target values from the test set?
  5. Temporal Order: For time series, did I use random splits or include future information?
  6. Post-Event Features: Are any features only available after the outcome occurs?
  7. Cross-Validation: Did I preprocess before setting up CV folds?
  8. Production Reality: At prediction time, can I actually compute this feature with available data?
The "Prediction Time" Test

For every feature and preprocessing step, ask:

When the model is deployed and receives a new data point at time T:
- Can I query this value from the database?
- Can I compute this statistic using only data available before time T?
- Does this feature require knowing the outcome I'm trying to predict?

If any answer is "NO", you have data leakage.

Examples

Example 1: Detecting Normalization Leakage

Input: Code that normalizes data before train-test split

python
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_normalized = scaler.fit_transform(X)
X_train, X_test, y_train, y_test = train_test_split(X_normalized, y, test_size=0.2)

Leakage Detection:

  • ❌ LEAKAGE DETECTED: fit_transform on entire dataset X
  • The scaler learned mean and std from the full dataset, including test set
  • At prediction time, you won't have access to test set statistics
  • Test set performance is artificially inflated

Corrected Code:

python
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
scaler = StandardScaler()
X_train_normalized = scaler.fit_transform(X_train)  # Fit on train only
X_test_normalized = scaler.transform(X_test)  # Transform using train statistics
Example 2: Detecting Missing Value Imputation Leakage

Input: Code that fills missing values with global mean

python
df = pd.read_csv('data.csv')
df['income'].fillna(df['income'].mean(), inplace=True)
X_train, X_test, y_train, y_test = train_test_split(df, y, test_size=0.2)

Leakage Detection:

  • ❌ LEAKAGE DETECTED: Global mean computed on entire dataset
  • The mean includes test set values
  • At prediction time, you won't know the test set mean
  • This inflates test performance

Corrected Code:

python
df = pd.read_csv('data.csv')
X_train, X_test, y_train, y_test = train_test_split(df, y, test_size=0.2)
train_mean = X_train['income'].mean()  # Compute mean from training data only
X_train['income'].fillna(train_mean, inplace=True)
X_test['income'].fillna(train_mean, inplace=True)  # Use training mean for test
Show full SKILL.md (290 more words)Show less
Example 3: Detecting Time Series Leakage

Input: Stock price prediction with random split and rolling average feature

python
df['rolling_avg_7d'] = df.groupby('stock')['price'].rolling(7, center=True).mean()
X_train, X_test = train_test_split(df, test_size=0.2, random_state=42)

Leakage Detection:

  • ❌ LEAKAGE DETECTED - Multiple Issues:
    1. center=True in rolling window uses future prices (t+3 days) to compute feature at time t
    2. Random split on temporal data means using "tomorrow's data" to predict "yesterday"
    3. At prediction time, you can't access future prices

Corrected Code:

python
# Fix 1: Use backward-looking rolling window (no center=True)
df = df.sort_values(['stock', 'date'])
df['rolling_avg_7d'] = df.groupby('stock')['price'].rolling(7, min_periods=1).mean().reset_index(0, drop=True)

# Fix 2: Time-based split instead of random
split_date = '2024-01-01'
X_train = df[df['date'] < split_date]
X_test = df[df['date'] >= split_date]
Example 4: Detecting Target Encoding Leakage

Input: Category encoding using target mean from full dataset

python
category_target_mean = df.groupby('category')['target'].mean()
df['category_encoded'] = df['category'].map(category_target_mean)
X_train, X_test, y_train, y_test = train_test_split(df.drop('target', axis=1), df['target'])

Leakage Detection:

  • ❌ LEAKAGE DETECTED: Target encoding uses test set labels
  • The encoding directly uses the answer (target) from test set
  • This is severe leakage - you're literally using test labels as features
  • At prediction time, you don't know the target value

Corrected Code:

python
X_train, X_test, y_train, y_test = train_test_split(df.drop('target', axis=1), df['target'])
# Compute target mean only from training data
train_df = X_train.copy()
train_df['target'] = y_train
category_target_mean = train_df.groupby('category')['target'].mean()

X_train['category_encoded'] = X_train['category'].map(category_target_mean)
X_test['category_encoded'] = X_test['category'].map(category_target_mean)
Example 5: Detecting Post-Event Feature Leakage

Input: Predicting customer churn using "number of retention calls" as feature

python
features = ['account_age', 'monthly_spend', 'support_tickets', 'retention_calls_count']
X = df[features]
y = df['churned']

Leakage Detection:

  • ❌ LEAKAGE DETECTED: Post-event feature
  • "retention_calls_count" only exists AFTER customer shows churn signals
  • Company only makes retention calls to customers who are already churning
  • This is a consequence of churn, not a predictor
  • At prediction time (before churn), this feature doesn't exist or is always 0

Corrected Code:

python
# Remove post-event features, use only pre-event features
features = ['account_age', 'monthly_spend', 'support_tickets', 'login_frequency',
            'feature_usage_decline', 'payment_delays']
X = df[features]
y = df['churned']

Leakage Severity Levels

CRITICAL (Model is completely invalid):

  • Target encoding with test labels
  • Post-event features
  • Using future data in time series

HIGH (Significantly inflated performance):

  • Preprocessing before train-test split
  • Feature selection on full dataset
  • Global imputation statistics

MEDIUM (Moderate performance inflation):

  • Incorrect cross-validation setup
  • Subtle temporal leakage in feature engineering

References

  • references/leakage-patterns.md: Comprehensive catalog of leakage patterns
  • references/temporal-leakage.md: Time series specific leakage issues
  • references/detection-strategies.md: How to detect leakage in existing code

Maintenance

  • Sources: ML best practices, Kaggle competition lessons, production ML experience
  • Last updated: 2026-01-19
  • Known limits: Cannot detect domain-specific leakage without context (e.g., whether a feature is post-event in a specific business context)

© foryourhealth111-pixel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in bundled/skills/ml-data-leakage-guard of foryourhealth111-pixel/Vibe-Skills.

  • SKILL.md
  • references/detection-strategies.md
  • references/index.md
  • references/leakage-patterns.md
  • references/temporal-leakage.md

Open the folder on GitHubat commit ddcaa2a

Compare with similar skills

ML Data Leakage Guard next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Data Leakage Guard compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Data Leakage Guard this skillforyourhealth111-pixel/Vibe-Skills3.6k—~3.4kAutomated safety check: PassApache-2.0
Senior Data ScientistRaidriar7170/hermes-skilleval1255 repos~1.4kAutomated safety check: PassMIT
Time Series Analytics Useropen-edge-platform/edge-ai-libraries171—~3.1kAutomated safety check: PassApache-2.0
Aeon Time Series Machine Learningdavila7/claude-code-templates33k13 repos~2.6kAutomated safety check: PassMIT
Setup Timescaledb Hypertablestimescale/pg-aiguide1.9k—~4.7kAutomated safety check: PassApache-2.0
Data Scientistdavila7/claude-code-templates33k8 repos~2.6kAutomated safety check: PassMIT

Similar skills

  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 5 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    171 GitHub stars~3.1k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Aeon Time Series Machine Learning

    davila7/claude-code-templates

    Guides time series machine learning with the aeon toolkit: classification, regression, clustering, forecasting, anomaly detection, segmentation and similarity search.

    33k GitHub starsUsed in 13 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Setup Timescaledb Hypertables

    timescale/pg-aiguide

    A skill your agent uses when creating database schemas or tables for Timescale, TimescaleDB, TigerData, or Tiger Cloud, especially for time-series, IoT, metrics, events, or log data.

    1.9k GitHub stars~4.7k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Data Scientist

    davila7/claude-code-templates

    Expert data scientist for advanced analytics, machine learning, and statistical modeling.

    33k GitHub starsUsed in 8 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Longbridge Quant

    helsome/folio

    Quantitative strategy frameworks: pairs trading/cointegration, volatility regime strategies, seasonality/calendar effects, multi-factor models (IC/IR), factor research and screening, correlation…

    271 GitHub starsUsed in 1 repo~1.6k tokens
    Data & AnalyticsAuto-check passed

More from foryourhealth111-pixel/Vibe-Skills

All 81 skills in this repo
  • Market Research Reports

    foryourhealth111-pixel/Vibe-Skills

    Produces long consulting-style market research and industry reports covering market sizing, competitive landscape, market entry and investment theses.

    3.6k GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check: notes
  • Academic Venue Templates

    foryourhealth111-pixel/Vibe-Skills

    Supplies venue-specific LaTeX templates and formatting rules for journals, conferences and posters, and checks a manuscript against page limits and submission requirements.

    3.6k GitHub stars~3.9k tokensUpdated 1 mo ago
    Auto-check: notes
  • Digital Brain

    foryourhealth111-pixel/Vibe-Skills

    This skill should be used when the user asks to "write a post", "check my voice", "look up contact", "prepare for meeting", "weekly review", "track goals", or mentions personal brand, content…

    3.6k GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Smart File Writer

    foryourhealth111-pixel/Vibe-Skills

    Diagnoses why a file write failed (permissions, disk space, path length, locks, read-only mounts) before retrying, instead of repeating the same call blindly.

    3.6k GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Automated Video Studio

    foryourhealth111-pixel/Vibe-Skills

    Turns footage, audio and a storyboard plan into a finished short video with FFmpeg jump-cuts, subtitle burn-in and a final polish pass.

    3.6k GitHub stars~838 tokensUpdated 1 mo ago
    Auto-check passed
  • Citation Management

    foryourhealth111-pixel/Vibe-Skills

    Turns DOIs, PMIDs and arXiv IDs into clean BibTeX, searches Google Scholar and PubMed, and checks and deduplicates a reference list.

    3.6k GitHub stars~7.6k tokensUpdated 1 mo ago
    Auto-check: notes

Questions about ML Data Leakage Guard

What does ML Data Leakage Guard do?

Detects and prevents data leakage in machine learning and mathematical modeling. ML Data Leakage Guard is an agent skill from foryourhealth111-pixel/Vibe-Skills. Detects and prevents data leakage in machine learning and mathematical modeling.

When should I use ML Data Leakage Guard?

ML Data Leakage Guard fits situations like: tasks that involve Machine learning; tasks that involve Data cleaning; tasks that involve Database schema design.

How do I install ML Data Leakage Guard in Claude Code?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill ml-data-leakage-guard -a claude-code`. Or copy the skill folder (bundled/skills/ml-data-leakage-guard in foryourhealth111-pixel/Vibe-Skills) into .claude/skills/ml-data-leakage-guard in your project. Claude Code loads it when a task matches its description.

How do I install ML Data Leakage Guard in Codex?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill ml-data-leakage-guard -a codex`. Or copy the skill folder (bundled/skills/ml-data-leakage-guard in foryourhealth111-pixel/Vibe-Skills) into .agents/skills/ml-data-leakage-guard in your project. Codex loads it when a task matches its description.

Can I use ML Data Leakage Guard in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill ml-data-leakage-guard -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-data-leakage-guard, .gemini/skills/ml-data-leakage-guard, .github/skills/ml-data-leakage-guard and .opencode/skills/ml-data-leakage-guard in your project.

What does ML Data Leakage Guard need to run?

SKILL.md names no scripts, command-line tools or credentials: ML Data Leakage Guard is instructions for the agent only. Our summary lists: Python 3.

Does ML Data Leakage Guard access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is ML Data Leakage Guard safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does ML Data Leakage Guard use?

ML Data Leakage Guard is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Data Leakage Guard use?

About 3.4k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.

What are the alternatives to ML Data Leakage Guard?

Skills that share tags, products or a category with ML Data Leakage Guard: Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Time Series Analytics User (open-edge-platform/edge-ai-libraries, 171 stars), Aeon Time Series Machine Learning (davila7/claude-code-templates, 33k stars) and Setup Timescaledb Hypertables (timescale/pg-aiguide, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Data Leakage Guard?

foryourhealth111-pixel (a GitHub user) maintains it in foryourhealth111-pixel/Vibe-Skills, which has 3,627 GitHub stars. The repository holds 81 skills in this directory. The repository was last updated on August 31, 2026.

Source: foryourhealth111-pixel/Vibe-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.