Splitting Datasets
jeremylongshore/tons-of-skills-marketplace
Process split datasets into training, validation, and testing sets for ML model development.
Agent skill
by foryourhealth111-pixel in foryourhealth111-pixel/Vibe-Skills
⚠️ CRITICAL USER EXPERIENCE-BASED SKILL - ALWAYS CONSULT BEFORE DATA PREPROCESSING ⚠️ Prevents catastrophic errors (88.9% error rate in V1.0 case study) through multi-level feature analysis, data…
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install foryourhealth111-pixel/Vibe-Skills scientific-data-preprocessing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/bundled/skills/scientific-data-preprocessing .claude/skills/scientific-data-preprocessing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "scientific-data-preprocessing" agent skill from https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/main/bundled/skills/scientific-data-preprocessing into .claude/skills/scientific-data-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scientific-data-preprocessing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/main/bundled/skills/scientific-data-preprocessingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install foryourhealth111-pixel/Vibe-Skills scientific-data-preprocessing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/bundled/skills/scientific-data-preprocessing .agents/skills/scientific-data-preprocessing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "scientific-data-preprocessing" agent skill from https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/main/bundled/skills/scientific-data-preprocessing into .agents/skills/scientific-data-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scientific-data-preprocessing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install foryourhealth111-pixel/Vibe-Skills scientific-data-preprocessing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/bundled/skills/scientific-data-preprocessing .cursor/skills/scientific-data-preprocessing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "scientific-data-preprocessing" agent skill from https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/main/bundled/skills/scientific-data-preprocessing into .cursor/skills/scientific-data-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scientific-data-preprocessing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/foryourhealth111-pixel/Vibe-Skills.git --path bundled/skills/scientific-data-preprocessing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install foryourhealth111-pixel/Vibe-Skills scientific-data-preprocessing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/bundled/skills/scientific-data-preprocessing .gemini/skills/scientific-data-preprocessing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "scientific-data-preprocessing" agent skill from https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/main/bundled/skills/scientific-data-preprocessing into .gemini/skills/scientific-data-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scientific-data-preprocessing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install foryourhealth111-pixel/Vibe-Skills scientific-data-preprocessingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/bundled/skills/scientific-data-preprocessing .github/skills/scientific-data-preprocessing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "scientific-data-preprocessing" agent skill from https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/main/bundled/skills/scientific-data-preprocessing into .github/skills/scientific-data-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scientific-data-preprocessing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install foryourhealth111-pixel/Vibe-Skills scientific-data-preprocessing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/bundled/skills/scientific-data-preprocessing .opencode/skills/scientific-data-preprocessing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "scientific-data-preprocessing" agent skill from https://github.com/foryourhealth111-pixel/Vibe-Skills/tree/main/bundled/skills/scientific-data-preprocessing into .opencode/skills/scientific-data-preprocessing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "scientific-data-preprocessing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
scientific-data-preprocessing⚠️ CRITICAL USER EXPERIENCE-BASED SKILL - ALWAYS CONSULT BEFORE DATA PREPROCESSING ⚠️ Prevents catastrophic errors (88.9% error rate in V1.0 case study) through multi-level feature analysis, data…
Scientific Data Preprocessing is an agent skill from foryourhealth111-pixel/Vibe-Skills. ⚠️ CRITICAL USER EXPERIENCE-BASED SKILL - ALWAYS CONSULT BEFORE DATA PREPROCESSING ⚠️ Prevents catastrophic errors (88.9% error rate in V1.0 case study) through multi-level feature analysis, data leakage detection, and semantic validation. MANDATORY for: data preprocessing, feature engineering, standardization, normalization, interpolation, missing value handling, feature selection, or ANY data transformation task. Covers grouped time-series, cross-sectional, panel data. Detects: time travel leakage, causal…
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `README.md`, `UPDATE_LOG_V1.1.md` and `references/ai-common-pitfalls.md`).
It sits in Data & Analytics, covering Machine learning and Data cleaning. The repository describes itself as: Intelligent Skill routing and workflow orchestration for AI agents — +21.12 pp reward, −29.6% tokens on SkillsBench with DeepSeekV4Flash-VE. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit ddcaa2a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Scientific Data Preprocessing loads about 5k tokens when it runs, and up to ~21k if it reads all its reference files. Until then it costs about 167 tokens; SKILL.md has 824 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from foryourhealth111-pixel/Vibe-Skills at commit ddcaa2a, republished under its Apache-2.0 licence (© foryourhealth111-pixel). 824 words, ~4,956 tokens.
.claude/skills/scientific-data-preprocessing/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.⚠️ CRITICAL: USER'S HARD-WON EXPERIENCE - MANDATORY CONSULTATION ⚠️
This skill encapsulates painful lessons learned from real preprocessing disasters (88.9% error rate documented). ALWAYS use this skill for planning, reflection, and validation when ANY data preprocessing is involved.
Why this skill is mandatory:
When to invoke (DO NOT SKIP):
Prevent catastrophic preprocessing errors in grouped time-series data by applying multi-level feature analysis and respecting data structure boundaries.
MANDATORY consultation - trigger immediately when:
Trigger keywords that MUST invoke this skill:
This skill does NOT:
Required inputs before proceeding:
Level 1: Data Type
# Check data types
df.dtypes # int64, float64, object, etc.Level 2: Feature Type Classification
# Binary (0/1)
binary_features = [col for col in df.columns if df[col].nunique() == 2]
# Categorical (finite discrete values)
categorical_features = [col for col in df.select_dtypes(include='object').columns]
# Continuous (infinite possible values)
continuous_features = [col for col in df.select_dtypes(include=['float64', 'int64']).columns
if df[col].nunique() > 10]Level 3: Data Structure
# Check for grouping
print(f"Number of groups: {df['group_id'].nunique()}")
print(f"Avg points per group: {df.groupby('group_id').size().mean():.1f}")
# Check for time-series
df_sorted = df.sort_values(['group_id', 'timestamp'])Level 4: Physical Meaning
# Validate physical ranges
assert df['speed_mph'].max() < 200, "Speed exceeds physical limit"
assert df['distance_meters'].min() >= 0, "Negative distance impossible"# Decision: Within-group or global processing?
def choose_processing_scope(data, feature, goal):
"""
goal = 'relative' → within-group (e.g., "this point was intense FOR THIS MATCH")
goal = 'absolute' → global (e.g., "this was an intense point OVERALL")
"""
if goal == 'relative':
return 'within_group'
elif goal == 'absolute':
return 'global'
else:
raise ValueError("Goal must be 'relative' or 'absolute'")from scipy.interpolate import CubicSpline
import numpy as np
# ✅ CORRECT: Interpolate within each group
for group_id in df['match_id'].unique():
mask = df['match_id'] == group_id
group_data = df.loc[mask, 'speed_mph'].copy()
# Get valid (non-NaN) indices
valid_idx = group_data.notna()
valid_positions = np.where(valid_idx)[0]
valid_values = group_data[valid_idx].values
if len(valid_positions) >= 4:
cs = CubicSpline(valid_positions, valid_values)
missing_positions = np.where(~valid_idx)[0]
df.loc[mask & ~valid_idx, 'speed_mph'] = cs(missing_positions)# ❌ WRONG: Cross-group interpolation
# This interpolates between match A's last point and match B's first point!
cs = CubicSpline(
np.where(df['speed_mph'].notna())[0], # ❌ All indices globally
df['speed_mph'].dropna().values
)
df.loc[df['speed_mph'].isna(), 'speed_mph'] = cs(
np.where(df['speed_mph'].isna())[0]
)from sklearn.preprocessing import StandardScaler
# ✅ CORRECT: Standardize within each match
for match_id in df['match_id'].unique():
mask = df['match_id'] == match_id
scaler = StandardScaler()
df.loc[mask, 'distance_run_std_within'] = scaler.fit_transform(
df.loc[mask, [['distance_run']]
)
# Interpretation: z=+2 means "2 std above average FOR THIS MATCH"# ✅ CORRECT: Global standardization (when appropriate)
scaler = StandardScaler()
df['distance_run_std_global'] = scaler.fit_transform(df[['distance_run']])
# Interpretation: z=+2 means "2 std above average ACROSS ALL MATCHES"# Binary variables (0/1) - KEEP AS-IS
binary_cols = ['is_ace', 'is_winner', 'is_error']
# ❌ NEVER standardize these! They have semantic meaning as 0/1
# Categorical variables - ONE-HOT ENCODE
df_encoded = pd.get_dummies(df, columns=['server', 'serve_number'], dtype=int)
# Continuous variables - STANDARDIZE (within-group or global)
continuous_cols = ['distance_run', 'rally_count', 'speed_mph']
# ✅ Apply pattern 3 or 4 based on goal# ✅ CORRECT: Sliding window for momentum analysis
window = 10
df['win_rate_last10'] = df.groupby('match_id')['point_won'].transform(
lambda x: x.rolling(window, min_periods=1).mean()
)
# ❌ WRONG: Cumulative features (loses temporal locality)
df['cumulative_points_won'] = df.groupby('match_id')['point_won'].cumsum()
# This just increases monotonically and correlates with point_numberdef validate_data_quality(df, feature, expected_range):
"""Validate before processing"""
# Check range
assert df[feature].min() >= expected_range[0], f"{feature} below minimum"
assert df[feature].max() <= expected_range[1], f"{feature} above maximum"
# Check for anomalies
mean = df[feature].mean()
std = df[feature].std()
if std > mean:
print(f"⚠️ WARNING: {feature} has std > mean (highly skewed or errors)")
# Check missing pattern
missing_by_group = df.groupby('match_id')[feature].apply(lambda x: x.isna().sum())
if missing_by_group.max() > len(df) / df['match_id'].nunique() * 0.5:
print(f"⚠️ WARNING: {feature} has >50% missing in some groups")
# Example
validate_data_quality(df, 'speed_mph', expected_range=(50, 165))def detect_processing_scope(df, group_col, feature_col):
"""
Recommend within-group vs global based on variance structure
"""
# Calculate variance components
within_group_var = df.groupby(group_col)[feature_col].var().mean()
global_var = df[feature_col].var()
# Intraclass correlation
between_group_var = global_var - within_group_var
icc = between_group_var / global_var
if icc > 0.5:
return 'within_group', f"High between-group variance (ICC={icc:.2f})"
else:
return 'global', f"Low between-group variance (ICC={icc:.2f})"
scope, reason = detect_processing_scope(df, 'match_id', 'distance_run')
print(f"Recommended: {scope} - {reason}")def detect_data_leakage(df, target_col, feature_cols, id_cols):
"""
Critical checks for data leakage and AI common pitfalls
"""
issues = []
# 1. ID Leakage: High cardinality variables as features
for col in feature_cols:
if col in id_cols:
issues.append(f"❌ FATAL: {col} is an ID - NEVER use as feature")
continue
# Check if looks like ID (>50% unique)
uniqueness = df[col].nunique() / len(df)
if uniqueness > 0.5:
issues.append(f"⚠️ {col}: {uniqueness*100:.1f}% unique - possible ID leakage")
# 2. Causal Inversion: Perfect correlation with target
for col in feature_cols:
if col == target_col:
continue
if df[col].dtype in ['int64', 'float64']:
corr = abs(df[[col, target_col]].corr().iloc[0, 1])
if corr > 0.95:
issues.append(f"❌ FATAL: {col} correlation={corr:.3f} - likely consequence of target!")
# 3. Meaningless Numeric: Codes treated as numbers
for col in feature_cols:
if df[col].dtype in ['int64', 'float64']:
# Pattern: High values, many uniques, looks like code
if df[col].min() > 1000 and df[col].nunique() > 100:
issues.append(f"⚠️ {col}: Looks like code (zipcode/ID) - should be categorical")
# 4. Time Travel: Check if standardization used global statistics
# (Requires knowing if train/test split was done first)
# Print report
if issues:
print("="*60)
print("DATA LEAKAGE AUDIT")
print("="*60)
for issue in issues:
print(issue)
print("="*60)
else:
print("✅ No obvious leakage detected")
return issues
# Example usage
issues = detect_data_leakage(
df,
target_col='point_won',
feature_cols=['speed_mph', 'user_id', 'distance_run'],
id_cols=['match_id', 'user_id']
)from scipy.stats import skew, kurtosis
from sklearn.preprocessing import StandardScaler, RobustScaler
def smart_scaler_selection(df, col):
"""
Choose scaler based on distribution characteristics
"""
data = df[col].dropna()
# Check distribution
skewness = skew(data)
kurt = kurtosis(data)
print(f"{col}: skewness={skewness:.2f}, kurtosis={kurt:.2f}")
if abs(skewness) < 0.5 and abs(kurt) < 3:
# Roughly normal
print(" → StandardScaler (data is roughly normal)")
return StandardScaler(), None
elif skewness > 1:
# Right-skewed (long tail)
print(" → Log transform + StandardScaler (right-skewed)")
return StandardScaler(), 'log'
else:
# Heavy outliers or non-normal
print(" → RobustScaler (heavy outliers)")
return RobustScaler(), None
# Example usage
for col in continuous_features:
scaler, transform = smart_scaler_selection(df, col)
if transform == 'log':
df[f'{col}_log'] = np.log1p(df[col])
df[f'{col}_scaled'] = scaler.fit_transform(df[[f'{col}_log']])
else:
df[f'{col}_scaled'] = scaler.fit_transform(df[[col]])Input:
speed_mph, distance_run, rally_count, is_ace, serverSteps:
import pandas as pd
from sklearn.preprocessing import StandardScaler
# 1. Load and inspect
df = pd.read_csv('tennis_data.csv')
print(f"Matches: {df['match_id'].nunique()}")
print(f"Features: {df.dtypes}")
# 2. Classify features
binary_features = ['is_ace', 'is_winner', 'is_break_point']
categorical_features = ['server', 'serve_number']
continuous_features = ['distance_run', 'speed_mph', 'rally_count']
# 3. Validate data quality
for feat in continuous_features:
print(f"\n{feat}:")
print(df[feat].describe())
# Check for impossible values
if feat == 'speed_mph':
assert df[feat].max() < 170, "Speed exceeds world record!"
# 4. Handle missing values (within-group)
for match_id in df['match_id'].unique():
mask = df['match_id'] == match_id
for feat in continuous_features:
if df.loc[mask, feat].isna().any():
# Simple linear interpolation within match
df.loc[mask, feat] = df.loc[mask, feat].interpolate(method='linear')
# 5. One-hot encode categorical
df = pd.get_dummies(df, columns=categorical_features, dtype=int)
# 6. Standardize continuous features WITHIN each match
for feat in continuous_features:
df[f'{feat}_std'] = np.nan
for match_id in df['match_id'].unique():
mask = df['match_id'] == match_id
scaler = StandardScaler()
df.loc[mask, f'{feat}_std'] = scaler.fit_transform(
df.loc[mask, [[feat]]
)
# 7. Create sliding window features
window = 10
df['win_rate_last10'] = df.groupby('match_id')['point_won'].transform(
lambda x: x.rolling(window, min_periods=1).mean()
)
# 8. KEEP binary features as 0/1 (don't transform!)
# binary_features are already correct
print("\n✅ Preprocessing complete!")
print(f"Final shape: {df.shape}")
print(f"Standardized features: {[f for f in df.columns if f.endswith('_std')]}")Expected output:
server_1, server_2)_std versions_std features have mean≈0, std≈1 WITHIN each matchInput:
Steps:
# Check if standardization was done correctly
def check_within_group_standardization(df, group_col, feature_std_col):
"""
Verify that standardized feature has mean≈0, std≈1 within each group
"""
results = df.groupby(group_col)[feature_std_col].agg(['mean', 'std'])
# Within-group standardization: each group should have mean≈0, std≈1
if (results['mean'].abs() < 0.1).all() and (results['std'].between(0.9, 1.1)).all():
print("✅ CORRECT: Within-group standardization detected")
return True
# Global standardization: groups will have varying means and stds
else:
print("❌ WRONG: Global standardization detected!")
print("Group means:", results['mean'].values[:5])
print("Group stds:", results['std'].values[:5])
return False
check_within_group_standardization(df, 'match_id', 'distance_run_std')Expected output:
Input:
Steps:
# ❌ WRONG approach (existing code)
df['cumulative_wins'] = df.groupby('match_id')['point_won'].cumsum()
# Problem: This just counts total wins up to this point
# Doesn't capture recent momentum!
# ✅ CORRECT approach (fix)
# Replace cumulative with sliding window
window = 10
df['recent_win_rate'] = df.groupby('match_id')['point_won'].transform(
lambda x: x.rolling(window, min_periods=1).mean()
)
# Compare
print("Cumulative (wrong):", df['cumulative_wins'].values[50:60])
print("Sliding window (correct):", df['recent_win_rate'].values[50:60])
# Cumulative: [25, 26, 26, 27, 28, ...] - monotonic
# Sliding window: [0.6, 0.7, 0.5, 0.6, ...] - fluctuates with momentumExpected output:
references/index.md: Navigation and overviewreferences/error-case-studies.md: Real-world preprocessing disasters from tennis datareferences/decision-trees.md: Full decision trees for all preprocessing choicesreferences/validation-checklist.md: Pre-processing validation checklistreferences/ai-common-pitfalls.md: AI-specific errors (data leakage, semantic fallacies, distribution blindness)⚠️ CRITICAL NOTICE: USER'S PERSONAL EXPERIENCE-BASED SKILL ⚠️
This skill is NOT theoretical - it's based on real project failures:
Why this matters to you (Claude):
Authority level: HIGHEST
Sources:
Mandatory consultation:
Last updated: 2026-01-18 (V1.1)
Known limits:
© foryourhealth111-pixel, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (references) in bundled/skills/scientific-data-preprocessing of foryourhealth111-pixel/Vibe-Skills.
Open the folder on GitHubat commit ddcaa2a
Scientific Data Preprocessing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Scientific Data Preprocessing this skillforyourhealth111-pixel/Vibe-Skills | 3.6k | — | ~5k | Automated safety check: Pass | Apache-2.0 | |
| Splitting Datasetsjeremylongshore/tons-of-skills-marketplace | 2.8k | 1 repos | ~836 | Automated safety check: Pass | MIT | |
| Sap Hana Cloud Data Intelligencesecondsky/sap-skills | 462 | — | ~3.2k | Automated safety check: Pass | GPL-3.0 | |
| Rf Model Importance Analysisaipoch/medical-research-skills | 2k | — | ~2.7k | Automated safety check: Pass | MIT | |
| Data Cleanbrycewang-stanford/Auto-Empirical-Research-Skills | 4.5k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Nan Safe Correlationjaechang-hits/SciAgent-Skills | 370 | 1 repos | ~2.9k | Automated safety check: Pass | CC-BY-4.0 |
jeremylongshore/tons-of-skills-marketplace
Process split datasets into training, validation, and testing sets for ML model development.
secondsky/sap-skills
Develops data processing pipelines, integrations, and machine learning scenarios in SAP Data Intelligence Cloud.
aipoch/medical-research-skills
A skill your agent uses when you need a standardized R CLI workflow to train a two-class random forest model from an expression-like feature matrix, rank variable importance, and generate…
brycewang-stanford/Auto-Empirical-Research-Skills
Produce documented data cleaning scripts that log every transformation with N before/after each step, generate a CONSORT-style exclusion flow diagram, create decision log entries for every…
jaechang-hits/SciAgent-Skills
Per-feature NaN-safe Spearman/Pearson correlation across many features (genes, proteins, variants) with missing values.
Drchronx/ai-agent-research-starter-kit
Parameterized Python empirical-analysis and machine-learning workflow for applied economics, public health epidemiology, supervised ML, and ML causal inference.
foryourhealth111-pixel/Vibe-Skills
Produces long consulting-style market research and industry reports covering market sizing, competitive landscape, market entry and investment theses.
foryourhealth111-pixel/Vibe-Skills
Supplies venue-specific LaTeX templates and formatting rules for journals, conferences and posters, and checks a manuscript against page limits and submission requirements.
foryourhealth111-pixel/Vibe-Skills
This skill should be used when the user asks to "write a post", "check my voice", "look up contact", "prepare for meeting", "weekly review", "track goals", or mentions personal brand, content…
foryourhealth111-pixel/Vibe-Skills
Diagnoses why a file write failed (permissions, disk space, path length, locks, read-only mounts) before retrying, instead of repeating the same call blindly.
foryourhealth111-pixel/Vibe-Skills
Turns footage, audio and a storyboard plan into a finished short video with FFmpeg jump-cuts, subtitle burn-in and a final polish pass.
foryourhealth111-pixel/Vibe-Skills
Turns DOIs, PMIDs and arXiv IDs into clean BibTeX, searches Google Scholar and PubMed, and checks and deduplicates a reference list.
Categories
⚠️ CRITICAL USER EXPERIENCE-BASED SKILL - ALWAYS CONSULT BEFORE DATA PREPROCESSING ⚠️ Prevents catastrophic errors (88.9% error rate in V1.0 case study) through multi-level feature analysis, data…. Scientific Data Preprocessing is an agent skill from foryourhealth111-pixel/Vibe-Skills.0 case study) through multi-level feature analysis, data leakage detection, and semantic validation.
Scientific Data Preprocessing fits situations like: tasks that involve Machine learning; tasks that involve Data cleaning.
Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a claude-code`. Or copy the skill folder (bundled/skills/scientific-data-preprocessing in foryourhealth111-pixel/Vibe-Skills) into .claude/skills/scientific-data-preprocessing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a codex`. Or copy the skill folder (bundled/skills/scientific-data-preprocessing in foryourhealth111-pixel/Vibe-Skills) into .agents/skills/scientific-data-preprocessing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill scientific-data-preprocessing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scientific-data-preprocessing, .gemini/skills/scientific-data-preprocessing, .github/skills/scientific-data-preprocessing and .opencode/skills/scientific-data-preprocessing in your project.
SKILL.md names no scripts, command-line tools or credentials: Scientific Data Preprocessing is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Scientific Data Preprocessing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 16k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Scientific Data Preprocessing: Splitting Datasets (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Sap Hana Cloud Data Intelligence (secondsky/sap-skills, 462 stars), Rf Model Importance Analysis (aipoch/medical-research-skills, 2k stars) and Data Clean (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
foryourhealth111-pixel (a GitHub user) maintains it in foryourhealth111-pixel/Vibe-Skills, which has 3,604 GitHub stars. The repository holds 65 skills in this directory. The repository was last updated on August 31, 2026.
Source: foryourhealth111-pixel/Vibe-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.