Agent skill

Walk Forward Validation

by agiprolabs in agiprolabs/claude-trading-skills

Walk-forward validation framework for trading strategies and ML models with time-series-aware splits, overfit detection, and regime-aware validation

MITAuto-check passedData & Analytics

Install Walk Forward Validation

skills CLI
$ npx skills add agiprolabs/claude-trading-skills --skill walk-forward-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agiprolabs/claude-trading-skills walk-forward-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agiprolabs/claude-trading-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/walk-forward-validation .claude/skills/walk-forward-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
walk-forward-validation
GitHub stars
410
Token cost
~2.2k tokens
SKILL.md length
802 words
Files
6 (incl. scripts, references)
Skills in repo
68
Repo updated
First seen
Licence
MIT

At a glance

Walk-forward validation framework for trading strategies and ML models with time-series-aware splits, overfit detection, and regime-aware validation

  • Works in 4 steps: Lookahead bias — Random splits let the… → Autocorrelation — Adjacent observations… → Regime dependence — Markets shift… → …
  • Tasks that involve Machine learning
  • SKILL.md covers Why Standard Cross-Validation…, Walk-Forward Framework, Purging and Embargo and Combinatorial Purged…, plus 5 more sections
  • Runs Python scripts from its folder

What it does

Walk Forward Validation is an agent skill from agiprolabs/claude-trading-skills. Walk-forward validation framework for trading strategies and ML models with time-series-aware splits, overfit detection, and regime-aware validation

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/methodology.md`, `references/overfit_detection.md` and `references/practical_guide.md`).

It sits in Data & Analytics, covering Machine learning, Trading and backtesting and Forecasting and time series. The repository describes itself as: 68 trading, DeFi, and quantitative finance Agent Skills. Works with Claude Code, Cursor, Codex, Gemini CLI, and 30+ other tools. The licence is MIT.

When your agent uses it

  • Tasks that involve Machine learning
  • Tasks that involve Trading and backtesting
  • Tasks that involve Forecasting and time series

Example prompts

  • “/walk-forward-validation”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Lookahead bias — Random splits let the model train on future data and predict past data, artificially inflating performance.
  2. Autocorrelation — Adjacent observations are correlated. A random split that puts Monday in test and Tuesday in train leaks information.
  3. Regime dependence — Markets shift between regimes. A model trained on a bull market and tested on a bull market tells you nothing about…
  4. Label overlap — If labels are computed over windows (e.g., 24h forward return), adjacent train/test samples share label computation…

What it can do on your machine

Read from SKILL.md and the folder at commit 981e1d7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Walk Forward Validation loads about 2.2k tokens when it runs, and up to ~7.6k if it reads all its reference files. Until then it costs about 43 tokens; SKILL.md has 802 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from agiprolabs/claude-trading-skills at commit 981e1d7, republished under its MIT licence (© agiprolabs). 802 words, ~2,184 tokens.

Download SKILL.mdSave it as .claude/skills/walk-forward-validation/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
walk-forward-validation
description
Walk-forward validation framework for trading strategies and ML models with time-series-aware splits, overfit detection, and regime-aware validation

Walk-Forward Validation

Walk-forward validation framework for trading strategies and ML models. Standard cross-validation (k-fold, random splits) fails catastrophically for financial time series because it introduces lookahead bias and ignores autocorrelation. This skill covers proper time-series validation techniques including rolling and expanding windows, purged cross-validation, combinatorial purged cross-validation (CPCV), and overfit detection metrics.

Why Standard Cross-Validation Fails

Standard k-fold CV assumes data points are independent and identically distributed (IID). Financial time series violate both assumptions:

  1. Lookahead bias — Random splits let the model train on future data and predict past data, artificially inflating performance.
  2. Autocorrelation — Adjacent observations are correlated. A random split that puts Monday in test and Tuesday in train leaks information.
  3. Regime dependence — Markets shift between regimes. A model trained on a bull market and tested on a bull market tells you nothing about bear market performance.
  4. Label overlap — If labels are computed over windows (e.g., 24h forward return), adjacent train/test samples share label computation periods, leaking information.

Walk-Forward Framework

Rolling Window (Fixed Train Size)

The train window has a fixed size and slides forward in time. This is preferred when you believe older data is less relevant (common in crypto).

Window 1: [===TRAIN===][=TEST=]
Window 2:    [===TRAIN===][=TEST=]
Window 3:       [===TRAIN===][=TEST=]

Parameters:

  • train_size: Number of bars/days in the training window
  • test_size: Number of bars/days in the test window
  • step_size: How far to advance between folds (often equals test_size)
Expanding Window (Growing Train)

The train window starts at the beginning and expands forward. This uses all available historical data, which helps when data is scarce.

Window 1: [==TRAIN==][=TEST=]
Window 2: [====TRAIN====][=TEST=]
Window 3: [======TRAIN======][=TEST=]

Parameters:

  • min_train_size: Minimum training samples before first fold
  • test_size: Fixed test window size
  • step_size: How far to advance between folds
Choosing Between Them
FactorRollingExpanding
Data recencyPrioritizes recent dataUses all history
Regime changesBetter adapts to new regimesMay dilute recent regime
Sample sizeFixed, may be smallGrows over time
Crypto preferencePreferred for < 6mo horizonsBetter for regime-stable models

Purging and Embargo

Purging

Remove training samples whose labels overlap with the test set's time range. If a label is computed as the 24h forward return starting at time t, any training sample where t + 24h extends into the test period must be purged.

python
def purge_train_indices(
    train_idx: list[int],
    test_start: int,
    label_horizon: int,
    timestamps: list[int],
) -> list[int]:
    """Remove train samples whose label windows overlap test period."""
    test_start_time = timestamps[test_start]
    return [
        i for i in train_idx
        if timestamps[i] + label_horizon < test_start_time
    ]
Embargo

Add a buffer gap between the end of training and start of testing to account for serial correlation that purging alone does not eliminate.

[===TRAIN===][--EMBARGO--][=TEST=]

Typical embargo sizes:

  • 1-minute bars: 60–240 bars (1–4 hours)
  • 5-minute bars: 12–48 bars (1–4 hours)
  • Hourly bars: 6–24 bars (6–24 hours)
  • Daily bars: 2–5 bars (2–5 days)
  • Crypto rule of thumb: Embargo >= 2x the label computation horizon

Combinatorial Purged Cross-Validation (CPCV)

CPCV (Lopez de Prado, 2018) generates all possible train/test combinations from N groups while maintaining temporal ordering. This produces far more test paths than standard walk-forward, enabling statistical tests for overfitting.

Key properties:

  • Splits data into N contiguous groups
  • For each combination of k test groups, the remaining N-k groups form the training set
  • Applies purging and embargo at each train/test boundary
  • Produces C(N, k) backtest paths (e.g., N=6, k=2 gives 15 paths)

See references/methodology.md for the full CPCV algorithm and formulas.

Show full SKILL.md (298 more words)Show less

Overfit Detection

Deflated Sharpe Ratio (DSR)

The observed Sharpe ratio must be adjusted for:

  • Number of strategies tested (multiple testing)
  • Non-normality of returns (skewness, kurtosis)
  • Length of the backtest
python
import numpy as np
from scipy.stats import norm

def deflated_sharpe_ratio(
    observed_sr: float,
    num_trials: int,
    backtest_length: int,
    skewness: float = 0.0,
    kurtosis: float = 3.0,
) -> float:
    """Compute the probability that observed SR > 0 after deflation.

    Args:
        observed_sr: Annualized Sharpe ratio of the selected strategy.
        num_trials: Number of strategies tested (including discarded ones).
        backtest_length: Number of return observations.
        skewness: Skewness of returns.
        kurtosis: Excess kurtosis of returns.

    Returns:
        p-value (probability SR is genuinely > 0).
    """
    sr_std = np.sqrt(
        (1 - skewness * observed_sr + (kurtosis - 1) / 4 * observed_sr**2)
        / (backtest_length - 1)
    )
    # Expected max SR under null (Euler-Mascheroni approximation)
    euler_mascheroni = 0.5772156649
    expected_max_sr = norm.ppf(1 - 1 / num_trials) * (
        1 - euler_mascheroni
    ) + euler_mascheroni * norm.ppf(1 - 1 / (num_trials * np.e))
    dsr = norm.cdf((observed_sr - expected_max_sr) / sr_std)
    return dsr

A DSR below 0.95 suggests the observed performance is likely due to overfitting across the trials tested.

Probability of Backtest Overfitting (PBO)

PBO uses CPCV to measure the fraction of backtest paths where the in-sample optimal strategy underperforms the median out-of-sample. A PBO above 0.50 indicates more-likely-than-not overfitting.

See references/overfit_detection.md for complete derivations and implementation details.

Crypto-Specific Considerations

  1. Shorter windows: Crypto regimes change faster than equities. A 90-day rolling window may be more appropriate than 252 days.
  2. 24/7 markets: No weekends or holidays to account for, but funding rate resets (every 8h on perps) create microstructure effects.
  3. Survivorship bias: Many tokens delist. Validation must include delisted tokens or at minimum acknowledge this limitation.
  4. Liquidity regime shifts: A token's liquidity profile can change dramatically (new CEX listing, liquidity mining end). Train/test splits should ideally not straddle major liquidity events.
  5. Data availability: Many tokens have < 1 year of data. Expanding windows with small min_train_size may be necessary.

Practical Window Sizes for Crypto

Strategy TimeframeTrain WindowTest WindowEmbargo
Scalping (1-5min)3-7 days1 day2-4 hours
Intraday (15min-1h)14-30 days3-7 days12-24 hours
Swing (4h-daily)30-90 days7-14 days2-5 days
Position (daily-weekly)90-180 days30 days5-10 days

Quick Start

python
from walk_forward import WalkForwardValidator, WalkForwardConfig

config = WalkForwardConfig(
    train_size=90,
    test_size=14,
    step_size=14,
    window_type="rolling",
    embargo_size=3,
    purge_horizon=1,
)

validator = WalkForwardValidator(config)
for fold in validator.split(price_data):
    model.fit(fold.train_X, fold.train_y)
    predictions = model.predict(fold.test_X)
    fold.record_performance(predictions, fold.test_y)

results = validator.aggregate_results()
print(f"OOS Sharpe: {results.oos_sharpe:.3f}")
print(f"Train/Test Sharpe ratio: {results.sharpe_ratio_ratio:.2f}")

Files

References
  • references/methodology.md — Walk-forward theory, window types, purging, embargo, CPCV algorithm with formulas
  • references/overfit_detection.md — Deflated Sharpe ratio, probability of backtest overfitting, multiple testing corrections
  • references/practical_guide.md — Window size selection for crypto, regime considerations, common validation mistakes
Scripts
  • scripts/walk_forward.py — Walk-forward validation engine with rolling and expanding windows; --demo mode with synthetic data
  • scripts/overfit_detector.py — Deflated Sharpe ratio and PBO computation; --demo mode with synthetic backtest results

© agiprolabs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/walk-forward-validation of agiprolabs/claude-trading-skills.

  • SKILL.md
  • references/methodology.md
  • references/overfit_detection.md
  • references/practical_guide.md
  • scripts/overfit_detector.py
  • scripts/walk_forward.py

Open the folder on GitHubat commit 981e1d7

Compare with similar skills

Walk Forward Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Walk Forward Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Walk Forward Validation this skillagiprolabs/claude-trading-skills410—~2.2kAutomated safety check: PassMIT
Longbridge Quanthelsome/folio2701 repos~1.6kAutomated safety check: PassMIT
Forecastingericrisco/rsc-harness174—~2.8kAutomated safety check: PassMIT
Options Spread Conviction EngineLeoYeAI/openclaw-master-skills2.2k—~5.7kAutomated safety check: NotesMIT
QuantMind Training Config Generatorqusong0627/QuantMind1.7k—~1.5kAutomated safety check: PassAGPL-3.0
Machine Learning Trading StrategyHKUDS/Vibe-Trading35k—~3.2kAutomated safety check: PassMIT

Similar skills

  • Longbridge Quant

    helsome/folio

    Quantitative strategy frameworks: pairs trading/cointegration, volatility regime strategies, seasonality/calendar effects, multi-factor models (IC/IR), factor research and screening, correlation…

    270 GitHub starsUsed in 1 repo~1.6k tokens
    Data & AnalyticsAuto-check passed
  • Forecasting

    ericrisco/rsc-harness

    A skill your agent uses when projecting history forward — sales, demand, units, revenue, signups, traffic — into a defensible number with an error band: method by data shape, rolling-origin…

    174 GitHub stars~2.8k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Options Spread Conviction Engine

    LeoYeAI/openclaw-master-skills

    Multi-regime options spread analysis engine with quantitative rigor.

    2.2k GitHub stars~5.7k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check: notes
  • Turns a plain-language model training request into a validated QuantMind training config file that can be imported from the Model Training page.

    1.7k GitHub stars~1.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Trains scikit-learn models with walk-forward validation on features from OHLCV data to predict return direction and turn the predictions into trading signals.

    35k GitHub stars~3.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Quant Statistical Methods

    HKUDS/Vibe-Trading

    Guides your agent through unit-root, cointegration, GARCH, bootstrap and regression-diagnostic tests on financial time series, using a tested helper module.

    35k GitHub stars~4k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from agiprolabs/claude-trading-skills

All 68 skills in this repo
  • Backtrader

    agiprolabs/claude-trading-skills

    Event-driven backtesting with bar-by-bar execution, complex order types, multiple analyzers, and custom indicators

    410 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Birdeye API

    agiprolabs/claude-trading-skills

    Solana token market data via Birdeye — prices, OHLCV, trades, token metadata, security checks, and trader activity

    410 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Coingecko API

    agiprolabs/claude-trading-skills

    Broad crypto market data from CoinGecko covering 13,000+ tokens.

    410 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Cointegration Analysis

    agiprolabs/claude-trading-skills

    Cointegration testing for pairs trading using Engle-Granger, Johansen, and rolling stability analysis

    410 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Copy Trading

    agiprolabs/claude-trading-skills

    Wallet evaluation, monitoring, and copy-trade strategy design for Solana DEX trading

    410 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Correlation Analysis

    agiprolabs/claude-trading-skills

    Cross-asset correlation analysis including rolling correlation, hierarchical clustering, tail dependence, and regime-dependent correlation

    410 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Walk Forward Validation

What does Walk Forward Validation do?

Walk-forward validation framework for trading strategies and ML models with time-series-aware splits, overfit detection, and regime-aware validation. Walk Forward Validation is an agent skill from agiprolabs/claude-trading-skills.

When should I use Walk Forward Validation?

Walk Forward Validation fits situations like: tasks that involve Machine learning; tasks that involve Trading and backtesting; tasks that involve Forecasting and time series.

How do I install Walk Forward Validation in Claude Code?

Run `npx skills add agiprolabs/claude-trading-skills --skill walk-forward-validation -a claude-code`. Or copy the skill folder (skills/walk-forward-validation in agiprolabs/claude-trading-skills) into .claude/skills/walk-forward-validation in your project. Claude Code loads it when a task matches its description.

How do I install Walk Forward Validation in Codex?

Run `npx skills add agiprolabs/claude-trading-skills --skill walk-forward-validation -a codex`. Or copy the skill folder (skills/walk-forward-validation in agiprolabs/claude-trading-skills) into .agents/skills/walk-forward-validation in your project. Codex loads it when a task matches its description.

Can I use Walk Forward Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agiprolabs/claude-trading-skills --skill walk-forward-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/walk-forward-validation, .gemini/skills/walk-forward-validation, .github/skills/walk-forward-validation and .opencode/skills/walk-forward-validation in your project.

What does Walk Forward Validation need to run?

Going by SKILL.md and its folder, Walk Forward Validation needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Walk Forward Validation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Walk Forward Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Walk Forward Validation use?

Walk Forward Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Walk Forward Validation use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.4k tokens, read only when the agent opens those files.

What are the alternatives to Walk Forward Validation?

Skills that share tags, products or a category with Walk Forward Validation: Longbridge Quant (helsome/folio, 270 stars), Forecasting (ericrisco/rsc-harness, 174 stars), Options Spread Conviction Engine (LeoYeAI/openclaw-master-skills, 2.2k stars) and QuantMind Training Config Generator (qusong0627/QuantMind, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Walk Forward Validation?

agiprolabs (a GitHub user) maintains it in agiprolabs/claude-trading-skills, which has 410 GitHub stars. The repository holds 68 skills in this directory. The repository was last updated on September 3, 2026.

Source: agiprolabs/claude-trading-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.