Agent skill

Backtest Expert

by tradermonty in tradermonty/claude-trading-skills

Expert guidance for systematic backtesting of trading strategies.

MITAuto-check passedBusiness, Finance & HR

Install Backtest Expert

skills CLI
$ npx skills add tradermonty/claude-trading-skills --skill backtest-expert -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tradermonty/claude-trading-skills backtest-expert --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tradermonty/claude-trading-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/backtest-expert .claude/skills/backtest-expert && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
backtest-expert
GitHub stars
3k
Used in
5 other repos
Token cost
~2.2k tokens
SKILL.md length
988 words
Files
7 (incl. scripts, references)
Skills in repo
74
Repo updated
First seen
Licence
MIT

At a glance

Expert guidance for systematic backtesting of trading strategies.

  • Works in 6 steps: State the Hypothesis → Codify Rules with Zero Discretion → Run Initial Backtest → …
  • Validating quantitative trading strategies
  • SKILL.md covers Core Philosophy, When to Use This Skill, Prerequisites and Workflow, plus 6 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Backtest Expert is an agent skill from tradermonty/claude-trading-skills. Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/failed_tests.md`, `references/methodology.md` and `scripts/evaluate_backtest.py`).

It sits in Business, Finance & HR, covering Trading and backtesting. The repository describes itself as: Claude Code skills for equity investors and traders — market analysis, technical charting, economic calendars, screeners, and trading strategy development. The licence is MIT.

When your agent uses it

  • Validating quantitative trading strategies
  • Asks about backtesting
  • Strategy validation
  • Robustness testing

Example prompts

  • “beating ideas to death”
  • “/backtest-expert”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. State the Hypothesis
  2. Codify Rules with Zero Discretion
  3. Run Initial Backtest
  4. Stress Test the Strategy
  5. Out-of-Sample Validation
  6. Evaluate Results

What it can do on your machine

Read from SKILL.md and the folder at commit eab8d5c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Backtest Expert loads about 2.2k tokens when it runs, and up to ~5.9k if it reads all its reference files. Until then it costs about 116 tokens; SKILL.md has 988 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from tradermonty/claude-trading-skills at commit eab8d5c, republished under its MIT licence (© tradermonty). 988 words, ~2,163 tokens.

Download SKILL.mdSave it as .claude/skills/backtest-expert/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
backtest-expert
description
Expert guidance for systematic backtesting of trading strategies. Use when developing, testing, stress-testing, or validating quantitative trading strategies. Covers "beating ideas to death" methodology, parameter robustness testing, slippage modeling, bias prevention, and interpreting backtest results. Applicable when user asks about backtesting, strategy validation, robustness testing, avoiding overfitting, or systematic trading development.

Backtest Expert

Systematic approach to backtesting trading strategies based on professional methodology that prioritizes robustness over optimistic results.

Core Philosophy

Goal: Find strategies that "break the least", not strategies that "profit the most" on paper.

Principle: Add friction, stress test assumptions, and see what survives. If a strategy holds up under pessimistic conditions, it's more likely to work in live trading.

When to Use This Skill

Use this skill when:

  • Developing or validating systematic trading strategies
  • Evaluating whether a trading idea is robust enough for live implementation
  • Troubleshooting why a backtest might be misleading
  • Learning proper backtesting methodology
  • Avoiding common pitfalls (curve-fitting, look-ahead bias, survivorship bias)
  • Assessing parameter sensitivity and regime dependence
  • Setting realistic expectations for slippage and execution costs

Prerequisites

  • Python 3.9+ (for evaluation script)
  • No API keys required
  • No external data dependencies — metrics are user-provided

Workflow

1. State the Hypothesis

Define the edge in one sentence.

Example: "Stocks that gap up >3% on earnings and pull back to previous day's close within first hour provide mean-reversion opportunity."

If you can't articulate the edge clearly, don't proceed to testing.

2. Codify Rules with Zero Discretion

Define with complete specificity:

  • Entry: Exact conditions, timing, price type
  • Exit: Stop loss, profit target, time-based exit
  • Position sizing: Fixed $$, % of portfolio, volatility-adjusted
  • Filters: Market cap, volume, sector, volatility conditions
  • Universe: What instruments are eligible

Critical: No subjective judgment allowed. Every decision must be rule-based and unambiguous.

3. Run Initial Backtest

Test over:

  • Minimum 5 years (preferably 10+)
  • Multiple market regimes (bull, bear, high/low volatility)
  • Realistic costs: Commissions + conservative slippage

Examine initial results for basic viability. If fundamentally broken, iterate on hypothesis.

4. Stress Test the Strategy

This is where 80% of testing time should be spent.

Parameter sensitivity:

  • Test stop loss at 50%, 75%, 100%, 125%, 150% of baseline
  • Test profit target at 80%, 90%, 100%, 110%, 120% of baseline
  • Vary entry/exit timing by ±15-30 minutes
  • Look for "plateaus" of stable performance, not narrow spikes

Execution friction:

  • Increase slippage to 1.5-2x typical estimates
  • Model worst-case fills (buy at ask+1 tick, sell at bid-1 tick)
  • Add realistic order rejection scenarios
  • Test with pessimistic commission structures

Time robustness:

  • Analyze year-by-year performance
  • Require positive expectancy in majority of years
  • Ensure strategy doesn't rely on 1-2 exceptional periods
  • Test in different market regimes separately

Sample size:

  • Absolute minimum: 30 trades
  • Preferred: 100+ trades
  • High confidence: 200+ trades
5. Out-of-Sample Validation

Walk-forward analysis:

  1. Optimize on training period (e.g., Year 1-3)
  2. Test on validation period (Year 4)
  3. Roll forward and repeat
  4. Compare in-sample vs out-of-sample performance

Warning signs:

  • Out-of-sample <50% of in-sample performance
  • Need frequent parameter re-optimization
  • Parameters change dramatically between periods
6. Evaluate Results

Questions to answer:

  • Does edge survive pessimistic assumptions?
  • Is performance stable across parameter variations?
  • Does strategy work in multiple market regimes?
  • Is sample size sufficient for statistical confidence?
  • Are results realistic, not "too good to be true"?

Decision criteria:

  • ✅ Deploy: Survives all stress tests with acceptable performance
  • 🔄 Refine: Core logic sound but needs parameter adjustment
  • ❌ Abandon: Fails stress tests or relies on fragile assumptions

Use the evaluation script for a structured, quantitative assessment:

bash
python3 skills/backtest-expert/scripts/evaluate_backtest.py \
  --total-trades 150 \
  --win-rate 62 \
  --avg-win-pct 1.8 \
  --avg-loss-pct 1.2 \
  --max-drawdown-pct 15 \
  --years-tested 8 \
  --num-parameters 3 \
  --slippage-tested \
  --output-dir reports/

The script scores across 5 dimensions (Sample Size, Expectancy, Risk Management, Robustness, Execution Realism), detects red flags, and outputs a Deploy/Refine/Abandon verdict.

Key Testing Principles

Punish the Strategy

Add friction everywhere:

  • Commissions higher than reality
  • Slippage 1.5-2x typical
  • Worst-case fills
  • Order rejections
  • Partial fills

Rationale: Strategies that survive pessimistic assumptions often outperform in live trading.

Seek Plateaus, Not Peaks

Look for parameter ranges where performance is stable, not optimal values that create performance spikes.

Good: Strategy profitable with stop loss anywhere from 1.5% to 3.0% Bad: Strategy only works with stop loss at exactly 2.13%

Stable performance indicates genuine edge; narrow optima suggest curve-fitting.

Show full SKILL.md (371 more words)Show less
Test All Cases, Not Cherry-Picked Examples

Wrong approach: Study hand-picked "market leaders" that worked Right approach: Test every stock that met criteria, including those that failed

Selective examples create survivorship bias and overestimate strategy quality.

Separate Idea Generation from Validation

Intuition: Useful for generating hypotheses Validation: Must be purely data-driven

Never let attachment to an idea influence interpretation of test results.

Common Failure Patterns

Recognize these patterns early to save time:

  1. Parameter sensitivity: Only works with exact parameter values
  2. Regime-specific: Great in some years, terrible in others
  3. Slippage sensitivity: Unprofitable when realistic costs added
  4. Small sample: Too few trades for statistical confidence
  5. Look-ahead bias: "Too good to be true" results
  6. Over-optimization: Many parameters, poor out-of-sample results

See references/failed_tests.md for detailed examples and diagnostic framework.

Output

  • reports/backtest_eval_<timestamp>.json — structured evaluation with per-dimension scores, red flags, and verdict
  • reports/backtest_eval_<timestamp>.md — human-readable report with dimension table, key metrics, and red flag details

Resources

Methodology Reference

File: references/methodology.md

When to read: For detailed guidance on specific testing techniques.

Contents:

  • Stress testing methods
  • Parameter sensitivity analysis
  • Slippage and friction modeling
  • Sample size requirements
  • Market regime classification
  • Common biases and pitfalls (survivorship, look-ahead, curve-fitting, etc.)
Failed Tests Reference

File: references/failed_tests.md

When to read: When strategy fails tests, or learning from past mistakes.

Contents:

  • Why failures are valuable
  • Common failure patterns with examples
  • Case study documentation framework
  • Red flags checklist for evaluating backtests

Critical Reminders

Time allocation: Spend 20% generating ideas, 80% trying to break them.

Context-free requirement: If strategy requires "perfect context" to work, it's not robust enough for systematic trading.

Red flag: If backtest results look too good (>90% win rate, minimal drawdowns, perfect timing), audit carefully for look-ahead bias or data issues.

Tool limitations: Understand your backtesting platform's quirks (interpolation methods, handling of low liquidity, data alignment issues).

Statistical significance: Small edges require large sample sizes to prove. 5% edge per trade needs 100+ trades to distinguish from luck.

Discretionary vs Systematic Differences

This skill focuses on systematic/quantitative backtesting where:

  • All rules are codified in advance
  • No discretion or "feel" in execution
  • Testing happens on all historical examples, not cherry-picked cases
  • Context (news, macro) is deliberately stripped out

Discretionary traders study differently—this skill may not apply to setups requiring subjective judgment.

© tradermonty, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in skills/backtest-expert of tradermonty/claude-trading-skills.

  • SKILL.md
  • references/failed_tests.md
  • references/methodology.md
  • requirements.txt
  • scripts/evaluate_backtest.py
  • scripts/tests/conftest.py
  • scripts/tests/test_evaluate_backtest.py

Open the folder on GitHubat commit eab8d5c

Used in 5 other repositories

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in tradermonty/claude-trading-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Backtest Expert next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Backtest Expert compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Backtest Expert this skilltradermonty/claude-trading-skills3k5 repos~2.2kAutomated safety check: PassMIT
Tushare Datazillionare/zillionare3212 repos~2.3kAutomated safety check: PassNone
Tradingview MCPatilaahmettaner/tradingview-mcp5k—~1.3kAutomated safety check: PassMIT
Digital Oraclekomako-workshop/digital-oracle875—~5.9kAutomated safety check: PassMIT
Polyclawchainstacklabs/polyclaw3591 repos~2kAutomated safety check: PassApache-2.0
Markdownfacioquo/stock-indicators-dotnet1.2k—~812Automated safety check: PassApache-2.0

Similar skills

  • Tushare Data

    zillionare/zillionare

    面向中文自然语言的 Tushare 数据研究技能。用于把“看看这只股票最近怎么样”“帮我查财报趋势”“最近哪个板块最强”“北向资金在买什么”“给我导出一份行情数据”这类请求,转成可执行的数据获取、清洗、对比、筛选、导出与简要分析流程。适用于 A 股、指数、ETF/基金、财务、估值、资金流、公告新闻、板块概念与宏观数据等研究场景。

    321 GitHub starsUsed in 2 repos~2.3k tokens
    Business, Finance & HRAuto-check passed
  • Tradingview MCP

    atilaahmettaner/tradingview-mcp

    AI Trading Intelligence — live prices, 30+ technical indicators, backtesting (6 strategies), walk-forward overfitting detection, trade logs, equity curves, licensed news sentiment (Marketaux), and…

    5k GitHub stars~1.3k tokensUpdated 2 days ago
    Business, Finance & HRAuto-check passed
  • Digital Oracle

    komako-workshop/digital-oracle

    Answer prediction questions using market trading data, not opinions.

    875 GitHub stars~5.9k tokensUpdated 2 mo ago
    Business, Finance & HRAuto-check passed
  • Polyclaw

    chainstacklabs/polyclaw

    Trade on Polymarket via split + CLOB execution. An agent skill from chainstacklabs/polyclaw.

    359 GitHub starsUsed in 1 repo~2k tokens
    Business, Finance & HRAuto-check passed
  • Markdown

    facioquo/stock-indicators-dotnet

    Format and lint Markdown in this repository against GitHub Flavored Markdown and its markdownlint-cli2 configuration — headers, lists, code fences, callouts (VitePress containers on docs-site pages…

    1.2k GitHub stars~812 tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Openmobius Skill

    MobiusQuant/OpenMobius-skill

    Provides multi-school trading Q&A, chart/OHLCV analysis, annotation, and fresh-market workflows covering ICT/SMC, ChanLun, Wyckoff, Price Action, Order Flow, VSA, and Elliott Wave.

    695 GitHub stars~7.2k tokensUpdated 1 mo ago
    Business, Finance & HRAuto-check passed

More from tradermonty/claude-trading-skills

All 74 skills in this repo
  • Technical Analyst

    tradermonty/claude-trading-skills

    This skill should be used when analyzing weekly price charts for stocks, stock indices, cryptocurrencies, or forex pairs.

    3k GitHub starsUsed in 4 repos~4.6k tokens
    Auto-check passed
  • Theme Detector

    tradermonty/claude-trading-skills

    Detect and analyze trending market themes across sectors. An agent skill from tradermonty/claude-trading-skills.

    3k GitHub starsUsed in 2 repos~4.9k tokens
    Auto-check passed
  • Trader Memory Core

    tradermonty/claude-trading-skills

    Track investment theses across their lifecycle — from screening idea to closed position with postmortem.

    3k GitHub starsUsed in 2 repos~4.3k tokens
    Auto-check passed
  • Edge Strategy Reviewer

    tradermonty/claude-trading-skills

    Critically review strategy drafts from edge-strategy-designer for edge plausibility, overfitting risk, sample size adequacy, and execution realism.

    3k GitHub starsUsed in 1 repo~988 tokens
    Auto-check passed
  • Sector Analyst

    tradermonty/claude-trading-skills

    This skill should be used when analyzing sector rotation patterns and market cycle positioning.

    3k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Stanley Druckenmiller Investment

    tradermonty/claude-trading-skills

    Druckenmiller Strategy Synthesizer - Integrates 8 upstream skill outputs (Market Breadth, Uptrend Analysis, Market Top, Macro Regime, FTD Detector, VCP Screener, Theme Detector, CANSLIM Screener)…

    3k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed

Questions about Backtest Expert

What does Backtest Expert do?

Expert guidance for systematic backtesting of trading strategies. Backtest Expert is an agent skill from tradermonty/claude-trading-skills. Expert guidance for systematic backtesting of trading strategies.

When should I use Backtest Expert?

Backtest Expert fits situations like: validating quantitative trading strategies; asks about backtesting; strategy validation; robustness testing.

How do I install Backtest Expert in Claude Code?

Run `npx skills add tradermonty/claude-trading-skills --skill backtest-expert -a claude-code`. Or copy the skill folder (skills/backtest-expert in tradermonty/claude-trading-skills) into .claude/skills/backtest-expert in your project. Claude Code loads it when a task matches its description.

How do I install Backtest Expert in Codex?

Run `npx skills add tradermonty/claude-trading-skills --skill backtest-expert -a codex`. Or copy the skill folder (skills/backtest-expert in tradermonty/claude-trading-skills) into .agents/skills/backtest-expert in your project. Codex loads it when a task matches its description.

Can I use Backtest Expert in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tradermonty/claude-trading-skills --skill backtest-expert -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/backtest-expert, .gemini/skills/backtest-expert, .github/skills/backtest-expert and .opencode/skills/backtest-expert in your project.

What does Backtest Expert need to run?

Going by SKILL.md and its folder, Backtest Expert needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Backtest Expert access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Backtest Expert safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Backtest Expert use?

Backtest Expert is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Backtest Expert use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.8k tokens, read only when the agent opens those files.

What are the alternatives to Backtest Expert?

Skills that share tags, products or a category with Backtest Expert: Tushare Data (zillionare/zillionare, 321 stars), Tradingview MCP (atilaahmettaner/tradingview-mcp, 5k stars), Digital Oracle (komako-workshop/digital-oracle, 875 stars) and Polyclaw (chainstacklabs/polyclaw, 359 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Backtest Expert?

tradermonty (a GitHub user) maintains it in tradermonty/claude-trading-skills, which has 2,973 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 5, 2026.

Source: tradermonty/claude-trading-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.