A skill your agent uses when running, interpreting, or designing backtests on Superior Trade — anything about backtest windows, trade-count thresholds, exit-reason mix, parameter sweeps…

MITAuto-check passedBusiness, Finance & HR

Install Backtesting

skills CLI
$ npx skills add Superior-Trade/superior-skills --skill backtesting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Superior-Trade/superior-skills backtesting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Superior-Trade/superior-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/backtesting .claude/skills/backtesting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
backtesting
GitHub stars
214
Used in
1 other repo
Token cost
~2.9k tokens
SKILL.md length
1,429 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when running, interpreting, or designing backtests on Superior Trade — anything about backtest windows, trade-count thresholds, exit-reason mix, parameter sweeps…

  • Works in 2 steps: Walk-forward to a similar pair. Pick a… → Defer the deployment recommendation. A…
  • Designing backtests on Superior Trade — anything about backtest windows
  • SKILL.md covers The trade-count bar (sample…, Pick a backtest window that…, Don't optimize against the… and The 3-variant parameter sweep, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Backtesting is an agent skill from Superior-Trade/superior-skills. Use when running, interpreting, or designing backtests on Superior Trade — anything about backtest windows, trade-count thresholds, exit-reason mix, parameter sweeps, walk-forward validation, zero-trade diagnosis, compute-cost estimation, or "is this backtest result trustworthy?". Pair with the relevant strategy skill, such as mean-reversion or breakout.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Business, Finance & HR, covering Trading and backtesting. The repository describes itself as: Open agent skills and tool schemas for Superior Trade — build, backtest, and deploy trading strategies on Hyperliquid. The licence is MIT.

When your agent uses it

  • Designing backtests on Superior Trade — anything about backtest windows
  • Trade-count thresholds
  • Exit-reason mix
  • Parameter sweeps

Example prompts

  • “is this backtest result trustworthy?”
  • “/backtesting”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Walk-forward to a similar pair. Pick a sibling perp with longer history (e.g. test the strategy on BTC if BTC-AAPL is too new), then…
  2. Defer the deployment recommendation. A 14-day backtest is too thin to recommend live capital, period. Tell the user that.

What it can do on your machine

Read from SKILL.md and the folder at commit 9d41db5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Backtesting loads about 2.9k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 1,429 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Superior-Trade/superior-skills at commit 9d41db5, republished under its MIT licence (© Superior-Trade). 1,429 words, ~2,907 tokens.

Download SKILL.mdSave it as .claude/skills/backtesting/SKILL.md (or your agent's skills folder).
name
backtesting
description
Use when running, interpreting, or designing backtests on Superior Trade — anything about backtest windows, trade-count thresholds, exit-reason mix, parameter sweeps, walk-forward validation, zero-trade diagnosis, compute-cost estimation, or "is this backtest result trustworthy?". Pair with the relevant strategy skill, such as mean-reversion or breakout.
metadata.version
0.1.0
metadata.updated
2026-05-07

Backtesting Best Practices

Read ../../references/unified-runtime.md for the shared API lifecycle. Submit every run with POST /runtime/backtests; creation queues it, so there is no separate start request. Use the framework and venue fields defined by GET /openapi.json and the relevant venue skill. This page is about the judgment calls — picking a window that means something, telling signal from noise in the result, and knowing when to give up vs. iterate.

The trade-count bar (sample size first)

Trade count is the single most important number on a result page. Look at it before PnL, before Sharpe, before win rate.

Trade countVerdict
< 30Coincidence, not a strategy. Don't promise anything; widen entries or extend window.
30-50Marginal. Sharpe is noisy. Treat results as directional, not numeric.
50-200Useful. Sharpe / profit factor start to mean something.
200+Statistical confidence. Now you can compare variants on micro-differences.

Watch for the trap: backtests with 5-10 trades and a 100% win rate. They look like world-beaters and almost always disintegrate live. The strategy is too selective — every signal is a coin flip you've cherry-picked, not a repeatable edge. Widen the entry threshold, lengthen the window, or accept that there's no statistical signal here.

Pick a backtest window that means something

A great backtest over the wrong window is a great fiction.

The window should answer: "if I had deployed this strategy on day one of this window, what would have happened?" — not "what's the prettiest curve I can fit?"

Cover at least one regime change

Pure bull, pure bear, sideways chop — your window should include at least two of the three. A 90-day backtest in a one-direction market is a 90-day cherry-pick. A momentum strategy that prints +50% over a +60% trending window has told you nothing about itself; it's just measured beta.

Useful default windows
  • 6 months of 1h data, or
  • 18 months of 1d data

Less than that and you're really looking at noise. More than that and Hyperliquid's history may not cover the pair (HL was launched in 2023; many alts have < 12mo of data).

When you can't get a long window

Some HIP3 / new-listing pairs only have a few weeks of data. Two pragmatic responses:

  1. Walk-forward to a similar pair. Pick a sibling perp with longer history (e.g. test the strategy on BTC if BTC-AAPL is too new), then assume the result transfers ±20%.
  2. Defer the deployment recommendation. A 14-day backtest is too thin to recommend live capital, period. Tell the user that.

Don't optimize against the test set (the cardinal sin)

If you tune parameters and validate on the same window, your backtest is a souvenir, not a forecast. This is the single most common reason strategies that look brilliant on paper die in the first week of live trading.

The honest workflow:

  1. Pick parameters from theory or a single observation, NOT from a sweep.
  2. Run the backtest.
  3. If results are bad, change the idea, not the parameters.

If you do sweep, walk forward to a different period or pair before promoting the winner. A parameter that wins on Q1 BTC and also wins on Q1 ETH (out-of-sample for the second pair) is real. A parameter that wins on Q1 BTC and gets re-validated on Q1 BTC is theatre.

Red flag

If the user's "strategy" started as a 5-variant sweep and now they want to deploy the winner — the result on the test window is overstated. Tell them to either run an out-of-sample check or accept they're deploying on optimistic numbers.

The 3-variant parameter sweep

When the user wants to find the right parameter (e.g. RSI threshold, ATR multiplier, lookback length), run 3 variants in one batch, not iterative single backtests. The shape of the 3-result table tells you what to do:

VariantConfigPnL%TradesSharpeMax DD
ARSI < 25…………
BRSI < 30…………
CRSI < 35…………

Read the shape:

  • All 3 profitable → pick the best Sharpe (not best PnL — PnL rewards luck on small samples). Parameter is robust; proceed to walk-forward / deployment with confidence.
  • 1-2 profitable → pick the winner, but flag that the param is sensitive. Either (a) walk-forward to a 2nd pair, or (b) one tighter sweep around the winner.
  • All 3 unprofitable / < 10 trades each → the idea doesn't work on this pair. Don't sweep again. Switch pair or strategy.
  • Monotonic edge (PnL strictly improves A → B → C) → the best variant is at the edge of the grid. Run one more variant past it (e.g. RSI < 40). Don't run another full 3-grid; just extend by one.

Read the exit-reason mix

After the backtest completes, look at the breakdown of how trades ended. A healthy strategy ends through a mix of:

  • Take-profit / minimal_roi
  • Signal-based exit (populate_exit_trend)
  • Trailing stop
  • Time-based timeout (custom_exit)

Each one tells a different story. Distortions diagnose specific bugs:

DistortionMeaning
90% stoploss hitsStop is too tight relative to the strategy's natural noise. Widen, or accept this strategy is in the wrong volatility regime for this pair.
90% time-based timeoutsExit signal does nothing. Either the exit conditions are too restrictive, or there's no real edge — you're just holding until the timer rings.
100% take-profit hitsminimal_roi is the strategy. The exit signal isn't earning its keep — could be removed or tightened.
50/50 stop vs profitHealthy. Strategy is choosing actively in both directions.
Show full SKILL.md (546 more words)Show less

Zero-trade diagnosis

Zero trades from a backtest means the entry condition never fired. Five common causes, in order of likelihood:

  1. Indicator threshold too strict (e.g. RSI < 30 rarely triggers on BTC 15m).
  2. startup_candle_count larger than available data — the indicator stays NaN forever.
  3. Pair has no data for the requested timerange.
  4. Wrong pair format for trading mode (BTC/USDC for futures will silently produce no fills).
  5. Multi-condition entry that's effectively false AND false — each condition individually rare, joined by AND rarer still.
The escalation rule (don't burn cycles)
  • 1st zero-trade backtest: Diagnose. Widen the threshold, drop one of the AND conditions, or extend the window. Next attempt should be a 3-variant sweep, not a single guess — gives sensitivity in one shot.
  • 2nd consecutive all-zero result: STOP. Don't run a third. Tell the user: "Two consecutive attempts returned zero trades. The strategy needs a fundamentally different approach." Suggest a different signal (mean-reversion vs momentum, BB vs RSI), a different timeframe, or a different pair.
  • Never run 3+ all-zero attempts in a row on the same idea. Wastes compute and user trust.

Compute cost (estimate before submitting)

Backtest run time scales with candle count. Estimate before submitting so you can set the user's expectations:

candles ≈ (days × 24) / timeframe_hours × number_of_pairs
Candle countBehavior
< 100KSubmit normally, poll every 10s.
100K–500KWarn user "may take a few minutes", use exponential polling.
> 500KWarn user "could take 10+ minutes", longer poll intervals.

Reference points:

  • 1 pair × 15m × 90 days = 8.6K candles (fast)
  • 3 pairs × 5m × 90 days = 78K candles (fast)
  • 12 pairs × 1m × 90 days = 1.5M candles (slow — warn upfront)

Always allow the user to proceed. Just set expectations. Don't block on size.

Exponential polling cadence

When polling backtest_status on long runs:

  • Polls 1-3: every 10s
  • Polls 4-6: every 20s
  • Polls 7-9: every 30s
  • Polls 10+: every 60s

This prevents hitting the 20-step tool limit on million-candle backtests.

What headline numbers to trust (and not)

MetricTrust atNotes
Total tradesAlways look first.Below 30 = ignore everything else.
Total profit %Useful for ranking, weak for forecasting.Large windows + small per-trade edge can produce big PnL from luck.
Win rateOK above 50 trades.A 60% WR on 8 trades is one good week, not an edge.
Sharpe ratioAbove 1.0 = good, above 2.0 = excellent.But fragile under 50 trades; don't quote "Sharpe 2.0" off a 12-trade run.
Profit factorGross gains / gross losses. > 1.3 is the practical floor.Penalizes hidden tail losses better than Sharpe.
Max drawdown> 20% is risky for retail-sized accounts.A 5% Sharpe-1 strategy with 30% DD is unrunnable for most users.
Avg holdingSanity check — does it match the strategy's intent?A "scalp" with 12h avg holding is misnamed; a "swing" with 5min holding likewise.

Walk-forward (the honest validation)

Before recommending live deployment, run the strategy on out-of-sample data at least once:

  1. Out-of-sample period: take the most recent 30 days that weren't in your tuning window. Re-run unchanged. Numbers should be in the same ballpark — not better, not dramatically worse.
  2. Out-of-sample pair: run on a sibling pair (ETH if you tuned on BTC). Allow ±30% degradation; anything beyond that means the parameters were pair-specific not regime-specific.

If both walk-forwards survive, you have a defensible recommendation. If either falls apart, you have a backtest, not a strategy.

© Superior-Trade, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/backtesting of Superior-Trade/superior-skills.

Open the folder on GitHubat commit 9d41db5

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Superior-Trade/superior-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Backtesting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Backtesting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Backtesting this skillSuperior-Trade/superior-skills2141 repos~2.9kAutomated safety check: PassMIT
Tushare Datazillionare/zillionare3192 repos~2.3kAutomated safety check: PassNone
Tradingview MCPatilaahmettaner/tradingview-mcp5k—~1.3kAutomated safety check: PassMIT
Digital Oraclekomako-workshop/digital-oracle870—~5.9kAutomated safety check: PassMIT
Fintoolsecond-state/fintool3161 repos~5.9kAutomated safety check: PassNone
Polyclawchainstacklabs/polyclaw3601 repos~2kAutomated safety check: PassApache-2.0

Similar skills

  • Tushare Data

    zillionare/zillionare

    面向中文自然语言的 Tushare 数据研究技能。用于把“看看这只股票最近怎么样”“帮我查财报趋势”“最近哪个板块最强”“北向资金在买什么”“给我导出一份行情数据”这类请求,转成可执行的数据获取、清洗、对比、筛选、导出与简要分析流程。适用于 A 股、指数、ETF/基金、财务、估值、资金流、公告新闻、板块概念与宏观数据等研究场景。

    319 GitHub starsUsed in 2 repos~2.3k tokens
    Business, Finance & HRAuto-check passed
  • Tradingview MCP

    atilaahmettaner/tradingview-mcp

    AI Trading Intelligence — live prices, 30+ technical indicators, backtesting (6 strategies), walk-forward overfitting detection, trade logs, equity curves, licensed news sentiment (Marketaux), and…

    5k GitHub stars~1.3k tokensUpdated today
    Business, Finance & HRAuto-check passed
  • Digital Oracle

    komako-workshop/digital-oracle

    Answer prediction questions using market trading data, not opinions.

    870 GitHub stars~5.9k tokensUpdated 2 mo ago
    Business, Finance & HRAuto-check passed
  • Fintool

    second-state/fintool

    Financial trading CLIs — spot and perp trading on Hyperliquid, Binance, Coinbase, OKX.

    316 GitHub starsUsed in 1 repo~5.9k tokens
    Business, Finance & HRAuto-check passed
  • Polyclaw

    chainstacklabs/polyclaw

    Trade on Polymarket via split + CLOB execution. An agent skill from chainstacklabs/polyclaw.

    360 GitHub starsUsed in 1 repo~2k tokens
    Business, Finance & HRAuto-check passed
  • Markdown

    facioquo/stock-indicators-dotnet

    Format and lint Markdown in this repository against GitHub Flavored Markdown and its markdownlint-cli2 configuration — headers, lists, code fences, callouts (VitePress containers on docs-site pages…

    1.2k GitHub stars~812 tokensUpdated today
    Business, Finance & HRAuto-check passed

More from Superior-Trade/superior-skills

All 31 skills in this repo
  • Deposit Qr

    Superior-Trade/superior-skills

    A skill your agent uses when a user needs a QR code or wallet payment URI to fund a Superior-managed EVM wallet on a specific chain before using Lighter, Polymarket, Hyperliquid, or other Superior…

    214 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Superior Trade Auth

    Superior-Trade/superior-skills

    A skill your agent uses when an agent needs to register a Superior Trade account, verify an email OTP, configure x-api-key authentication, or recover from missing or invalid credentials before using…

    214 GitHub stars~950 tokensUpdated 28 days ago
    Auto-check passed
  • Aerodrome

    Superior-Trade/superior-skills

    A skill your agent uses when creating, validating, backtesting, deploying, sizing, or troubleshooting Aerodrome/Base spot trading strategies through the Superior Trade API, especially Freqtrade…

    214 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • External Deposit

    Superior-Trade/superior-skills

    A skill your agent uses when a user wants to fund or bridge from an external wallet through a third-party UI such as Relay, especially when asking for MetaMask Mobile QR codes, bridge links…

    214 GitHub stars~1.7k tokensUpdated 28 days ago
    Auto-check passed
  • Polymarket

    Superior-Trade/superior-skills

    A skill your agent uses when the user wants to trade, research, or backtest Polymarket prediction markets through Superior Trade — finding markets by slug or event URL, placing a single immediate…

    214 GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • Superior Trade

    Superior-Trade/superior-skills

    A skill your agent uses when a user wants to start trading with Superior Trade and does not have everything set up yet — getting an API key, creating and funding a trading account, choosing a venue…

    214 GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed

Questions about Backtesting

What does Backtesting do?

A skill your agent uses when running, interpreting, or designing backtests on Superior Trade — anything about backtest windows, trade-count thresholds, exit-reason mix, parameter sweeps…. Backtesting is an agent skill from Superior-Trade/superior-skills.".

When should I use Backtesting?

Backtesting fits situations like: designing backtests on Superior Trade — anything about backtest windows; trade-count thresholds; exit-reason mix; parameter sweeps.

How do I install Backtesting in Claude Code?

Run `npx skills add Superior-Trade/superior-skills --skill backtesting -a claude-code`. Or copy the skill folder (skills/backtesting in Superior-Trade/superior-skills) into .claude/skills/backtesting in your project. Claude Code loads it when a task matches its description.

How do I install Backtesting in Codex?

Run `npx skills add Superior-Trade/superior-skills --skill backtesting -a codex`. Or copy the skill folder (skills/backtesting in Superior-Trade/superior-skills) into .agents/skills/backtesting in your project. Codex loads it when a task matches its description.

Can I use Backtesting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Superior-Trade/superior-skills --skill backtesting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/backtesting, .gemini/skills/backtesting, .github/skills/backtesting and .opencode/skills/backtesting in your project.

What does Backtesting need to run?

SKILL.md names no scripts, command-line tools or credentials: Backtesting is instructions for the agent only.

Does Backtesting access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Backtesting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Backtesting use?

Backtesting is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Backtesting use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Backtesting?

Skills that share tags, products or a category with Backtesting: Tushare Data (zillionare/zillionare, 319 stars), Tradingview MCP (atilaahmettaner/tradingview-mcp, 5k stars), Digital Oracle (komako-workshop/digital-oracle, 870 stars) and Fintool (second-state/fintool, 316 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Backtesting?

Superior-Trade (a GitHub organization) maintains it in Superior-Trade/superior-skills, which has 214 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on September 10, 2026.

Source: Superior-Trade/superior-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.