Your backtest says you found an edge. These eight tests check whether the evidence actually holds up. Upload your export from TradingView, MetaTrader 5, or any platform that writes CSV. The tests come from peer-reviewed research and run on your actual trades. Tell us how many versions you tried and we take that back out of the result, because the more you tested, the better the best one looks for no reason at all.
of strategies that win in backtesting underperform on data they never saw. Measured with combinatorial cross-validation, the same method behind our overfitting test.
t-statistic a Sharpe ratio needs before it counts as statistically meaningful. Most backtests never check this.
average drop in Sharpe ratio once commission, spread, and slippage are priced in at realistic levels.
real strategies of known outcome, from a live edge to two killed systems. The engine graded all six in the exact order their real-world verdicts demanded.
Run a strategy in TradingView, QuantConnect, or NinjaTrader and you get numbers. A Sharpe ratio, an equity curve, a win rate, a max drawdown. Those numbers look like answers. They're not.
No platform tells you where the numbers came from. A genuine edge produces a good backtest. So does a strategy fitted too tightly to one stretch of history. The printouts look identical. One makes money live. The other blows up.
Telling them apart takes statistical tests most traders have never heard of, let alone run. Overfitting probability. Sharpe ratios discounted for how many ideas you tried. Out-of-sample decay. Hedge funds run these as a matter of routine. Retail platforms run none of them.
Each test targets one specific way a backtest can mislead you. Together they answer the question no platform asks. Is this result real, or did it just fit the data it was built on?
Did your strategy learn something real, or did it memorize the past? We split your backtest into 12,870 combinations of build and test periods, then check whether the winner in each build window keeps winning out-of-sample. If it usually doesn't, your result was fitted to the data, not discovered in it.
−35 points if PBO exceeds 0.60We shuffle the order of your trades 10,000 times and ask one question. Could pure luck have produced this result? If 30% of random orderings beat your actual strategy, your returns came from which trades landed when, not from your logic.
−25 points if p-value exceeds 0.10We split your backtest in time order. The first 70% is the training window, the last 30% is the test. Then we compare Sharpe ratio, win rate, and profit factor across the two. A strategy that earns Sharpe 2.1 in training and 0.4 in testing has a problem you want to find before your account does.
−35 points if ratio falls below 0.20A Sharpe of 1.9 from your first idea and a Sharpe of 1.9 from your 500th attempt are not the same evidence. This test discounts your Sharpe for how many configurations you tried, the shape of your returns, and the bias that comes from keeping only the best performer.
Checks significance at the 95% levelWe label every stretch of your backtest as trending, ranging, or volatile, then check where the money came from. If 80% of your profit sits in three months of 2021, you have a strategy for one kind of market. It may stop working the day conditions change.
−20 points if 70%+ of profit is concentratedA Sharpe of 2.0 over 30 trades proves nothing. That little data cannot separate a real edge from a lucky streak. We calculate the minimum number of trades your Sharpe ratio needs before it can be trusted, and tell you how far short you are.
−30 points if below the minimum requiredMost backtests assume clean fills and near-zero costs. Live trading charges commission, spread, and slippage on every trade. We model those for your asset class and show what your Sharpe ratio and returns look like after real execution costs.
Flags when cost drag exceeds 40%A number from 0 to 100 and a grade from A to F. Every deduction is named and shown, so you can check the arithmetic yourself.
From TradingView, open the Strategy Tester panel and select "Download data as XLSX." From MetaTrader 5, save the Strategy Tester report. Anything else, QuantConnect, NinjaTrader, or your own Python, exports a CSV of trades. Upload the file as is. Don't open or edit it first.
Eight statistical tests run on your actual trade data. That includes 10,000 Monte Carlo shuffles and 12,870 cross-validation splits for the overfitting check. Every backtest in our six-strategy calibration run, from 164 to 1,111 trades, finished in under 30 seconds.
A number from 0 to 100, a letter grade, and a plain-English readout of every test with the fixes spelled out. Not "consider improving," but "cut the parameter count from 6 to 3." Every upload is saved to your library, so you can rerun after a change and see whether the evidence actually improved.
The score starts at 100. Each test that finds a problem deducts points. Every deduction is listed, with what we found and why it costs you.
A score of 100 does not mean the strategy will make money live. Markets change. The score tells you whether your backtest evidence holds up. It does not predict profit.
Most analysis tools are validated against synthetic data, fixtures built to pass. In August 2026 we ran the full OverfitCheck pipeline on six real strategies from a multi-year systematic-trading research program. Their true quality was already settled by live trading and years of forensic review. The engine was not told which was which.
It got them right. Both killed strategies graded F. The one live, validated edge graded B, and beat every strategy the research program had killed. Ranked by the engine's continuous outputs, the Monte Carlo p-value and the Probabilistic Sharpe Ratio, all six landed in the exact order their real-world verdicts demanded. Every audit finished in under 30 seconds.
The full report is public, including the engine's known limitations. Where a test cannot tell you anything, we say so.
You have months of TradingView strategies behind you. You know what a Sharpe ratio is. You're not sure yours means what it appears to mean. OverfitCheck runs the checks you haven't been running.
You're paying $200 to $500 per challenge attempt. Most failures aren't about raw profitability. The strategy breaks a rule nobody was tracking, or the backtest overstated what it would do live. Find out before you pay the fee, not after.
You run everything through code. Python, QuantConnect, Backtrader. You know these statistical tests exist. You've never gotten around to implementing them properly. OverfitCheck runs them correctly, in seconds, every time.
Three audits free, no card. The paid report is for a strategy somebody is selling you, where a verdict is worth having before money moves. Every tier runs the same eight tests on the same engine, so paying more cannot buy a friendlier answer.
The manual treatment. Cost forensics, session-integrity sweep, stress-test battery, survival sizing. For operators with live money at stake. Apply-only because it takes analyst hours, and we will tell you when your case does not need one.
We take no affiliate or referral money from any prop firm, broker, or EA vendor.
Four checks specced from the same research program that produced our calibration set. Listed so you can see where the product is going. Nothing here is for sale until it exists.
Does your backtest trade the session you think it trades? Timestamp forensics on the trade log. Catches daylight-saving drift and server-clock changes.
Spread measured against each trade’s own stop distance. Produces a per-era breakeven win rate and flags the years your strategy could not afford by construction.
Finds fixed-point parameters whose meaning drifted as the instrument grew. One backtest that is secretly two different strategies.
Thousands of resampled challenge paths against prop-firm rulesets, with honest cleared-payout accounting. P(pass) is not the number that matters.
OverfitCheck provides statistical analysis of backtest data only. Results do not constitute financial advice, investment recommendations, or a guarantee of future performance. Statistical robustness in historical testing does not predict live trading outcomes. All trading involves substantial risk of loss. The score and all test results are tools for your own informed decision-making. You are solely responsible for any trading decisions you make.