In-Sample vs Out-Of-Sample Testing: The Test Most Backtests Skip
A backtest on all your data proves almost nothing. Split it. Optimise in-sample, then face the out-of-sample verdict once. That’s the test most traders quietly skip.
You’ve got a beautiful equity curve. Ten years of data. One backtest. No obvious losses. Then you go live and the market behaves like it never got the memo.
This is a plain-English guide to in-sample vs out-of-sample testing.
The missing step was in-sample vs out-of-sample testing.
Most traders don’t skip it because they’ve got a better method.
They skip it because it’s the first time the system is allowed to actually say “no”.
Why one big backtest is basically a confidence trick
The usual ritual looks like this: pull every tick or bar you can get, throw the strategy at it, twiddle parameters until the equity line looks smooth, then declare victory.
That whole process happens on one continuous chunk of data.
You are judging the system on the same history you used to build it.
If that sounds suspiciously like marking your own homework, that’s because it is.
The market gives you a finite sequence of prices. You then ask, “What rules would have made this line go up?” and keep bending the rules until it does.
With enough parameters you can make almost any history look profitable.
But what you’ve usually built is not a trading strategy; it’s a bespoke explanation of that exact past, plus all its random quirks, plus every lucky wiggle you managed to capture by accident.
There’s a name for that: overfitting. I’ve written about it in detail here: /blog/what-is-overfitting-in-trading-and-how-to-spot-it.
In-sample vs out-of-sample testing is how you stop flattering yourself with that one perfect curve.
In-sample vs out-of-sample testing: the simple version
Strip the jargon out and it’s just a data split.
You deliberately keep some data back you’re not allowed to touch until the very end.
Step-by-step, it looks like this:
- Take all your historical data.
- Split it into two parts: in-sample and out-of-sample.
- Design and optimise your strategy only on the in-sample period.
- Freeze the rules.
- Run exactly those frozen rules once on the out-of-sample period.
If it collapses out-of-sample, the message is clear.
The important bit: no sneaking back to adjust after you’ve seen the out-of-sample report.
Once you do that, you’ve turned your out-of-sample into more in-sample.
That’s the most common way traders accidentally (or very deliberately) lie to themselves about robustness.
How much data should you use for each? The boring maths
People love exact formulas for this. There isn’t one.
There are trade-offs, and they’re boring, which is why they actually matter.
Think of each completed trade as one data point.
You need enough in-sample trades to design the rules and get a rough handle on expectancy, drawdown, and variance.
You then need enough out-of-sample trades that the verdict on those rules isn’t pure noise.
As a rough, hypothetical guide:
- If a system takes, say, 2 trades a week, 5–7 years of data might give you 500–700 trades total.
- You might use 60–70% of that for in-sample (300–500 trades).
- The remaining 30–40% (200–300 trades) sits as out-of-sample.
It’s not about the years, it’s about the number of trades and regimes.
You want both samples to see bull phases, bear phases, choppy periods, crashes, dull ranges – the lot.
If you’ve not thought about how many trades you actually need, that’s a separate problem; I break that out here: /blog/how-many-trades-do-you-need-to-test-a-strategy.
But the core point: sacrifice some data from the build process to pay for a clean test later.
What a good in-sample vs out-of-sample result really looks like
Here’s the part people get wrong.
They expect the out-of-sample equity curve to look as perfect as the in-sample one.
If it does, odds are you “accidentally” used the out-of-sample during optimisation, or the market was absurdly kind for a while.
A more realistic pattern:
- In-sample: stronger performance, smoother equity, higher profit factor.
- Out-of-sample: weaker but still positive performance, similar behaviour style.
The key word is similar.
Similar win rate range. Similar drawdown magnitude relative to returns. Similar pattern of runs and recoveries.
If in-sample is a gentle staircase up and out-of-sample is a cliff, that’s not “market changed”; that’s “the system never worked in the first place”.
Compare some basic stats side by side:
| Metric | In-sample | Out-of-sample | Comment |
|---|---|---|---|
| No. of trades | 400 | 220 | Both samples decently large |
| Win rate | 52% | 49% | Close enough to be noise |
| Average R multiple | +0.35R | +0.22R | Worse, but still positive |
| Max drawdown | 12% | 15% | Comparable risk profile |
| Profit factor | 1.6 | 1.3 | Reasonable degradation |
Those are hypothetical, but that pattern is what you’re aiming for.
Not identical.
Survivable.
The ugly version: when out-of-sample fails hard
Let’s outline the more common reality.
In-sample report looks outstanding. High profit factor. Tiny drawdowns. Fifty parameters.
Out-of-sample: flat to down, big equity zigzags, a handful of big losers wiping weeks of grind.
At this point the temptation is strong to “just tweak one thing”.
You nudge an entry filter, add a time-of-day rule, move a stop by a few pips on the trend days where it hurt the most.
Then you re-run the out-of-sample test and miraculously it’s better.
You’ve just done a hidden optimisation pass on the out-of-sample.
It has quietly turned in-sample.
If you really want to make honest use of a failed out-of-sample result, you have two options:
- Admit the system is not robust and bin it (the option nobody likes).
- Use the full data set again to rethink the logic, then create a completely fresh split and rerun the full procedure once.
Is that painful? Yes.
Is it cheaper than trading a curve-fit system live with real money? Also yes.
Walk-forward testing: in-sample vs out-of-sample on repeat
If you want to go further, you can cycle this process with walk-forward testing.
This is where the term out-of-sample validation often shows up.
Mechanically, a simple walk-forward might work like this (dates hypothetical):
- Optimise on 2012–2016 (in-sample 1), test on 2017 (out-of-sample 1).
- Optimise on 2013–2017 (in-sample 2), test on 2018 (out-of-sample 2).
- Optimise on 2014–2018 (in-sample 3), test on 2019 (out-of-sample 3).
You then stitch together the out-of-sample years (2017, 2018, 2019) into a single walk-forward equity curve.
Every point on that curve is performance on data that was unseen at the time of each optimisation window.
That makes it a harsher, and therefore more useful, test of robustness.
Done properly, walk-forward optimisation forces you to accept that edges evolve and degrade.
Done badly, it’s just a more complicated way to overfit everything.
Why in-sample vs out-of-sample matters more for automated trading
Automation doesn’t make a bad edge good.
It just executes your bad edge very, very faithfully.
With discretionary trading, you at least have a human in the loop.
They might pull risk when volatility explodes, or skip a setup into a central bank decision, or simply lose confidence and stop trading when it obviously isn’t working.
With automated trading, if you deploy a system that only ever looked good in-sample, it can keep grinding through that non-edge 24/5 until you intervene.
This is where people confuse speed with safety.
The maths doesn’t care whether a human or a server is clicking the button.
Your risk of ruin is still driven by expectancy, variance, position sizing, and drawdown tolerance.
If you’ve built that expectancy entirely on in-sample data, your risk model is fantasy.
You’re optimising for the past and measuring risk on the same dream.
If you haven’t already, read through the expectancy point properly here: /blog/what-is-trading-expectancy.
When you combine realistic expectancy with in-sample vs out-of-sample testing, you finally have a model that at least lives in the same universe as the market.
Why your position sizing needs out-of-sample reality
Here’s another place this bites: sizing.
People build a beautiful in-sample curve, then plug the metrics into a position sizing formula and get aggressive.
They might use fixed fractional sizing, or some cut-down Kelly, because a high in-sample profit factor makes the numbers look friendly.
But if that edge is inflated by curve-fit noise, you are betting too big on something that was never there.
The system doesn’t just underperform; it can blow up.
Position sizing works on real edges, not imagined ones.
If you want to see the difference between fixed fractional and fixed lot sizing, that’s broken out here: /blog/fixed-fractional-vs-fixed-lot-sizing-that-compounds.
The point in this context: use out-of-sample and, ideally, walk-forward results as the inputs to any risk model.
Base your max drawdown assumptions on the worst out-of-sample sequence, not the prettiest in-sample patch.
If you size off an equity curve that only ever existed in your optimiser, the market will correct you.
What does a sensible split look like in practice?
Let’s say hypothetically you’re building a trend-following system on EURUSD H1 with 10 years of data, and you get around 800 trades total.
One reasonable approach:
- Use the first 6–7 years (say 2014–2020) as in-sample. That’s 480–560 trades.
- Use the last 3–4 years (2021–2024) as out-of-sample. That’s 240–320 trades.
Within that in-sample block you can do your parameter sweeps, robustness checks, and sanity tests.
You look for parameter regions where performance is stable rather than sharp single-point peaks.
You test sensitivity: “If I move the stop by 10%, does it implode?”
Once you’re happy the logic makes sense and isn’t absurdly delicate, you freeze it.
Then you press run on the out-of-sample block and do not touch anything until the test finishes.
Then, and only then, you compare.
Using in-sample vs out-of-sample on faster products
If you’re working on gold (XAUUSD) or indices, you usually get more trades in fewer years.
That can help, because both the in-sample and out-of-sample chunks can contain hundreds of trades even if you only have, say, 5–7 years of decent tick data.
But there’s another issue: volatility.
Gold in particular moves fast, and with something like a 0.01 lot minimum on some accounts, the swings can be large relative to a small balance.
That makes it even more important that your equity curve assumptions come from data the system has not already been “trained” on.
If you size a gold system off a pipedream in-sample test, then trade it on a small account with a 0.01 lot minimum, the percentage swings can be much bigger than you expect.
Out-of-sample is where you find that out in the safety of a simulator, instead of with your actual capital.
The human problem: can you actually sit through the out-of-sample?
One last, awkward point.
Out-of-sample testing isn’t just for the spreadsheet.
It’s for your head.
A valid system will have losing streaks and drawdowns in out-of-sample, just as it will live.
If you look at the out-of-sample equity curve and think “I’d bail at that point”, that’s valuable information.
You’ve just learned something about your own risk tolerance before money is at stake.
There’s a separate deep dive on what sort of losing streaks are normal here: /blog/how-long-a-losing-streak-should-you-actually-expect.
The takeaway is simple: if the honest out-of-sample equity curve is psychologically untradeable for you, the system doesn’t fit you, even if the maths is fine.
The one-line summary: prove it on data the system hasn’t seen
In-sample vs out-of-sample testing is not optional ceremony.
It’s the only real defence you have against convincing yourself that hindsight equals edge.
Split the data.
Optimise in-sample.
Freeze the rules and face the out-of-sample verdict once.
If it holds up, you still haven’t “won”.
You just have a candidate system with some evidence behind it, which you then forward test in small size and keep under review.
Trading involves risk. You can lose some or all of your capital. No test, method, or technology removes that.
Start your free 14-day ArcisTrade demo →
Watch live automated systems running across FX, gold and indices — with winners and losers visible, and test accounts sitting right next to live ones.
No martingale. No grid. Just real curves you can study before you risk a penny.
P.S. Use what you’ve just read: treat the demo as your out-of-sample view of how automation actually behaves.
Common questions
What is the difference between in-sample and out-of-sample data?
In-sample data is the historical period you use to design and optimise a trading strategy. Out-of-sample data is a separate, unseen period you keep back and only test on once the rules are frozen. Performance on in-sample shows how well you fitted the past; performance on out-of-sample is a more honest check of whether there is a real edge.
How much data should I allocate to out-of-sample testing?
There is no fixed percentage, but a common approach is to allocate around 30–40% of your total data to out-of-sample testing, ensuring both sets contain enough trades and different market regimes. What matters most is having a large enough number of out-of-sample trades to reduce the impact of random noise on your verdict about the system.
Can I re-optimise after seeing poor out-of-sample results?
You can, but you must then treat the original out-of-sample period as contaminated. Once you adjust rules based on that performance, it effectively becomes more in-sample data. To keep the process honest, you would need to rebuild using the full history and then create a fresh in-sample and out-of-sample split for a new, single-pass test.
Is walk-forward testing better than a single in-sample vs out-of-sample split?
Walk-forward testing extends the idea by repeatedly optimising on a rolling in-sample window and testing on the next out-of-sample period, then stitching those out-of-sample segments together. Done carefully, it can give a stronger sense of robustness across time. Done carelessly, it just multiplies the opportunities to overfit, so the discipline about not peeking and not over-tuning still applies.