← NeuPortal blog

Your 30-Day Backtest Has 3,275 Windows and 110 Observations

By ·

Your 30-Day Backtest Has 3,275 Windows and 110 Observations

Here is a mistake that survives code review, passes every unit test, and quietly inflates the confidence of almost every backtest published on the internet.

You want to know how a 30-day holding period behaves. So you pull every daily candle Binance holds for BTCUSDT - 3,304 of them, reaching back to August 2017 - and run a 30-day window along the series. Start at day 1, measure the 30-day return. Start at day 2, measure again. Keep going.

You end up with 3,275 measurements. That feels like a serious sample. You compute a mean, a standard deviation, some quantiles, and you draw error bars around them. The error bars are narrow, because 3,275 is a big number and uncertainty shrinks with the square root of your sample size.

The error bars are wrong. Not slightly wrong. Wrong by a factor of about five.

The window that shares 29 days with its neighbour

Consider two consecutive measurements. The first covers days 1 through 30. The second covers days 2 through 31. They share 29 days out of 30.

They are not two observations of how the market behaves over a month. They are substantially the same observation, measured twice, with one day of new information between them. The third window shares 28 days with the first. The thirtieth window is the first one that shares nothing at all.

Statistics has a name for this. The samples are not independent, and nearly every formula you learned assumes they are. The standard error of a mean, the width of a confidence interval, the p-value in your significance test, the t-statistic on your Sharpe ratio - all of them take your sample size at face value and quietly assume each observation brought fresh information.

Overlapping windows break that assumption completely, and they break it in the direction that flatters you.

The arithmetic, on real history

Non-overlapping windows are the honest count. If you have T daily candles and you want h-day returns, you get T divided by h genuinely independent observations. Everything above that number is the same data counted again.

Measured on Binance daily history as of 2 September 2026:

| Symbol | Candles | Span | 30-day windows | Independent | Inflation | |---|---|---|---|---|---| | BTCUSDT | 3,304 | 9.0 years | 3,275 | 110 | 30x | | ETHUSDT | 3,304 | 9.0 years | 3,275 | 110 | 30x | | BNBUSDT | 3,223 | 8.8 years | 3,194 | 107 | 30x | | SOLUSDT | 2,214 | 6.1 years | 2,185 | 73 | 30x |

The inflation factor is not a coincidence or an artefact of these particular assets. It is exactly h. A 7-day horizon inflates your sample sevenfold. A 30-day horizon inflates it thirtyfold. The longer the horizon, the more aggressively you are counting the same months over and over.

For the one-day horizon the factor is one. Daily windows do not overlap, so a daily band built on 3,304 candles really does rest on 3,304 observations. This is the quiet reason short horizons are so much more defensible than long ones, and it has nothing to do with markets being easier to predict tomorrow than next month.

What this does to your error bars

Uncertainty around an estimate falls with the square root of the sample size. So if you claim 3,275 observations when you have 110, you have understated your uncertainty by the square root of thirty - a factor of about 5.5.

Read that again in practical terms. Your confidence interval is five and a half times narrower than the data supports. A band you believe covers half the outcomes might cover far more, or far fewer, and you have no way to tell from inside the backtest. Your Sharpe ratio comes with a t-statistic computed on a sample size that does not exist. A strategy that looks significant at the 1% level might not clear 20%.

Nothing in your code will warn you. The arithmetic is all correct. The data is real. The only broken thing is the assumption underneath, and assumptions do not throw exceptions.

The yearly horizon is where it stops being subtle

Push the same arithmetic out to a one-year holding period and the illusion becomes almost comic.

| Symbol | 365-day windows | Independent years | Inflation | |---|---|---|---| | BTCUSDT | 2,940 | 9 | 327x | | ETHUSDT | 2,940 | 9 | 327x | | BNBUSDT | 2,859 | 8 | 357x | | SOLUSDT | 1,850 | 6 | 308x |

Two thousand nine hundred and forty windows. Nine independent years.

You can compute a mean annual return from those 2,940 numbers and it will look authoritative to four decimal places. What you actually have is nine observations, and those nine are not even nine draws from a stable process - they include a halving cycle or two, one regulatory era ending and another beginning, and at least one market structure that no longer exists.

Nine observations does not constitute a distribution. It is a small pile of episodes carrying a standard deviation. Any annual band drawn from it is a guess wearing a lab coat, and the honest response is not to widen the error bars. It is to decline to publish the horizon at all.

This is why a serious forecasting system hands you a day, then a week, maybe a month, and refuses to go further. Not because longer horizons are uninteresting, but because the data to support them does not exist yet and no amount of resampling conjures it into being.

Why the mistake is so easy to make

It is worth being sympathetic here, because this is not a beginner's error. It shows up in published papers.

The pull toward overlapping windows is real and rational. Crypto history is short. Non-overlapping 30-day windows on Solana give you 73 observations, and 73 feels uncomfortably thin when you are trying to estimate a tail. Sliding the window feels like extracting more information from limited data rather than manufacturing it.

And overlapping windows genuinely are useful - for estimating the centre of a distribution. Your median and mean get more stable as you add overlapping samples, because those estimators are not badly hurt by correlation between observations. The damage is concentrated in exactly one place: anything that quantifies uncertainty. Standard errors, confidence intervals, significance tests, error bars.

So the rule is not "never slide a window". The rule is that sliding a window improves your estimate of the centre and does nothing whatsoever for your estimate of how uncertain that centre is.

What to do instead

Four habits, in rising order of effort.

**Report the effective sample size, not the row count.** If a chart of 30-day returns rests on 110 independent observations, put 110 on the chart. This single change kills most of the overconfidence, because the number is visibly small and readers calibrate accordingly.

**Compute intervals on non-overlapping windows.** Estimate the centre however you like, then step back and derive your uncertainty from the independent subset only. It is a smaller number and a more honest one.

**Use a block bootstrap rather than a plain one.** Resampling individual days destroys the serial structure that made your windows dependent in the first place, which puts you right back where you started. Resampling contiguous blocks preserves it.

**Let the data decide which horizons you offer.** If a horizon does not have enough independent history to support a band, that is not a formatting problem to be solved with wider bounds. It is a signal to drop the horizon.

The part that survives

Strip out the statistics and one idea remains.

A backtest cannot tell you how confident to be in a backtest. Every number it produces comes from the same finite slice of history, and the more aggressively you resample that slice, the more certain everything looks while nothing new has been learned. The row count grows. The information does not.

The only thing that adds genuinely independent observations is time passing, with predictions committed in advance so they cannot be revised afterwards. Nine years of Bitcoin gives you nine independent years no matter how finely you slice it. The tenth one arrives on schedule, not by resampling.

That is the uncomfortable arithmetic underneath every long-horizon claim you will read this week. Before you trust one, ask how many independent observations it rests on. The answer is usually smaller than the chart implies, and often small enough to change your mind.

Our own board applies exactly this test to itself. Each call is sealed ahead of its resolution and marked once the outcome lands, failures kept on the page. The board sits at neuportal.ai/experiment. It exists to be audited, not admired.

*Educational content - not financial advice.*