A May 2026 preprint by Carlos Baquero of the University of Porto reviewed 23 Bitcoin price-prediction papers selected for methods, genuine out-of-sample evaluation, or influence, and reached a sobering conclusion. Across the peer-reviewed record, no model demonstrated durable superiority over the appropriate naive benchmark at one-to-six-month horizons spanning several market regimes. The review itself remains awaiting peer review.
The weakness is structural rather than specific to any technique. Naive forecasting works because financial prices are persistent: a model predicting $100,100 tomorrow when Bitcoin trades at $100,000 today can produce a tiny percentage error while learning almost nothing about direction. Francesco Puoti, Fabrizio Pittorino, and Manuel Roveri reached a similar result, comparing 12 approaches (ARIMA, Prophet, random forests, XGBoost, LSTM networks, and N-BEATS) across five major cryptocurrencies at one-day, seven-day, and 30-day horizons. Simple naive models consistently produced better forecasts than every complex alternative tested.
Why it matters
Bitcoin offers enormous quantities of data, but the number of independent market cycles is still small. Millions of minute bars keep repeating observations from the same 2018 bear market or the same 2020 liquidity shock. A rule calibrated to the retail-led 2017 cycle encountered a different derivatives structure in 2021, and spot ETFs created another route for capital and price discovery in 2024. Each era supplies historical data from a version of the market that no longer exists in quite the same form, the non-stationarity problem at the heart of every quantitative forecasting failure.
Backtest overfitting compounds the issue. Trying more model variations raises the odds of finding an excellent historical result through chance, and selecting the winner while hiding the failed attempts produces what David Bailey and his co-authors formalized as the probability of backtest overfitting. Among the peer-reviewed papers Baquero examined, none evaluated the same approach across multiple non-overlapping holdout windows covering different regimes. A single chronological split offers little protection because a researcher can train through 2020 and evaluate in 2021, producing an apparently out-of-sample result that owes much of its performance to a single bull market.
Market impact
The result has direct implications for the valuations retail traders actually use.
Frequently asked questions
-
What did the University of Porto preprint conclude about Bitcoin price forecasting?
Carlos Baquero's May 2026 review examined 23 peer-reviewed Bitcoin prediction papers selected for methods, influence, or genuine out-of-sample evaluation. It found none demonstrated durable superiority over the appropriate naive benchmark at one-to-six-month horizons across several market regimes.
-
Why do naive forecasts beat complex models on Bitcoin?
Financial prices are persistent, so a model predicting $100,100 tomorrow when Bitcoin trades at $100,000 today produces a tiny percentage error while learning almost nothing about direction or return. Puoti, Pittorino, and Roveri confirmed this empirically: simple naive models beat ARIMA, Prophet, random forests,…
-
What is backtest overfitting and why does it matter for crypto models?
Backtest overfitting, formalized by David Bailey and co-authors, describes the elevated probability of finding an excellent historical result through chance when researchers try many model variations. Selecting the winner while hiding failed attempts produces apparent out-of-sample performance that does not survive…
-
Do stock-to-flow and Metcalfe models predict Bitcoin price?
Alexander Shelton's 2024 peer-reviewed examination found stock-to-flow and Metcalfe variables explained Bitcoin returns in-sample but offered limited or zero predictive ability out of sample. Once time effects entered the stock-to-flow regression, its statistical force disappeared. The stock-to-flow model diverged…
-
What would an honest Bitcoin forecasting standard require?
Baquero's review outlines several requirements: publish the naive benchmark beside every model, report every market regime separately rather than as a single aggregate error, include trading costs, release public code and data for reproducibility, and disclose how many model variations were attempted, since that…
CryptoSlate