Crowding and options: three pre-registered tests, no edge
A model of crowded crypto derivatives, a volatility rule borrowed from currency research and its mirror image. All three were written down and hashed before the data was touched. None produced an edge; the third now runs as a live paper test.
Key findings
- A crowding model failed. Gradient-boosted trees on Bitcoin open interest, funding, perpetual premium and taker flow, retrained monthly and tested out of sample over 2022–2026, earned +8.8% a year of alpha with a t-statistic of 0.84. The pre-registered bar was 2.0. The model did not rank the next 72 hours at all.
- Buying volatility after an inverted short end lost about 20% per trade. When 7-day implied volatility traded above 30-day, a weekly at-the-money straddle bought at real prices returned −19.5% on Bitcoin (79 trades) and −21.3% on Ether (55 trades). The rule was taken from currency research; in crypto it picks the weeks in which options are most expensive.
- The mirror image could not be tested on free data. Selling a defined-risk iron butterfly in those weeks needs four legs filled by real trades, which happened in only 15 of the 50 required weeks. It now runs as a live paper test on Deribit quotes, every Friday from 9 October 2026, for at least a year.
These three tests were designed after the rest of this site was finished, in a dialogue with a second AI model (ChatGPT), which proposed the hypotheses, the features and the success criteria; we checked the data, argued with parts of the design and implemented it. Each protocol was written down before any number was computed. Its SHA-256 hash and a UTC timestamp went into a log, every later change was a dated addendum made before the results, and each test was run exactly once.
Test 1: crowded derivatives and the next 72 hours
Hypothesis. Open interest, funding, the premium of the perpetual future over the index and aggressive (taker) flow tell whether the Bitcoin derivatives market is crowded. A crowded market should have a lopsided spot return over the following 72 hours, beyond what plain Bitcoin exposure explains.
Design. One decision a day at 00:00 UTC from data known by then, execution an hour later on spot, a fixed 72-hour hold. Eighteen features fixed in advance (returns and realised volatility over 1 and 3 days, open-interest changes, funding level and z-score, premium means and z-score, taker imbalance, volume z-score, two interactions, and Ether’s return and funding). An XGBoost regression with hard-coded settings, no tuning, retrained every month on all earlier rows whose outcome was already known. Long Bitcoin when the predicted 72-hour return is above 1.5%, otherwise cash. Costs of 0.30% per side: the exchange fee plus slippage.
The bar. Daily strategy returns regressed on Bitcoin’s daily returns, with Newey–West errors. The alpha needs a t-statistic above 2.0, must stay positive at 0.75% round-trip costs and after leaving out any calendar year, and there must be at least 60 trades. We lowered the bar from ChatGPT’s first proposal of 3.0 because a 21-month test cannot reach it: t > 3 over 1.75 years needs an information ratio of about 2.3. Instead we lengthened the test to 4.75 years by walking forward.
| Criterion | Required | Result | |
|---|---|---|---|
| t-statistic of alpha, 0.60% round trip | > 2.0 | 0.84 | fail |
| Alpha at 0.75% round trip | > 0 | +5.8% a year | pass |
| Trades | ≥ 60 | 94 | pass |
| Alpha without any single calendar year | > 0 | at least +1.6% a year | pass |
| 2022–2026, out of sample | Strategy | Buy and hold | Hold at the same 21% exposure |
|---|---|---|---|
| Annual return | 13.0% | 13.0% | 4.8% |
| Sharpe ratio | 0.59 | 0.49 | 0.49 |
| Maximum drawdown | −46% | −67% | −19% |
Does the crowding model rank the next 72 hours?
Mean BTC spot return over the next 72 hours, by quintile of the model's out-of-sample prediction, 2022–2026.
Show data · values in %
| Prediction quintile, lowest to highest | mean BTC return over the next 72 hours |
|---|---|
| Q1 | 0.41 |
| Q2 | 0.38 |
| Q3 | -0.12 |
| Q4 | -0.28 |
| Q5 | 0.55 |
The strategy beat 99% of random entry schedules with the same number of trades and the same costs, but random timing pays the full costs with little market exposure, so that comparison mostly says the entries were not random. As a forecast the model was empty: the correlation between its predictions and the outcomes was 0.01, and the quintiles above have no order. Selection skill is possible, money after costs is not shown. By the protocol this is a failure, with no second run.
Test 2: buying the straddle when the short end inverts
Hypothesis. A 2025 study of currency options (Kostakis, McBride, Sarno and Wang) found that an inverted implied-volatility term structure, with short-dated volatility above longer-dated, comes before large moves. If that holds for crypto, a long straddle bought in those weeks should pay.
Design. Every Friday from 2019 to 2026. The signal uses only option trades before 07:00 UTC: the amount-weighted median implied volatility of at-the-money calls and puts expiring in 7 days and in about 30 days. When the 7-day value is higher, buy the 7-day at-the-money straddle at the first buyer-initiated trades between 07:00 and 10:00, so the bid–ask spread is in the price, and hold it to expiry. The payoff uses the exchange’s official delivery price. Fees are 0.03% of the underlying per leg and 0.015% at delivery, both capped at 12.5% of the option’s value. Success needed a bootstrap p-value below 0.05, a positive mean under a 5% worse fill, at least 50 trades and a result better than buying the straddle every week.
The free trade data was too thin for the first version of the rules (4 possible trades). Before computing any return we fixed a ladder of three looser variants and took the first that reached 60 signal weeks with fills. That was the loosest one.
| Bitcoin | Ether (replication) | |
|---|---|---|
| Trades in signal weeks | 79 | 55 |
| Mean net return per straddle | −19.5% | −21.3% |
| Median | −50% | −48% |
| Share of profitable straddles | 30% | 25% |
| Bootstrap p-value (mean > 0) | 0.995 | 0.992 |
| Every Friday with a fill, mean | +1.0% (253) | −3.7% (190) |
Buying the weekly straddle after an inverted short end
Mean net return per weekly at-the-money straddle, bought at real buyer-initiated Deribit prices and held to expiry, after fees.
Show data · values in %
| weeks with 7-day IV above 30-day IV (79 BTC, 55 ETH trades) | every Friday with a valid fill (253 BTC, 190 ETH) | |
|---|---|---|
| BTC | -19.50 | 1.00 |
| ETH | -21.30 | -3.70 |
The loss in signal weeks, year by year (BTC)
Mean net return of the weekly straddle in weeks with an inverted short end, by calendar year. Labels show the number of trades.
Show data · values in %
| BTC, signal weeks | |
|---|---|
| 2019 (6) | -39.70 |
| 2020 (2) | -63.90 |
| 2021 (8) | -34.30 |
| 2022 (12) | -15.30 |
| 2023 (16) | 1.90 |
| 2024 (9) | 2.70 |
| 2025 (15) | -22 |
| 2026 (11) | -40.30 |
The slope did not predict the size of the next week’s move either: the rank correlation with the absolute move was −0.05 for Bitcoin and −0.01 for Ether. An inverted short end simply means short-dated options are expensive, and the move that follows does not pay for them. The currency result does not carry over.
Test 3: selling it instead, and why it runs live
The obvious next idea is to sell in those weeks. Bitcoin’s history could no longer serve as evidence for it, because the idea came from Bitcoin’s own result. So we wrote the protocol on Bitcoin’s result, froze it, and only then opened Ether: sell the at-the-money call and put, buy wings 10% away to cap the loss, and measure the return on the maximum loss. ChatGPT argued against a naked short straddle because its tail risk is open-ended, and we agreed.
The test could not run. With real trades instead of quotes, all four legs filled within 30 minutes in only 40 Fridays, and in only 15 of them was there a signal; the protocol required 50. We declared it infeasible and adapted nothing.
It now runs forward as a paper test on Deribit’s live order book, with a new protocol frozen before the first observation:
- every Friday at 07:00 UTC from 9 October 2026, Bitcoin and Ether;
- the signal from live mark implied volatility, the short legs filled at the best bid, the wings at the best ask, settlement at the delivery price;
- at least 52 weeks, preferably 104; the first 26 are only a check that it works;
- judged on cumulative return, mean return per unit of risk and the worst 5% of weeks;
- stopped early only if the cumulative loss reaches 10 times the risk per trade.
Its results will appear on the live page.
Caveats
- Trade prints, not quotes. The option tests use executed trades. A buyer-initiated trade is a reasonable proxy for the ask a small buyer would pay, but it is not an order-book snapshot, and the thin data forced the loosest pre-set variant.
- Bitcoin and Ether are not independent. Their volatilities move together, so the Ether replication is a weaker check than a second market would be.
- A data gap in the crowding test. The exchange archive has no taker data from late December 2021 to mid-May 2022, so the model did not trade in those months, as the protocol required. The two Ether features with the same gap were dropped before the first fit.
- Researcher degrees of freedom. These hypotheses were chosen by people who had already seen the rest of this site. Pre-registration protects each test, not the choice of which tests to run.
- The fee changes. From 5 October 2026 the European spot exchange charges a flat 0.25% per side at the entry tier, so the crowding test’s 0.30% per side is close to what it would pay today.