Methods
How the studies on this site were run: sample splits, cost models, simulators, statistics, safeguards against look-ahead, and the rules we held ourselves to.
The rules we followed
- Costs are always included. Every simulator charges fees, spread and commission per trade or per unit of turnover, and currency tests that hold overnight also pay the broker’s swap mark-up for every night.
- Trade on the next bar. A signal computed on a bar’s close is acted on in the next bar, minute or second, and unit tests prove it.
- Lag every non-price series by its publication delay. Thresholds use only past data.
- Split before searching. Training, then validation, then a hold-out opened once for finalists chosen in advance. When a hold-out had already been seen, we say so.
- Correct for multiple testing and compare with chance: Benjamini–Hochberg across whole families of tests, random rules, random entries and placebo days or levels.
- Use robust standard errors (Newey–West) whenever returns overlap or are autocorrelated.
- Pre-register new tests with a cost-based decision rule. A significant t-statistic alone never counts.
- One final test per hypothesis. The vault scripts and the ladder test run once.
- Report negative results. Every configuration’s result is kept, and almost all of them are failures.
- Demo accounts only. Live servers are refused in code.
- Production must equal research. A unit test checks that the deployed model reproduces the research model.
- Write down what was chosen after seeing results instead of presenting it as out-of-sample.
Sample splits
| Study | Training | Validation | Hold-out |
|---|---|---|---|
| Kronos forecasts | pre-trained weights; fine-tuning on data before Aug 2025 | Aug 2025 | forecasts from 1 Sep 2025 to 31 Aug 2026 |
| Classic crypto and FX rules, dips, patterns | 2020–2023 | 2024 | 1 Jan 2025 – 31 Aug 2026, opened once |
| Calendar, weather and Moon hypotheses | 2012–2019 | 2020–2023 | 2024 – Sep 2026 |
| Indicator atlas and combinations | to end of 2023 | – | 2024 – Sep 2026 |
| FX intraday and cross-pair grids | 2016–2022 | 2023–2024 | 2025–2026, finalists only |
| FX tick scalping and exits | first two thirds of six months of ticks | – | last third |
| FX weekly positioning and trend | before 2016 | – | 2016–2026 |
| FX 11-predictor lab | G10, 2001–2013 | – | 2014–2026, emerging markets as a control |
| FX daily mean reversion re-test | – | – | 2000–2015, never used by the atlas |
| Direction model | data to end of 2022 | – | 2023 – Sep 2026, plus 9 altcoins never used in selection |
| Stablecoin signal | no split; stability across sub-periods | – | the live demo run since 28 Sep 2026 |
| Order-book imbalance | Mar – Jun 2026 | – | Jun – Sep 2026 |
| Depth ladder test | pre-registered, evaluated once on at least 24 covered hours | – | – |
The first-round hold-out was protected in code: the research scripts only compute training and validation, and a separate script with hard-coded finalists opened the hold-out once. In the currency grids the hold-out statistics were computed for every configuration but read for only one pre-declared hypothesis, so there the protection was discipline, not code.
Costs
| Market | Cost model |
|---|---|
| Crypto spot | 0.25% taker or 0.10% maker per side; research presets add slippage of 0.02% (BTC, ETH) to 0.10% (altcoins) per side |
| Crypto, 1-minute dip tests | maker 0.10%, taker 0.25%, 0.10% slippage on market orders, borrow cost for shorts, and limit orders that fill only when the price trades 0.1% through them |
| Indicator atlas | 0.1% per side for crypto, 0.5 bp per side for currencies and gold, per unit of position change |
| FX, 1-minute grids | spread by instrument and hour of day, measured from tick samples, plus 1 bp commission per round trip |
| FX, tick scalping | real bid and ask quotes on a 1-second grid plus 0.6 pip commission (1.1 pip in the exit study) |
| FX, daily to monthly | 0.3–3 bp of turnover depending on the market, plus the broker’s swap mark-up for every night held |
| FX, order-book tests | 250 ms delay on entry and exit, 0.6 pip commission |
Simulators
- Bar backtester: the target position from a bar’s close is held over the next bar, and costs are charged on every change of position, including the first entry.
- Indicator lab: numba state machines for every rule, with the same next-bar convention and per-trade accounting.
- 1-minute currency simulator: entries at the next bar’s open, longs at the ask and shorts at the bid, the stop checked before the target within a bar, gaps through the stop filled at the open, trailing stops from the previous bar’s best price, and a forced exit before the rollover.
- 1-minute dip simulator: tranches as limit orders that need a trade-through, no take-profit in a bar that added a tranche, and the stop wins ties.
- Tick simulator: a 1-second grid of bid and ask quotes, signals at the end of a second and entries in the next.
- Order-book simulators: tick level with a 250 ms delay, passive fills for limit orders and markouts at 30 and 120 seconds.
Statistics
- Newey–West t-statistics with the lag length matched to the overlap of the return windows.
- Overlap-corrected event tests for pattern events, with events at the same time pooled across instruments.
- Benjamini–Hochberg false-discovery control across whole families of tests, with a Bonferroni step inside a rule when it was tested on several portfolios.
- Random baselines: 20 random position rules and always long in the indicator atlas, random-direction entries in the reward-to-risk test.
- Placebos: Friday the 13th among the calendar hypotheses, round-number levels against levels shifted by half a step, month-end currency flows against random days.
- Decile and tercile analyses with cut-offs taken from the training period only.
- Bootstrap methods were not used.
How many hypotheses
| Study | Configurations |
|---|---|
| Indicator atlas (162 rules × 20 instruments × 3 timeframes) | 9,720 |
| Classic rules and portfolios | 4,032 |
| 1-minute dip and pump rules (288 × 12 coins) | 3,456 |
| Candlestick patterns | 1,584 |
| FX tick scalping grid | 1,536 |
| Chart patterns and “smart money” concepts | 1,230 |
| FX intraday patterns | 788 |
| FX cross-pair lead-lag | 654 |
| Reward-to-risk ratios | 144 |
| FX tick exit rules | 94 |
| Crypto exits, volatility targeting, stablecoin audit, indicator combinations | 84 |
| Direction model | 70 |
| FX daily, weekly and monthly studies | 52 |
| Kronos run × horizon cells | 43 |
| Calendar, weather and Moon hypotheses | 40 |
| Macro regime filters | 18 |
| Order-book tests | 12 |
| Total | ≈ 23,600 |
One configuration is one parameterised rule on one instrument, pooled test or portfolio, counted once however many periods it was evaluated on. A few hundred more tests were printed to the console without a result file and are not counted. Nothing in the large searches passed its multiple-testing gate.
Safeguards against look-ahead
- Unit tests that fail on leakage. A signal that sees the next bar must produce a Sharpe ratio above 20, while a same-bar signal on a random walk must not profit. Every strategy and all 183 indicator and baseline rules must produce identical positions on truncated data. Random signals must lose exactly their costs.
- Publication lags in days: Wikipedia 1, CoinMetrics 1, Fear & Greed 1, Fed balance sheet 2, Treasury account 2, reverse repos 1, M2 56, real yields 1, financial conditions 6, CFTC positioning 4.
- Past-only thresholds. The stablecoin threshold excludes the current day, and the flow of a day is traded on the next day.
- Research and production parity. The deployed direction model must reproduce the research composite on more than 2,500 days, to twelve decimal places.
- Pre-registration. The depth ladder test and its decision rule were written down before any data was examined, and three clarifications were added before its first real run: the rollover hours are excluded, hours are counted as covered time without gaps, and terciles are formed by rank.
Known in-sample choices
- The live Bitcoin bot’s Parabolic SAR exit was chosen from about 40 exit variants on the same 2020–2026 data it is evaluated on.
- The 40% volatility target was chosen after seeing its results, so it was not deployed.
- The hour-of-day currency family took each hour’s direction from the training data, so its training statistics are in-sample.
- The indicator combinations were designed after the full-sample diagnosis of the indicators.
Data vintages and provenance
All histories were downloaded in September 2026. Macro series such as financial conditions and money supply are revised, stablecoin aggregates change their coverage, and on-chain providers revise their attribution of exchange addresses, so the data we used is not always the data that was visible at the time. The crypto universes contain coins that are listed today, so delisted coins are missing. The project has no commit history, and file timestamps are the only record of when each result was produced.
Tooling
Python 3.12 with pandas, numpy, numba, statsmodels, scipy and TA-Lib, Parquet caches for every download, PyTorch on one consumer GPU for the Kronos model, and 270 unit tests. The live bots use the cTrader Open API for currencies and the exchange’s REST API for crypto.
How this was made
The research code, the tests and this write-up were produced by one person working with an AI coding assistant. Every number on this site comes from the project’s own code and result files, and every chart has a data table below it. The code is not published at the moment.