How we decide
whether something is real.

Almost everything that looks like a trading edge is luck wearing a convincing costume. This page is about how to tell the difference — explained from scratch, with no assumed knowledge.

LONA has tested roughly 145 trading rules. One is still standing. That ratio is not pessimism, it is what honest testing does to ideas. The rest of this page is the method that killed the other 144, one rule at a time, with the real numbers it produced.

THE PROBLEMRandomness is a very good actor

Give a hundred people a coin each and ask them to flip it five times. Three or four of them will get five heads in a row. If you then interview those three about their technique, they will have one — people always do. Nothing about their method was real, but the evidence in front of you looks exactly like skill.

Trading research is that, with money attached. You try an idea on historical prices, and some ideas will have made money by chance alone. The question is never "did this make money in the test?" — something always does. The question is "would a coin have done as well?"

Here is the first real example from our own work. We hunted for a second strategy across 26 different ideas and picked the best one. It scored 0.94 on the measure we use.

Best of 26 real ideasBest of 26 real ideas: Sharpe 0.94 — what we found0.94what we foundBest of 26 random onesBest of 26 random ones: Sharpe 1.11 — what noise alone manages1.11what noise alone managesSHARPE RATIO — HIGHER LOOKS BETTER
Finding 011. The score is a “Sharpe ratio” — return divided by how bumpy the ride was. Higher looks better.
⚠ Read that again

If you generate 26 completely meaningless ideas and keep the best one, it typically scores 1.11. Our best real idea scored 0.94worse than what pure noise usually manages.

Not "not quite significant". Below the coin. We threw all 26 away.

This is the single most important idea on this page: the more things you try, the better your best result has to be before it means anything. A score of 0.94 found by testing one idea and a 0.94 found by testing twenty-six are completely different pieces of evidence, even though the number is identical.

RULE ONECount what you actually have

More data sounds like more certainty. It usually isn't, because most rows of data are repeating what the row before them already said.

On 12 August our recorder captured 158,517 snapshots of Bitcoin's order book in a single day. That sounds like an enormous sample. It isn't — each snapshot is almost identical to the one a fraction of a second earlier, so they are not independent pieces of evidence.

158,517 rows of order-book datawhat it is actually worth4,723 independent observations — 3%
Finding 026. Bitcoin's order book updates constantly and says nearly the same thing every time.

Once you account for how much each row repeats the last, those 158,517 rows carry about 4,723 genuinely independent observations — three per cent. A confidence figure calculated on the raw row count would be five times too confident.

Our first five findings died of exactly this mistake in a different costume: six weeks of data read as though it were a large sample. It was one market mood, sampled very finely.

RULE TWOInclude the ones that died

Imagine judging how good a school is by interviewing its graduates. You would never hear from anyone who dropped out. Every answer you get is from someone the process already selected.

Trading data has exactly this hole. Many crypto markets are launched, trade for a while, collapse and get removed from the exchange. If your historical data only contains the markets that still exist today, you have quietly deleted the failures — and any strategy tested on it will look better than it was.

✓ What we do instead

Our main dataset is 742 markets, and 148 of them are dead. They are left in, with all 97,044 days of price they produced before they died. The strategy has to survive owning things that later went to zero, because in real life it would have.

There is a sting in the tail, and it is a good example of why measuring beats assuming. We measured this bias twice. On one three-year sample, removing the dead names made the strategy look 6.2 points a year better. On a different 6.6-year sample it made it look 2.9 points worse. Opposite directions.

The spread of that measurement turned out to be larger than the measurement itself, which means there is no such thing as "the" survivorship correction — and two of our earlier findings had been subtracting one as though there were. Both were corrected.

RULE THREEBuild a test that can fail

Before trusting any result, we deliberately break the test to check the test still notices. The main one is blunt: let the strategy cheat.

We feed it tomorrow's prices today. If a strategy that can see the future doesn't return something absurd, then the machinery is broken and every honest number it produced is meaningless too.

ControlResultWhat it proves
Let it see tomorrow's price+164,956% The engine works. If this ever looked reasonable, everything else is void.
Hold all 20 markets, alwaysloses money The edge is in choosing, not in owning crypto.
Shuffle the choices at random, 200×10% beat the real one On that dataset, luck closes the gap one time in ten — so we do not use it as evidence.

That last row is the point of doing controls at all. It was our own result, and the control said it wasn't good enough. So we reported it as not evidence.

RULE FOURWrite down what you'll accept, first

Here is how self-deception actually happens. You test an idea. It doesn't quite work. You try a slightly different version. Also not quite. On the ninth attempt something works, and by then you have honestly forgotten there were eight others.

The fix is to write down, before running anything: what is being tested, how it will be measured, and exactly what result counts as a pass. Then you cannot move the goalposts, because they are already in writing.

Here is the most recent one, and it is the best example on this page — because the rule stopped a result I wanted to keep.

The question

Our strategy buys a market when it hits a 50-day high, then holds it for five days. Why five? An earlier study noticed, in passing, that holding for two days looked better — but noticing something in a grid of results is not evidence, it is fishing. So we wrote the test down first and ran it on a completely different dataset.

Binance spot, 2017–20Binance perps, 2020–26Hyperliquid, 2023–26
0.60.81.01.21.41.62d3d5d7d10d15d20dHOW MANY DAYS WE HOLDHyperliquid 2023-26 — 2-day hold: Sharpe 1.23Hyperliquid 2023-26 — 3-day hold: Sharpe 0.94Hyperliquid 2023-26 — 5-day hold: Sharpe 1.08Hyperliquid 2023-26 — 7-day hold: Sharpe 0.85Hyperliquid 2023-26 — 10-day hold: Sharpe 0.85Hyperliquid 2023-26 — 15-day hold: Sharpe 0.73Hyperliquid 2023-26 — 20-day hold: Sharpe 0.70HyperliquidBinance perps 2020-26 — 2-day hold: Sharpe 1.48Binance perps 2020-26 — 3-day hold: Sharpe 1.43Binance perps 2020-26 — 5-day hold: Sharpe 1.21Binance perps 2020-26 — 7-day hold: Sharpe 1.15Binance perps 2020-26 — 10-day hold: Sharpe 0.89Binance perps 2020-26 — 15-day hold: Sharpe 0.88Binance perps 2020-26 — 20-day hold: Sharpe 0.83Binance perpsBinance spot 2017-20 — 2-day hold: Sharpe 1.50Binance spot 2017-20 — 3-day hold: Sharpe 1.39Binance spot 2017-20 — 5-day hold: Sharpe 1.58Binance spot 2017-20 — 7-day hold: Sharpe 1.39Binance spot 2017-20 — 10-day hold: Sharpe 1.46Binance spot 2017-20 — 15-day hold: Sharpe 1.53Binance spot 2017-20 — 20-day hold: Sharpe 1.25Binance spot
Findings 029 and 031. Three datasets — and the important thing is which two of them overlap in time.

The two that agreed were not independent

The blue and orange lines agree almost perfectly. The ranking of the seven hold lengths matches at 0.95 out of 1, both peak at two days, both slope steadily down. When we first ran this we wrote that down as “the same mechanism showing up twice”.

It wasn't. Look at the dates. Hyperliquid covers 2023–26. Binance perps cover 2020–26. Every single day of the Hyperliquid sample sits inside the Binance one. Two exchanges, one period, mostly the same coins — two windows onto one history. Of course they agreed. We had checked the venues were different and never checked the years were.

The green line is Binance spot, 2017–2020 — a period neither of the others contains. It is the only genuinely independent test on the chart, and it disagrees. Its best hold is five days, not two, and the tidy downward slope is gone:

comparedoverlap in timeagreement
Binance perps vs Hyperliquidcomplete+0.95
Spot vs Binance perpsnone+0.25
Spot vs Hyperliquidnone+0.45
⚠ The lesson, which cost us our best-looking result

Two datasets that overlap in time are one dataset. Before treating agreement as proof, check the calendars, not the logos.

Counting datasets is the wrong unit. Count independent years. This project has five sources and roughly two eras — which is far less evidence than five sources sounds like.

Nothing about this changes the verdict below; the two-day hold was already not adopted. What it removes is the sentence we were proudest of. That check had been available for a day and nobody thought to run it until somebody asked whether a third dataset would help.

And the effect on risk is dramatic. Here is the worst fall from a peak in each year, for both versions:

-0%-10%-20%20202020: 5-day hold, worst fall -14%2020: 2-day hold, worst fall -13%-13% vs -14%20212021: 5-day hold, worst fall -15%2021: 2-day hold, worst fall -12%-12% vs -15%20222022: 5-day hold, worst fall -25%2022: 2-day hold, worst fall -8%-8% vs -25%20232023: 5-day hold, worst fall -19%2023: 2-day hold, worst fall -10%-10% vs -19%20242024: 5-day hold, worst fall -15%2024: 2-day hold, worst fall -10%-10% vs -15%20252025: 5-day hold, worst fall -9%2025: 2-day hold, worst fall -9%-9% vs -9%20262026: 5-day hold, worst fall -10%2026: 2-day hold, worst fall -3%-3% vs -10%WORST FALL FROM A PEAK, EACH YEAR2-day5-day
The two-day hold fell less in every single year. In 2022 — a brutal year for crypto — it lost 3% where the five-day version lost 18.8%.
⚠ And we did not adopt it

The written test required two things to pass. One passed decisively. The other — the main one — failed by a whisker: it needed a number above zero and got −0.037.

So the answer is no. Not "no, but". Just no.

Adopting it anyway would mean the written rule was decoration — and every earlier idea we rejected under the same rule would retroactively mean nothing, because the rule would only bind when it agreed with us.

A rule you wrote down is only worth something on the day it stops you. That day was 15 August 2026.

RULE FIVEPrefer new time to new analysis

When a result is borderline, the tempting move is to analyse the same data harder — more variations, more slices, more statistics. This almost never helps, because the uncertainty comes from how much independent data you have, and re-reading the same years does not create more of it.

We once ran 70 versions of our strategy on the same three years. All 70 were positive, which sounds like overwhelming evidence. It is one observation read 70 ways.

✓ What actually moves the needle

Only new time or new information. Testing the same rule on another exchange over the same years would produce a nearly identical answer — it is the same market history wearing a different logo.

So we went looking for genuinely older data, twice. What we found is the next section, and it is not the happy answer.

That is why the most valuable thing this project does now is wait. Every day, it writes down what it would have traded, before knowing what happens. Nobody can search that record, because it has not happened yet. It needs months, not cleverness.

RULE SIXKnow when you have run out of history

Rule five says prefer new time. The obvious next move, then, is to go and buy more of it — find an exchange whose data starts earlier than ours. We went looking, and the answer was not the one we wanted.

We found the best-preserved archive in crypto. Bitfinex has daily prices going back to March 2013, seven years earlier than most of our data, and unlike almost everyone else it still serves the markets that died — we tested fifty delisted pairs and all fifty came back complete.

Then we counted what was actually in those seven years.

The 20 biggest coins, market-wideBitfinex
0510152020 — what the rule needs20142015201620172018MARKETS LIQUID ENOUGH TO TRADE — $1m A DAYour archive startsThe 20 biggest coins, January 2014 — 4 markets over $1m a dayThe 20 biggest coins, January 2015 — 4 markets over $1m a dayThe 20 biggest coins, January 2016 — 2 markets over $1m a dayThe 20 biggest coins, 2016 — 5 markets over $1m a dayThe 20 biggest coins, January 2017 — 13 markets over $1m a dayThe 20 biggest coins, 2017 — 20 markets over $1m a dayThe 20 biggest coins, January 2018 — 20 markets over $1m a dayThe 20 biggest coinsBitfinex, 2013 — 0 markets over $1m a dayBitfinex, 2014 — 1 market over $1m a dayBitfinex, 2015 — 1 market over $1m a dayBitfinex, 2016 — 1 market over $1m a dayBitfinex, 2017 — 7 markets over $1m a dayBitfinex, 2018 — 14 markets over $1m a dayBitfinex
Findings 032 and 033. The upper line is capped at 20 by how it was measured — we could only read the twenty biggest coins in each snapshot. That cap is the point: the rule needs twenty, so the question is simply whether the twenty biggest were even tradable.

On 3 January 2016, exactly two coins on earth traded a million dollars a day. Five by that summer. Thirteen by the start of 2017. Our rule ranks a list of markets and buys the best twenty of them — and twenty liquid markets did not exist anywhere until around the middle of 2017.

✓ What this actually means

Our earliest archive starts in August 2017. That is within weeks of the earliest date this test could exist at all. There is no hidden shelf of older data we have been too lazy to fetch. The ceiling belongs to the market, not to us.

Which is a more useful sentence than “we could not find more data”, because it tells you to stop looking and start waiting.

One dataset would have been a disaster

Poloniex was the biggest home for small coins in 2016 and 2017 and looked like the obvious answer. We tested thirty of its markets that have since been delisted. Twenty-nine of them have been deleted outright — no listing, no prices, nothing.

A backtest on Poloniex would have produced a beautiful 2017 and would have been measuring one thing only: which coins are still alive in 2026. That is rule two's mistake, handed to you gift-wrapped. It is now written down as a rule of its own: an exchange that does not keep its dead is not a weaker source, it is a disqualified one.

What we used Bitfinex for instead

There was still a job worth doing. One stretch of our evidence — August 2017 to 2019, which contains the whole 2018 crash and is the most informative period we own — had only ever been seen through one exchange. If that result were a quirk of one venue's fees or listings, nothing we had would have noticed.

So we ran the identical rule on Bitfinex over the identical years. Getting the list of markets took 17,576 separate requests, one for every possible three-letter ticker, because Bitfinex's own published list contains none of its dead ones. It found 268 markets, 204 of them long gone.

Same years, two exchangesReturn/yrScoreBuy all 20 instead
Binance · 2017–2020+35.7%1.28+2.2%, −91% fall
Bitfinex · 2017–2020+49.2%1.29−0.5%, −92% fall

Two exchanges, two sets of customers, two fee schedules. The scores land at 1.28 and 1.29.

And here is the check that makes that worth anything. The obvious objection is that it is the same years, so it must be the same coins. We measured it instead of arguing about it: the two exchanges' traded lists overlap by a quarter. Fifteen markets were tradable on Bitfinex and never on Binance. Three quarters of the two universes are different.

The prize was 2018. Every previous look at that crash had too few markets to choose between — below our own minimum, so we had to write “this data cannot settle it” and mean it. Bitfinex's 2018 clears the minimum, and it says: the rule lost 4.4% in the year bitcoin fell 71%. BitMEX, which shares no markets with either exchange, measured the same crash at −5.3% while bitcoin fell 73.4%. Two independent looks, one point apart.

⚠ And the part that keeps it honest

That Bitfinex result scores 1.79 on confidence — below 2, so within reach of luck on its own. Its worth is the agreement, not the number. Two of its four years are still too thin to count. And Bitfinex from 2021 onwards is weak — on a median of six markets, too few to be evidence either way. That last line is published because we wrote down in advance that we would publish it, whichever way it came out.

It also does not add a new era. It is a second view of years we already had. By rule five's own logic, our count of genuinely independent periods is still two.

THE RESULTWhat survived all of that

One rule. Buy a market when it closes above every price of the previous 50 days, hold it five days, size it at a twentieth of the account, otherwise sit in cash. Roughly 16% invested on an average day.

Tested onSpanReturn/yrWorst fallConfidence
Binance · 742 markets6.6 yr+32.4%−28%2.4
BitMEX · bitcoin only10.25 yr+27.7%−54%2.39
OKX · 356 markets6.6 yr+26.1%−35%2.1
Hyperliquid · 232 markets3.0 yr+20.9%−14%1.5

“Confidence” is a t-statistic. Below 2 means the result is within reach of luck; above 2 starts to be worth something. Note the bottom row: the three-year sample is the least convincing, and it was the first one we had.

Two things this does not mean

It does not mean +32% a year. Trend-following on crypto is one of the most heavily studied ideas in finance, so some of that number is the field's selection rather than ours. A realistic live expectation is meaningfully lower.

It does not mean a comfortable ride. Only 47% of trades are up after their first day. The strategy loses slightly more often than it wins — its edge is that the wins are bigger. Anyone watching will see red on most days while it behaves exactly as measured, and there is roughly a one-in-eight chance of a losing year.

IN SHORTThe whole method, in eight lines

  1. Ask whether a coin would have done as well — and correct for how many things you tried.
  2. Count independent observations, not rows.
  3. Leave the failures in the data.
  4. Build controls that are supposed to fail, and check they do.
  5. Write down the pass mark before you look.
  6. Prefer new time over new analysis.
  7. Know when the history has run out — and never take it from a source that deleted its failures.
  8. When the rule and the result disagree, the rule wins.

None of this was designed in advance. Every line exists because a result died on it — and several of them exist because a result we had already published turned out to be wrong.

LONA · research only, not advice · every figure on this page comes from a numbered finding in the research log, and several of them are corrections to earlier findings in the same file.