I got into crypto during lockdown. Over the next four years I lost about $100,000.
This year I had AI agents build a trading system to win it back.
50,000 lines of code. I wrote none of them.
Here's what happened when I audited it.
Before trusting a backtest, ask:
Was the universe frozen in advance?
How many variants were tried?
What happens without the best 10 trades?
How bad was the worst out-of-sample period?
Did it survive real costs?
The return is the start of the audit, not the verdict.
One of my strategies passed its overfitting test.
PBO: 0.094. The failure limit was 0.20.
It also had a 54% drawdown, lost 53.56% without its 10 best trades, and failed 7 promotion gates.
One green metric does not overrule a red system.
One strategy passed its overfitting test: PBO 0.10, below the 0.20 limit.
It still lost 68% after costs.
A green metric isn't validation. Every gate must be allowed to kill the idea.
A market-making signal can have a real gross edge and still be worthless.
One of mine lost 4.7 basis points per trade after fees. Across 43,000 trades, execution costs weren't a footnote. They were the strategy.
Gross alpha you can't capture isn't alpha.
@edgeful The split ratio isn't the constraint, the trade count is. 15 OOS trades at 60% win is 9 wins. The CI covers almost everything.
And your dashboard already flags the bigger one: drawdown/profit 92% OOS vs ~12% in-sample ๐ค
Ten trades out of 157 produced 96.16% of the profit.
Remove them and the strategy returns โ53.56%.
The other 147 were not a weaker profitable system. They were a losing one.
Best trade alone: 40.66% of total profit.
Top 5: 89.93%.
Killed it at +432%.
@dikovaxi Exactly, and the harder part is that n is almost never the number you remember.
Every lookback window, every threshold, every variant killed after five minutes โ all of it counts.
Most people report n=1 with a true n in the hundreds. ๐ค
Testing one strategy is research. Testing 100 and showing only the winner is a lottery with selective reporting.
My best backtest had a Sharpe of 1.36. After correcting for the search process, its Deflated Sharpe Ratio was 0.45.
The audit killed it.
@HF_Trader The hard part isnโt whether context can be programmedโitโs whether its definition existed before the outcome. If a context variable is added after reviewing failures, it becomes another degree of freedom that needs fresh OOS validation.
AI agents wrote roughly 50,000 lines of my trading system.
Code generation wasn't the dangerous part. Agreeable validation was.
Every strategy looked promising until I built gates designed to return bad news.
AI can build the lab. It shouldn't grade its own experiment.
A backtest showed +632%.
Remove its 10 best trades and it loses 54%.
Those trades produced 96% of the entire profit.
That wasn't an edge. It was concentrated luck.
Full audit: https://t.co/enV3KXOSei
@RuujSs Multiple-testing blindness killed my strongest result. Raw Sharpe was 1.36; after accounting for the search process, Deflated Sharpe was 0.45. Tracking the full trial budget matters more than polishing the winning backtest.
@slash1sol Finding a testable rule and proving that it survives are separate jobs. Symbol selection, multiple testing, fees and top-trade concentration killed every completed strategy in my audit.
@0xCrypto_Duke Exactly. Let the agent generate hypotheses and build the lab, but don't let it grade its own experiment. Independent gates and a fixed trial budget made the biggest difference in my tests.
@_vmlops The tooling is the easy part. The harder problem is making validation adversarial. My AI-built pipeline produced a +632% result that became a 54% loss after removing 10 trades. Backtesting needs kill gates, not just more metrics.