@marcelpociot@typesafeai on a historical backtest benchmarking i ran i did just fine with around ~3.5% profits and > 85% no trade decisions, however there is some contradictory decisions you can see on my video on youtube here
Jev AI Trading Model Tested: 300 Decisions, Real Benchmark Results
I ran TypeSafe’s new Jev model on 300 real trading decisions.
NVDA. JPM. XOM. Bitcoin. Gold.
85% of the time it said: stay flat.
The 46 trades it did take: 23 wins, 23 losses. +3.82%.
Then the report got ugly.
Every single short came back “not applicable” for entry, take-profit, and holding period.
Trades with no internal contradictions: +6.36%
Trades that contradicted themselves: −2.54%
Highest-confidence trades were worse.
I replayed every answer in a controlled backtest lab. Full method, every cutoff, every contradiction — in the video.
Which model should I put through the same 300-decision gauntlet next?
https://t.co/gitO1pUXWG
I just finished a 300-decision trading benchmark with Jev — and the result is much more interesting than the +3.82% return.
5 markets.
300 point-in-time decisions.
46 actual trades.
85% of the time: NO TRADE.
Results:
• +3.82% net return
• 50% win rate
• 1.06 profit factor
• 30 longs / 16 shorts
• 23 winners / 23 losers
At first glance, +3.82% doesn’t look spectacular.
But that’s not what caught my attention.
Jev is fundamentally different from the LLMs I’ve tested in this trading arena.
It doesn’t generate a long analysis and then convert that text into JSON.
It directly returns typed, probabilistic decisions.
That makes it extremely interesting for something like real-time algorithmic trading:
market state → structured decision → execution
without needing a large natural-language reasoning layer in between.
And there’s another fascinating result:
Jev stayed flat on 254 / 300 decisions.
It wasn’t constantly trying to invent a trade.
That kind of selectivity is something I’ve struggled to get consistently from general-purpose frontier models.
The benchmark also uncovered contradictions inside Jev’s answers — including cases where it traded despite:
• low setup probability
• recommendations to stay flat
• missing stop-loss / take-profit answers
Which is exactly why I’m interested in this experiment.
The model isn’t “solving trading.”
But it may represent a very different architecture for building machine-native decision systems.
I’m going much deeper into the results next.
Particularly:
confidence calibration
decision consistency
latency
abstention / NO_TRADE behavior
contradiction rates
comparison against frontier LLMs
This one surprised me.
Jev may be more interesting as a decision engine than as a chatbot.
Going to upload a performance video really soon!
thank @CompleteSkeptic for this great model
@theo On a 100$ plan one prompt perfect 12% of weekly limits!
Yet to be discovered on another task that usually consume 4-6% of Soal at max, going to update this with the final result and increase
@Henryf1w Not so true, as its consume more tokens for lower quality that's the main reason, Its a very good Jump however not the Fable 5 comparable yet