@Mnilax For the name of God stop using AI to build your misleading posts or at least read it before posting. I bet you use some Chinese model to write it anyways. "saves up to 98% on tokens ur GPT Astra and Opus 5.5" - What are you even talking about?
Did my OSS-Jev called Verdict, beat Official Typesafeai 's Jev ?, and Laya (the best performing OSS Jev) on their own benchmark?
When I first tested Verdict 1.0(my earlier attempt of oss-jev) earlier this week, it scored 26.10% accuracy. Uniform random guessing on that split is 26.90%, meaning the model was literally worse than a coin toss.
I dug into the failure and traced it to an option-pointer misalignment. I scrapped that setup and rebuilt Verdict 2.0 with ModernBERT-base (149.6M parameters), symmetric permutation-KL loss to eliminate option-order bias, and strict case hashing to keep the test set clean.
I just finished evaluating on the 2,000 held-out enterprise decisions from LocalLLaMA/typed-decisions.
1. Top-1 Accuracy: 77.10%
Ahead of both Laya's 421M ModernBERT-large (76.60%) and TypeSafe AI's Jev (72.70%). A 149.6M base encoder outperforming a model with 2.8x the parameters, on less than half the memory.
2. Solved the "Soft-Label Trap"
The benchmark scores Brier against soft human panel distributions and scores ECE against hard right or wrong. Those two objectives pull in opposite directions. Laya answers both with a single probability vector and lands at 0.2140 ECE.
Verdict 2.0 splits the output into two heads:
β’ Distribution head: 0.0636 Brier, 0.1513 ECE, 29% below Laya
β’ Confidence head: 0.0144 ECE (1.44%) and 0.7664 AUROC
3. Permutation Stability: 4.76% Flip Rate
Kev(another better oss jev version) reported a 7.41% flip rate when option positions were shuffled. Verdict 2.0 drops this to 4.76% across 2,918 permutations via symmetric KL training.
4. Production Gating with Selective Classification
Calibrated confidence enables safe automation thresholds:
β’ At 80% coverage: accuracy rises to 83.44%
β’ At 60% coverage: accuracy reaches 89.00%
β’ At 50% coverage: accuracy reaches 91.40%
Deterministic code can auto-route routine decisions and escalate low-confidence cases to humans.
5. Runs in the Browser via WebGPU
At 149.6M parameters it runs locally in browser tabs and edge runtimes, with no cloud dependency and no inference fees.
I Will be writing a full technical Article on the dual-channel calibration math, the option-pointer wiring fix, and training dynamics. If you want to catch the breakdown when it drops, drop a follow.
Link for the post-trained model and github repo are in the comments