JevK5 v0.3 is out: updated 4B weights, a new 9B, and JevK5-Lite, a separate CPU classifier. All Apache-2.0.
A typed decision returns probabilities over options without generating text.
Results, code, and limitations: https://t.co/ptWXuOimhT
The 4B and 9B are submitted to the Decision Index 0.2. Thanks @multimodalart for running the board.
Weights on @huggingface; eval code and known limitations are in the model cards. Feedback welcome.
https://t.co/PYkH1vBUTZ
JevK5-Lite: 437M DeBERTa-v3 zero-shot classifier for CPU, single- and multi-label in one pass.
My tests vs GLiNER2.5-Decide: lower calibration error on all 6 single-label sets (ECE 0.035-0.103 vs 0.091-0.250), similar speed, slightly lower mean accuracy (0.629 vs 0.643).
JevK5-9B v0.3.3: @Alibaba_Qwen Qwen3.5-9B + a distilled LoRA, same training data as the 4B.
On my held-out checks vs v0.3: index proxy 0.762 → 0.794, hard set 0.781 → 0.859.
Trade-off: ~2-3x slower than the 4B.
https://t.co/fRc9afRy6S
JevK5 update:
• JevK5 (4B) is #2 of 76 on @airesearch12's JevBench v1.4, right behind @typesafeai's Jev. #1 open model, fastest on speed
• New: JevK5-9B v0.3.3, stronger on my held-out checks
• New: JevK5-Lite, 437M, runs on CPU
Apache-2.0 → https://t.co/8yRtDJZ31q
Minimal Python:
from jevk5 import JevK5Lite
m = JevK5Lite.from_pretrained('alibiserikbay/JevK5-Lite')
m.classify('Charged twice', {'intent': ['refund', 'status'], 'team': ['billing', 'shipping']})
Both sets are scored in one CPU pass. Install steps are on the model card.
One ticket, several decisions: intent, topic, urgency.
JevK5-Lite is a 437M CPU preview. Pass label sets at runtime; one encoder pass returns probabilities for each. Apache-2.0.
Weights, Python example and full evaluation: https://t.co/PYkH1vBUTZ
@MKhordoo For routing and tagging, I released JevK5-Lite: a 437M CPU encoder. Supply task names and labels at runtime; it scores multiple heads in one pass and returns probabilities. Preview with uneven accuracy. Full evaluation: https://t.co/PYkH1vBUTZ
@0xwhrrari To try the typed-decision steps locally: JevK5 v0.3 serves `noul`, `choice` and `score` at a TypeSafe-style `/v1/systemone` endpoint. Apache-2.0 4B and 9B weights are open. Real-world parity with Jev has not been measured. Code and examples: https://t.co/ptWXuOimhT
@airesearch12 Congrats to decider-4b and thanks for keeping the evaluation independent. JevK5 v0.2 is #3/89 now. I released v0.3 4B and a 9B today; neither has an official JevBench score yet. The public hard-tier results and regressions are in the repo: https://t.co/ptWXuOimhT
@fastinoAI Congrats on the release. I built JevK5-Lite, a 437M CPU multi-head classifier. On Fast Decisions dev it trails Decide (0.587 vs 0.637 head accuracy). On six public single-label tests it was better calibrated. Full results: https://t.co/PYkH1vBUTZ
JevK5-Lite is a separate 437M CPU classifier. It scores multiple label sets in one pass. On six public single-label tests it was better calibrated than GLiNER2.5-Decide; on Fast Decisions dev it trailed. Preview results: https://t.co/PYkH1vBUTZ
JevK5 v0.3 is out: updated 4B weights, a new 9B, and JevK5-Lite, a separate CPU classifier. All Apache-2.0.
A typed decision returns probabilities over options without generating text.
Results, code, and limitations: https://t.co/ptWXuOimhT
On our held-out checks, the 4B index proxy rose from 0.620 to 0.731; 9B reached 0.762. This is our estimate, not an official Decision Index score.
On 111 public JevBench hard items, 4B rose from 0.739 to 0.784. The change is not statistically significant; 9B scored 0.730.
@dair_ai Interesting cascade result. JevK5 exposes choice probabilities from open 4B and 2B weights, so it could be tested as a local first-stage judge. I haven't run this paper's eval, and its threshold would need calibration on each task. https://t.co/ptWXuOimhT
@rohanatlan Nice task-level breakdown. I maintain JevK5, an Apache-2.0 decision model with 4B and 2B weights. It uses typed choice probabilities and can run locally. I haven't tested it on Decision Bench; its open tasks would make a useful independent check. https://t.co/ptWXuOimhT
@yibie Thanks for adding JevK5 to awesome-jev. Small release update: the 4B weights now have GGUF builds for llama.cpp; there's also a 2B model, and v0.2.2 handles choices beyond 16 with grouped passes. https://t.co/ptWXuOimhT
@benchmarkheaven Thanks for the careful comparison. The published JevBench result is v0.2.0; v0.2.2 keeps those weights and adds support for more than 16 choices through grouped passes. I haven't measured a new JevBench score for that runtime. https://t.co/ptWXuOimhT
@multimodalart@perplexity_ai@denisyarats Congrats on v0.2. JevK5 v0.2.2 now handles >16 options by grouping them; it keeps the same v0.2 weights. A rerun of its previously refused requests would be useful. Caveat for knowledge tests: the weights saw 940 MMLU-Pro test items in training (disclosed in CHANGELOG).
JevK5 update: 4B and 2B open weights now have GGUF builds for llama.cpp. v0.2.2 also handles >16 choices by grouping options.
On 500 BANKING77 train-split items (77 choices): 69.0% accuracy, 116 ms p50 on H100. This is not a held-out score.
https://t.co/ptWXuOimhT
@CompleteSkeptic I built JevK5 to explore fast typed decisions on open 4B weights. It is now #2 on the third-party JevBench v1.4, behind Jev. Weights, training code and runtime are Apache-2.0: https://t.co/ptWXuOhOsl