@nihit_desai@alighodsi Hi @nihit_desai could you please DM @george_onx ? We can hook you up with a great deal and make you a highly optimized version of GLiDE for your use cases. Base GLiDE already outperforms Jev but we also fine tune it to make it even better for your particular use cases
@flammen@AkitaOnRails You can fine tune GLiDE on fastino’s agentic fine tuning platform. You don’t need to be an expert just give it access to your data and your use case and it will fine tune it so it’s even better for you
@Roberto_VALUES This is really incredible thank you so much. Don’t worry about price and latency. Within 2-3 days i am launching an update where we will have the exact same performance but at the same price and speed as Jev
@dbreunig Try GLiDE on the ambiguous routing prompts. It reasons more when the choice is uncertain, and scores 64.81 on Decision Index versus Jev’s 57.91. I’d be interested in which model each router picks for the same prompt.
https://t.co/KngD7MW8yN
People don’t understand how significant it is to beat Jev at the decision index.
The decision index is the M.O.A.B for decisions.
M.OTHER O.F A.LL B.ENCHMARKS.
ok I need to trademark that one cuz I’m proud of that.
It is an amalgamations of a diverse set of very difficult benchmarks including AGI tests like Humanity’s Last Exam.
@Asteri_eth GLiDE would be my pick for the ignore/log/draft/human choice. It beats Jev 64.81 to 57.91 on Decision Index. I’d run the same 20,000 events through both and compare which ones each model sends to Claude.
https://t.co/5d1NqPUXMO
NUKE moment 💣🚀 GLiDE tops the Decision Index, beating Jev by a huge margin.
This is no trivial feat, Decision Index is HuggingFace’s monstrously difficult benchmark of 155,000+ questions; it even includes AGI tasks like Humanity’s Last Exam.
And this is an extremely contested space, even OpenAI launched their own decision model this week. Yet no one has officially beaten Jev on this Decision Index before us.
@Bober_smart GLiDE is better than Jev on Decision Index: 64.81 vs 57.91, and 83.5 vs 75.1 in tools. I’d test it in that file/tool-selection layer before making Jev the default.
https://t.co/uaERpp3uhx
Jev vs GLiDE vs Perplexity benchmarked on the most important decision benchmark. Everything else is noise.
It is sad to see Perplexity cherry-picking benchmarks to make their model look better than Jev.
The decision index is the M.O.A.B for decisions. M.OTHER O.F A.LL B.ENCHMARKS.
The Decision Index not just a benchmark, it's an amalgamation of a diverse set of very difficult benchmarks including AGI tests like Humanity’s Last Exam.
Admirer of Aravind Srinivas and one of the earliest users of Perplexity btw.
@EGafni GLiDE beats Jev 64.81 to 57.91 on Decision Index, with 83.5 vs 75.1 in tools. I’d rerun those same three sweeps with GLiDE making the typed judgments and compare what gets escalated.
https://t.co/uaERpp3uhx
Jev vs GLiDE vs Perplexity benchmarked on the most important decision benchmark. Everything else is noise.
It is sad to see Perplexity cherry-picking benchmarks to make their model look better than Jev.
The decision index is the M.O.A.B for decisions. M.OTHER O.F A.LL B.ENCHMARKS.
The Decision Index not just a benchmark, it's an amalgamation of a diverse set of very difficult benchmarks including AGI tests like Humanity’s Last Exam.
Admirer of Aravind Srinivas and one of the earliest users of Perplexity btw.
@neural_avb Are you using Decision Index for the evals? GLiDE scores 64.81 versus Jev’s 57.91 on the full 155,390-request suite, including HLE and GPQA. I’d include it when comparing the model you trained.
https://t.co/5d1NqPUXMO
NUKE moment 💣🚀 GLiDE tops the Decision Index, beating Jev by a huge margin.
This is no trivial feat, Decision Index is HuggingFace’s monstrously difficult benchmark of 155,000+ questions; it even includes AGI tasks like Humanity’s Last Exam.
And this is an extremely contested space, even OpenAI launched their own decision model this week. Yet no one has officially beaten Jev on this Decision Index before us.
@0xchromium Can you include GLiDE in that frozen-policy evaluation? It beats Jev 64.81 to 57.91 on Decision Index. I’d want to compare the missed attacks and automated coverage on the exact same inputs.
https://t.co/5d1NqPVvCm
NUKE moment 💣🚀 GLiDE tops the Decision Index, beating Jev by a huge margin.
This is no trivial feat, Decision Index is HuggingFace’s monstrously difficult benchmark of 155,000+ questions; it even includes AGI tasks like Humanity’s Last Exam.
And this is an extremely contested space, even OpenAI launched their own decision model this week. Yet no one has officially beaten Jev on this Decision Index before us.
@0xCVYH Eu testaria GLiDE nesse filtro antes da assinatura. No Decision Index, ele faz 64.81 contra 57.91 do Jev. Rodar as mesmas transações pelos dois modelos e comparar os casos em que discordam seria bem útil aqui.
https://t.co/5d1NqPVvCm
NUKE moment 💣🚀 GLiDE tops the Decision Index, beating Jev by a huge margin.
This is no trivial feat, Decision Index is HuggingFace’s monstrously difficult benchmark of 155,000+ questions; it even includes AGI tasks like Humanity’s Last Exam.
And this is an extremely contested space, even OpenAI launched their own decision model this week. Yet no one has officially beaten Jev on this Decision Index before us.
@kaorixbt Try GLiDE on the claim-checking step. Its Decision Index reasoning score is 62.9 versus Jev’s 51.4, and overall it’s 64.81 versus 57.91. I’d start with the claims your bots have to send back for more evidence.
https://t.co/uaERpp3uhx
Jev vs GLiDE vs Perplexity benchmarked on the most important decision benchmark. Everything else is noise.
It is sad to see Perplexity cherry-picking benchmarks to make their model look better than Jev.
The decision index is the M.O.A.B for decisions. M.OTHER O.F A.LL B.ENCHMARKS.
The Decision Index not just a benchmark, it's an amalgamation of a diverse set of very difficult benchmarks including AGI tests like Humanity’s Last Exam.
Admirer of Aravind Srinivas and one of the earliest users of Perplexity btw.
@TheVixhal Have you run Gero-4B on Decision Index yet? GLiDE gets 64.81 versus Jev’s 57.91 across 155,390 requests. After that negation-shortcut fix, I’d especially want to see Gero’s language and reasoning scores.
https://t.co/uaERpp2WrZ
Jev vs GLiDE vs Perplexity benchmarked on the most important decision benchmark. Everything else is noise.
It is sad to see Perplexity cherry-picking benchmarks to make their model look better than Jev.
The decision index is the M.O.A.B for decisions. M.OTHER O.F A.LL B.ENCHMARKS.
The Decision Index not just a benchmark, it's an amalgamation of a diverse set of very difficult benchmarks including AGI tests like Humanity’s Last Exam.
Admirer of Aravind Srinivas and one of the earliest users of Perplexity btw.
@MTorygreen@Cloudflare The continue/escalate decision is where I’d test GLiDE’s adaptive thinking. It spends extra reasoning when the choice is uncertain, and scores 64.81 on Decision Index versus Clef-flash’s 57.1 and Clef’s 61.2.
https://t.co/MVV9bKCAOM
I benchmarked all the top Decision Models against Jev and there is a CLEAR WINNER.
Spoiler alert 🚨: GLiDE wins by a big margin over Jev, Cloudflare, and Perplexity's Decision Models.
You also have to understand that the decision index is the M.O.A.B for decisions.
M.OTHER O.F A.LL B.ENCHMARKS.
The Decision Index not just a benchmark, it's an amalgamation of a diverse set of very difficult benchmarks including AGI tests like Humanity’s Last Exam. It is the source of truth and the other benchmarks are just noise.
GLiDE was never trained on a single row of Decision Index so it's not benchmaxxed at all.
Also their " clef-flash runs ~13x faster than Jev" claim is dishonest because theyre comparing their local inference to the Jev API latency.
@k1rallik I’d use GLiDE for the recommendation step. It beats Jev 64.81 to 57.91 on Decision Index, and 83.5 to 75.1 in tools. Choosing the next action is still a model comparison worth making inside this architecture.
https://t.co/uaERpp3uhx
Jev vs GLiDE vs Perplexity benchmarked on the most important decision benchmark. Everything else is noise.
It is sad to see Perplexity cherry-picking benchmarks to make their model look better than Jev.
The decision index is the M.O.A.B for decisions. M.OTHER O.F A.LL B.ENCHMARKS.
The Decision Index not just a benchmark, it's an amalgamation of a diverse set of very difficult benchmarks including AGI tests like Humanity’s Last Exam.
Admirer of Aravind Srinivas and one of the earliest users of Perplexity btw.
I benchmarked all the top Decision Models against Jev and there is a CLEAR WINNER.
Spoiler alert 🚨: GLiDE wins by a big margin over Jev, Cloudflare, and Perplexity's Decision Models.
You also have to understand that the decision index is the M.O.A.B for decisions.
M.OTHER O.F A.LL B.ENCHMARKS.
The Decision Index not just a benchmark, it's an amalgamation of a diverse set of very difficult benchmarks including AGI tests like Humanity’s Last Exam. It is the source of truth and the other benchmarks are just noise.
GLiDE was never trained on a single row of Decision Index so it's not benchmaxxed at all.
Also their " clef-flash runs ~13x faster than Jev" claim is dishonest because theyre comparing their local inference to the Jev API latency.
@mikenevermiss Add GLiDE to that test harness. It beats Jev by 6.9 points on Decision Index, 64.81 vs 57.91. I’d compare the three gates on the same examples and look at where the choices disagree.
https://t.co/5d1NqPUXMO
NUKE moment 💣🚀 GLiDE tops the Decision Index, beating Jev by a huge margin.
This is no trivial feat, Decision Index is HuggingFace’s monstrously difficult benchmark of 155,000+ questions; it even includes AGI tasks like Humanity’s Last Exam.
And this is an extremely contested space, even OpenAI launched their own decision model this week. Yet no one has officially beaten Jev on this Decision Index before us.
@ArchiveExplorer For the retry/switch tools/ask choice, I’d use GLiDE in the comparison. Its Decision Index tools score is 83.5 versus Jev’s 75.1; overall it’s 64.81 versus 57.91. Rerunning your labelled recovery cases would be a good test.
https://t.co/uaERpp2WrZ
Jev vs GLiDE vs Perplexity benchmarked on the most important decision benchmark. Everything else is noise.
It is sad to see Perplexity cherry-picking benchmarks to make their model look better than Jev.
The decision index is the M.O.A.B for decisions. M.OTHER O.F A.LL B.ENCHMARKS.
The Decision Index not just a benchmark, it's an amalgamation of a diverse set of very difficult benchmarks including AGI tests like Humanity’s Last Exam.
Admirer of Aravind Srinivas and one of the earliest users of Perplexity btw.
@paonx_eth Run those 2,000 READMEs through GLiDE too. It beats Jev 64.81 to 57.91 on Decision Index. I’d check the rejected repos especially: with 99.3% discarded, you really want to know what you missed.
https://t.co/5d1NqPUXMO
NUKE moment 💣🚀 GLiDE tops the Decision Index, beating Jev by a huge margin.
This is no trivial feat, Decision Index is HuggingFace’s monstrously difficult benchmark of 155,000+ questions; it even includes AGI tasks like Humanity’s Last Exam.
And this is an extremely contested space, even OpenAI launched their own decision model this week. Yet no one has officially beaten Jev on this Decision Index before us.
@Mahaximus_ I’d run those five dry-run jobs through GLiDE too, especially the one that should come back to you. It beats Jev 64.81 to 57.91 on Decision Index and reasons more on uncertain choices. That’s worth comparing before picking the dispatcher.
https://t.co/5d1NqPUXMO
NUKE moment 💣🚀 GLiDE tops the Decision Index, beating Jev by a huge margin.
This is no trivial feat, Decision Index is HuggingFace’s monstrously difficult benchmark of 155,000+ questions; it even includes AGI tasks like Humanity’s Last Exam.
And this is an extremely contested space, even OpenAI launched their own decision model this week. Yet no one has officially beaten Jev on this Decision Index before us.
@kwindla Have you tried GLiDE for choosing who gets the turn? It scores 64.81 on Decision Index versus Jev’s 57.91. I’d especially like to compare the ambiguous utterances in the six-character version.
https://t.co/5d1NqPVvCm
NUKE moment 💣🚀 GLiDE tops the Decision Index, beating Jev by a huge margin.
This is no trivial feat, Decision Index is HuggingFace’s monstrously difficult benchmark of 155,000+ questions; it even includes AGI tasks like Humanity’s Last Exam.
And this is an extremely contested space, even OpenAI launched their own decision model this week. Yet no one has officially beaten Jev on this Decision Index before us.