With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds.
Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️
To learn more point your agent at our GitHub here: https://t.co/uNojQBFIZR
Thank you for all the support! Our first official post about our work went pretty far!
Stay tuned. We have a lot more to share with you all that will help improve your local AI experience.
We're just getting started. 🦾
With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds.
Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️
To learn more point your agent at our GitHub here: https://t.co/uNojQBFIZR
With today's release of Jovian Judgement R37, DeepSeek V4.1 Flash has full support with blazing speeds.
Decodes of 500 tok/s and prefills of 20k tok/s for single stream workloads on 4 RTX 6000 Pros. ⚡️
To learn more point your agent at our GitHub here: https://t.co/uNojQBFIZR
500 toks/s at home 🤯
The thing that bothers me the most about local AI, and too few talk about this, is the fact I've already dropped obscene amounts of money and I just simply don't have enough vram.
I have two RTX 6000 pros and it's pretty rare I can find a model that I actually want to run.
It seems like four is the minimum honestly. This is gonna end up costing me over $30,000 more.
It's true that smaller models are getting way better, and I'm not discounting these relatively tiny models
But at the end of the day when there is a significantly better larger model, you're always going to want to run that.
I'm getting serious fomo.
Anyway, these guys are cracked.
It's Friday already?! 😱
Time flies when everything is accelerating.
We're talking about AI security, the latest in local and much more.
Panel: @philipbankier@LLMJunky@secemp9
Guests: @ZackKorman@BrandonMusicKy@msuiche
Join us at 4:00 pm ET!
https://t.co/a21BggLohK
Now trying to get deepseek v4.1flash to run on my rtx pro 6ks. Really excited to try this one. Following a deliberate process to make sure there are no max logit error and zero relative L2., and cosine similarity is appropriate. Running in eager. Nowtime for graphs and MTP lol
This is an excellent read on why more companies are investing in local AI, but not to leave the cloud. Instead to change their relationship with the cloud.