@yacineMTB 1T under 250 vram at over 100 tk/s, smarter than sol max.
The current MoE architecture execution is a toy compared to what it can and will achieve
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
This is a one prompt page I did long time ago: https://t.co/LIgPhykvzf I've full website, research blogs etc... But just stopped working on it, hopefully one day I'll retake it.
I've some of my own models running, but didn't have enough time (or the tokens) to finish up the inference pipelines optimizations plus many of the other puzzle pieces
It works, I've theorized it in my own project a year ago, and you can also go much further than this, I've been just stupidly dealing with other stuff instead of putting 100% of my time on my lab.
There is even crazier stuff you can do with the MoE architecture which you're going keep seeing.
the router, context engines, post training, expert freezing, expert packing, expert labeling, cost-aware gating, predictive expert routing, dynamic expert counts, expert tiering, hot/cold expert sets, expert prefetching, hierarchical routing, expert swapping, delta experts, expert specialization, sparse channel execution, adaptive KV compression, band-aware attention, hardware-aware scheduling, modular expert updates.
People are massively underestimating how far this architecture goes.
Has happened to me on specific times (I might not be at your level of course).
But I switch to a mentality of giving joy to those around me, it gets me more involved in anything I do with others (fam/friends...).
My current problem is doing mundane things like, hey car registration expired, or you need insurance, or there is internet issues call the Internet company, lease company needs this, AC is having trouble call someone to fix it. You need to do laundry....
The daily life issues, for those am basically mentally exhausted.
for your first question, (me?) around 90% of the time.
There are 0 issues with inference speed at the implementation level, only cost, nothing we can do, frontier is frontier for a reason.
And even then, most people would rather buy groceries in a Lambo than pretend they love the Toyota.
Intelligence is the eval because intelligence is the objective. I don't care about doing the same work faster. I care about solving problems I couldn't solve before or would take an enormous amount of research, study and work.
If you look at these models/harness as daily task/jira solvers, or your benchmark is "can it solve my daily tasks cheaper?" you're optimizing for labor, not true capability.
Garlic I agree, but unlike to cut my onions, with a sharp enough knife, there is no crying.
Same for peppers, I cut them both (and used different types), based on what I am cooking, but only cook a few times a month since it's a huge time waster.