PhD applications will be starting soon! Time to start drafting your SOP.
Sharing my SOP from last cycle with annotations and general takeaways. Hope this helps people applying this cycle!
https://t.co/ZG0gKSZ1ov
PhD applications will be starting soon! Time to start drafting your SOP.
Sharing my SOP from last cycle with annotations and general takeaways. Hope this helps people applying this cycle!
https://t.co/ZG0gKSZ1ov
PhD applications will be starting soon! Time to start drafting your SOP.
Sharing my SOP from last cycle with annotations and general takeaways. Hope this helps people applying this cycle!
https://t.co/ZG0gKSZ1ov
Hi, I'll be presenting ELT (https://t.co/MhiFVDeMNu) on 12th September at #ECCV2026
📍MalmöMässan Exhibit Hall, Poster 141
🕦 10:30AM - 12:30PM
Excited to discuss looped transformers for visual generation, and in general efficiency & adaptivity in large models.
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
I was an intern under Seb in 2020. Unfortunately, the allegations of unscrupulous behavior is 100% believable. I am glad that the mask is off publicly. I really hope he doesn't weasel out of this.
https://t.co/Pe6e6cK920
@matiks_play sent me a hoodie for 200 days streak or something, bro I have opened the site just once in my lifetime. Thanks for the hoodie though, one size larger would help :)
@55CncAe the std enc-dec models are trained jointly, but Mostik is composing two independently trained models (basically cross-family decoder only models) that too by just training the bridge and freezing the models.
but yeah functionally it does look like an extremely asymmetric enc-dec
Mostik reminds me of Tandem Transformers (ICML 2024). Tandem trains the small drafter using xent and distillation, which feels more natural to me than freezing both models and training only the bridge.
Tandem also repeatedly refreshes the context using large model, making it more resistant to small model drift. Though large model still processes every generated token, but efficiently in parallel blocks.
Conceptually, Mostik's showcased setup is close to the no-refresh limit of Tandem: one large model prefill, followed by a full decode from a 4B model conditioned on that initial strong signal.
Its frozen cross-family setup is interesting though, but the final generation is still fundamentally driven by the small 4B model. This may make it difficult to close the performance gap with the large model, especially on long-horizon tasks.
meet @mostik_ai!
what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can.
everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning?
we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched.
how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance.
we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst
WIRED has the first external account of the company and the work: https://t.co/tP8nItCsDl
full writeup, the setup, and all the numbers: https://t.co/C9NZ5vtV1V
I think they'll naturally converge towards Tandem Transformers (https://t.co/LJ2X1PLBsJ) to close the performance gap.
step1: Train the smaller model too along with bridge
step2: Add distillation loss along xent loss
step3: Add layer-wise bridge
step4: Periodically update the context using large model (processing tokens efficiently in blocks)
step5: Add specdec
and your product is ready!
what the helly is this...
not spec dec, not disagg, but some lobotomized franken-glueing where the glued models are frozen? why? if you have access to latents, you probably also have access to model weights. freezing weights seem completely unnecessary, not to mention what kind of cursed deployment hell would need this kind of setup...
I am super excited to share this educational video that I had GPT-6 Astra create from a single prompt: “Create a 5 minute educational video about T cells,” cells that I had devoted most of my life studying!
I also had Astra to use the Remotion plugin for the video generation. Astra also used Imagegen to create the visuals, it also suggested to use HeyGen for narration (have not used that before) and then assembled the entire educational video on its own one-shot, including what to say, how to explain it!
The result is so incredibly well made that, narration, the animations are all so perfectly created that even with 35 years of experience studying T cells, I really don’t think I could have explained this any better myself! The quality is also better than anything I had seen before, Astra is far ahead of all models now!
I’ve also been creating much longer and more advanced videos on T cells, immunology, and related topics, and I’m now planning to build an entire video lecture series and put it all on a dedicated website.
I am just having the time of my life making these, it actually makes me emotional to be able to magically create these with few prompts!