Several new methods to shape the latent distributions of autoencoders have popped up recently.
They are often compared against the traditional VAE setup, where a KL penalty encourages the latents to be Gaussian. 🧵👇
(1/10)
Hello World. MAI image is our first image generation model, following MAI voice and MAI text.
Congrats team!
Oh, and if you are a strong engineer who wants to make this model climb to Number 1, please send your CV to [email protected]
Meet our third @MicrosoftAI model: MAI-Image-1
#9 on LMArena, striking an impressive balance of generation speed and quality
Excited to keep refining + climbing the leaderboard from here!
We're just getting started.
https://t.co/33BiNfIjPg
Stoked to get to talk to @lexfridman + my homie @dylan522p for 5+ hours to try and get to the bottom of what is actually happening in AI right now.
DeepSeek R1 & V3, China v US, open vs closed, decreasing hype, datacenters, everything in between...
🚀 what a fun whirlwind week
New paper: What happens once AIs make humans obsolete?
Even without AIs seeking power, we argue that competitive pressures will fully erode human influence and values.
https://t.co/efbmjqcTIy
with @jankulveit@raymondadouglas@AmmannNora@degerturann@DavidSKrueger 🧵
We reproduced DeepSeek R1-Zero in the CountDown game, and it just works
Through RL, the 3B base LM develops self-verification and search abilities all on its own
You can experience the Ahah moment yourself for < $30
Code: https://t.co/UcGKN2SVGj
Here's what we learned 🧵
Most AI researchers I talk to have been a bit shocked by DeepSeek-R1 and its performance.
My preliminary understanding nuggets:
1. Simple post-training recipe called GRPO: Start with a good model and reward for correctness and style outcomes. No PRM, no MCTS no fancy reward models. Basically checks if the answer is correct. 😅
2. Small models can reason very very well with correct distillation post-training. They released a 1.5B model (!) that is better than Claude and Llama 405B in AIME24. Also, their distilled 7B model seems better than o1 preview. 🤓
3. The datasets used are not released, if I understand correctly. 🫤
4. DeepSeek seems to be the best at executing Open AI's original mission right now. We need to catch up.
This can be big. Google unveils the successor to the Transformer architecture
"We present a new neural long term memory module that learns to memorize historical context and helps an attention to attend to the current context while utilizing long past information. We show that this neural memory has the advantage of a fast parallelizable training while maintaining a fast inference."
"From a memory perspective, we argue that attention due to its limited context but accurate dependency modeling performs as a short term memory, while neural memory due to its ability to memorize the data, acts as a long-term, more persistent, memory. Based on these two modules, we introduce a new family of architectures, called Titans, and present three variants to address how one can effectively incorporate memory into this architecture."
"Our experimental results on language modeling, common sense reasoning, genomics, and time series tasks show that Titans are more effective than Transformers and recent modern linear recurrent models."
"They further can effectively scale to larger than 2M context window size with higher accuracy in needle in haystack tasks compared to baselines."
Qwen released a 72B process reward model (PRM) on their recent math model. A good chance it's the best PRM openly available for reasoning research. We like Qwen.
@HaiperGenAI API is now available on @ComfyUI with the features:
🟢 Key Frame Conditioning
🖼️ Image2Video
📽️ Text2Video
🏞️ Text2Image
Try out: https://t.co/Nt5niaKf0I
Haiper API: https://t.co/xGrpjKcGx2
#Haiper#ComfyUI
🚀 Haiper 2.5: Enhanced Mode is here.
Take control like never before with Keyframe Conditioning Timeline, letting you customize every frame to perfection.
Sharper. Smoother. Simply revolutionary.
Catch us at https://t.co/QXrGpak5Ab and take your creativity to the next level. 🌟
#AI #EnhancedMode #Haiper2_5
Don't sleep on @HaiperGenAI
They just released their v2.5 model and it passes the brain worm test - both in t2v prompt adherence and not blocking what I want visualized. Plus the resolution is quite good on initial tests. Now go eat some turkey 🦃
🚀 Introducing Haiper 2.0: Text-to-Image Like Never Before! 🚀
Unleash your creativity with sharper, more realistic visuals at lightning speed. Whether you’re a creator or a brand, Haiper 2.0’s Text-to-Image feature makes transforming ideas into images effortless. Ready to see the magic? 🪄
Whaaa, https://t.co/ZQrl5yl9Dk is now more than just #AIvideo?! YUP.
Check out these 8 wild AI text-to-images, then try it out for free at https://t.co/QqFfBhqzxK! #AIimages