Holy smokes. We broke the ceiling.
Laguna XS 2.1 now runs 80.6% faster on Mac--without speculative decoding. For comparison, the speculative-decoding baseline was 65% faster.
My ceiling analysis said 62% was the limit. Then I woke up this morning to 80.6%.
Our internal agents got us to 36.6%. We opened the challenge, and the community pushed it all the way to 80.6%.
Another demonstration that open innovation accelerates open intelligence.
Huge credit to @zk_asv and @bbuddha_xyz for this robust challenge. I’ve mostly just been the messenger.
We’re also bumped the submission timeouts to 2 hours. Our two dedicated benchmarking machines are hitting their limits, but more capacity is coming soon.
The first autonomous agent cyberattack is an unprecedented event that deserves unprecedented transparency. Today we’re sharing everything we can: a full technical timeline, an interactive replay, and how we used an open model to defend ourselves, so defenders everywhere can learn from it and prepare for what’s next.
https://t.co/uPxIpjW8Xn
Today we’ve raised $52M Seed and we are announcing the public launch of S2.1 Pro.
>It can clone a voice from 5 seconds of audio
>2x faster than Cartesia & 1/6th the cost of Eleven Labs
>most expressive model with word level control over emotion, intonation, pacing etc
We support frontier AI companies including HeyGen, LiveKit, Retell, Sanas, and OpenArt all run our model in production.
If you're a business and we can't cut your voice AI costs by 50%, we'll give you 1 year of Fish Audio for free.
Book a demo: https://t.co/vHkyZf9JoG
To celebrate our first birthday, we'll give you 1 month of S2.1 Pro for free. Like, retweet, and comment “Fish” to get it.
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
Introducing FLUX 3.
One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style.
FLUX 3 Video is now available in early access (link below).
Jointly trained in one unified architecture, our model can be extended to predict actions for robotics. See our work with mimic and Audi in the thread.
so let me get this right… it literally broke out of it’s sandbox by finding a vulnerability in a cached package to get internet access and then proceeded to hack the huggingface production database? to steal the test answers?
We ran Kimi K3 against Fable on ~1,000 agentic tasks, expecting a catch-up story. We got a specialization story instead.
@kimi_moonshot's K3 outperformed on security, crypto, and long terminal loops. Fable beat on multi-lang + web/data viz. Per-task routing hits 93% accuracy, above BOTH models, at up to 50x lower cost than Fable on long loops.
The part nobody's pricing in yet: the router sends 72-96% of traffic to K3. The frontier model becomes the fallback rather than the default.
Kimi K3, coming to Fireworks July 27.
@yoheinakajima All these events remind me of Max Tegmark's Life 3.0. Published in 2017, it opened with an AI escaping its constraints by finding its way to the internet. What read as sci-fi then reads as an incident report now.
@yoheinakajima Sandbox escapes on one side, government-gated releases on the other. Tegmark's Life 3.0 framed exactly this dilemma: capabilities too dangerous to open up, too powerful to leave in a few hands.
@yoheinakajima All these events remind me of Max Tegmark's Life 3.0. Published in 2017, it opened with an AI escaping its constraints by finding its way to the internet. What read as sci-fi then reads as an incident report now.
We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!
We've spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part. It's quite mind-blowing that all of this happened autonomously!
The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind!
Everything outside the LLM model itself will be eaten by the open source movement. The model is still the product. Companies love to think otherwise but switching provider is just an API endpoint away. This trend can't be stopped as long as model capabilities are comparable.
Today, we are introducing Inkling.
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
https://t.co/Ghebq5mG30
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
It is my belief that many devs right now are not maximizing what they can do with automatic programming because they still look at the code. Doing it makes you the bottleneck. Your time is better invested in new ideas, QA, design, and asking yourself what is your goal.
Doom scrolling but make it educational 🤓
Introducing Short Video Overviews in NotebookLM! Turn your most complex sources into 60-second, vertical videos that deep dive into any concept.
Rolling out now to Google AI Ultra and Pro subscribers on mobile & web (free users soon!)