🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta!
🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇
🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex!
Check out the configuration details in our official API docs: https://t.co/smCwQZMeiq
Seedance 2.5 in a realistic style, generated with Dreamina (2nd try, 20+ ref images). Result was ok, but overloading the refs likely hurt the output. Still way more reliable than 2.0! Quality improves with fewer refs, but you trade off accuracy for what you want to see.
Seedance 2.5 generated with Dreamina (first try with 30 ref images). Result was almost exactly what I wanted! I included too many elements so it might not look great, but 2.5 is way more reliable than 2.0 with a good storyboard.
If I have to choose only one AI right now, I'll go with OpenAI because it can be used directly with Hermes Agent. Plus, since it's cheaper than DeepSeek V4 Pro and Gemini 3.1 Pro Preview, it’s definitely my new favorite model for Hermes Agent! 😁
major price cuts today:
*80% drop for GPT-5.6 Luna, now $0.20 per million input tokens and $1.20 per million output
*20% drop for GPT-5.6 Terra, to $2/$12
*GPT-5.6 Sol gets Fast mode in the API, up to 2.5x the speed for 2x the price, same intelligence
Today, we open-source PrunaVAED, a faster drop-in replacement decoder for video generation with LTX-2.3!
- 𝗛𝗼𝘄 𝗲𝗳𝗳𝗶𝗰𝗶𝗲𝗻𝘁 𝗶𝘀 𝗶𝘁? ~1.7-2.1x faster decoder & ~50% lower peak VRAM
- 𝗛𝗼𝘄 𝗴𝗼𝗼𝗱 𝗶𝘀 𝗶𝘁? Near-original visual quality & no changes to the latent encoding
- 𝗛𝗼𝘄 𝗱𝗼𝗲𝘀 𝗶𝘁 𝘄𝗼𝗿𝗸? We pruned, trained & distilled video decoders to maximize efficiency of the LTX-2.3 series from @Lightricks
Want to try the fast open-source integration? Check it here on Diffusers: https://t.co/gJLged4ezV
Want to try the fastest video generation endpoint? Check p-video here on API: https://t.co/fbzM9gs7s8
Introducing FLUX 3.
One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style.
FLUX 3 Video is now available in early access (link below).
Jointly trained in one unified architecture, our model can be extended to predict actions for robotics. See our work with mimic and Audi in the thread.
🚨 Alibaba just launched Qwen-Image-3.0 and it looks ridiculously capable
The official improvements include:
• Up to 4.5K-token prompts
• Legible text as small as 10px
• Native rendering across 12 languages
• More than 100 supported visual styles
• Complex newspapers, exam papers and interfaces generated in one pass
• Entire 3×3 grids containing nine detailed infographics
There are no independent benchmark results yet, but the examples Qwen have shown are incredible, I can't tell if some of the images are AI or real, they may have just built one of the most useful image models available.
The Chinese labs are moving unbelievably fast.
Do you think this beats the competition in image gen?
NVIDIA Nemotron just dropped an audio-native model that hears the world, not just words 🔥
transcription, translation, sound recognition & audio Q&A, TTS and full speech-to-speech, hears, thinks, and talks back natively
open weights, 2B and 30B
▶️ https://t.co/jUv9HVSfr3
We’re rolling out three new models to make AI agents faster, smarter, and cheaper at scale:
🔵 Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost.
🔵 Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks like processing documents and agentic search.
🔵 Gemini 3.5 Flash Cyber: A cybersecurity model built to find and patch critical software vulnerabilities.
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:https://t.co/YRvcGdB9Bv
China:https://t.co/PKMUNwUuRp
Today we released Nemotron 3 Embed 8B and it reached #1 overall on RTEB 🏆
RTEB benchmarks retrieval accuracy across real-world tasks. Better retrieval gives agents more relevant context, helping improve response accuracy.
ICYMI: Google just open-sourced GNM ✨ (and its amazing)
> a parametric 3D head model with 253 identity + 383 expression parameters, built from real 3D scans
so impressed by the level of control🔥
Try it → https://t.co/8tGLANS2Ci
Kimi K3 is insane🤯
It built an Animal Crossing-style game, and it generated a fully playable experience with the cozy aesthetic, interactions, and gameplay loop in a single shot.
Open-weight models are starting to rival the best closed models for game generation. The pace of progress is getting ridiculous. 🌱🎮
No sound department? No problem.
Our new Foley LoRA built on LTX-2.3 adds synced sound design to silent footage. Footsteps, impacts, materials, and ambience layered in to match the action. No music bed, no dialogue, ready to drop straight into your mix.
Meet kbd-1.0-codex-micro, built with @work_louder.
Map the buttons and joystick to your workflow, and keep your pinned chats in view.
Get yours before stock returns 410.
We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions!
Here is a breakdown of what’s being fixed and updated in this release: 🧵👇