As I said before, but I will be nicer this time, Arena needs to overhaul the benchmarks for image and video models and stop these useless benchmarks. They need to test the models on complex prompts and measure model censorship, colour and texture quality, and usefulness. We are in August 2026; this is disrespectful to everyone.
We did it, finally... Codex & ChatGPT desktop, now on Linux.
Thanks for waiting and you can cancel that MacBook order if you got impatient. It’s that good.
👀
🎬 Global Launch: Dreamina Seedance 2.5 is now live!
Our most cinematic video model yet has arrived, and Dreamina AI is your official platform. Step into a new era of precision control and long-take storytelling:
· Native 30s videos in a single generation
· A new interactive editing experience: reshape any part of the frame
· Long video mode: up to 3 minutes, with consistency that holds throughout
· Dreamina AI plugins for Maya and Blender, built for film-grade production pipelines
· Up to 50 multimodal references, now with support for 3D white models and green-screen footage — bringing AI generation into real production workflows
· More true-to-life lighting and shadows, richer motion detail, and precise timestamp control down to 1 second
Dreamina Seedance 2.5 is more than a model. It's an upgrade to what you can create. Now open to Dreamina AI subscribers.
🕐 Available now across Southeast Asia, the Middle East, Africa, Europe, and South America. More countries and regions coming soon.
✨ Watch @yingjian355352 in action with the new model
⏳ Limited time: save up to 32% on credits per generation
🎁 Repost and comment – first 300 users get 1,000 credits in 12 hours
#DreaminaSeedance25 #DreaminaAI
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
I switched after Anthropic disabled my account the day after identifying me as Chinese—an experience that felt discriminatory and deeply dehumanizing. OpenAI’s GPT-5.6 Sol has been the opposite: generous usage, excellent quality, and consistently reliable. Easy choice.
Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched?
Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens!
https://t.co/8mU93eA13i
I am super excited to share that I launch a weekly Video Model Journal Club. Every week we pick one paper and go deep, i.e. video generation, world models, physical reasoning, diffusion, flow matching, and everything in between.
This Friday, we will have Yilun Du @du_yilun from @Harvard giving us a talk on Embodied Reasoning with World Models in person at @moonlake - really grateful for Fan-yun Sun @sunfanyun, Charlotte @xia_char and Shin @shinshin_oob for hosting.
Register for in-person via Luma: https://t.co/nquwdXfaKc
#video #AI #SF
🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length.
🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.
🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice.
Try it now at https://t.co/GCdiMzk1Dl via Expert Mode / Instant Mode. API is updated & available today!
📄 Tech Report: https://t.co/drlDrxkYtp
🤗 Open Weights: https://t.co/T13Y8i7SDM
1/n
Anthropic is guilty of stealing training data at massive scale and has had to pay multi-billion dollar settlements for their theft. This is just a fact.
Company that trained on everyone's data without asking is upset that someone trained on its data without asking
2026 is the year of open source for a reason
This is the secret to Seedream4.0's success:
They use a custom RL model framework,
RewardDance: Reward Scaling in Visual Generation
Max derivative to 26B model.
Just published my blog site along with a new blog "Go with the Flow" - I've been diving deep into flow-based models over the past few months, and this is the first part where I break down how they work internally. I have covered topics like Normalizing Flows, Flow Matching, Conditional and Marginal Probability Paths and Vector Fields, Rectified Flows, Optimal Transport and Reflow.
Link to the blog in the next tweet!
Some further thoughts on the idea of "thinking with images":
1) zero-shot tool use is limited -- you can’t just call an object detector to do visual search. That’s why approaches like VisProg/ViperGPT/Visual-sketchpad will not generalize or scale well.
2) visual search needs to be a native, end-to-end component within multimodal LLMs. This was our focus in V*, but two yrs ago, we didn’t realize how powerful RL would become, so we had to stick with SFT to train detection heads. It worked, but it was slow and kind of a pain.
3) however, when the tools are simple and low-level -- say, basic python image processing functions, rather than a faster R-CNN -- they can be incorporated directly into the end to end system. With RL at scale, these simple tools become visual primitives you can mix and match to build scalable visual skills.
4) we should keep identifying these visual primitives. It’s definitely not just simple image processing functions; think about video and 3D.
5) finally, I think most traditional visual recognition models are dead. As the great @inkynumbers said, they’re parsers (https://t.co/E9zzVkOkA1). But vision itself isn’t dead. i think it’s more alive and exciting than ever.