🎨 Meet Qwen-Image-3.0 — the third generation of our foundational image generation model.
If 1.0 was about "Precision," and 2.0 added "Variety, Completeness, Beauty & Authenticity," then 3.0 comes down to a single word: Real (实).
Three dimensions of "Real":
📰 Rich Content — prompts up to 4.5k tokens. One-pass generation of complex layouts: newspapers, storyboards, exam papers — even a 3×3 infographic grid or picture-in-picture-in-picture UIs.
🔬 Authentic Details — text legible down to 10px, full LaTeX paper pages, pores, hair strands & near-photographic skin texture.
🌏 Deep Knowledge — native rendering in 12 languages, 100+ art styles, realistic UIs (web / games / livestreams), plus world knowledge & live web retrieval.
Not just "good-looking" — genuinely useful. Image generation as a real productivity tool for design, content, education & e-commerce.
Go create 🏃🎨
💬Qwen Chat: https://t.co/941HmITJ2W
📝Blog: https://t.co/5mnS4uI9Ar
Introducing FLUX 3.
One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style.
FLUX 3 Video is now available in early access (link below).
Jointly trained in one unified architecture, our model can be extended to predict actions for robotics. See our work with mimic and Audi in the thread.
ReChannel
A new way to extract depth, normals, matting, and segmentation from a single image using a frozen text-to-image DiT — with only 33K trainable parameters per task.
This is Segment Anything for #GaussianSplatting running on the web... but it comes with a caveat.
This is SplatEdit, a powerful #3DGS editing software built on top of SuperSplat.
Last week, it added a segmentation selection feature powered by Segment Anything, running locally in the browser. Just upload a splat, click on the "mug" icon and select entire objects in the scene.
The problem is that everything after the object is also selected (as you can see in the video).
To fix that, just use the brush tool to subtract the unwanted portion of the scene (hold CTRL while brushing).
I feel we are getting closer and closer to having more editing control on these amazing fuzzy clouds 😉
GenCeption from Google DeepMind
A single feed-forward vision model that matches specialist models on depth, surface normals, camera pose, keypoints & segmentation
CT/MRI画像から3DCGを高速に作成できるKaloLumenの無料デモ版をご提供します。医療関係者・研究者で以下のスペックのPCを持つ方はお気軽にDMください。このソフトは教育・研究用途で診断・診療には使えません。 OS: Windows 11 Home or Pro
CPU: 12th Gen intel® Core™ i7-12700Hと同等以上
RAM: 16GB以上
GPU: NVIDIA® GeForce RTX™ 3050, VRAM 6 GBと同等以上
Playing with getting trellis2cpp working e2e on my 16gb card. They recommend using 24gb, but depending on how long you are willing to wait, it appears to work even on 8gb. @rms80
Cool data!
https://t.co/11IcdK9FJP runs in your browser
Just drag and drop your colmap dataset with 3dgs ply
You could also host data on Huggingface for free or using your private bucket.
"PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation"
TL;DR: a minimalist pixel-space diffusion transformer predicts dense 3D point maps directly from a single image, removing the need for VAEs and hybrid architectures while producing sharper geometry.
We added bending and tapering to superquadric reconstruction / decomposition.
➡️ big step forward in reconstruction quality
Check out SuperFlex for compact & explicit 3D object/scene representations. @eccvconf#ECCV
More details: https://t.co/ZVaZQznVAL
Does 3D reconstruction have to be complex?
We answer this question with PointDiT (#ICML2026): a minimalist pixel-space Diffusion Transformer without bells and whistles.
We show that a plain ViT can estimate dense 3D point maps by operating directly on raw patches. No hybrid ViT+Conv architectures, no lossy VAEs, no complicated training losses. (1/5) 🧵👇
Super cool new sun direction relighting LoRA for Flux Klein 🌞
> Had my agent build a ZeroGPU demo for it using our new space builder skill for LoRAs from huggingface/skills
https://t.co/rn5nz8bdJg