ml-headsup
Large-Scale High-Quality 3D Gaussian Head
Reconstruction from Multi-View Captures
Apple
https://t.co/mpiW75a1Ql
Abstract
We propose HeadsUp, a scalable feed-forward method for reconstructing high-quality 3D Gaussian heads from large-scale multi-camera setups. Our method employs an efficient encoder-decoder architecture that compresses input views into a compact latent representation. This latent representation is then decoded into a set of UV-parameterized 3D Gaussians anchored to a neutral head template. This UV representation decouples the number of 3D Gaussians from the number and resolution of input images, enabling training with many high-resolution input views. We train and evaluate our model on an internal dataset with more than 10,000 subjects, which is an order of magnitude larger than existing multi-view human head datasets. HeadsUp achieves state-of-the-art reconstruction quality and generalizes to novel identities without test-time optimization. We extensively analyze the scaling behavior of our model across identities, views, and model capacity, revealing practical insights for quality-compute trade-offs. Finally, we highlight the strength of our latent space by showcasing two downstream applications: generating novel 3D identities and animating the 3D heads with expression blendshapes.
Roomform
Parses indoor point cloud scans into structured, editable scenes — an open RoomPlan API-style pipeline for any registered point cloud (laser scans, ARKit exports, RGB-D reconstructions): walls and openings, classified oriented object boxes, per-object segmented points, and evidence-gated object meshes.
I'm releasing Roomform, an open-source alternative to the RoomPlan API.
It ingests point cloud scans and extracts geometry including walls, doors, windows, objects. No rectangular-room or Manhattan layout assumptions.
The model also infers structure behind occluded areas like furniture, cabinets, and unscanned corners.
home designers, this one is for you.
A rough sketch can now become a complete home design. Draw the footprint, list the rooms you need, and get detailed floor plans, a 3D model, and an exterior rendering.
AI can now reconstruct the 3D layout of an entire building from images alone.
It's called PolyLayout.
Most systems map one room at a time and stitch the results together. PolyLayout reconstructs connected indoor spaces all at once, so the whole floor plan holds together instead of breaking at every doorway.
→ Rebuilds multiple connected rooms in 1 pass
→ Combines neural networks with geometric reasoning
→ Stays precise across different camera setups
100% open-source.
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
TL;DR: Qwen-3D unifies multi-view inputs in a shared 3D space using depth and camera poses. With 3D RoPE and a dense mask decoder, it supports spatial reasoning, grounding, segmentation, and VQA while outperforming prior 3D LMMs and retaining 2D capability.
https://t.co/zyinrSnV6V
2D is what the camera sees.👀
3D is what the world is. 🌍
Today, we’re proudly open-sourcing QuerySplat — an open-source feed-forward 3D model that turns a few unposed images into a navigable 3D scene in seconds (run on 4090🪶).
Your input: a few unposed images
Our output: a sharp, navigable 3D scene — generated in seconds! 🚀
🧠 organize scenes with 3D queries, not pixel-bound Gaussians
📷 no camera poses or manual calibration required
⚡ reconstruct new scenes in a single forward pass
✨ preserve sharp geometry and high-frequency appearance
🏆 achieve state-of-the-art results on DL3DV-Evaluation
🔓 code and model weights are now open-source
From pixels to scenes, we are making 3D generation scalable.
Try our free iOS app — it’s a lot of fun! 😆
🔗 https://t.co/VaguGSnRqZ
Try it on the Web: 🔗 https://t.co/3ZrCuGZCKS
Paper:
🔗 https://t.co/gMHB9b5vjR
Code & weights:
🔗 https://t.co/ofbNZMjtcF
Project page:
🔗 https://t.co/GRkQm4oTIh
#AIGC #3DGS #AI #Computervision #3D #artist #3dartist #3d
Finally, the Bathroom Configurator is live! 🚀
• Custom walls, doors , windows etc
• Vanities, basins, taps, toilets , showers & more.
• Wall lights & bathroom accessories
• Random Scene Generator
• Texture Wheel for easy customization
Try it now:
https://t.co/ntMJn2C1Wr
While creating a tool to create simple line drawings illustrations for a client in order to overlay some airflow dataviz, I may gone a bit too far and created something a bit more advanced 😅
This studio apartment is a 3D Gaussian splat - and the collision that makes it walkable is just 21KB. 🤯
Auto-generated voxel collision, perfectly approximating the gaussians. Watch me toggle it on and off 👇
[1/2]
Been building a system that turns point clouds and scans into reconstructed, editable environments.
The goal isn’t just visual fidelity. Walls, floors, ceilings, and objects should remain separate so users and agents can understand and modify the scene.
holy sht.. you can now control computers with your thoughts
someone at WAIC 2026 is playing Black Myth: Wukong using only a brain-computer interface
the system flashes visual targets at different frequencies, your brain generates unique EEG signals in response, which are detected by a headset and translated into in-game controls
staff say it only takes around 5 minutes to calibrate
we’re watching the next generation of human-computer interaction in real time
Apple Maps is having one of the craziest glow-ups I've seen👀
Gaussian Splats already look incredible
Dynamic lighting would make this feel straight out of a next-gen game
Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
CVPR 2026
https://t.co/Eh9p0KYlal
Abstract
Stereo foundation models achieve strong zero-shot generalization but remain computationally prohibitive for real-time applications. Efficient stereo architectures, on the other hand, sacrifice robustness for speed and require costly per-domain fine-tuning. To bridge this gap, we present Fast-FoundationStereo, a family of architectures that achieve, for the first time, strong zero-shot generalization at real-time frame rate. We employ a divide-and-conquer acceleration strategy with three components: (1) knowledge distillation to compress the hybrid backbone into a single efficient student; (2) blockwise neural architecture search for automatically discovering optimal cost filtering designs under latency budgets, reducing search complexity exponentially; and (3) structured pruning for eliminating redundancy in the iterative refinement module. Furthermore, we introduce an automatic pseudo-labeling pipeline used to curate 1.4M in-the-wild stereo pairs to supplement synthetic training data and facilitate knowledge distillation. The resulting model can run over 10× faster than FoundationStereo while closely matching its zero-shot accuracy, thus establishing a new state-of-the-art among real-time methods.
zoxilsi studio https://t.co/mmWRUSfyyv
A browser based mesh gradient design tool built on WebGL no installs no sign up just open and create
Tired of boring static backgrounds wanted a fast premium feeling tool that works right in the browser
♡ Features
● Live WebGL mesh gradient editor with fluid silky color blending
● OKLab color space for perceptually accurate colors with no muddy transitions
● 100 plus hand crafted presets including Glass Mosaic Silk Aurora and Cyberpunk
● Real time effects like grain blur chromatic aberration vignette and glow
● Geometric overlay patterns with tiles hexagons dots waves and stripes
● Export to 4K PNG animated video CSS SVG or React code
● Animate with drift and hue flow with one click
● Dark and light mode keyboard shortcuts and full undo redo history
● Fully open source and deployable on Vercel in one click
Tech
Next.js 15 React Three Fiber Three.js Zustand Framer Motion Tailwind
Open source
https://t.co/i9WQUQW3CP