Neural renderers shine in CUDA but break in game engines. MeshSplatBench reveals the exact pipeline flaws — like mismatched normals and occlusion handling — that tank fidelity on real hardware, so your 3DGS assets ship clean 🎯
📄 https://t.co/raB7xOE0uc
Novel view synthesis from sparse views often passes errors through a lossy rendering or 3D bridge. RoGe cuts that out: end-to-end implicit reconstruction and generation conditioned directly on the input images
📄 https://t.co/zJUvB6PVaL
We share H3-World 🌍
The first to turn MiniMax-H3 itself into world model.
No new action module. We directly convert H3’s pretrained language understanding into world control.
Only 8K samples + 0.199% Trainable Params.
📄 https://t.co/Bn4M4JwEJs
💻 https://t.co/sT44pUHAl8
Unstructured view sampling leaves gaps in large-scene digitization — coverage holes kill consistency. InceptionGS blends reconstruction with generation to repair missing regions without quality loss elsewhere.
📄 https://t.co/4S52PyyhiJ
Clever bit is decoupling structure & appearance diffusion priors — this lets DualDiff3D enforce view consistency without blurring geometry, crucial for reliable 3DGS from sparse inputs like Steam Frame’s dual-camera setup. https://...
↪️ @SadlyItsBradley https://t.co/gon0fVOfQ4
Arcturus Industries, who built the Steam Frame’s camera tracking system, is launching an official Color Passthrough Module
Dual 32 megapixel sensors, 10 bit HDR capture, 64mm IPD
Realtime environment mapping
Spatial Video recording
Gaussian Splatting support with tools
Camera-controlled video gen breaks on large baselines—attention logits explode. MeRoPE keeps them norm-preserving with multi-frequency rotary phases and epipolar priors, stable at any physical scale
📄 https://t.co/jEOZRSTwuu
🌐 https://t.co/dyaUo75iA3
Furry objects render beautifully in NeRFs but crumble in real-time graphics. This method replaces volumetric blobs with explicit, anti-aliased line segments for hair and fur that rasterize, shade, and simulate properly.
📄 https://t.co/QnpJqk3KVV
🌐 https://t.co/9w8F8a54SH
4D world models stuck on sequential video-then-reconstruction? Streaming4D couples block-wise video generation with incremental 3D updates — geometry evolves online with the stream, cutting latency without sacrificing fidelity
📄 https://t.co/Trndk8aDcG
The clever bit is Lucida’s joint optimization of parse-generate-place, breaking the circular dependency where each step needs clean input from the prior one. That’s what makes it practical for cluttered real captures, unlocking editable...
↪️ @qin_minghan https://t.co/t1h9sHxaQU
Everyone benchmarks embodied AI in sims that look nothing like your living room.
OVER just closed the loop: robot + VLM navigating INSIDE a real-world Gaussian Splat — pose, render, decide, move, repeat.
Capture once, get a training world for free.
https://t.co/3qdC6DhVLw
A Gaussian Splat can become a world where Robots and AI agents can act.
In our latest OVER Research experiment, we placed a robot inside a real-world 3D capture, with a VLM making decisions based on what it sees.
At every step, the robot holds a pose in the reconstruction, gets a newly rendered view of the environment, takes an action, moves, and sees the world again from its new position.
Why does this matter?
Because 3D captures can become more than reconstructions to explore. They can become environments where embodied AI and robots can navigate, act, be evaluated and eventually train across real-world spaces at scale.
Capture a place once. Then turn it into a world where AI and Robots can act.
The full experiment, including what we discovered once we actually put the loop to the test: https://t.co/vR0uURk6YW
Stylizing a 3DGS scene meant choosing between weak VGG transfer or per-view diffusion drift. DReSG grounds diffusion proposals in 3D: attention-guided residuals, filtered across views, absorbed into one stable scene 🎨
📄 https://t.co/L7xD30KRga
🌐 https://t.co/StdjiVqp3d
Novel view synthesis can hallucinate shapes or blur depth for unseen parts. ReconSplat uses latent diffusion with 3D Gaussian features to produce photorealistic views and sharp, consistent depth maps.
📄 https://t.co/mPlPH0TU8j
🌐 https://t.co/uvpT7yHasq
The clever bit is KISS-GS reparameterising 3D Gaussians as 9 fixed-resolution images, unlocking standard codec compatibility. This bypasses the need for custom entropy models, making 3DGS compression instantly deployable on existing hardwa...
↪️ @NarenBao https://t.co/FSuJ0fLmvW
Training large radiance fields like 3DGS on a single GPU? VRAM blows up with scene size. ABCD reformulates training as block coordinate descent over spatial partitions—peak VRAM becomes O(1) with <5% PSNR loss 🔥
📄 https://t.co/bTO6zI4fiC
🌐 https://t.co/vXsfiWkvPn
3DGS gives stunning renderings but bloats models with useless semi-transparent Gaussians. Fast and Compact kills them early—a Polarized Opacity Prior forces primitives to commit: opaque or gone 🔥
📄 https://t.co/qzKQubr9qX
Global SfM scales to huge datasets but a few ambiguous image pairs collapse the camera graph. Robust Global SfM prunes bad edges via internal consistency checks in reliable subgraphs — clean poses, no artifacts.
📄 https://t.co/u9YFqkbO5n
🤯 Turn one image into an entire explorable 3D world? SpatialCrafter does it with a global 3D proxy — no more drift or hallucination in generated scenes.
📄 https://t.co/Zly5lVGNEu
🌐 https://t.co/Mb1JifafF6
The clever bit is fusing depth-derived geometric priors via inter/intra-patch fusion, not just adding depth as a channel. This lets DINOcular align semantic features from DINO with metric 3D structure, unlocking spatial reasoning for e...
↪️ @hermannsblum https://t.co/YhKVELt7wg
new result that I am excited about: A self-supervised visuospatial representation with better 3D awareness.
If you have RGB-D inputs (robots usually have), our model makes use of the depth to give you better semantic and geometric features than DINO.
📄 https://t.co/zOWWzEBIHw