Introducing Boogu-Image-0.1 ๐จ
Open-source unified model family for image understanding + generation โ Base, Turbo, Edit, Edit-Turbo.
Just 208M images, ~$400K compute โ yet rivals open SOTA & approaches Nano-Banana-Pro & GPT-Image-2.
Apache 2.0 ๐งต๐
๐ https://t.co/9ec1rLMv7n
Excited to introduce #TruckDrive ๐ at #CVPR2026: a new long-range driving dataset built specifically for long-range truck autonomy, where safe braking and anticipatory planning demand perception hundreds of meters ahead, far beyond existing robotaxi datasets.
๐ฆ TruckDrive includes:
๐น 475K samples, with 165K densely annotated frames
๐น Benchmarks for end-to-end driving, tracking, planning, depth estimation, and up to 1,000m for 2D detection and 400m for 3D detection ๐๐ฏ
๐ฐ๏ธ A purpose-built long-range sensor suite:
๐ธ 7 long-range FMCW LiDARs (range + radial velocity)
๐ธ 3 high-res short-range LiDARs
๐ธ 11ร 8MP surround cameras for short and long-range๐ท
๐ธ 10ร 4D FMCW radars ๐ก
โ ๏ธ Key finding: current state-of-the-art models break down at long range
๐ with 31% to 99% drops on 3D perception tasks beyond 150m. TruckDrive exposes a long-range generalization gap that current architectures and training signals are not closing yet - a benchmark for the next generation of long-range highway autonomy research ๐
๐ Project and Data: https://t.co/fNzDCbGQRQ
Fun work together with @torc_robotics led by Filippo Ghilotti, Edoardo Palladin, Samuel Brucker, Adam Sigal, and Mario Bijelic.
Congrats to Linhan for his first NeurIPS paper.
DC-Gaussian could reconstruct high-quality 3D Gaussain Splatting via shitty cameras shooting through windshield (car dash cameras)๐.
Codes have been released.
See project page here: https://t.co/F6DMTOAg6q
Thanks for the tweet.
We improve 3D Gaussain splatting for dash cam videos that contain serious reflections and obstructions.
With Dash cam videos, it's possible to leverage massive driving data.
Visit our page at https://t.co/F6DMTOzIgS
Codes will be soon.
Since Sora is out, I have been thinking about our role in academia. One thing we can do at school is fast prototyping with very talented students, showing the potential, the possibility. Of course, the future will always be scaling up.
Seeing and Hearing
Open-domain Visual-Audio Generation with Diffusion Latent Aligners
Video and audio content creation serves as the core technique for the movie industry and professional users. Recently, existing diffusion-based methods tackle video and audio generation separately, which hinders the technique transfer from academia to industry. In this work, we aim at filling the gap, with a carefully designed optimization-based framework for cross-visual-audio and joint-visual-audio generation. We observe the powerful generation ability of off-the-shelf video or audio generation models. Thus, instead of training the giant models from scratch, we propose to bridge the existing strong models with a shared latent representation space. Specifically, we propose a multimodality latent aligner with the pre-trained ImageBind model. Our latent aligner shares a similar core as the classifier guidance that guides the diffusion denoising process during inference time. Through carefully designed optimization strategy and loss functions, we show the superior performance of our method on joint video-audio generation, visual-steered audio generation, and audio-steered visual generation tasks.
Code for Neural Spline Fields is out! We use burst image fusion to see through occlusions and remove reflections.
We just released the code here: https://t.co/frpg3ex38K
Comes with a fun set of notebooks and tutorials!
Ultra-thin flat cameras are possible with nanophotonic optics! Excited to share recent work that shrinks the entire optical stack down to a 700-nanometer thick layer of optics on the sensor cover glass.
We will present this paper later this week at #SIGGRAPHAsia2023 !
Paper: https://t.co/B0frFesaqg
Super fun work with @praneethchk, @JipengSun2 from @Princeton, and our friends from UW Arka Majumdar, and Johannes E. Frรถch.
Splines instead of Gaussians ๐ Introducing Neural Spline Fields, which can see through occlusions!
We learn to represent a stack of misaligned captures as a multi-layer image sandwich. Then you can extract your favorite layer to remove occlusions, reflections, or even your own shadows from the scene!
Paper and Code: https://t.co/7hToSxgqPu
Fun work by @_ilya_c , @DavidShustin , @RuyuYan00, and Chenyang Lei.