@Finumus1 xhigh seems like a very good model as well, straightforward and not too much talking, straight to business.
overall, very impressive results.
Introducing Kaleido💮 from @AIatMeta — a universal generative neural rendering engine for photorealistic, unified object and scene view synthesis.
Kaleido is built on a simple but powerful design philosophy:
3D perception is a form of visual common sense.
Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models.
Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures.
It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding.
Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable.
Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings.
We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details.
🔗 Explore more results and paper: https://t.co/fOcssVKQiW
📢 Hiring a PhD intern in Generative AI @Meta (London) to work with my team on video generation.
If you’ve worked on diffusion / flow matching (plus ideally LLMs/VLMs) and have strong engineering skills, apply here: https://t.co/HxYgEVBa7T
24-week internship preferred
We’re releasing access to Speech-1, our conversational speech model designed specifically for customer phone calls (8khz telephony) to reduce call drop rates.
Announcing -
WebGPU Puzzles:
Learn GPU Programming in Your Browser
It's a web app that let’s you practice writing GPU compute kernels using WebGPU - runs 100% in the browser, locally on your GPU (even tiny integrated laptop GPUs).
As many of you know, over the past few months I have been sharing Prompt Engineering resources in different forms. I have now compiled them all into a cohesive publication and uploaded to arxiv: https://t.co/7TZgGF67lj