Thank you for the retweet!
Our dataset, PixelProse contains descriptive and dense captions for over 16M images through the Google Gemini Vision model!
We carefully curate images from 3 different sources, filter for CSAM, and provide additional filters and metadata.
Forget about all the captioning datasets you've tried before!
PixelProse is a captioning dataset of 16M image-caption pairs, with less toxicity and higher details ✨
https://t.co/xYrMOjsyzU
⛷️Here’s my entry for the fast generative model olympics🥇
The Sphere Encoder is an autocoder so powerful that it produces high quality images quickly and without diffusion.
At training time, we learn an encoder that maps natural images uniformly onto the surface of a sphere. At inference time, we sample a random vector from the sphere, and a decoder makes it into an image.
Guardrails with custom polices are hard for models trained on safety and harm-related datasets. But what if you trained a guardian model on arbitrary rules?
Introducing DynaGuard, a guardian model for custom policies: https://t.co/oPWOZstRUQ
Introducing MORSE-500
🌐 https://t.co/drjCHyEgjJ
500 scripted videos that stress-test six reasoning skills — beyond math, beyond static pics, built to get harder.
Key Features:
🚀 Fresh & Portable
🎯 Diverse Categories
👁️ Pure Visual Cues
📈 Scalable Difficulty
Dive in 🧵
Attention sinks in LLMs are weird. There’s ~20% of heads that don’t seem to do anything.
Do these heads matter? Turns out that if we get rid of them, benchmark scores don’t change.
Looking at the reviews in ICML, I am noticing more and more that some reviewers are assuming knowledge or rumors that may or may not exist in industry labs. This isn't great for open research
My team at @GoogleDeepMind is hiring.
If you are passionate about robust ML, the provenance of synthetic media, and the trustworthiness of data, consider applying: https://t.co/vWlfLPI4sN
Ok, so I can finally talk about this!
We spent the last year (actually a bit longer) training an LLM with recurrent depth at scale.
The model has an internal latent space in which it can adaptively spend more compute to think longer.
I think the tech report ...🐦⬛
New open source reasoning model!
Huginn-3.5B reasons implicitly in latent space 🧠
Unlike O1 and R1, latent reasoning doesn’t need special chain-of-thought training data, and doesn't produce extra CoT tokens at test time.
We trained on 800B tokens 👇
📢 We're hiring a Postdoctoral Associate to research 3D scene reconstruction, novel view synthesis, and inverse rendering. Join our team and contribute to cutting-edge projects in computer vision!
🔗 https://t.co/lN7CWRqOa3