🔴 ¡PRÓXIMO GEMINI 4 ARGON!
Google acaba de anunciar la que será su próxima generación de modelos Gemini 4, con nomenclatura nueva (similar a lo hecho por Anthropic y OpenAI) donde Argon es su nuevo modelo frontera!
Y en varios benchmarks lo logra, superando a Astra y Fable 🔥
One intriguing thing about SiamJEPA: its attention maps become surprisingly object-centric 👀
We are still investigating why this happens, but the patterns are remarkably clear.
Preliminary results below 👇
Introducing Diffusion Controller, a lightweight steering damper that precisely steers image generation for better prompt alignment without breaking stability. Read the blog to learn how it boosts image quality without breaking baseline stability → https://t.co/ejCnuqoTy1
Today we’re introducing Gemini 4 Argon.
It delivers frontier performance in complex workflows across real-world software engineering, knowledge work, and cybersecurity defense with an industry-leading 1M token output limit.
5 years ago I was working on LLMs because I believed in them. And yet I didn't think we'd get where we are today without a major paradigm shift. I definitely didn't expect things to move so fast.
I was wrong.
It's astonishing to see some of the very vocal people who were far more skeptical than I was insist that they were right all along, because "hey, today's models do a lot more than predict the next token."
Like, sure, coding agents can run commands in a terminal. But they do it by predicting the next token. What in the world did you expect?
No one was claiming that LLMs would solve every problem without any way to interact with their environment. But the layer on top is very thin, and tool use isn't a new paradigm. "Giving BERT a Calculator" came out 7 years ago.
Science is about updating your beliefs based on evidence. Admitting that your view has changed doesn't make you a worse scientist. It makes you a better one.
This obsession with "I've been saying the same thing for X years and I was right all along" is so tiresome, and sets a terrible example for the general public and the new generation of scientists.
Ollama now supports Jev-like decision models all locally.
Use decision models like Nimble for tasks like ticket triaging, model routing, and content moderation.
ollama pull nimble
Here’s Nimble playing Ollama racer through the new local /v1/systemone API by making decisions in real-time. 🏎️
Es habitual ver cómo en ocasiones se inflan las capacidades de algunos modelos opensource para elevarlos a lo que los modelos privados pueden ofrecer.
Lo extraño es ver que quien lo hace sea la misma Anthropic queriendo demostrar que GLM 5.3, de la empeza china Zai, ya está al nivel de su afamado y temido Mythos Preview!
Por ahora demuestran que esto ocurre con las capacidades de ciberseguridad (justo las que más temen los de Anthropic). Y obvio que ese reconocimiento a los modelos GLM viene motivado por querer hacer sonar las alarmas ante el riesgo de los modelos abiertos, el gran terror de Dario Amodei.
JEPA Learns What the Mask Leaves Unrecoverable
TL;DR: Why do block masks work so well in JEPA? Not because of their shape itself, but because they hide information that low-level interpolation cannot recover. Experiments on images and videos show that both unrecoverability and context reachability matter.
https://t.co/IjtzaC8AXz
🔥Welcome 🔺Simplex Diffusion Models🔻!
Standard discrete diffusion models suffer from "information collapse" because they discard uncertainty at intermediate steps by sampling categorical tokens. (1/5)
¿Buscas una alternativa de código abierto a Jev?
Laya hace lo mismo y puedes entrenarlo con tus datos.
✓ Open Source + Corre en local
✓ Clasifica con score, noul y choices
✓ Versión en inglés y otra multilenguaje
→ https://t.co/nKDGpFrGqJ
Target representation matter for denoising models (e.g., RAE).
But what about the Source?
Can we learn source for better generation?
Flow matching doesn't require Source to be Gaussian noise.
Meet CSFM: end-to-end source learning for T2I.
🎉 NeurIPS 2026
https://t.co/ngsS3rRq54
NVIDIA drop a 3B vision-language model for fast, high-quality visual grounding in real time accurately.
- Parallel box decoding
- 10× faster than Qwen3-VL
- Trained on 138M queries/785M boxes
- GUI, OCR, and document layout, dense detection
- Open source
Useful for computer-use agents and Physical AI,
FreeBridge: Variational Schrödinger Bridges for Cellular Transition Dynamics
TL;DR: Models cellular perturbations beyond endpoints by constraining intermediate trajectories to remain biologically plausible. Uses Schrödinger Bridges in latent space and achieves strong results on BBBC021, RxRx1, and JUMP.
https://t.co/HaQebTsiXj
Join us on Oct 10 if you’re a world model lover in SF! I will share our ICML paper on representation learning for WMs and recent work AdaJEPA. Excited to attend my first reading club 📚!
🚀 We introduce Neural Theorizer (NEO) — a new type of world model that learns to theorize the world from observation, without language or LLM supervision.
Selected as an ICML 2026 oral presentation — 0.7% of submitted papers.
The paper asks:
"What does it mean to understand the world and build a world model?"
Today’s world models are often trained to predict the future: the next frame, next latent state, or next observation.
But is prediction enough?
We argue that a world model should be a theory-building system: one that discovers reusable primitives, composes them into executable explanations, and transfers those explanations to novel phenomena.
NEO is our first step toward this vision — a World Theory Model that learns explicit, compositional theories from raw observation.
This work was led by my wonderful students: Doojin Baek*(@doojin_a_baek), Gyubin Lee* (@gyubin0521), Junyeob Baek (@JunyeobB), and Hosung Lee (@HosungLee_).
For more details, take a look at the paper — and if you’re attending ICML, let’s talk there!
📄 arXiv: https://t.co/TGMXLLfzP7
🌐 Project page: https://t.co/aLJywp8rfq
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family.
It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
🧵We discovered a new phenomenon in speculative decoding models that we call attention drift.
As the drafter generates tokens, its attention moves from the "sink" onto its own recently-generated tokens.
Fixing the underlying issue recovered up to 2× acceptance length, but why?
Learning a Flow to Self-Supervised Representations
TL;DR: FBDM brings Flow Matching to self-supervised representation learning via spherical conditional velocity regression, removing the need for an adversarial critic while matching DM performance and training 1.48–1.83× faster.
https://t.co/rS5raYa1nY
Can we build world models purely in the latent space, without pixel prediction?
We present Contrastive World Models - we train the latent states to maximize mutual information with future observations, without any decoders / pixel reconstruction.
Contrastive World Models enables
⚡ Substantially more robust representations
🤖 More efficient training by removing the pixel decoder entirely
🌍 A general approach with minimal assumptions [1/n]