We live in a 𝕏 bubble.
Only 2% of us households pay for AI. Comparatively:
- 25% of households pay for SiriusXM
- 55% of pay for cloud storage
- 91% pay for at least one streaming service
You are so early
Thrilled to announce HalluWorld 😵💫 🌎 , a controlled benchmark + methodology to verifiably measure how LLM agents hallucinates while in an environment.
Now accepted at #NeurIPS2026 Evals & Datasets!!
We start with the thought that Reality is vast, dynamic, fractally unfolding.
W/o a stable reference and controllable complexity, verifiably measuring and inducing situations that lead to Hallucination a.k.a deviation from observed Reality gets tough.
This is even more so if we want to go from single turn Q/A setting --> Agents hallucinating while solving a task once placed in an environment.
HalluWorld puts models inside environments where we know exactly what is true at every step: Gridworlds 🕹️ , Chess ♟️, and Terminal 👩💻 tasks.
What's more?
Having such environments makes things neatly modifiable to am up support different levers of complexity, think canyons, signboards and long tunnels (Gridworlds) ; arbitrarily complex game situations or mods (Chess) inter alia --> This gives us ways to craft situations that amp up cognitive pressure on the agent
This is inspired and in vein of the bAbi methodology of creating questions with knowable answers from a world you can create by @jaseweston et al from 2015.
Come check out more at https://t.co/7fIKlxoh6y , or read our arXiv at https://t.co/9ZseW9PC0h !
It was enriching and fun working on this with @_emliu , @karanps, @stevenyfeng, @_michael_yu_, @ZhuofuTao and @shocheen from @LTIatCMU , @PatronusAI , @stanfordnlp and @OhioStateCSE over the past few months!
Incredibly proud of our team who got 4 papers accepted to NeuroIPS 🚀
Including @getdarshan's single author paper "Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL""
https://t.co/4GF238DNe1
To spotlight how much of an exemplary achievement that is, only 1% accepted papers are single author
"impressive" metrics:
- $$$ raised
- growing headcount
- how many tokens you've burned
ACTUALLY impressive metrics:
- if you hopped on a quick call today
- how many quick calls you hopped on
- how you made people feel when you hopped on a quick call
Australia has been hacked.
'And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia's extreme concern about this incident. And I also expressed my disappointment that it took the company way too long to inform the government what had occurred, and the nature of the way that that notification occurred as well was unacceptable.'
my cto said we have to let more people know that we are doing deep model alignment research.
so hi guys, psa we're doing tons of models alignment research.
Introducing Halo, the best framework for post-training of open-source models.
Halo delivers up to 2.8x the throughput of stock TRL with less peak memory, while models stay in their native HuggingFace format.
Star us on GitHub: https://t.co/3mAiUljdrN
I’m surprised people are piling on Noam here.
I thought this point was pretty clear: given a capable enough actor, sandbox guarantees become increasingly difficult.
I encourage people to watch the entire episode.
OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change
"But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI."
"So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar."
"You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient."
"There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors."
"One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate."
_________
Link and more key quotes from OpenAI's safety related conversations: https://t.co/uGBDtpmLBj
frontier model keep crushing public coding benchmarks with every release yet public sentiment seems to be that agents are frustrating and hard to use.
How can these models be so capable but still be so hard to use, What’s going on?
I find it almost disorienting to look at this.
The fastest-growing segment for frontier labs is coding, and it has absolutely destroyed the notion we built over 60+ years of what it means to be a SWE. It’s the biggest jump in abstraction since the compiler (i believe compiler went mainstream somewhere in the late 1950s or 60s)
Claude Code went GA only 16 months ago so this whole thing really only started 16 months ago. And coding is still 3x smaller than trading and 15x smaller than robotics.
imagine what will happen once robotics really takes off.
frontier labs in one way or another need to also become data labs. data labs are also frontier labs. as intelligence scale the methods through which we produce data to train models will become increasingly sophisticated
you might be the first person talking about the data over the architecture! 🥲
we consider ourselves a data research lab! the vast vast vast majority of research was on making data that is truly general (ala a cognitive core) and 100% of our data is synthetic (but not the type of crap that is just spit out from an LLM obviously)