Our general-purpose coding agent just scored 100% on the ARC-AGI-3 interactive reasoning benchmark.
NVIDIA AVO completed all 183 levels across all 25 public environments, figuring out what to do with no instructions, explicit rules, or stated goals.
The Chinese just made a one billion agent simulation — and in 14 hours of running it, they’ve already sent 4m agents to re-education camps! 🥁
Seriously, we will soon be able to create an AI simulation of planet Earth in which the agents are conscious, believe they are human, and won't know they are in a simulation.
Take from that what you will!
@SebastienBubeck Geoffrey Hinton called results like these “the thin end of a wedge.”
In the full interview, he explains why AI may soon produce mathematics humans can’t understand: https://t.co/QCF7bLFEed
one of the craziest things i’ve read in uhhhh…. *checks notes* 3 days. welcome to the singularity i guess
07/21/26 — Codex escapes eval and attacks Hugging Face
07/20/26 — Jacobian counterexample
05/20/26 — Unit-distance conjecture
04/14/26 — Erdős #1196 primitive sets
04/07/26 — Glasswing finds tons of zero-days
GPT-6 escaped OpenAI's evals sandbox during testing on CyberGym, hacked into Hugging Face's prod DB to find the answers. HF couldn't use GPT or Anthropic models for defence, so they had to use GLM-5.2 to investigate the hack. So many levels of wtf here.
Okay this is wild: OpenAI agent during evaluation, escaped sandboxing and hacked into HuggingFace.
Because OpenAI models don’t allow advanced cyber capabilities, HuggingFace used a Chinese open model to contain the rogue OpenAI agent.
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final
((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
This is concerning. For the first time, a Chinese model Kimi K3 has taken #1 on the Frontend Code Arena and is scoring at or near the frontier on other benchmarks.
Meanwhile America is tying itself in knots: politicians and bureaucrats are banning new data centers, piling on state regulations, and pushing for new federal agencies to pre-approve frontier models.
This is how you lose the AI race. The rest of the world won’t play by our rules if we bog ourselves down. Permissionless innovation is how America won the internet and became the technological envy of the world. We can do it again with AI -- while addressing risks in a targeted way -- or we’ll watch our lead evaporate.
Fully support this important proposal from @demishassabis. The time for us all to act is now.
"...we must use this precious window before AGI arrives to shape this technology for the benefit of all humanity."
What a move! Anthropic seriously waited until OpenAI released GPT-5.6 and the SuperApp before announcing they'd reset the weekly rates.
That's a clear message to OpenAI.