Meet Mistral Large 4, aka Le Chonk.
• 1T parameters, natively multimodal. 49B active.
It is the best open weights model from US or Europe on aggregated benchmarks.
• State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding.
• Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure.
• Available to all via API today. Working with cybersecurity partners privately.
Open weights release end of October.
Announcing Gemini 4 Argon, our new frontier model.
Argon is built to sustain deep reasoning across complex, long-horizon workflows and delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.
We’re also expanding the model’s output token limit to an industry-leading 1M tokens.
Argon is currently rolling out to a set of trusted cyber defenders in the Fairwind Program, with broader availability as soon as possible.
if you’re in b2b
PAY ATTENTION to what Rippling is doing with their outbound
i’d bet money these guys are booking at least 500-1,000+ demos/mo with these cold emails
We asked ten Claude Opus 5.5 agents to devise a faster shortest-path algorithm and prove it in Lean. Within 15 hours, they produced C-HD: a formally verified improvement over the published bounds.
Oxford researchers argue that LLMs can never invent anything.
It is mathematically impossible.
They published a paper called “Theory Is All You Need" and it argues against the claim that computational models can generate genuine novelty or new knowledge.
They analyzed the limits of generative ai, and the results are a brutal reality check for the idea that ai will replace human decision making under uncertainty.
Here is why AI is stuck and human cognition wins:
backward-looking vs forward-looking.. llms are probability machines that look backward at existing data. human cognition is forward-looking and capable of generating genuine novelty. human cognition operates theoretically "top-down" rather than "bottom-up" from data.
the "data-belief asymmetry".. the researchers use the invention of "heavier-than-air flight" to illustrate this concept. an ai relies on data-based prediction, which is largely imitative. humans, however, use theory-based causal logic that allows them to hold beliefs that go beyond existing data.
the intervention gap.. humans don't just process information; we use theory to practically "intervene" in the world. we engage in directed experimentation to generate entirely new data. ai-based models are theory-free and place primacy on existing data and prediction.
tldr?
AI uses a probability-based approach to knowledge and ia largely imitative. It can process data and make predictions, but human cognition relies on theory-based causal reasoning.
The decades-old analogy comparing human minds and computers to mere "input-output" devices is fundamentally flawed.
GPT-6 Astra is here.
We hope it will begin to enable a new generation of entrepreneurship, scientific discovery, and building.
We believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more.
It took us some extra time to ensure that we could meet the safety and alignment standards required for this capability level, but we think you’ll find it worth the wait.
It scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI 3, and 100% on ExploitBench.
congrats to the team - new paper from the epileptic seizure forecasting group i worked with
good reminder that eval/data design can matter more than model choice
going to start posting here again
spent the last few years doing a phd in ai and building things on the internet. now mostly deep in agents + ai systems
will share more of what i'm building / learning here
(1/7) New research: how can we understand how an AI model actually works? Our method, SPD, decomposes the *parameters* of neural networks, rather than their activations - akin to understanding a program by reverse-engineering the source code vs. inspecting runtime behavior.