One of the AI trends I'm following is recurrent reasoning. The main idea is to let the model do the reasoning in its own continuous latent space and avoid the inefficiencies of CoT reasoning, which forces the model to reason one token at a time and in a discrete space.
We're seeing different forms of it (perhaps the most popular being the looped transformer that got a lot of attention after the announcement of GPT-6 Astra).
BDH-CQ by @pathway_com combines the benefits of latent reasoning with in-context learning, which was previously only seen in autoregressive transformer models.
The result is an AI model that can reason on novel tasks based on the given examples while being very compute efficient. The model set a new Pareto frontier on ARG-AGI-1, reaching 29.5% at $0.0007 per task with 150M parameters. And the researchers behind the model tell me that the prospects for scaling the model and working on other modalities are very positive.
Worth following.
UndoBench is currently #14 Trending on @askalphaxiv
If you’re interested in our other work, I’ll be speaking at @aiDotEngineer New York, Oct 12–14 — happy to connect: https://t.co/UrNEDD6H2b
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
https://t.co/7N6TPlft1P
Both American and Chinese open model devs liking this tweet hello I see you and thanks for agreeing with my passive observation - healthy competition is good :) More positivity in the open community pls
so uh bros ?
n² → n^1.9992 and n³ → n^2.9995 sounds like nothing.
It's the difference between "nobody has found a way" and "there is no way." Twenty years of work falls with this
pretty cool work.
last week we published our new work (animation bench) where we showed how frontier models perform poorly on motion reconstruction. seems like the thesis is just getting validated across niche domains.
https://t.co/ceSYxXhXQC
This is exactly the right direction.
Auto-review being free is a huge win. Let agents actually work for hours without constantly asking me for permission, while a second agent quietly catches the genuinely risky stuff.
More autonomy, less babysitting, without throwing safety out the window.
This is how agents become actually useful.
How should you approach the presence of AI slop in training data?
Abdullahi Dattijo offers actionable insights based on extensive testing he's conducted, leveraging three different approaches. https://t.co/R0do2DX6sp
Congratulations to @reflection_ai
decent start , but nowhere near the hype. gets mogged by most of its competition, half of the benchmark is NR, but a good start.
🚨 AI can now legally write prescriptions in the US. No human doctor involved.
Nolla Health says it's the first organization in the US (and possibly the world) to get regulatory approval for an AI to issue initial prescriptions.
The first use case is acne, in Utah.
You should be able to use smarter models in your agent.
You should be able to use cheaper models in your agent.
You should be able to use local models in your agent.
You should be able to use a different model for every job.
You should be able to switch models in the middle of a conversation.
You should be able to choose what data your agent has.
You should be able to choose what it remembers and what it forgets.
You should be able to choose what computer your agent runs on.
You should be able to run it on a machine that never touches the internet.
You should be able to run it while your laptop is closed.
You should be able to reach it from the apps you already use.
You should be able to choose how your agent thinks.
You should be able to read the prompt it runs on.
You should be able to see every tool it has, and turn off the ones you don't want.
You should be able to teach it something once and never explain it again.
You should be able to know what it did, and when, and why.
You should be able to export your agent.
You should be able to read the code.
You should be able to change the code.
Your agent should be yours.
🚨 First Week of October news : fable edition
we got gemini 4 argon , fable 5.5 routing , codex updates , and many more watch this news video
brain rot edition fully coded in JS there fights will give you news
Games development already had its “AI transition”. Casey Muratori @cmuratori, programmer and performance nerd:
“The licensable engine thing kind of was our AI transition already, unfortunately. I regret to inform you that the news is probably not that positive.
There are some definite positive things that happen early on because it opens up the ability to make games to people who could not have marshaled the technical staff necessary to produce a competitive engine. You get some really cool games coming from some sources that just simply wouldn't have been able to do it. Thumbs up.
The problem is that it rapidly accelerates into this nasty scenario where you have massive numbers of releases. We're at the point where, I want to say, Steam games are in the tens of thousands per year. It's so massive that there is no way that your game will be organically noticed anymore, period.”
Reflection’s first open-weight model is reportedly close. Nvidia-backed, aiming to compete with the top Chinese open models.
The US could use a few more serious open-weight models and a few fewer $200 subscriptions lol
Huge release from EverMind AI.
Raven is an open-source multi-agent system that builds a separate harness for each model and domain, evolves those harnesses from their failures, and coordinates them through a host agent.
The host agent splits a goal into a task graph and sends each part to a model-harness pair for research, code, design, or on-call work.
Harness components are diagnosed, mutated, and recombined, and a candidate replaces the current harness only after passing a statistical check.
With DeepSeek-V4-Flash, the research harness scores 69.3% on BrowseComp and 60.0% on Humanity's Last Exam, against at most 62.4% and 43.3% for two other harnesses on the same model.
Its curated skill library raises SkillsBench Pass@1 from 9.2% with no skills to 22.6%.
Paper: https://t.co/FGWHgSug42
Chat with Paper: https://t.co/BsNe1SGsPr
With the arrival of Jev, practitioners can now choose between generation-focused models and decision-making-based ones.
@nhu_hoang offers a thoroughly tested comparison of both approaches in her excellent technical deep dive. https://t.co/CIiLK5FgRr