AI can increasingly make progress on open problems like Navier-Stokes. Yet defining a new problem still relies on human research taste.
Introducing ScholarCatalyst, a benchmark built from AI researchers’ firsthand accounts of what inspired their work.
Thank you, IROS 2026.
This week marked a special moment for Dexmate. You witnessed the launch of our new website and new identity, and we got to share this new chapter with the global robotics community in Pittsburgh.
Being a California based company, it’s great to visit Pittsburgh and meet with builders and developer from all around the world. We love meeting with the community and talking about the endless possibilities we can build on Vega.
Thank you to everyone who stopped by and spent time with us. We’re excited for what we’ll build together next.
#IROS2026 #Dexmate #PhysicalAI #Robotics #Humanoid
This is DoorDash Air.
As a pilot, I can tell you that you wouldn’t build a plane and then ask if it can carry cargo. So we built @DoorDash Air from the ground up, designed around what 13 years of deliveries have taught us: what people order, how much it weighs, and how far it has to travel.
Our drones are as silent as a passing car and early tests show that we’ve gotten flight delivery times in under ~5 mins on average, so you can get your @chipotletweets & @popeyes (our first partners) and local favorites right to you, fast. The first deliveries will be in NorCal.
We proved autonomy works on the ground. Now see how it works in the sky.
Kalshi’s temperature markets beat the strongest individual public weather forecast in six of seven US cities, according to a new study.
Researcher Alexander Crosier compared forecasts drawn from market prices with public weather models across 7,590 city-days.
One hour into trading, the market’s forecast error was about 10% lower than the National Blend of Models. Its lead continued overnight into the day it was predicting.
Traders could already read those public forecasts. The finding suggests prices were adding information beyond what any single forecast offered.
That makes weather markets worth watching even without placing a bet.
The September 21 preprint covers data through August 12. In the chart, lower error means a better forecast: purple is Kalshi, green is the public model.
I compared every single AI personal assistant on the same task:
Grok Bot: 7 min 40 sec
Meta Muse: 4 min 36 sec
Instinct: 14 min +
Claude Cowork: 6 min 25 sec
Human (me) : 37 sec
Cerebras: 22 sec
yeah officer so he was using astra, and i told him he should've used fable, because fable writes more mergeable code, you know how gpt models love their tests? i also saw him on the terminal, we have guis now, officer have you heard about t3ode? it's the all in one interface for using your codex and claude subscription
NEVER thought i would hear this from jeff dean
he thinks we can compress chip design from 2 years to 3 months with RL + new EDA tooling
essentially just by a specialized auto research loop for hardware!
BREAKING: Terence Tao and 24 other Fields Medalists just signed a letter telling AI companies they're destroying mathematics
>headlines: AI solves famous math problems
>25 greatest mathematicians respond today
>"we are witnessing a general threat to intellectual work"
>ai labs are treating math problems like a benchmark you can brute force
>“solutions are announced in a rush, leaving no time for a proper writeup and citing relevant previous work of others"
>they're talking about openai
>"this raises severe attribution and plagiarism questions"
>they're definitely talking about openai
>openai offered to put math professor name on navier-stokes proof
>condition: abandon his coauthor bc he works at anthropic
>"the most precious resources of our profession are students and ideas"
>for the labs the most precious resource is GPUs
A Severe Misalignment of AI in Mathematics
I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why.
AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly.
Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet.
I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going.
Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards.
I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development.
I wrote up more thoughts here on my decision and what I hope changes: https://t.co/doX17mrHYq