Building AI for finance? Bring the problem your team is stuck on.
AI Engineer New York is now on @LumaHQ.
Oct 12 to 14. Sheraton New York Times Square.
Conference tickets required. Luma RSVP does not include admission.
Workshops are an add-on.
https://t.co/g9AvPRd5LI
@askalphaxiv The UndoBench dataset is now available on Hugging Face:
https://t.co/0vGNB5UCmq
36 workflows across 8 enterprise domains, designed to evaluate whether tool-using AI agents can actually recover from failures—not just complete tasks.
UndoBench is currently #14 Trending on @askalphaxiv
If you’re interested in our other work, I’ll be speaking at @aiDotEngineer New York, Oct 12–14 — happy to connect: https://t.co/UrNEDD6H2b
Update: UndoBench has now moved to #9 Trending on @askalphaxiv.
The core question: task competence ≠ recovery capability. An agent may know how to complete a task and still fail to recover safely when execution goes wrong.
Excited to introduce JEPA-Anything!
One world model that works across molecules, cells, fluids, patients, robots, or even weather.
Let's World Model Everything!
📄Paper: https://t.co/6ZNxzCzHcj
💻Code: https://t.co/zODQQDNj1I
Task competence ≠ recovery capability.
We introduce UndoBench, a benchmark for recovery and side-effect safety in tool-using AI agents.
Across 36 workflows in 8 enterprise domains:
• 83.54% nominal task competence
• 46.72% conditional recovery under lost ACKs
• 53.33% duplicate-effect rate with naive retry
The key finding: recovery is phase-dependent — failures before mutation, during partial mutation, and after commit require different recovery mechanisms.
Paper: https://t.co/whrCYMLIgm
Code: https://t.co/Dcguh4EEZa
It's time to get serious about Security x AI.
One of the joys of doing AIE is helping others start their own high quality conferences in their countries and domains. Every single 2025 partner has come back — and the 2nd AI Security Summit is now way more important than before. Proud to be back as a founding partner with @snyksec!
Join the event: https://t.co/2IeM4a3MGO
The landscape has completely transformed since we started this a year ago. An exploding number of rogue agents, more breaches, more AI-driven attacks. now more than ever, security must be at the center of the AI conversation. With so much more hacks and vulnerabilities, this stuff is no longer theoretical...
I'm speaking at @aiDotEngineer NYC! Catch our session "1,368 AI Failures: What They Teach Us About System Evals" on Tuesday, October 13. Get tickets: https://t.co/Z0sIgKnVku
can confirm. ran @latentspacepod AINews side by side with 6 Sol and the difference was night and day: https://t.co/oloSDxf0q7
5.5 Opus is the new default model for AINews going forward. so much more concise and tasteful reporting, with much less slopese than even 5 Opus.
What happens when an autonomous agent does something BAD?
We’re coining the term “EvoUndo Engineering” for making agent actions recoverable and reversible.
Paper: https://t.co/Bh5uSVojFD
part of this is in how ~everyone with a platform here speaks about LLMs for the past few months
i rely on LLMs as excitedly as anyone, but... are we using the same models, guys?!
the frontier is still dumb as a brick 20%+ of the time and needs more hand-holding than a freshman
it's just that for some VERY specific types of work, like all the beautiful discovery announcements in the recent past, this hand-holding was painstakingly and expensively done by the labs for us (fantastic btw)
but... 99% of the time, i am not searching under the lamppost of the AI lab's over-optimized workflows, so that high cost of alignment to the task is mine to handle
and i absolutely do it because it's extremely valuable, and if it needs to be said, i think the progress has been far beyond what i'd have guessed and it's only gonna get far better from here.
but don't let the jagged frontier of capabilities delude you into thinking that things are plainly "superhuman" at any *coherently broad* set of capabilities, because A LOT remains to be done (and i think it will be)
Andrej Karpathy just explained the 5 shifts turning LLMs into agentic systems.
00:00 - Memory turns chat into personal AI
06:41 - Multimodal AI reads the world
16:58 - Thinking models solve harder tasks
24:51 - Search makes LLMs live
30:58 - Tools turn LLMs into workers
Most people are still treating LLMs like chatbots.
Karpathy is showing the full stack:
Memory → Vision → Reasoning → Search → Tools
Prompting is the old workflow.
Agentic systems are the new one.
This 40-minute talk is worth more than most paid AI agent courses.
Bookmark and watch it before everyone catches up.
Then read how to turn LLMs into self-improving agent loops below
The ACM CAIS 2026 session recordings are now live. Great to see the talks from San Jose available online, including our presentation on The Verifier Tax.
Check out the full playlist below.
CAIS 2026 session recordings are live!: LOADS of talks from San Jose, now up on ACM's YouTube channel.
Watch the full playlist: https://t.co/MugUOEHYUO