@signulll I use it + Antigravity and have had more success with it than Claude Code. If you build a good harness Gemini is a lot stronger imo and a lot cheaper per task.
@cicatriz@julianboolean_@ylecun You are not recovering any of the wasted work $$$ going down a bad branch. You are not recovering any of your db records if they get wiped. While there does exists some probability r of an a token trigger an IRRECOVERABLE error branch, the generic p is still wasted compute.
@cicatriz@julianboolean_@ylecun As tasks become agentic and run longer, the risk of missteps increase with each additional token and the cost of a misstep increases. The advantage of models that understand fundamentals is that when something violates that you can detect earlier and prune early.
@cicatriz@julianboolean_@ylecun That is like asking how useful are safety measures on a nuclear reactor if they don’t trigger frequently? It’s long tail risk you want to avoid.
Furthermore, modern LLMs struggle on ARC-3 for the exact reason this model addresses. I’ve been in the space since 2015.
@julianboolean_@ylecun The problem is simple. You have a fixed probability P that any given “step” takes you OUTSIDE the tree of recoverable decision (in this case think of an LLM deleting your entire DB). The longer the task runs, you keep rolling this dice but as some point it will make a mistake.
I did research on @deepseek_ai recent papers and the implications for the R2 model later this month. I think Deepseek has built a ~10x more efficient model and they will offer a better AI than OpenAI at a similar price or a similar one at a cheaper price.
https://t.co/a2e9mG8sAY
@jhuntermav@deepseek_ai Now, [Deepseek] is accelerating the launch of the successor to January's R1 model, according to three people familiar with the company.
Deepseek had planned to release R2 in early May but now wants it out as early as possible, two of them said, without providing specifics.