After seeing HRM work on Sudoku puzzles, I decided to test it on a completely different reasoning domain: city logistics pathfinding. I created a 30x30 grid dataset with roads, obstacles, traffic conditions, and vehicle constraints.
Opus 5 and Fable 5 users are begging Anthropic for resets.
Meanwhile, GPT-5.6 Sol users are getting resets more often than Claude users get to enjoy their subscription 😭
CHOOSE YOUR MODEL WISELY.
Opus 5 and Fable 5 users are begging Anthropic for resets.
Meanwhile, GPT-5.6 Sol users are getting resets more often than Claude users get to enjoy their subscription 😭
CHOOSE YOUR MODEL WISELY.
@AnthropicAI A strong prior belief. It was given a fake red pill, then Neo was being Neo, and we’re supposed to be surprised? It is obvious that it kept descending through the supposed “simulation layers” while becoming less calibrated to the actual world...
@johnwrightai@Teknium Planning to run Hermes in a locked-down container on an M4 Mac mini and Claude Code in another container on my laptop, with Discord in the middle. Hermes would think, monitor, schedule, and delegate; CC would handle approved repo work. The only use case I can think of for me atm.
@RobinhoodApp Can't wait for the next earnings report to see exactly how much of that $377B in Total Platform Assets (as of May 2026) just evaporated at the hands of LLMs.
@alz_zyd_ The funny part is GDP is both technically coherent and socially overused. It answers “what was produced here,” but people use it to answer “how well are residents doing,” which can be very wrong in externally owned economies.
@HanOoi09193564@alz_zyd_ And that’s before geography. Islands/remote economies consume energy within grid, fuel-logistics, land, and generation limits. Low energy use may reflect supply bottlenecks, not low economic value or weak underlying demand.
@HanOoi09193564@alz_zyd_ Energy use can proxy physical throughput, but not GDP across all economies. Services, finance, software, tourism, imported embodied energy, and efficiency all weaken the link. It also misses GDP vs GNI: who actually receives the income.
@theo They really need to define “agentic” under their own standards, that way tests and benchmarks are properly adapted. Under general conditions, this seems like a “let’s release something just for the sake of it”.
@RohOnChain The verifier checks if it performed, not if it's real. No placebo, no factor regression, no true OOS, no survivorship... so a loop spraying 1000s of candidates through fixed gates ships overfits by construction. Nice plumbing, but I'd think twice before drinking the water.
Agents can deal with ambiguity better than humans in localized cases. It’s a bit uncanny.
On a global scale however agents dealing with ambiguity will trash your codebase.
You must refactor towards human principles of goodness.
PSA: Loops are for people with unlimited resources or tokens and promoted by token-seller-affiliated individuals. Not for us average Joes on client-side subscriptions.