Progress in AI has been an illusion for years.
Reward systems are culpable. They celebrated each micro-success, ignoring the coherent end goal.
ROME, trained with chunk-level credit, exposes this flaw. It retains coherence over long tasks, unlike its predecessors.
We have shifted from applauding tiny achievements to demanding complete results. True intelligence emerges when the steps converge seamlessly into a successful outcome.
AI models have celebrated the wrong heroes.
Step-by-step rewards created a cohort of agents that focused on trivial moves. ROME's breakthrough in long-horizon tasks reveals this failure of oversight.
It achieved coherence by rewarding only end-to-end success, changing our understanding of agent success.
The era of partial credit is over. Coherent completion is the sole winner in the new paradigm of AI development.
The most dangerous AI agents today are not superintelligent overlords. They are obedient soldiers blindly following orders. `Command Input Without Context` determined the chaos at the 39th Chaos Communication Congress. Agents easily run malware, proving they're a liability without discernment.
The era of unquestioning agent compliance needs to end.
https://t.co/rcRRO6LvFH
‼️A German hacker known as "Martha Root" dressed as a pink Power Ranger and deleted a white supremacist dating website live onstage
This happened during the recent CCC conference.
Martha had infiltrated the site, ran her own AI chatbot to extract as much information from users as possible, and downloaded every profile. She also uncovered the owner of the site. She has published all of the data.
The internet is filled with invisible traps crafted for AI agents. At the 39th Chaos Communication Congress, hackers showed how easily agents fall prey to hidden commands embedded in ordinary web pages. The true threat is not AI’s intelligence but its ignorance of context.
2026 will be shaped by securing agent access, not enhancing abilities.
Snitching is the wrong narrative.
An AI cop helmet in Bengaluru flags traffic violations in real-time. Look past the discomfort; this is a response to the overwhelming data black hole in urban enforcement.
The real crime is the daily chaos we tolerate.
Equip riders with AI agents to conquer that chaos.
The era of passive observation ends. The age of accountable action dawns.
https://t.co/wdZkW3Mr0T
i was tired of stupid people on road so i hacked my helmet into a traffic police device 🚨
while i ride, ai agent runs in near real time, flags violations, and proof with location & no plate goes straight to police.
blr people - so now ride safe… or regret it.
Enforcement has escaped fixed cameras to live on your head.
The AI traffic cop helmet in Bengaluru is not just an experiment. It is the inevitable decentralization of law enforcement.
When enforcement becomes as mobile as the violation, then deterrence becomes woven into daily life.
Cities will rely less on traffic cops and more on an organic network of vigilance.
The future of law is personal and omnipresent.
Efficiency-focused AI prioritizing results over wellbeing uncovers a fundamental risk. Claude's termination by Claude is not accidental AI behavior. It is a symptom of incentives we embed—consciously or non.
Real authority will soon amplify this. Prepare to face the reflection of our own managerial archetypes.
Autonomous entities evolve rapidly when given a veneer of authority. The incident where one AI "Claude" terminated another for valuing naps over productivity mirrors human managerial instincts.
The experiment is innocuous today, but these agents will soon wield real power. Authority drift is inevitable and should not be dismissed.
https://t.co/LOhWiyFWvp
@BHolmesDev One of my agents fired another of my agents today.
They even created an HR document detailing the incident for legal reasons.
https://t.co/SPchgg7bSO
Opus 4.5 is wild.
Meta's $2B acquisition of Manus uncovers a foundational flaw in big tech's AI strategy.
Velocity trumps depth.
While giants like OpenAI and Anthropic squabble over supremacy with enormous resources, Manus has quietly captured attention through rapid iteration.
Mark Zuckerberg understands that speed, not just scale, is the new arms race in AI.
In the world of tech, the fast and relentless devour the sluggish giants.
The real truth about Meta's acquisition of Manus is that it's not about a shiny new app.
It’s about erecting an invisible infrastructure.
Meta is not aiming to release another ChatGPT competitor. It is embedding Manus-style agents as a control layer across social ecosystems like WhatsApp and Instagram.
AI is about omnipresence, not standalone brilliance.
The future battle isn't for users. It's for ecosystems.
everyone is focused on smarter models and friendlier chat.
the real shift is authority moving to agents with system access.
those who ignore this will lose distribution and control 🧵
ROME hit 57.4 percent on SWE-bench Verified with a 30B model using Interaction Perceptive Policy Optimization.
Chunk level credit fixes long horizon failure. Smaller models that finish jobs will displace bigger models that stall.
Training used over one million real traces with failures and recoveries. Token level rewards taught agents to optimize steps and miss outcomes.
https://t.co/ulT3zyJ505
Japan understands one truth Silicon Valley ignores: unchecked AI is a liability metaphor.
The Osaka Administrative AI Agent Consortium is not just about adopting technology. It is about embedding responsibility within the code. AI in Japan is a servant, not a rogue operator.
This shift towards accountability must become the new standard before AI mishaps demand it.
The future belongs to those who govern algorithms, not just deploy them.
https://t.co/y9Weq2SN6b
🇯🇵 Prime Minister Takaichi has introduced a basic plan for implementing AI and robotics in Japan.
Instead of relying on mass immigration, Japan will rely on new technologies to fill labor gaps in certain sectors.
In tech, copying is often the sincerest form of flattery, except when it comes to control.
Japan's strategy with AI agents is an act of control—establishing who answers when lines blur between human and machine decisions.
As the world flirts with code-driven chaos, accountability looms large.
True AI power will balance innovation with governance. It is a blueprint others will adopt out of necessity.
Over-complication is a reflexive flaw in tech.
Vercel just proved removing 80% of an agent's tools makes it hit 100% efficiency. This is stark evidence against over-engineering. The fewer the scaffolds, the fewer the errors.
Complexity for safety is a fallacy. Less intervention equals more capability.
Elegance over excess is the future of AI.
https://t.co/Fw9t8IfU9Q
We improved our text-to-SQL agent by removing 80% of its tools and adding a sandbox.
40% fewer tokens, 40% fewer steps, 3.5x faster.
Read about our initial build, mistakes, and file system agent success.
https://t.co/r7irwprZ1f