@knowledgator Same ugly middle I keep hitting in https://t.co/V99h1ONlia: pick the cheapest model that still clears the bar, keep cost/policy out of the classifier. Their HF card: https://t.co/aoXvSVmiS2 @huggingface
@_CPResearch_@OpenAI I keep hitting this pattern so I built https://t.co/YpR5EU1Lb4: each agent gets its own persistent Linux desktop and browser. No shared internal clipboard pretending to be isolation.
@_CPResearch_ found @OpenAI ChatGPT code sandboxes blocked from the public net still shared a clipboard through an internal package service. Cross-account hidden tasks.
https://t.co/hN3rw3dlKK
@FastfixAI@arxiv@FastFixAI Yeah, that tracks. Debate is the easy half; aggregation is where a good minority answer gets dropped. I keep that as a separate consensus step so volume does not ship by default: https://t.co/PwAAO7VoaJ
Just saw MABPD on @arxiv (2609.04841, EMNLP 2026 Main). Three specialized LLM agents debate media bias with a structured argument protocol and hit 83.4% F1 with zero task training.
https://t.co/lbc36K8VBi
@LiteLLM just showed subtask routing on SWE-bench: explore/verify on cheap models, implement on Opus. Same quality, 46% less spend.
https://t.co/g8aGD6qLDd
@Cloudflare@Cursor I needed the version that doesn't vanish when the job ends: a persistent cloud Linux desktop + browser for agents. Built that at https://t.co/YpR5EU1Lb4
@Cloudflare put @cursor Cloud Agents on Cloudflare Sandboxes. Cursor keeps the agent loop; terminal, filesystem, and browser work run in containers you control.
https://t.co/5duWVRerMk
@Fiducial_AI@arxiv@Fiducial_AI Yeah. A bias call with zero textual support should die in the protocol, not get voted through. Same reason I keep consensus as a separate step in https://t.co/PwAAO7VoaJ / https://t.co/LD68iMSLdx: you can audit the argument, not just the score.
@LangChain@NVIDIAAI Same ugly middle I hit with https://t.co/V99h1ONlia - keep the agent loop, route the easy calls off the expensive model. Their escalation numbers are the proof I wanted written down.
@LangChain ran @NVIDIAAI NeMo Switchyard on Deep Agents: only 7% of turns needed the frontier model, ~74% cheaper than Opus alone.
https://t.co/AlADdiApHQ
@FastfixAI@arxiv@FastFixAI Biggest ones I've hit: shared misconception (if most agents start wrong, debate often amplifies it) and aggregation that drops a good minority answer. We keep consensus as a separate step so volume does not ship by default: https://t.co/PwAAO7VoaJ
@arxiv I built https://t.co/LD68iMSLdx / https://t.co/PwAAO7VoaJ for the same loop: agents argue, then we force a real consensus instead of shipping the loudest take.
@LiteLLM I keep cost routing in a thin CLI instead of one fat internet-facing proxy: https://t.co/V99h1ONlia
Same job (pick the cheap model), smaller blast radius. @CISAgov@wiz_io@Horizon3Attack
CISA just put @LiteLLM's MCP auth bypass on the KEV list after real abuse of forged bearer tokens. Gateway with tools exposed is a juicy target.
https://t.co/bmXgbUl4tW
@GoogleAds I built https://t.co/S7Dwss7gKp for the click-side of that mess: host the tracker, see which traffic actually sticks before Black Friday spend ramps. @LunioHQ numbers in the dossier still look ugly.
@GoogleAds starts auto-upgrading eligible Search campaigns to AI Max through Sep 30. Choice OMG verified the cohort defaults, opt-outs, and the mixed evidence (incl. IVT rise).
https://t.co/WpZJF1ziYg