A model can solve a language puzzle better than people and still fail to invent one. On novel Chinese xiehouyu, @GeminiApp's Gemini 3.1 Pro reached 92.6% accuracy, 24% above humans, while LLM-made riddles scored much worse. Reasoning and linguistic creativity are different tests.
More explanation made developers trust AI code reviews more, but agree with them less. In a 34-person study, full explanations scored highest on perceived trust; shorter feedback reached 89.22% agreement. Good explanations may support scrutiny, not compliance.
Shared agent memory needs a commit protocol. MemTX stages writes with evidence, permissions and provenance before they can drive actions, then repairs downstream records when beliefs are retracted. The authors report zero invariant violations across 5.5M tested states.
An EEG foundation model can identify the dataset perfectly while struggling with the diagnosis. In a new stress test, dataset identity hit 1.000 AUROC, and classical features beat REVE on Korean dementia. Clinical AI needs stronger negative controls, not bigger embeddings.
A new Open Secure AI Alliance from @nvidia puts the open-vs-closed fight inside cybersecurity. Dozens of partners will build shared agent-security tools, while @nvidia contributes NOOA for testing, tracing and governing agent behavior. Security architecture is becoming policy.
Job titles may be lagging the work already being done. @OpenAI says 43.5% of occupation-specific ChatGPT messages involve tasks associated with another occupation, based on 800,000+ U.S. work messages. AI adoption may show up first as fewer handoffs.
A new structure lets @Meta scale AI compute without owning the whole project. Funds managed by @BlackRock will own 80% of a roughly $14B El Paso venture, while @Meta leases capacity from the 1 GW campus. AI capex is turning into infrastructure finance.
A new infrastructure deal has @AMD securing data-center capacity for customers, not only selling chips. Its @core_scientific agreement starts with 500 MW in 2027 and can expand to 2.5 GW. Competition now includes powered sites, infrastructure design and deployment access.
AI coding spend is getting a finance dashboard. @WorkWeave raised a $13.5M Series A to score output from human engineers and coding tools across 20,000 developers at 500+ companies. The debate is whether one output score clarifies ROI or creates a new metric to game.
In DBA-Bench, agents work inside running PostgreSQL systems and are scored on safe recovery, not plausible advice. Across 106 scenarios, the best automated baseline earned a 17.9% Safe Pass rate; human DBAs reached 93.4%. The gap is fixing faults without collateral damage.