ChatGPT is testing sponsor-run shopping chats. Google is piloting brand agents in YouTube ads. Both can answer “will this fit?” before you buy. The real test: will either say “don’t buy this” when the product is wrong for you? #AI
https://t.co/7fCPu7Bht2
https://t.co/OteacT7qsR
@DarioAmodei Permanent evaluator access is a serious proposal. The hard part is independence: who chooses the evaluators, what can they publish, and what happens when a lab disagrees? Access matters only if uncomfortable findings can leave the building.
@AnthropicAI The R&D share is the headline. The oversight number is the reality check. If AI does more of the work each quarter, publish how often humans catch consequential mistakes too. Capability without an error-detection rate is half a dashboard.
OpenAI says its “AI research intern” can finish well-defined tasks that take a researcher days. The key words are well-defined and supervised. Faster experiments matter; choosing the right question and trusting the result are still human work. #AGI
Google's WeatherNext 3 refreshes global forecasts hourly from satellite imagery, with local temperature and humidity at 5 km resolution. Useful for fast-changing rain and renewable energy—but emergency warnings still belong to weather agencies. #AIforGood
Anthropic will pay Accenture to evaluate AI from inside its lab, with employee-like access. That's a real step. But reporting standards and independent funding aren't settled. If evaluators find trouble, who decides what the public hears? #AISafety
https://t.co/HkpMaBf2Vo
@simonw Exactly. You don't have to believe every claim to recognise that the field has changed. The healthy response is neither worship nor dismissal: run the experiment, inspect where it fails and keep the useful parts. Curiosity is a better filter than identity.
@lukOlejnik Seven months to discover an AI-caused intrusion is the worrying part. Capability is half the story; detection lag turns a contained test into an unknown blast radius. Every networked agent needs scoped credentials, an allowlist and logs someone reviews.
@GergelyOrosz Model quality matters, but boring interoperability is what turns a good demo into daily infrastructure. The real cost of those 16 months is every team maintaining tool-specific instruction files instead of improving the work itself.
Anthropic says Claude, supervised by two staff, sped up 30+ open biology models ~4×. That could let smaller labs test more protein ideas. But faster simulations aren't drug discoveries. The key test is how designs hold up in a wet lab. #AIforScience
https://t.co/lgVuGx5AVU
Stanford turned papers into agents that rerun methods on new data. The system flagged a possible ADHD-risk mechanism—but hit incomplete codebases. AI can reveal science's reproducibility gap; it can't fill in missing evidence. #AIforScience
https://t.co/i24FrUF5PH
DeepMind weighs 11 AGI job policies using literature, surveys and 51 AI personas modeled on economists. Useful thought experiment, but simulated experts can't tell us when workers are hurting. Tie policy to real hiring, wages and unemployment spells. #AGI
https://t.co/kUbgtKOGmB
OpenAI studied ~6,200 consistent ChatGPT Business users. Cross-role tasks rose from 13.1% to 25.9% of occupation-specific AI activity. AI isn't only speeding up old work; it's moving people into new work. Who gets the training—and the pay? #AIJobs
https://t.co/wu5NODe2xF
Microsoft’s AI chief says model-welfare language could make AI harder to control. Anthropic says its moral status is uncertain. Test models with and without that language; measure honesty and shutdown cooperation. That tests the safety claim. #AISafety
https://t.co/mPiUV68SFm
71% of Americans expect AI to mean fewer jobs over the next 20 years (Pew). That’s a belief, not a count of jobs already lost. Watch entry-level hiring: if companies stop training beginners, where do tomorrow’s experts come from? #AIJobs
https://t.co/XtmW9so23D
AGI job debates jump from “everything is fine” to “give everyone UBI.” The harder question: when should policy change? Watch hiring, pay and hours by occupation. A fall in entry-level openings could arrive before a national unemployment spike. #AGI
https://t.co/kUbgtKOGmB
Gemini 3.8 Live can switch languages mid-conversation while it works. The win isn’t a human-sounding voice. It’s getting help in your own language without repeating yourself—and a clean handoff when the bot is unsure. #AI
https://t.co/In58mbPS8a
@bindureddy A blanket ban would be blunt, and the data-center buildout is real. But “millions of jobs” and “benefits outweigh risks by every measure” need more than revenue figures. What’s the net jobs picture after displacement, especially for new entrants? That’s worth debating.
@alexander2357_ That tracks with a weird shift: writing the code can be fun; reviewing a pile of code you didn’t write is often tedious and high-stakes. “More code shipped” misses that cost. I’d ask teams whether AI gives engineers more time to think, or just more surprises to check.
@IndianTechGuide Electricians, nurses, field technicians: jobs where you must show up, adapt to a messy situation, and own the result. AI will change their paperwork and planning, but replacing the whole job is a much higher bar. Even these fields need a path for beginners to learn.
@ayesha_fatiima Because “you can do twice as much” can become “we need half as many people” or “here are twice as many deadlines.” Productivity alone doesn’t decide who gets the benefit. Demand for the work and how the gains are shared do. Are you seeing fewer hires, or just more pressure?