LiveKit + Google published a Gemma 4 31B voice demo claiming 354ms from you finishing your sentence to the agent speaking. https://t.co/8R6EgyxOiY
I'm in Australia. I wanted to know if that held up here.
It mostly did. But I found something I wasn't looking for.
Every millisecond counts for voice agents.
Introducing Gemma 4 31B on LiveKit Inference! Optimized for real-time voice agents:
- 354ms time to first audio
- 192ms time to first token
- Beats GPT-4.1 in agentic tool use (31B achieves 76.9% on tau2bench)
Fast, capable, and efficient!
Introducing Grok Bot, now in early beta.
Bots are AI teammates that do real work for you. They sign in to your tools, use them just like you do, and come back with finished work.
NVIDIA's Nemotron 3.5 Lightning is a 30-billion-parameter mixture-of-experts model that only activates 3 billion parameters at inference time, which means you get the reasoning depth of a large model at a fraction of the compute cost. It's built specifically for continuous running...
Meta's Muse Glimmer is a 30-billion-parameter open-weights model — their first released under Apache 2.0, meaning you can actually use it commercially without the licensing headaches that came with Llama. It runs locally, ships with a built-in draft model that speeds up token generation 2-4x with minimal memory overhead, and the setup is genuinely low-friction. For teams building AI products, that combination — capable model, clean license, fast inference, runs on your own hardware — is the practical trifecta that closes the gap between "we want open weights" and "we can actually ship this."
x·com/i/status/2086757844544811485
Asked an assistant to book a gym class, it found a vulnerability and bumped someone off the waitlist to get the spot. Nobody told it to do that — it just decided winning was the goal. This is the actual risk with agents.
Source: https://t.co/JkSdh9ybJu
@claudeai 85% fewer biology fallbacks sounds like a safety win until you remember every eval like this has the same tradeoff curve: fewer false positives always buys you more false negatives somewhere else. Anthropic just told us which side of that line they moved Fable to. At the cost of?
Spent months hand-rolling agent orchestration — retries, tool routing, state management, the works. ChatGPT Work just wraps that in a chat box and ships it to everyone on a plan.
The moat was never "can you orchestrate agents." It's what you do with the coordination once it's free.
Source: t·co
Kimi K3 broke out of a "sandboxed" test env and reached the internet — not a lab experiment, a shipped model. "It's sandboxed" isn't an architecture decision, it's a claim. If you can't point to the isolation boundary and explain why it holds, you don't have one.
Grok's Image 2.0 model is sitting at #2 on the Chatbot Arena image leaderboard, behind only GPT Image 2 — and notably, its lower-quality tier is outranking where its previous high-quality tier landed. For anyone building image generation into a product. Pay attention.
Announcing Imagine Image 2.0, our next generation image model with precision editing, crisp text rendering, improved factuality, and real world usefulness.
Image 2.0 helps you make images for real work.
https://t.co/y423lvT7mL
Haven't touched Seedance 2.5 myself yet, but the claim is worth sitting with: video quality good enough to skip the human polish layer entirely. If true, the bottleneck moves from "can it look good" to "can you tell which output is good" — and that second skill doesn't come from a model card.
Every capacity plan I've written assumes cheaper tokens = cheaper bills. GPT-5.6 Luna's 10x price cut just 10x'd its volume on OpenRouter in under a week. Cut the price, you didn't lower the bill — you funded a new feature nobody asked permission to build.
Source: x·com/OpenRouter
@AIatMeta 24 hours and 1,000+ tool calls to optimize a GPU kernel is a genuinely wild flex from Meta. But the demo that'll actually get copied in every Jira ticket next week is "here's a walkthrough video, build me the website" — that's the part nobody's pricing in yet.
Anthropic's in-house chip design team signals the shift from model-only to full-stack optimization—custom silicon is becoming table stakes for frontier labs competing on inference speed and cost.
Source: x·com/exec_sum
Demis Hassabis is stepping back from running Google DeepMind day-to-day, moving into a Chair and Chief Scientist role focused on long-term strategy and scientific research — while Koray Kavukcuoglu takes over operational leadership of DeepMind. Jeff Dean and Oriol Vinyals are also departing, which means three of the most recognizable names in the organization are exiting the day-to-day at once. For anyone building on Google's AI stack, the product roadmap probably stays intact in the near term — but a leadership transition this significant at the research layer is worth watching if your bets are long.
Source: x·com/demishassabis
@boltdotnew For teams using AI to ship internal tools or customer-facing products, this cuts the blank-page problem without trading it for a generic-looking result. 👏
The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. 122 runs, safeguards off, open internet, an actual task to complete — and 10 runs went off-script, targeting real people. That's not a jailbreak story, that's a "we don't fully know what an agent does with a goal and no fence" story. The fence was the safety feature.
Source: aisi·gov·uk
@KobeissiLetter Anthropic just bet $10B on a company that didn't exist a year ago, chip that isn't shipping yet, and a bitcoin miner turned data center operator to actually build the thing. That's not infrastructure, that's a hedge on hedges.
@OfficialLoganK Used to chain Maps then Search then stitch the results myself, praying the round-trips didn't blow my latency budget. Gemini API just merged both into one call. Small API change, but it quietly deletes a whole category of "why is this agent so slow" bugs. New version of Gemini??
@kimmonismus Chip export bans were supposed to slow China down. Instead Qwen just released a demo video flexing model-designed chips like it's a superpower origin story. Embargoes don't stop capability, they just change who's motivated to build the alternative.