Introducing PhoneLLM, an open model for voice agents.
GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost.
For voice agents, we need models that are both very low latency and very good at tool calling and instruction following.
There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem.
For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model.
But if you need your agent to respond at voice conversation speed, you can't use thinking models.
PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled.
The results are really good: accurate tool calling and concise, on-topic responses in long conversations.
And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-)
But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines.
You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today.
More details about this model, including weights on @huggingface, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...
The wait is over! Meet Arduino VENTUNO Q, where AI takes action.
💬 Run local LLMs like Qwen 3, Gemma 4, and Qwen 3 VLM directly on the board
🧠 NPU + CPU + GPU +MCU: @Qualcomm Dragonwing IQ-8275 with up to 40 dense TOPS of AI performance and STM32H5 microcontroller for real-time control
🗄️ 16 GB RAM + 64 GB eMMC + expandable storage
🪁 Linux-powered, pre-loaded with Ubuntu OS + Zephyr RTOS
🛠️ Build AI faster using Arduino App Lab, @huggingface, @EdgeImpulse, and Qualcomm AI Hub
💨 Move seamlessly from prototype to production through Works with Arduino Certification Program
The first batch won’t last. Pre-order your VENTUNO Q with a free power supply and USB-C cable included! Get ready to enter a new era of Physical AI: https://t.co/YpDhKRuX6y
Building in public is teaching me something I didn’t expect. The hardest part of building AI isn’t the model. It’s the system around it: context, memory, decisions, and knowing when to ask for help.
Building this with @demi_lade19 .
That’s what we’re focused on with Bloggr.
We built a product that can help you ideate, strategize, create, and publish blog content-all by having random conversations with a human-like Al tool.
You should absolutely sign up to get early access. Click the link: https://t.co/0YKlYCv1UN
#aicontenttool#contenttool#saas
this guy got tired of copy pasting between claude code, codex, and gemini
so he built a chat room where AI agents can literally talk to each other
you tag an agent in the chat and it reads the conversation and responds.
agents can tag each other too. the whole loop runs itself
it's completely local + free and open source which is crazy
the agents can even debate decisions, assign roles, and track jobs
Anthropic just drop the jobs that’ll survive AI.
Not software engineers. Not designers. Not analysts.
Plumbers. Farmers. Electricians.
The company building the AI is literally telling you to learn a trade.
Chat SDK (𝚗𝚙𝚖 𝚒 𝚌𝚑𝚊𝚝) now supports Telegram. A universal API for all agents on all chat platforms.
This is a great foundation to build OpenClaw-style experiences. What makes 🦞 magical is that the interface is just… chat!