Train in sim. Run it for real. Teach it new tricks. 🦆
This is the sim2real loop that powers Microduck.
RL stack open source here : https://t.co/mveM7y5fYa
Buy it here : https://t.co/Rmk8F4vpKd
Git : https://t.co/kcoCKdBi52
Join our community : https://t.co/TcX5uaf0yL
Introducing PhoneLLM, an open model for voice agents.
GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost.
For voice agents, we need models that are both very low latency and very good at tool calling and instruction following.
There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem.
For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model.
But if you need your agent to respond at voice conversation speed, you can't use thinking models.
PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled.
The results are really good: accurate tool calling and concise, on-topic responses in long conversations.
And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-)
But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines.
You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today.
More details about this model, including weights on @huggingface, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...
It's me again. I come bearing great news.
First of all, we have hit 20M active users for Codex some time this week. Second of all, this is cause for celebration and during the day we will credit every Codex and ChatGPT Work user with a BANKED reset that you can use at your own leisure. And we will have some other good news later too!
Now, on usage limits draining faster, while we're not seeing anything abnormal, we do take it incredibly seriously and there is an ongoing investigation. I will share if we do find anything and my below post is really a clarification on a specific pattern that we did see that I wanted to call out.
Go do something amazing today.
Claude Code can design now. The new /design skill (research preview) brings Claude Design's artboard workflow into the CLI and Desktop, built on artifacts.
Run /design to get editable artboards for your UI — pick one, tweak it, then have Claude implement it.
Okay, it's done
Cancelling Claude will probably be bad for my engagement as Claude complaint posts always do well, but it's the right thing to do
Goodbye, Claude, it was fun while it lasted
What you upload is what you get. Now in 3D.
One reference, one take. Watch the proportions, the part placement, and the surface detail against the input.
That's Meshy 7. 👩🎨
SL2T is our breakthrough sign language-to-text model powering new features for Deaf and hard of hearing users on @Android.
Starting with American Sign Language-to-English on Pixel 11, people can sign directly into Gboard and Live Transcribe instead of typing.