Voice AI has gotten remarkably good, but natural conversation remains a high bar. Small delays, awkward interruptions, or the wrong tone can quickly break the illusion, and adding vision and visual presence only raises the stakes.
In this episode, @smolix, co-founder and CEO of @boson_ai, explores the path from today’s voice agents to audiovisual agents and AI avatars. We discuss the technical tradeoffs behind real-time voice, including audio tokenization, latency, model size, and inference cost, as well as what changes when these systems can both see and be seen.
We also explore the role of emotional intelligence in AI, how agents can learn from human interactions, and what it will take to move beyond impressive demos toward interactions that actually feel natural.
🗒️ Full show notes: https://t.co/MhtLzKDnzr.
📖 CHAPTERS
===============================
00:00 - Introduction
01:36 - Challenges of Voice AI in Real-World Environments
06:12 - Why Voice Systems Are a Whole-System Challenge
12:50 - Trade-Off Between Size, Compute, and Affordability
13:53 - Boson AI’s Path to Audio
19:59 - Training Dataset
23:49 - Dealing with Noisy Data
28:41 - Model Training Approach
33:14 - Architectural Decisions: End-to-End vs. Multi-Threaded Systems
39:26 - MCP, Context Windows, and Inference Costs
41:04 - Agentic Systems for Voice
46:31 - Benchmarking
50:03 - Why EQ Matters for Voice and Video
53:43 - Recursive Self-Improvement and Personalization
01:00:03 - Brilliance, Interaction Improvement, and Wrap-Up
Only 1 hour left before our Generative AI study group starts! ⏳ Don’t miss out on insightful and thought-provoking conversations about AI agents, multimodality, test-time scaling, and more! See you in an hour!
#llms#generativeai#studygroup#TWIMLMeetup#TWIML#machinelearning #ml #ai #artificialintelligence
Head over to https://t.co/TDhwjxmtZ6 to register for tomorrow's session of our Generative AI study group. Explore image-text encoders, exchange latent reasoning approaches, dig into self-verification in reinforcement learning, and more! Meetup is at 8 am PT.
Be part of our TWIML community and join our Generative AI study group every Friday at 8 am PT. Dig into visual grounding techniques, explore Mixture-of-Modality models, discuss agent governance, and more! Feel free to invite a friend and register at https://t.co/4MmQy8rjg6. See you there!
1 hour left before this week’s meetup of our Generative AI study group starts! If you haven’t registered yet, head over now to https://t.co/TDhwjxmtZ6 and don’t forget to join the #generative-ai Slack channel. See you there!
Visit https://t.co/TDhwjxmtZ6 to register for tomorrow's Generative AI study group at 8 am PT. Discuss distributed inference pipelines, exchange ideas on reasoning-time scaling techniques, explore multimodal transformer models, and more.
Don’t miss out and hop in an hour to meet us in our Generative AI study group! Ask questions, share your insights, learn alongside fellow AI practitioners, and more! See you there!
Join us for another engaging conversation around open-source distributed inference frameworks, LVLM spatial understanding, AI observability for agentic systems, and more in our Generative AI study group! Visit https://t.co/9fyBVNIm8y to register and we’ll see you tomorrow at 8 am PT.
Join our weekly Generative AI study group every Friday at 8 am PT! Dig into geospatial reasoning, share strategies for patch-text alignment, discuss KV cache routing, and more! For registration, head over to https://t.co/TDhwjxmtZ6 to sign up.
As reasoning models consume more tokens and AI systems become more expensive to run, understanding what those tokens actually buy is becoming increasingly important. In this episode, @Stanford professor and Big Spin co-founder @ChrisGPotts joins us to discuss AI tokenomics and his research into “tokenflation”—the possibility that token usage is growing faster than the measurable value those tokens produce.
We explore how to measure the return on AI spending, why benchmarks alone provide an incomplete picture of model progress, and what inference-time scaling means for the economics of increasingly capable models. Chris also explains why expert AI users tend to get better results by challenging and iterating with models, how AI fluency affects outcomes, and why more efficient architectures could change the underlying economics. We also discuss DSPy, interpretability, the limits of today’s transformer architectures, and where Chris sees opportunities for more fundamental innovation in AI.
🗒️ Full show notes: https://t.co/SuX2rzHbc5.
📖 CHAPTERS
===============================
00:00 - Introduction
05:38 - Linguistics in the Age of Language Models
09:26 - Scale Limitations in NLP Research
12:54 - Challenging the Bitter Lesson Mindset
15:12 - Relationship Between Data, Mechanistic Interpretability, and Efficiency
17:03 - DSPy
21:32 - Prompt Optimization and Model Variability
24:35 - Tokenomics and the Rising Cost of AI
28:13 - Measuring Token Purchasing Power with a CPI
32:18 - Inference-Time Scaling
35:36 - Defining Value Across Different AI Tasks
38:33 - AI Value Creation
40:38 - Predicting AI Costs
42:25 - Tokenflation
46:22 - AI Fluency
50:12 - Key Lessons of AI Fluency Work
54:44 - Future Directions
Get ready to join us in 1 hour for our Generative AI study group! ⏳ Engage in the AI news discussion, bring your questions, exchange ideas, share experiences, and learn together! Meet you there!
Be part of tomorrow's discussion in our Generative AI study group and dig into long-horizon autonomous agents, explore prompt caching techniques, discuss sparse grouped-query attention (GQA), and more! The meetup starts at 8 am PT. Register today at https://t.co/9fyBVNIm8y and see you there!
Our Generative AI study group is open to everyone! Just register at https://t.co/4MmQy8rjg6 and join us every Friday at 8 am PT. Share RAG grounding techniques, discuss long-context transformer models, explore Cache-Augmented Generation (CAG), and more! We’re looking forward to you joining us!
In this episode, @jcjohnss, co-founder of @theworldlabs, joins us to discuss world models and the emerging field of spatial AI. We explore why many researchers see capabilities beyond language as an important frontier for AI, and what it means to build models that can understand, generate, and simulate the environments around them.
Justin explains the different approaches to world modeling, including explicit 3D representations and generative models, and why there is still no established recipe for building these systems. We also discuss World Labs’ Marble system, which can generate navigable 3D worlds from images and other inputs, the challenges of evaluating world models, and the role of simulation, planning, and action. Finally, Justin shares his vision for models that bring these capabilities together, supporting everything from interactive virtual environments to agents and robots that can operate in the physical world.
🗒️ Full show notes: https://t.co/7knA87Ue7g.
📖 CHAPTERS
===============================
00:00 - Introduction
03:19 - Defining World Models
08:10 - World Models as Theory Builders
11:37 - POMDPs and Agent–World Interaction
15:40 - Training Agents with Behavior Cloning
18:40 - Ground-Truth State and Learned State
23:34 - Explicit 3D vs. Implicit 3D
28:14 - Reconstruction vs. Generative World Modeling
30:38 - Gaussian Splat Anatomy and File Formats
33:52 - Why Gaussian Splats Work with Neural Networks
37:00 - Consistency by Construction and at Scale
40:31 - How Marble Generates 3D Worlds
44:25 - Training Data and Output Representations
47:48 - World Models as Renderers, Planners, and Simulators
51:30 - When Rendering and Simulation Overlap
56:28 - The Path to Unified World Models
59:30 - Architectures, Loss Functions, and Long Contexts
01:02:47 - Where to Learn More
Our Generative AI study group starts in 1 hour! Share noteworthy AI news and announcements, show and tell your personal projects, exchange valuable tips and tricks, and more! We’ll see you in a bit!
Secure your spot for tomorrow's session of our Generative AI study group by registering at https://t.co/TDhwjxmtZ6. Discuss vision-language agent workflows, share effective pruning techniques, explore graph neural networks, and more! We’ll meet you there tomorrow at 8 am PT.
Be part of our weekly Generative AI study group every Friday at 8 am PT. Exchange model merging techniques, discuss TPU orchestration, dig into interpretability frameworks, and more! Make sure to register at https://t.co/9fyBVNIm8y and join the #generative-ai Slack channel. See you there!
The conventional wisdom in AI is that the next breakthrough will come from more compute, more data, and larger models. But what if the next leap comes from somewhere else?
In this episode, @wellingmax—co-founder and CTO of @cusp_ai and professor at the University of Amsterdam—argues that physics may provide some of the ideas behind the next generation of AI systems.
We begin with CuspAI’s work using generative AI to design entirely new materials for semiconductors, batteries, carbon capture, and clean energy. Max explains how foundation models for chemistry, agentic workflows, simulation, and automated experimentation are dramatically accelerating the search for new materials and reshaping scientific discovery.
The conversation then broadens into a deeper question. Beyond giving AI new scientific problems to solve, can physics also teach us how to build better AI? Max explores surprising connections between machine learning and thermodynamics, why waves may become a new computational primitive for neural networks, and how concepts like symmetry breaking and statistical physics could inspire AI architectures beyond today’s scaling paradigm.
🗒️ Full show notes: https://t.co/LpjD9xIqBN.
📖 CHAPTERS
===============================
00:00 - Introduction
02:08 - From Equivariant Networks to AI for Science
04:55 - Founding CuspAI
07:30 - Using AI to Discover New Materials
09:03 - Designing Materials for Carbon Capture
13:20 - Partnerships and the CuspAI Business Model
14:59 - The End-to-End Materials Discovery Process
19:13 - Experiments and Self-Driving Labs
23:04 - Open-Source Molecular Dynamics on GPUs
25:23 - Foundation Models for Chemistry
30:20 - The Future of AI-Driven Materials Science
32:22 - Connecting Generative AI and Thermodynamics
39:29 - How Physics and Machine Learning Can Inform Each Other
44:53 - Waves as a New Primitive for Neural Networks
49:34 - Memory, Stability, and the Edge of Chaos
52:40 - Spontaneous Symmetry Breaking in Neural Networks
56:10 - The Two-Way Exchange Between AI and Physics
We’re just 1 hour away from our Generative AI study group! Share insights on single-vector retrieval methods, explore multimodal model alignment, discuss knowledge graphs, and more! See you there!