Sarvam Epoch | Chapter 1
Join us for the first edition of Sarvam Epoch, our flagship AI conference, in Bengaluru.
Across the two days, we will share new model updates, launch new products, present our latest research, and host hands-on technical workshops.
Builder Edition, 30 July | Enterprise Edition, 31 July
Register here: https://t.co/WD4z1oovEC
We're thrilled to announce that we have raised $234M in the first close of our $300M Series B at a $1.5B valuation.
@HCLTech and @BessemerVP have joined us in this round, alongside continued support from @khoslaventures and @peakxvpartners
For countries and companies, sovereign control on the AI stack is no longer an optionality. Sarvam will be the partner of choice for this aspiration. The capital allows us to accelerate our momentum towards this full stack of models, compute, and deployments.
A huge thank you to our customers, partners, investors, and the Sarvam team for your trust and belief in what we are building. We’re just getting started.
Read more: https://t.co/VmLtpnj8gx
Speaking this Thursday evening in SFO on all things model research and future plans @SarvamAI. Will also give a peak into the scale and impact of our deployed products. Bottomline: It's a generational opportunity to build a consequential Indian deep tech firm...
RSVP to join - https://t.co/kvZD24Eguz
I will be presenting our recent work on writing efficient GDN kernels on B200 GPUs at MLSys 2026 today (11 am–1 pm PDT)!
FlashInfer ran a kernel competition for B200 GPUs. Our team (@thepushkarp + me) won 🥇 1st place on the Gated Delta Net track (more details here: https://t.co/QnVVWWE6cr)
Do join if you are around.
#MLSys2026
Speaking tomorrow at Stanford about the opportunity to build deep tech in India. If you want to train models, build products, create population scale impact, or are just curious what we are up to then RSVP and show up - https://t.co/TLlNBp8ojJ
🚀 Building https://t.co/e16WEwjWW8 (with Shiven) - an AI platform that turns educational content into interactive, gamified learning experiences.
Goal: move beyond static assignments toward AI-compatible, engaging, feedback-rich, and integrity-aware learning. Looking for educators, students, collaborators, schools, and funding/partnerships.
🌐 Try it: https://t.co/41ZtdUid4r
📄 Paper: https://t.co/VTlRenwZrn
🎥 Demo: https://t.co/EKvNOam3XW
Would love to connect and collaborate on responsible AI in education.
Speech-native models like Moshi sound great and answer fast, but aren’t as smart as text LLMs. In our new paper, MoshiRAG, we show how Moshi can ask for advice from a text LLM or a knowledge base. The tricky part is how to do this in real time without adding latency. 🧵
New work with @AlecRad and @DavidDuvenaud:
Have you ever dreamed of talking to someone from the past? Introducing talkie, a 13B model trained only on pre-1931 text.
Vintage models should help us to understand how LMs generalize (e.g., can we teach talkie to code?). Thread:
🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length.
🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.
🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice.
Try it now at https://t.co/GCdiMzk1Dl via Expert Mode / Instant Mode. API is updated & available today!
📄 Tech Report: https://t.co/drlDrxkYtp
🤗 Open Weights: https://t.co/T13Y8i7SDM
1/n
I will be presenting @SarvamAI Vision in the @github constellations 2026 event this Saturday in Bengaluru. We will also be talking about our LLMs and Audio models with @sumanthd17@ManavSinghal157 and Aditya M. See you at the event!
New Anthropic research: Emotion concepts and their function in a large language model.
All LLMs sometimes act like they have emotions. But why? We found internal representations of emotion concepts that can drive Claude’s behavior, sometimes in surprising ways.
Computer use is now in Claude Code.
Claude can open your apps, click through your UI, and test what it built, right from the CLI.
Now in research preview on Pro and Max plans.
1/10 🚀 Qwen3.5-Omni is here! Scaling up to a native omni-modal AGI.
Meet the next generation of Qwen, designed for native text, image, audio, and video understanding, with major advances in both intelligence and real-time interaction.
A standout feature:
Audio-Visual Vibe Coding: Describe your vision to the camera, and Qwen3.5-Omni instantly builds a functional website or game for you.
Highlights:
Script-Level Captioning: Generate detailed video scripts with timestamps, scene cuts & speaker mapping.
SOTA Performance: Qwen3.5-Omni has secured 215 SOTA scores across various sub-tasks, matching the top-tier text/vision capabilities of the Qwen3.5 series.
Audio-Visual Understanding: From auto-segmentation to fine-grained script generation, it understands the relationship between characters and their environment like never before.
Seamless Interaction: With native API support for Semantic Interruption, voice conversations feel human-like and background-noise resistant.
Global Multilingual Mastery: Pioneering support for 74 languages in speech recognition and 29 languages in expressive speech generation, breaking down global communication barriers.
Autonomous Intelligence: Native support for WebSearch and complex Function Calling—the model now independently decides when to pull real-time data.
Qwen3.5-Omni is built to be the backbone of next-gen AI applications, empowering developers and users alike with true multimodal reasoning.
@wataru9871 After trying on few audios:
Worked really well for https://t.co/esqDIVTT7I
Mixed the speakers + 1 stream was almost blank for : https://t.co/cRjVN25rTD
Now trying out more on other languages