Applied AI researcher + 2x founder. Built @PortoAIHq. Ex @samsungresearch & @BoschGlobal. 15+ yrs from AI research to production. Building reliable agents.
In 2026 the industry has a new favourite phrase. Voice AI!!
But forgetting the real innovators .
Samsung Research in India was building this technology for the world. 🫡
Every week a new demo, a new fundraise, a new manifesto about how voice AI changes everything.
It does change everything. It also already happened once.
In November 2018, at the Samsung Developer Conference in San Francisco, the Bixby team opened the Capsule platform and Bixby Developer Studio to third-party developers. A capsule defined typed concepts and actions. The Bixby planner took a user utterance, derived an intent, dynamically generated an execution graph over the capsule's models, and executed it against real backends with transactional semantics.
Replace the word capsule with tool. That is MCP. That is the voice AI stack of 2026.
The Vision component, shipped on day one of Bixby in April 2017, did OCR, live translation, object recognition, and visual search as first-class inputs to the assistant. Multimodal voice assistants in 2026 are catching up to what was running on a Galaxy S8.
A significant portion of the pipeline ran on-device. ASR, wakeword, intent classification, input resolution. Years before edge AI became a fundraising slide.
The hard problems were on the table early. Cross-capsule planning, what people now call multi-agent orchestration. Deterministic execution on probabilistic input, the single biggest unsolved problem in voice AI in 2026 and most teams are just now hitting it. Long-tail intent coverage across hundreds of capsules and dozens of locales, a combinatorial nightmare current LLM voice agents avoid by narrowing scope. Trust calibration. When to confirm, when to execute, when to fall back to UI.
All of it. Named, fought, partially solved. In production. On hundreds of millions of devices. Eight years ago.
What happened next is a marketing story, not a technology story. Samsung scoped a platform as a phone feature. Pressed a hardware button into the chassis. Sold infrastructure as Siri. Journalists wanted Siri so they reviewed Siri, and when they did not find it, they wrote off everything underneath.
The technology never had the problem. The distribution was limited to Samsung Ecosystem .
This post is for the Bixby alumni and the engineers still inside, who shipped that system across hundreds of millions of devices and dozens of languages while the rest of the industry was still arguing about wake words. They named the failure modes the field would not encounter for another five years. They built infrastructure. The company sold it as a feature. The press judged it by the wrong benchmark. None of that diminishes what was built.
BTW, my team was called Voice AI. 🫡
As promised, I open-sourced Dhwani today and the beta is live: https://t.co/T6QU9UoV4Y
I built it mainly because I wanted a much better way to talk to AI agents. Typing long prompts, corrections and half-formed thoughts was becoming the bottleneck, so I made something simple for myself and kept using it.
It is still early and not on the App Store yet, so for now you need to download the file directly and install it manually. But the code is fully open. Use it, change it, fork it, improve it, or make your own version from it.
Would love to see what people do with it.
Everyone wants to build “the best AI dictation app.”
I think that framing is mostly bullshit.
Speech recognition wasn’t invented last month. Most of the technology behind these products has existed for years. What changes is how you combine it, package it, and what you build it for.
I built mine because I wanted a really good way to talk to AI agents.
And I realised there’s no magical IP here that I need to guard.
So I’m open-sourcing it today.
Take it. Edit it. Fork it. Make it yours. If you can make it better, contribute back.
That seems more valuable than pretending every useful piece of software needs a moat.
AI is making code generation cheaper and faster.
That makes formal verification increasingly interesting to me. If software can be generated at enormous scale, the next question is not only “does it run?” but “what can we actually guarantee about it?”
That curiosity took me into a completely new domain and eventually to Talos by @CajalResearch.
Talos is a WebAssembly interpreter written in Lean where the same semantics used to execute a program can also be used to reason about it and prove properties about its behaviour.
I spent the last few weeks going deep into WebAssembly semantics, Lean, validation, spec tests and weakest-precondition reasoning. That exploration turned into several contributions to the Talos validator.
But I kept thinking about another problem too: how do we make systems like this easier for the next person to understand?
So I spent a few days, and roughly $200 experimenting with animation, narration and tooling, building the visual explanation I wished existed when I started.
The Talos team has now made it part of the official README.
I’m increasingly convinced that good engineering has both sides:
make complex systems correct, and make complex systems understandable.
Still a lot for me to learn in this space, but the intersection of AI-generated software, executable semantics and formal reasoning feels incredibly worth exploring.
I am gonna say something risky:
Most of the FDEs I've seen are solving BS problems.
1. Existing enterprise workflows were designed around people, not agents.
2. Real AI gains require rebuilding the org chart.
3. Most execs have no incentive to do that, and most CEOs aren't in founder mode.
4. So FDEs automate the status quo, because that's what's easy to sell. But the status quo is stupid and will be obsolete.
5. If you FDE on critical workflows, 1) enterprises are really this dumb to hand off trade secrets? 2) if you know that, shouldn't you start a competing firm?
If you FDE on non-essential AND soon-obselete workflows, did you learn anything meaningful and repeatable?
6. If you are a smart young FDE, why spend your best years automating dying workflows and managing people who fear AI will take their jobs? Shouldn't you start an agent-native firm from scratch in that industry instead?
Analogy would be, instead of teaching everyone to be AI builders within a tech company, you as FDE build agents that auto-ingest meeting notes for PM, QA, front-end, back-end, devops, UX and help each of them write PRD, Figma, and meeting notes faster for their sync meeting, but the right solution should be to shrink roles, hire full-stack AI builders, and kill meetings. You gain lots of experience with your PRD/Figma/standup meeting sync agents but it's a useless skill.
I am not a FDE expert so I might oversimplify things. I am genuinely curious about this topic and welcome any critiques and comments.
Everyone is focused on making AI agents better at running experiments.
I think the bigger unlock is making experiments accumulate knowledge.
A failed run should eliminate a hypothesis.
A contradiction should change what gets tried next.
A result should become reusable context for every future agent.
Otherwise we’re just automating trial and error.
@dineshpaii Well thats what kiteMCP did by building the thing.. But i think it is just half story and there is so much to build for it to become thing .. waiting for zerodha 2.0 🫡
I’ve spent 15+ years taking AI, optimization and embedded systems from hypothesis to production, across research teams and two startups.
Now I’m focused on one question:
What makes an AI agent reliable enough to trust with real work?
I’ll publish the work here: evals, failure analyses, architecture, code, diagrams, open-source PRs, and lessons from systems that failed before they worked.
First investigation: can an agent improve its own harness without overfitting the eval?
I was going through an open-source repo today and had a small realization.
The interesting part wasn’t that the code was hard.
It was that I couldn’t quickly figure out what was unique about the repo, where the real gaps were, or where a contribution would actually move things forward.
We talk a lot about making software open source. But if you want more people, and increasingly agents, contributing, the repo also has to be designed for contribution.
Good onboarding should make one thing obvious: Here’s where your next hour can create value.
Open code is step one. Open contribution is the real unlock.
I started in embedded systems: kilobytes of memory, deterministic code, hard real-time constraints, and failures you could not hand-wave away.
Now I’m thinking about the opposite extreme: billion-parameter, probabilistic, agentic systems with access to real-world tools.
The interesting problem is the same:
How do you make powerful systems safe enough to trust?
The most fucked-up road in India is in Bangalore, Harlur Main Rd near Prestige Ferns. No road, no footpath, potholes.
Didn't get a cab for the last 2 hours. Went onto the road to find one on the go. No luck. Every auto is filled with passengers. Now government buses are operating to and fro.
These buses create more congestion than before and take 1 hour daily to get out of this 1.5 km stretch.
Was supposed to attend a research talk at Microsoft and hear the researchers. Could not make it.
We live in a country that would need 20 years to do a project.
What if there was an ambulance on this lane?
Infrastructure is such a basic need. @GBA_office@krishnabgowda
Support this man !! Let him help you build better India!
I wish I had millions to support @AadityaYuvraj mission !!
If you’re just fucking rich , just back the ecosystem in right way!!
Toxic is the most Amazing movie I have seen in recent times. @TheNameIsYash
So much witch hunt and manufactured negative PR to not let the man rule then big screen..
I hope he gets recognition in Hollywood.
Indian doesn’t deserve such passion.. gawar loag