Computer Vision Researcher | Alum @iiit_hyderabad | @BoschGlobal | Product Team @KLACorp | Dare, Dream and Do.| Laugh, Learn, Grow, Live and Be Happy 😃😊
I am happy the BCCI gets tax relief. Cricket is a great Indian success story and deserves encouragement.
But why should research labs pay GST on essential equipment and software? If we can make sports tax-friendly, we should make R&D tax-friendly too. Right thing will be to remove GST on essential research inputs and help build India’s future.
We grew from ~0 to $500M ARR, adding $250M last year alone while being EBITDA profitable.
1,300-word post on every growth tactic that worked for us:
1. What got us from 0-$400M ARR in the US works in every country.
2. Retention and Monetisation Hacks.
3. Localization > Translation.
4. Experiment with $1M internal seed checks.
5. Pivoting to only AI content production unlocked 100% ARR growth.
Pocket FM is like Netflix for audio-only dramas, with our own pool of one-person studios.
1. Acquisition Playbook for 0-$10M ARR in any country in 6 months
We run a 90 sec video trailer of an audio drama as an Ad and ask users to download the app if they're interested in the rest of the story.
Our core insight after spending >$100M on acquisition is: If the clickthrough rate (CTR) for an ad goes from 2% to 2.25%, our customer acquisition cost (CAC) decreases by ~ 30%.
We remodelled our system around this insight and built an AI-first 2.25% CTR ad manufacturing machine that works in every country.
For every new country, we take our hit shows -> use LLMs to extract and most intriguing moments -> Write a 5-min script combining all the best parts. The first 60 seconds has to have a hook every 5 seconds and needs to end with a crazy cliffhanger to force a download mid-scroll.
If a marketing video is not hitting our benchmarks (2.5% CTR and 55% 3-second through-play), a creative director gets involved to change the hook or cliffhanger to get the numbers there.
AI lets us make 1,000 Ads per show, and in total we do ~17.5k Ads per month. When a business creating scales from 1k to 10k Ads, the normal thing is for CAC to skyrocket. But with what I just shared, we 7-8x'd our User Acquisition budget without increasing our CAC materially.
It took us 8 months to figure this out, but then the timeline from 0-$10M in every country got shorter and shorter:
US revenue grew to $25M in 19 months. (US is now 78% of total)
Germany to 21M in 11 months.
France to $10M in 3 months.
2. Retention and Monetisation Hacks.
We knew we wanted to create an audio entertainment platform, but there was no standard format. For the first 2-3 years, we tried 10 different formats before landing on the winner.
After we got it right with audio drama (8-12min chapters, written for mobile fiction, serialized, episodes have strong hooks and end on cliffhangers), avg daily streaming time went from ~25mins to 150+.
Audio drama made us realize that Pocket FM was creating a whole new medium. There was no playbook for anything that we were doing. Everything had to be thought & built from scratch.
This is what we did for each major bullet:
- Monetization: Users have some free daily minutes to listen; then it's pay per episode. We also added the option for users to unlock episodes by watching ads, which is doing extremely well. Ads scaled from zero to a ~$90M run rate in 12 months.
- Engagement & production: You don't become obsessed with an app you open once a week. To have users engage daily, we make the next episode free every day. Also helps them build the habit.
- Discovery. Huge problem because people consuming Pocket FM enter the app, tap the show, and lock the screen. We fixed it by doing "playlists" of episodes. Once users's free minutes on a show are done, we ask users if they want to pay. If they say no, we play a new show, one where they haven't used their free mins. And we stitch episodes of different shows together that way, creating natural discovery.
3. Localization > translation.
78% of our revenue is still concentrated in the US. Localisation is fixing this: You might write a show for a Spanish audience where language, jokes, folklore, have a certain flavor; if you merely translate the show for, say, a Norwegian audience, that color is lost and hurts the show in the Norway. Listeners would relate more if the show was written by a Norwegian.
That's why, instead of translating, we localize shows to different regions. We use AI to take the spine of stories and adapt their whole cultural layer to the other country.
The results: a US show that was localized for a German audience had 50% higher retention than the translated version.
Localization + our user acquisition playbook led to:
- $10M+ ARR in France within 3 months
- $21M+ ARR in Germany within 11 months.
And this expands writers' addressable market.
We get messages of writers thrilled to have revenue coming from the US, India, EU, LatAm, without them doing much incremental work.
4. Experiment.
We give $1M checks to new internal initiatives, and the team has 12-18 months to prove their thesis. If they prove it, we double down.
This is how Pocket Saga came about, our AI video microdrama app. It's an AI video equivalent of Pocket FM. Same shows and structure, but as a vertical 2-min mobile video series.
We launched it two months ago and it's at ~$15M ARR.
Pocket's broader thesis is to help creators tell their stories to as many people as possible. Start with audio drama -> multiple languages or localize to diff countries -> microdrama -> more formats like movies, TV shows, and games.
Best of all is that writers get revenue streams not only from countries they wouldn't have tapped into, but also from formats they wouldn't have thought possible.
5. Pivoting to only AI content production unlocked 100% ARR growth.
In mid-2024, we pivoted to only AI content production.
Our growth took a very direct hit. Pocket was already at $200M ARR, growing 50% YoY, and we flatlined for a whole semester.
Six months later, the business exploded.
After the switch, we went from ~25k hours of content produced per year to over 2.5M hours, which are also higher quality. Because AI orchestrates the writing and more data is fed into it, more blockbusters come out.
Over 90 titles have $1M+ in lifetime earnings and 13 crossed $10M.
Quality control is done with LLM as a judge, and LLMs are quite tough.
Results are very encouraging:
- 12-month revenue retention went from 44% to 76%.
- In the past year, we added $250M in net new ARR.
__________________________________
Pocket Entertainment (Pocket FM + Pocket Saga) has become the largest AI entertainment platform.
We have the largest storytelling catalog with 770,000 titles, 550k creators, and 5.5B hours of playtime with minute by minute retention & engagement data. We are using all of this data to improve every aspect of our business.
Our Bet: In the next 3 years we will be able to produce Naruto-like series for $1000.
Netflix has to spend $17B to find 100 Blockbusters per year.
What happens when you can produce 1M high quality shows for $1B?
The streaming wars caused a $300B reallocation of market cap.
The AI entertainment wars may be a $1T+ reshuffling of market cap.
We are at a unique spot. Unlike other AI creator tools like Runway or Midjourney, we own both the supply and demand side of AI content.
We've attracted over 550,000 writers who are producing an annualized 2.5M hours of content every year. This pairs with 137B+ minutes streamed.
Because of this, growth is accelerating as we scale further.
We have multiple S-curves inflecting at the same time:
- AI produces better and more ads -> scale faster in new countries
- Writers + AI produce more shows -> more blockbusters
- Localisation -> more reach per blockbuster
- Audio to video format expansion -> larger TAM -> more creators
Every aspect of our business is designed to improve another, and everything is aligned so that the number of blockbusters continues to rise.
If you're interested in building at the intersection of ai, tech and entertainment, DM me.
I’ve been studying income, wealth, inequality and poverty for more than two decades.
Over that period I’ve also learnt to arrive, roughly, at unaccounted income and wealth from sources available in the public domain.
With AI, asking the right questions and drilling down further and further gives an approximate picture.
You can be approximate and still get the direction right. It is very difficult to be precise on these things in our country.
I’ve taken into account both accounted and unaccounted wealth.
Tamil Nadu has about 2.1 crore households.
This is a rough picture of wealthy families in the state, primary residence excluded.
70,000 households, 0.33% of the state’s households, are worth ₹10 crore or more (around $1 million, HNIs).
2,000 households, 0.01% of the state’s households, are worth ₹100 crore or more (around $10 million, decimillionaires).
900 households, 0.004% of the state’s households, are worth ₹300 crore or more (around $30 million, ultra HNIs).
85 households, 0.0004% of the state’s households, are worth ₹1,000 crore or more (around $100 million, ultra rich).
a new kind of engineer is showing up. they don't write prompts, they design how the AI works. it's called graph engineering, and the people who get it now are about to make everyone else look slow https://t.co/z1HhE6voRu
whoever leaked this has bigger balls than sense
Google Research and MIT ran the same agent jobs 260 different ways for Nature last month: they held the prompts, the tools and the compute budget identical and moved nothing but the wiring between the agents, and the same work swung from 70% worse than a single agent to 80.8% better, averaging out at 0.0%
i ran my own single agent against the task list first and it cleared 6 of 10 alone, already past the line where a crew starts subtracting
this is Graph Engineering, the layer that decides whether a crew is worth 80% more or 70% less, and it installs into the agent you already pay for:
- score your solo agent on the real task first: above roughly 45% success that study predicts zero to negative returns from any crew you put around it
- under that line, put one supervisor over the fan out: crews with no correction step amplified their own errors to 17.2x the single agent rate, supervised aggregation held it to 4.4x
- give every worker one output and let none of them read a peer's draft, so a wrong step reaches the supervisor instead of four other agents
- run the comparison again after every model upgrade, because a better model raises your baseline and a higher baseline is what makes a crew stop paying
- keep the single agent alive as the control, the only number that says the wiring is earning its calls
turns out the shape does not travel: the biggest win came off a finance task under one supervisor and the worst collapse off a planning task with independent agents
my position, and it is the arguable one: a crew is a bet on your own diagram, and the model you pick moves that bet less than one arrow does
bookmark this, the three moves that draw those arrows before you pay for one extra call are in the post below ↓
It’s one thing to talk about the future of India's urban mobility. It’s another to stand right next to it.
Honored to host Shri @sanjeevsanyal (Member, EAC-PM) at our facility. His visit deeply validates our physics-first solution to urban congestion. (1/2)
Message from Sonam :
20th JULY
आज़ादी का दूसरा आन्दोलन
भय मुक्त भारत, अन्याय मुक्त भारत
Freedom from injustice (Like paper leaks)
Freedom from Fear (my illegal detention)
India’s 2nd FREEDOM MOVEMENT
March to the Parliament
Please make it a big success
Sent through Gitanjali
from my illegal detention at Safdarjung
LLM Knowledge Bases
Something I'm finding very useful recently: using LLMs to build personal knowledge bases for various topics of research interest. In this way, a large fraction of my recent token throughput is going less into manipulating code, and more into manipulating knowledge (stored as markdown and images). The latest LLMs are quite good at it. So:
Data ingest:
I index source documents (articles, papers, repos, datasets, images, etc.) into a raw/ directory, then I use an LLM to incrementally "compile" a wiki, which is just a collection of .md files in a directory structure. The wiki includes summaries of all the data in raw/, backlinks, and then it categorizes data into concepts, writes articles for them, and links them all. To convert web articles into .md files I like to use the Obsidian Web Clipper extension, and then I also use a hotkey to download all the related images to local so that my LLM can easily reference them.
IDE:
I use Obsidian as the IDE "frontend" where I can view the raw data, the the compiled wiki, and the derived visualizations. Important to note that the LLM writes and maintains all of the data of the wiki, I rarely touch it directly. I've played with a few Obsidian plugins to render and view data in other ways (e.g. Marp for slides).
Q&A:
Where things get interesting is that once your wiki is big enough (e.g. mine on some recent research is ~100 articles and ~400K words), you can ask your LLM agent all kinds of complex questions against the wiki, and it will go off, research the answers, etc. I thought I had to reach for fancy RAG, but the LLM has been pretty good about auto-maintaining index files and brief summaries of all the documents and it reads all the important related data fairly easily at this ~small scale.
Output:
Instead of getting answers in text/terminal, I like to have it render markdown files for me, or slide shows (Marp format), or matplotlib images, all of which I then view again in Obsidian. You can imagine many other visual output formats depending on the query. Often, I end up "filing" the outputs back into the wiki to enhance it for further queries. So my own explorations and queries always "add up" in the knowledge base.
Linting:
I've run some LLM "health checks" over the wiki to e.g. find inconsistent data, impute missing data (with web searchers), find interesting connections for new article candidates, etc., to incrementally clean up the wiki and enhance its overall data integrity. The LLMs are quite good at suggesting further questions to ask and look into.
Extra tools:
I find myself developing additional tools to process the data, e.g. I vibe coded a small and naive search engine over the wiki, which I both use directly (in a web ui), but more often I want to hand it off to an LLM via CLI as a tool for larger queries.
Further explorations:
As the repo grows, the natural desire is to also think about synthetic data generation + finetuning to have your LLM "know" the data in its weights instead of just context windows.
TLDR: raw data from a given number of sources is collected, then compiled by an LLM into a .md wiki, then operated on by various CLIs by the LLM to do Q&A and to incrementally enhance the wiki, and all of it viewable in Obsidian. You rarely ever write or edit the wiki manually, it's the domain of the LLM. I think there is room here for an incredible new product instead of a hacky collection of scripts.
Andrej Karpathy (@karpathy) — co-founded OpenAI, led AI at Tesla, coined "vibe coding."
In 4 minutes he explains why software is changing - and why Claude Skills, MCP servers, and AI agents aren't hype anymore.
They're the foundation of how software gets built from now on.
Imo, worth every second (i've added subtitles)👇
Three days ago I left autoresearch tuning nanochat for ~2 days on depth=12 model. It found ~20 changes that improved the validation loss. I tested these changes yesterday and all of them were additive and transferred to larger (depth=24) models. Stacking up all of these changes, today I measured that the leaderboard's "Time to GPT-2" drops from 2.02 hours to 1.80 hours (~11% improvement), this will be the new leaderboard entry. So yes, these are real improvements and they make an actual difference. I am mildly surprised that my very first naive attempt already worked this well on top of what I thought was already a fairly manually well-tuned project.
This is a first for me because I am very used to doing the iterative optimization of neural network training manually. You come up with ideas, you implement them, you check if they work (better validation loss), you come up with new ideas based on that, you read some papers for inspiration, etc etc. This is the bread and butter of what I do daily for 2 decades. Seeing the agent do this entire workflow end-to-end and all by itself as it worked through approx. 700 changes autonomously is wild. It really looked at the sequence of results of experiments and used that to plan the next ones. It's not novel, ground-breaking "research" (yet), but all the adjustments are "real", I didn't find them manually previously, and they stack up and actually improved nanochat. Among the bigger things e.g.:
- It noticed an oversight that my parameterless QKnorm didn't have a scaler multiplier attached, so my attention was too diffuse. The agent found multipliers to sharpen it, pointing to future work.
- It found that the Value Embeddings really like regularization and I wasn't applying any (oops).
- It found that my banded attention was too conservative (i forgot to tune it).
- It found that AdamW betas were all messed up.
- It tuned the weight decay schedule.
- It tuned the network initialization.
This is on top of all the tuning I've already done over a good amount of time. The exact commit is here, from this "round 1" of autoresearch. I am going to kick off "round 2", and in parallel I am looking at how multiple agents can collaborate to unlock parallelism.
https://t.co/WAz8aIztKT
All LLM frontier labs will do this. It's the final boss battle. It's a lot more complex at scale of course - you don't just have a single train. py file to tune. But doing it is "just engineering" and it's going to work. You spin up a swarm of agents, you have them collaborate to tune smaller models, you promote the most promising ideas to increasingly larger scales, and humans (optionally) contribute on the edges.
And more generally, *any* metric you care about that is reasonably efficient to evaluate (or that has more efficient proxy metrics such as training a smaller network) can be autoresearched by an agent swarm. It's worth thinking about whether your problem falls into this bucket too.
American wafer fab equipment maker KLA Corporation is setting up a R&D centre in Chennai. They would be employing 3000 people. The jobs are of high quality requiring very strong skills.
Moneycontrol reports today that no one thought India could build its own chip, until Chennai startup Mindgrove did it.
The company specialises in microcontrollers and microprocessors, the tiny systems that power everything from smart locks to washing machines and biometric readers.
To quote Moneycontrol:
“In the old days, design and manufacturing were under the same entity,” says Shashwath. “But the world has since moved to a modular ecosystem…companies like NVIDIA, AMD, and Qualcomm design the chips, while foundries make them. We do the same…we design, get it manufactured, and sell it under our brand.”
That’s the simple version. But what they actually do, he says, is closer to writing software that becomes hardware.
“At the most basic level, what we do looks like writing software,” he explains. “But instead of compiling something that runs on a computer, we compile something that gives the computer itself a brain.”
It’s this translation from code to silicon that forms the soul of Mindgrove’s work.
Also when Trump is heavily discouraging fresh investments in India, Ford is reopening it's Chennai factory, not for selling cars in India, but to set up a next-generation powertrain facility that will exclusively produce advanced engines for global markets.
Tamil Nadu is truly progressing in advanced manufacturing.
🚨Semiconductor Major 🇺🇸KLA opens it's newest Global Capability Centre (GCC)- Only one in India at Chennai's DLF Downtown. The 300 Crs facility spread over 311,000 Sq.ft and supports 1,300 employees. News floating of possible 3,000 Crs expansion too ... #InvestInTN#GCC 🧑💻