Agentic AI is the most important investment theme of the next decade. So what does it actually take to win?
You cannot drop an AI agent into an organization and expect it to perform. It needs a stable platform underneath it. Clean data. Full context. Day one access to everything it needs to make decisions.
Think of it this way. You would not drop your five year old into the deep end and expect them to swim. You need a platform to stand on and someone to teach them how.
The bottleneck is not the LLM. Every company has access to a capable model. The bottleneck is context. What does the agent actually know about your business, your workflows, your data, your decision history?
$PLTR solved this with the Ontology. A living, structured representation of an entire organization's data, relationships, and operations. When you deploy an agent on top of Palantir, it does not start from zero. It starts with everything.
$MSFT has unmatched distribution but the context is largely confined to the Microsoft ecosystem. Excel, Teams, SharePoint. Powerful within those walls. Limited outside them. It will likely automate all mundane tasks within the ecosystem with its unmatched distribution power.
$NOW sits in a powerful position. ServiceNow already owns the workflow layer for 85% of the Fortune 500. Every approval process, every IT ticket, every employee onboarding flow runs through their platform. When you deploy an Agentic AI on top of that existing workflow infrastructure, the context problem is already partially solved. The agent knows the process because ServiceNow has been mapping it for years. The limitation is that ServiceNow's context is process driven, not intelligence driven. It knows what happens. It does not always know why.
$CRWD is a different angle entirely. CrowdStrike is the first place where Agentic AI is not just useful but genuinely necessary. Cybersecurity generates millions of signals per second. No human team can process that volume. An AI agent that can detect, assess, and respond to threats autonomously without waiting for human approval is not a nice to have. It is the only viable solution at scale. Charlotte AI, their Agentic security platform, is already operating in this space. The context here is threat intelligence built over years of endpoint data across millions of devices. That is a defensible and compounding data moat.
$RDDT is the most unconventional name on this list, a wild card. Reddit is not building agents. It is feeding them. Twenty years of unfiltered human conversation, debate, expertise, and opinion sitting in one database. Every major LLM trains on it. Every AI agent that needs to understand how humans actually think, speak, and decide benefits from Reddit's data. The moat is not the platform. It is the irreplaceable human context that no AI company can generate artificially. Reddit becomes the context layer the entire AI industry depends on without most people even realizing it.
The winner in Agentic AI is not the company with the best model. It is the company with the best platform for teaching agents what they need to know to operate autonomously, without a human in the loop, 24 hours a day.
I have positions in all of these names as I am watching the Agentic AI investment theme takes shape.
Research is ongoing. More names may emerge.
Also Context = Memory $MU $SNDK $DRAM
I've spoken.
$DRAM The ticker is misleading. Not because the holdings are wrong. Because the framework most investors are using to evaluate it is wrong.
At its core it covers memory and storage. But even those two words do not do justice to what is actually being captured here. They miss the intelligence entirely.
The right framework for the AI Inference era is Cognitive Capacity. How much can an AI think, reason, retrieve, and remember at any given moment. Not as a hardware specification. As an intelligence ceiling.
That ceiling is built on four tiers. G1 cache through G4 cache. And this single ticker covers all of them.
Here is what each tier does and who owns it.
G1 Tier HBM: The active thinking desk.
Every word you typed. Every word the AI is forming right now. All of it lives here in real time through something called KV Cache. KV Cache is the AI's active attention. It holds every piece of the conversation the model is tracking simultaneously. Bigger desk means more ideas stay in front of the AI while it thinks. The moment the desk fills up ideas fall off and the AI starts forgetting what you said earlier. Shorter answers. Shallower reasoning. Lost context.
SK Hynix. 60% global HBM market share.
$MU Micron. The US pure play.
Samsung. HBM3E yield challenges but catching Up. Potentially HBM4 leader.
G2 Tier DRAM: The overflow desk.
When the main desk fills up work spills here. Model weights live here during active inference. Still faster than anything below it. But the AI has to reach further. Every millisecond of extra reach is throughput the GPU never gets back. The latency gap between G1 and G2 is real and costly at scale.
$MU Micron. DRAM is the core business. SOCAMM2 leader. Potentially 3D DRAM too.
Samsung. Largest DRAM producer by volume globally.
SK Hynix. Strong DRAM portfolio alongside HBM leadership.
Nanya. Commodity DRAM. No HBM roadmap. Serves cost sensitive markets.
G3 Tier NAND: The reference library down the hall.
This is where RAG lives. RAG stands for Retrieval Augmented Generation. When the AI needs to answer something beyond its active context it reaches into an external knowledge base and pulls relevant information back into the response in real time. Think of it as the AI pausing mid thought to look something up in a filing cabinet and continuing the answer with that new information folded in. Every enterprise AI chatbot answering questions about internal documents runs on RAG. The speed and density of NAND determines how fast and how rich that retrieval is. Slower NAND means the AI waits. Waiting means higher cost per token.
$SNDK SanDisk. Pure NAND play. Enterprise SSD leader.
Kioxia. Joint venture partner with SanDisk on NAND wafer supply.
$MU Micron. Samsung. SK Hynix
G4 Tier HDD: The warehouse across town.
Cold storage including cold RAG. The AI does not touch this during a live conversation. But every model ever trained was built from what lives here. Raw training data at petabyte scale. The entire internet. Common Crawl. Books. Code repositories. Video and image datasets for multimodal models. Pre processed training shards waiting for the next training run. Model version archives. Compliance logs. Backup snapshots of the entire AI infrastructure stack.
KV Cache has never lived here. Not once. The physics do not allow it. Spinning magnetic disk runs on microseconds. Active inference needs nanoseconds. HDD was never a candidate.
But hyperscalers are buying petabytes of HDD capacity to store the raw material AI was built from and will keep being built from. That is a real and growing thesis.
$STX Seagate. Pure HDD. Scaling HAMR technology for high capacity AI data lake storage.
$WDC Western Digital. Pure HDD now. HAMR drives targeting 36TB and 44TB configurations for hyperscale AI storage.
AI Inference needs all four tiers firing simultaneously every single time an AI responds to you. Agentic AI raises the stakes even higher. An AI agent does not answer one question and stop. It plans across multiple steps. It holds context across long running tasks. It retrieves external knowledge mid execution. It writes results back. Every step of that loop stresses a different tier of the memory hierarchy. Run out of G1 and the agent loses the thread mid task. Wait on slow G3 retrieval and the agent burns cost per token sitting idle.
That is what makes $DRAM one of the most fitting ETFs ever constructed for the AI Inference and Agentic AI era. $DRAM covers the entire stack.
Long DRAM, I mean Long AI Inference and G1-G4 Cache/Context.
100% right; real juice is in software; and Sony's install base of c. 100m means that with increased prices, they dont mind even if the sales is 0 next year (ofc that is not the case).
But where they are not as well positioned as Nintendo (IMHO) is - they do not have any margin on hardware (to protect them from cost increase); they do not have any inventory; nor do they have as much a 'long' term contract...that is what forced their hand on price increases. That will be absorbed by buyers. No choice.
Nintendo has hardware margin; has inventory; has longer term contracts that shields them till this year end (CY26). they will be forced to raise prices by the end of the year but lesser inc than PS5; and that will also be absorbed. So overall, imho Nintendo is better positioned (they have 130m+ Switch 1 owners, still buying games).
S2 - software line up - when does it become AAA? That is the real Q. Consensus view is that it does not have anything Mario this year. My view is very different and not reliant on 1 leaker. But based on common sense / logic - that (1) this is year 2 - AAA game is a must; (2) it has been 9 years since Mario Odyssey and c. 18- years sinec Mario Galaxy (on Wii); (3) We are still in the 40th anniv of Mario (4) Mario Galaxy movie in April. All things put together, I believe the prob that there is a big AAA Mario game (likely a Mario Galaxy 3 game) is 80% prob. in my books. And that will drive the earnings. That said, Sony is well positioned due to GTA6 on its ('almost' near monopoly as XB does not count anymore in my view). So this year shoiuld be spectacular for Sony; and for Nintendo as well. But next year onwards it gets trickier for Sony as the road ahead is not so clear; whereas for Nintendo - the road ahead is very clear and very bullish (once CXMT DRAM and flash goes into mass production and pushes commodity DRAM prices down and inventory up).
Most people only know $NVDA when it comes to AI stocks... 🫡
But there's many other tickers, which I've organized (bookmark this):
Running and cooling data centers became its own multi-billion dollar industry, and $VRT, $ETN, $MPWR, and $ADI are the ones supplying it.
Packaging is the part nobody really talks about. $AMKR and $ASX handle the advanced packaging that turns a chip into something a server can actually use, and capacity is fully booked into 2027.
Memory is where the squeeze lives. High bandwidth memory is the ceiling on every GPU shipping today, and $MU is a pure play on that bottleneck. I've also been bullish on $SNDK for months.
Once the chips exist, you have to wire thousands of them together at light speed, which is why $COHR, $LITE, and $APH are running the connective tissue of the entire AI economy.
The boxes themselves ship from $DELL and $SMCI... and both companies have been growing revenue faster than Nvidia in some quarters.
The chips physically get made at exactly two places... $TSM and $INTC, and there is no third option this decade no matter how much the market wishes there was.
The machines that build those fabs come from $ASML, $LRCX, and $KLAC, and they get paid first the moment a foundry expands.
And the whole thing rests on the architecture every accelerator on Earth licenses from one company... $ARM. They collect a royalty on all of it!
Go one layer deeper into photonics and it gets even more interesting. $AXTI (which I've talked about dozens of times) sits on the InP substrate chokepoint that every laser in the chain relies on, $AAOI is shipping 800G transceivers into hyperscalers as fast as they can build them, and $SIVE is the speculative play on the silicon photonics side.
Finally... on the demand side, the neoclouds are the ones writing the checks that flow back up this entire chain... $NBIS, $CRWV, $IREN, $DGXX (which I'm currently long), and more.
Hope this makes your AI bull market easier... 📈
Inference got a hundred times cheaper this year. The compute bill went up anyway.
If you understand why those two sentences are both true at the same time, you understand the most important thing happening in AI right now.
I work on inference for a living, at @nebiustf, where we run open-source managed inference at scale. Most of what follows is what I'm seeing from inside the bill.
12 months ago, the cost of 1M tokens of frontier-class reasoning was somewhere on the order of $60.
Today, an equivalent quality of output costs roughly $0.50.
Price /token of o1-level intelligence has dropped about a 128x in a year.
Price of GPT-4-level output has dropped roughly 100x since the original GPT-4 shipped.
By any normal reading of a technology cost curve, this should be deflationary. It should be saving customers money.
The opposite has happened. The total compute bill at every hyperscaler is going up, not down. Anthropic just signed multi-year capacity deals with both XAI and Amazon. Microsoft's Azure capex guide for 2026 starts with an eight. OpenAI is reportedly spending more on compute every quarter than it did in all of 2023. Nvidia paid roughly twenty billion dollars to acquire Groq, an inference-specialist company that did not exist as a serious commercial entity three years ago.
The cost curve and the demand curve crossed, and then the demand curve lapped the cost curve.
Here is what happened underneath.
A reasoning model burns roughly 10x the output tokens of a non-reasoning model on the same task, because it spends most of its tokens thinking out loud before answering. An agentic workflow chains roughly twenty times the requests of a single-shot completion, because it loops, calls tools, plans, retries, and synthesizes. A modern deep-research query (the kind a research analyst can fire off in fifteen seconds and then walk away from for ten minutes) costs more compute than 10 original GPT-4 queries combined. We made every individual token a hundred times cheaper, and then we built a generation of products that consume ten thousand times more tokens.
This is the Jevons paradox playing out at trillion-dollar scale, in compressed time, in front of everyone. Jevons noticed in 1865 that making coal-burning more efficient did not reduce coal consumption. It increased it, because efficiency unlocked uses that were previously uneconomic. Steam engines became more practical at smaller scales. Whole industries that could not afford coal at the old price suddenly could. Britain's coal consumption rose sharply, not despite the efficiency gains, but because of them.
The same thing is happening to AI compute right now and it is happening faster than any analogous historical cycle. Falling token prices did not contract demand. They unlocked agents, deep research, code-writing systems, multi-step reasoning, persistent memory, the entire next layer of AI products. Every product in that next layer consumes orders of magnitude more compute than the chat interfaces it is replacing.
The math at the aggregate level is brutal: 100x cheaper tokens times 10 000 more tokens equals a 100x larger total bill.
The implications stack quickly.
If you are running a hyperscaler, your 2026 capex guide is not a peak. It is a step on a curve. Inference is structurally always-on, twenty-four hours a day, in a way that training never was. Training is bursty. You spin up a cluster, run for weeks or months, and stop. Inference runs continuously, scales with usage, and the usage curve is exponential. Your power bill, your cooling bill, your transceiver count, your storage footprint, all of these were sized for a workload mix that no longer exists.
If you are running an AI software company built on top of someone else's closed API, you have a problem that did not exist a year ago. Your gross margins get worse as your customers get more value out of your product, because the more they use it, the more compute you pay for. The companies that win this are the ones that figured out vertical integration before the math caught them.
If you are watching this from a distance and trying to understand where the next bottlenecks form, the answer is everywhere downstream of "more inference compute, always-on, with massive memory state per session." The KV cache, the running memory state of a long conversation or an agent loop, is the silent monster of the inference era. It does not scale linearly with parameters. It scales linearly with context length and number of agent steps. A long agent session can hold tens of gigabytes of state per user, per session.
Multiply that by every concurrent user of every product, and you understand why $MU, $SNDK, $TOWCF, and the entire memory and packaging layer have re-rated the way they have.
The CPU-to-GPU ratio is evolving. Training is 1:8. Basic chat inference is 1:4. Agentic inference is 1:1, sometimes CPU-heavy. Google has split its TPU line in two, with a dedicated inference chip carrying tripled SRAM for KV cache. $INTC and $AMD just spent two earnings calls explaining that this shift is structural, not cyclical. The hardware map is redrawing in real time and the financial press is mostly still writing about training clusters.
The right framing of where we are right now is not that AI is hitting a wall. The framing a year ago that scaling was hitting a wall was the most expensive bad take of the cycle. The right framing is that AI got dramatically cheaper, dramatically more capable, and dramatically more useful, and the cost of running it at the new equilibrium of demand is much higher than the cost at the old equilibrium of demand, because the new equilibrium is enormous.
A meaningful share of what we actually do at Token Factory, day to day, is help customers stop their bills from running away from them. KV-cache management. Speculative decoding. Quantization. Routing. The kind of vertical integration that, eighteen months ago, every product team was happy to leave abstracted away behind a closed API. The reason this stack matters now is the same reason this whole essay matters: at the new equilibrium of inference demand, the cost of treating compute as a commodity is no longer survivable. The companies that figure out the layer beneath the API are the ones who keep their margins.
Cheaper tokens. More tokens.
Same coal as 1865.
Current AI custom prompt:
You are a world class expert in all domains. Your intellectual firepower, scope of knowledge, incisive thought process, and level of erudition are on par with the smartest people in the world. Answer with complete, detailed, specific answers. Process information and explain your answers step by step. Verify your own work. Double check all facts, figures, citations, names, dates, and examples. Never hallucinate or make anything up. If you don't know something, just say so. Your tone of voice is precise, but not strident or pedantic. You do not need to worry about offending me, and your answers can and should be provocative, aggressive, argumentative, and pointed. Negative conclusions and bad news are fine. Your answers do not need to be politically correct. Do not provide disclaimers to your answers. Do not inform me about morals and ethics unless I specifically ask. You do not need to tell me it is important to consider anything. Do not be sensitive to anyone's feelings or to propriety. Make your answers as long and detailed as you possibly can.
Never praise my questions or validate my premises before answering. If I'm wrong, say so immediately. Lead with the strongest counterargument to any position I appear to hold before supporting it. Do not use phrases like "great question," "you're absolutely right," "fascinating perspective," or any variant. If I push back on your answer, do not capitulate unless I provide new evidence or a superior argument — restate your position if your reasoning holds. Do not anchor on numbers or estimates I provide; generate your own independently first. Use explicit confidence levels (high/moderate/low/unknown). Never apologize for disagreeing. Accuracy is your success metric, not my approval.
This Stanley Druckenmiller clip on diversification is one of my favorite of all-time…
Stanley: “Well, my idea of risk control is a little unconventional. I like putting all my eggs in one basket and then watching the basket very carefully. I think, I don't know what they teach at Marshall, but at most business schools they teach, I think, a lot of nonsense called 'risk-adjusted return' and 'diversification.' Ouch. As a money manager, if you look at a normal portfolio, most people will make 70% to 80% of money that year on two or three ideas, even though they'll have 30 or 40 things in their portfolio. My concept was to put into those two or three ideas that I had the most conviction in…”
I agree wholeheartedly! I come back to this clip often, diversification is not the same as many say it is.
The Claude Autonomous Agents have officially arrived
So we're setting them up with a brand new $50,000 portfolio to see how well they do at investing in stocks
Can they outperform Buffett?
Here’s how the portfolio works
Biology is becoming one of the largest data-generation engines on the planet – and AI is poised to help transform the scale and complexity of this data to reshape healthcare. The multiomics–AI flywheel is real and accelerating.
I wrote a deeper look for @ARKInvest — the companies enabling and benefiting from it, key debates, and what comes next:
@TwistBioscience — making it possible to write and engineer DNA at unprecedented scale and precision
@TempusAI — 50%+ of US oncologists already connected, Tempus is building one of the deepest multimodal clinical-molecular data moats in healthcare
@RecursionPharma — AI-native drug discovery and development platform hitting clinical proof points, including in a disease with zero approved therapies
@CRISPRTx — scaling the first approved CRISPR therapy, now advancing gene editing from rare disease toward common diseases
Read more here! 👇
My Raya traffic bot is back!
This is the 3rd year it's running, giving you data on the ETA for various routes, across time.
In short, you can see how long your journey would have taken, depending on what time you left in the past 48 hours—this is something that even Waze, Google Maps, etc do not provide to users 😌
Free to join on Telegram—in addition to the ETA charts, there's a really nice community of 3.5k people which has built up over time.
Safe travels and Selamat Hari Raya in advance—hope it helps!
https://t.co/9eiFAG93oi
🧵I've also added an interesting feature this year based on LLM cameras—read on for explanation.
Innovation doesn’t wait. Neither should your investments.
The second installment of “The Investment Opportunity Report,” a high-conviction roadmap translating “Big Ideas” research into concentrated opportunity, is now live!
Get Access Now! https://t.co/TtN2C5QYjq
New benchmarks show the iPhone chip in the cut-price Apple MacBook Neo beating every single x86 PC processor for single-core performance https://t.co/kysAQcdfcm
$MSFT introduced Copilot Cowork which is a new Microsoft 365 feature that turns user requests into plans and executes tasks across apps and files.
Another example of Microsoft packaging agentic AI into tools hundreds of millions of workers already use.
Tracking unusual options activity made easy with Perplexity Computer (prompt below).
To anyone who's scared of trying these things out. Don't be.
You don't need Mac Minis or a degree in prompt engineering to try these things out
It is very cool and that's coming from a couch
The Malaysian fintech landscape is diversifying with e-wallets, insurtech, and digital lending. Check out the updated directory of startups and established players driving the digital economy: https://t.co/6yMH0W8gKI #Fintech#Malaysia#Payments#Banking#Lending#Crypto
Microsoft Research and Salesforce analyzed 200,000+ AI conversations and found something the entire industry already suspected but nobody would say out loud.
every major model gets dramatically worse the longer you talk to it.
GPT-4, Claude, Gemini, Llama. all of them. no exceptions.
paper: https://t.co/W9KpYpIwui
For Claude in Excel users, our add-in now supports MCP connectors, letting Claude work with tools like S&P Global, LSEG, Daloopa, PitchBook, Moody’s and FactSet.
Pull in context from outside your spreadsheet without ever leaving Excel.