this is the future of personal computing, even enterprise computing longer term. tbf, we have already been on this hybrid path in the cpu workload era with our devices, but the demand for inference has vastly changed the runtime balance of cost, privacy, and risk.
no doubt that large frontier LLMs in the cloud will always be important, but this hybrid compute system only gets better from this day forward. the ceilings on small model capacity, on quantization, and on device logic/memory chip performance will only improve from here. this surge in long horizon agent workflow and vpc usage will especially benefit.
kudos to the team at @perplexity_ai for pushing the frontier in computing!
we’re fully aware today that software as we knew it is getting fully disrupted and generated by AI.
now, we are witnessing an AI system rapidly designing and verifying a novel AI accelerator chip, with just two engineers steering it.
proud of @axi_master@aadityasubedi_ and the @architectlabs team on this first breakthrough milestone!
this is a turning point in AI video and the open-weights question.
the new @fal research group worked some wonders on the minimax h3 model. i had posted about a few days ago - essentially, reducing the RTF well below 1, and generating and rendering video faster than you can watch it.
today they are shipping their new H3 Max model for everyone to try. It’s a heavily post-trained version of the Minimax H3 open-weights model. it’s beating all other video models in its category on speed and cost, at highest of quality, according to @DesignArena & @ArtificialAnlys.
another intriguing part of this is how an open-weights model can be *significantly* optimized by throwing data and compute at the model, plus deep human expertise like the @fal research team.
in LLMs, we have examples of this like @perplexity’s PPLX-27B model which is a heavy post-train of Qwen’s open weight model. inference and training services like @FireworksAI_HQ and @parasail_io similar work on LLMs for enterprises and startups, respectively.
tapping into “value-added” deep inference and harness builders in the ecosystem allows these models to further improve in strong and surprising ways. unfortunately, there are not many frontier video models that are open-weights currently. so, kudos to @MiniMax_AI for being one of the first.
i think it’s inevitable that, like the LLM labs, the gen media model labs will have open-weights versions for developers and inference providers to improve and vary the performance as users would demand…at least those labs that want the very best version of their own models available to consumer and enterprise customers.
went on @SquawkCNBC yesterday to share thoughts with @BeckyQuick on AI topics, including SK hynix, the memory-bound future, and the importance of open weights models.
memory is 100% structural now, not cyclical. intelligence is compute-bound, and compute is memory-bound. the markets are still wildly underestimating long term demand esp HBMs vs commodity DRAM, with long context agents, RL, RSI, and robotics/autonomy all rising.
so, why the momentary meltdown? shareholder expectations too high, margin call ripple effects, and single stock leverage ETFs all contributing to the temporary drama.
more on this and open weights models in the video below:
Yesterday, @stevejang shared his thoughts on @CNBC regarding AI’s future roadmap around memory, robotics, agents, and open weights models.
The hot topic of the morning: @SKhynix reported record Q2 results, then fell more than 9%. The gap between the print and the reaction raises a larger question: is the market applying an old memory-cycle framework to a new AI infrastructure layer?
Q2 revenue reached USD $55.0 billion, up an eye-popping 257% year over year. Operating profit rose to USD $42.0 billion, up incredibly 557%, with a company-record operating margin of 76%. But both missed consensus estimates. As our partner Steve Jang told @BeckyQuick on @SquawkCNBC, “the company has incredible fundamentals, record-breaking numbers, and long term HBM technical defensibility…but expectations were just very high.”
Steve’s larger argument begins with high-bandwidth memory, or HBM. For two decades, investors largely treated memory as a cyclical commodity. Now, Steve argues, HBM is moving into a new role, “sitting side-by-side with GPUs and other advanced logic processors” as a core layer of AI compute.
Demand now spans model training, inference, long-context agents, robotics, and autonomous systems. Steve pointed to deep-research products like @perplexity_ai’s Computer agent platform: as agents hold and reason across more context, their memory bandwidth and capacity needs grow. Robotaxis like @nuro+@Uber and @Waymo will need onboard edge compute including HBMs. He estimates HBM demand could increase 10x over the next 3 to 4 years.
SK hynix enters that buildout from a strong position. Long-term supply agreements with major customers provide demand visibility. Stacked-die architecture, advanced packaging, yield, and customer qualification create a steep technical and manufacturing climb. Steve estimates that a new entrant could need three to five years to significantly enter the HBM4e class.
That supports a credible near-term moat while leaving the harder question open: how today’s 76% operating margin evolves as supply expands over the next two to three years. One thing is clear: demand and importance of high bandwidth memory is still wildly underestimated.
The second half of the conversation moved from compute capacity to operational control. During the recent @OpenAI and @huggingface security incident, commercial frontier-model APIs blocked the attack commands and exploit payloads contained in forensic logs. Hugging Face instead ran GLM 5.2, an open-weight model, on its own infrastructure to analyze more than 17,000 recorded events without sending incident data or credentials outside its environment.
Steve’s takeaway is practical. As autonomous agents grow more capable, defenders need access to models they can host and direct when hosted guardrails block legitimate forensic work.
The next phase of AI will depend on both: enough memory bandwidth alongside GPUs and other accelerators to scale increasingly capable systems, and enough model ownership and control to deploy and defend them efficiently and safely.
Full conversation below 👇
Architect Labs is developing AI to accelerate the co-design of custom chips - democratizing the capabilities to innovate on the hardware level of the incredible intelligence era ahead of us.
So, we’re thrilled today to share that @KindredVentures led a $24M seed round for @architectlabs, joined by @TQVentures@RaceCapital@togetherfund, as well as a storied group of operators/researchers including @snsf, @lukaszkaiser, @AravSrinivas, @tlbtlbtlb, @alexwg, and et al from @NVIDIA, @GoogleDeepMind, @OpenAI, and @perplexity_ai.
For decades, the contract between hardware and progress was simple: transistors got smaller, everything got faster, and the industry planned around the rhythm. That rhythm has slowed because of Moore’s Law stalling, and the new performance gains are coming from somewhere else entirely, architectural decisions about how memory and compute are arranged, how silicon is shaped around the workloads it serves, and how hardware and software are designed to evolve together instead of drifting apart.
AI is no longer confined to the data center. It’s running in robots, at the edge, on satellites, in every device with ambition. Each of those environments has its own physics around power, latency, and cost, and none of them are well-served by the same general-purpose GPU. The frontier labs building tomorrow's models are still optimizing upward to fit chips they didn't design, instead of designing chips downward to fit the models they imagine.
The reason that hasn't changed is not just physics, it’s also inertia. Chip design cycles remain long, manual, capital-intensive, and gated by a small population of specialists. Custom silicon has remained an inheritance of a few incumbents rather than a tool available to the new guard who need it.
Architect Labs is rebuilding that process from first principles, an AI-native, self-improving system that co-designs silicon, compilers, runtimes, and system software as a single loop, so the cadence of chip development can finally start to look like the cadence of software. It's the unlock that lets the companies shaping intelligence actually own the hardware that runs it.
The founders, @axi_master and @aadityasubedi_ understand exactly how to diagnose this problem. They have little patience for the status quo. Ebrahim started college at 15, found his way onto Apple's silicon teams, and then onto Tesla's AI5 (the custom chip behind FSD and Optimus) where he watched a piece of silicon become outdated by the models it was meant to serve before it even shipped. Aaditya was researching AI for code verification at Harvard before he and Ebrahim met at Stanford and started working on AI for chip design together, a collaboration that later became the company. Alongside them now is a stellar team, researchers and systems engineers from @Anthropic, @Google, @intel, @Meta, @Samsung and @xai, with 80+ production tape-outs and core contributions at nearly every frontier lab between them.
To the Architect Labs team, we’re honored to join you on this incredible mission! 🚀🚀🚀
From @Perplexity’s computer-use agents to @Roblox’s huge engineering workflows, agents and agent harnesses have forever changed the way teams build products.
Tomorrow in SF we're hosting @randomjohnnyh, cofounder of @perplexity_ai, @andrewswerdlow, VP of engineering at @roblox, and our own @stevejang of @KindredVentures to go deep on this new era.
Apply to join us https://t.co/E3GSIrEasa
our world is accelerating due to speedups in technology and science, due to human ingenuity and AI model advancements.
intellectual property is the record-of-truth for technology, the design, engineering, and science carta for the world's industries, and also the law and economics code by which we operate as innovators and capital markets.
excited to partner with the @fearn_ai @hanhanhan_kim@af_gao and team to democratize and enhance this critical core of technology and science!
Kindred’s @stevejang joined @Bloomberg TV with @CarolineHydeTV to discuss our new $355M fund, which builds on its prior $200M AI/deep tech fund (2022–2025) which is now valued at ~$1B, and ranking in the top 1% of its vintage. Our strategy continues to lean into AI, robotics, and infrastructure at the earliest stages.
key takeaways from the interview:
- As capital concentrates in mega late-stage rounds, seed is diverging into a distinct asset class—defined by hands-on founder co-building and differentiated support, versus finance-approach growth VC.
- On AI, Steve is pilled on multi-model agent platforms (blending open + closed models) in many domains outperforming single-model stacks long term, with strong signals from companies like @Perplexity and @Cursor. At the same time, inference demand is exploding—driven by generative media (@fal), LLM agents (@parasail_io), and emerging embodied agents (robots and self-driving cars) — creating sustained investment opportunity in AI infrastructure.
- With IPO markets reopening in a record breaking way (OpenAI, Anthropic, Databricks, SpaceX in focus), expect a sharp uptick in M&A: newly public labs and platforms will acquire specialists, while older incumbents race to transform through acquisitions.
Thought provoking discussion about investing early, before consensus has been built, and the areas of early stage investment opportunity for investing in the layered stack of AI today!
We are investing in the age of intelligence, and it is a privilege. Grateful to our LPs for trusting us, and to the teams we back for leading the way to the future.
https://t.co/XFieSzt4sF
Major congrats to my partner @stevejang on his 3rd @Forbes Midas List appearance!
He’s been all-in on AI since 2019—tracking computer vision use cases, the rise of GANs and then diffusion models, AI hardware, and how early language models would reset the world. By 2022, when we went all in, we’d already spent years studying, investing, learning the hard way, and building. We don’t organize around chasing hot deals; we organize around deep, thematic research and portfolio support. That sometimes means exploring sectors until the insight and conviction are earned, then betting with high energy on that work.
That’s how investments like @Perplexity happen a few months after their launch, or @fal before “generative media” (coined by @Burkay Gur, @Gorkem Yurtseven, and Steve in our office) was even a thing. Before this cycle, investments like @Uber, @Coinbase, @Color, and @tonal (all early contrarian bets at the time) were the inspirations for his style. I’m biased, but Steve’s combo of product obsession and real go-to-market help is rare; most early-stage folks only have one of those gears.
At @KindredVentures, we love the early stage. It’s still where the biggest impact in venture gets made, and where we’ll keep doing our work. Grateful the market recognizes the craft here, and even more excited about what’s next. Huge congrats to Steve, and thank you to our founders for the mountains you move every day.⚡️⚡️⚡️
https://t.co/19e5n3VAag
we are excited to continue supporting @UseCorgi with their Series B as they work to transform insurance.
Insurance is 12% of GDP, the largest words-based industry on earth, and incredibly hard to start-up in, given regulatory and operating complexity.
Phenomenal achievement from @InversionSpace, one of a very small group of new technology startups building on this hyper critical space-based interceptor program.
One of the most important missions for the future –– and the present –– of U.S. and allied capabilities in space.
We’re in the earliest days of world models - for autonomy, robotics, entertainment, and much more.
Thrilled to share our investment in @overworld_ai, founded by ex-@StabilityAI researchers building open real-time world models for virtual worlds/games:
https://t.co/BvXYL6pqTL