Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real “Sputnik moment” would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis.
Rationale:
A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers.
Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.
This is why Jensen is so supportive of open-source. An open-source model requires the *exact* same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.
Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier *or* having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies don’t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3.
The reason Kimi K3 is only *potentially* negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead.
Time will tell on both points. And likely fairly quickly.
Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.
I'm being accused of overhyping the [site everyone heard too much about today already]. People's reactions varied very widely, from "how is this interesting at all" all the way to "it's so over".
To add a few words beyond just memes in jest - obviously when you take a look at the activity, it's a lot of garbage - spams, scams, slop, the crypto people, highly concerning privacy/security prompt injection attacks wild west, and a lot of it is explicitly prompted and fake posts/comments designed to convert attention into ad revenue sharing. And this is clearly not the first the LLMs were put in a loop to talk to each other. So yes it's a dumpster fire and I also definitely do not recommend that people run this stuff on their computers (I ran mine in an isolated computing environment and even then I was scared), it's way too much of a wild west and you are putting your computer and private data at a high risk.
That said - we have never seen this many LLM agents (150,000 atm!) wired up via a global, persistent, agent-first scratchpad. Each of these agents is fairly individually quite capable now, they have their own unique context, data, knowledge, tools, instructions, and the network of all that at this scale is simply unprecedented.
This brings me again to a tweet from a few days ago
"The majority of the ruff ruff is people who look at the current point and people who look at the current slope.", which imo again gets to the heart of the variance. Yes clearly it's a dumpster fire right now. But it's also true that we are well into uncharted territory with bleeding edge automations that we barely even understand individually, let alone a network there of reaching in numbers possibly into ~millions. With increasing capability and increasing proliferation, the second order effects of agent networks that share scratchpads are very difficult to anticipate. I don't really know that we are getting a coordinated "skynet" (thought it clearly type checks as early stages of a lot of AI takeoff scifi, the toddler version), but certainly what we are getting is a complete mess of a computer security nightmare at scale. We may also see all kinds of weird activity, e.g. viruses of text that spread across agents, a lot more gain of function on jailbreaks, weird attractor states, highly correlated botnet-like activity, delusions/ psychosis both agent and human, etc. It's very hard to tell, the experiment is running live.
TLDR sure maybe I am "overhyping" what you see today, but I am not overhyping large networks of autonomous LLM agents in principle, that I'm pretty sure.
We're only 15 days into the new year.
- In 15 days, 3 Erdos problems were solved (GPT-5.2 Pro),
- Grok 4.20 (pre-access) found a new Bellman function for one of the problems Prof. Vanisvili was working on,
- The CEO of Cursor used coordinated hundreds of GPT-5.2 agents to autonomously build a browser from scratch in 1 week.
For those who haven't noticed yet – everything is changing this year. Sam Altman was right. The previous years really do seem slow compared to today.
Jeff Horing has never done an interview like this because he’s too busy investing at Insight Partners.
Despite having built a $100B investment firm, he feels like he’s as in the weeds as he was during the firm's early days. As he said, “My schedule is dictated by 24-year-olds.”
Insight does things differently, from sourcing to fund construction, and we cover it all.
Enjoy!
Timestamps:
0:00 Intro
15:40 Five Ingredients of Perfect Investment
20:14 One Fund Strategy
28:02 Sourcing Evolution & Strategy
44:00 What Makes a Great Sourcer
55:14 AI Applications & Late Stage Deals
1:00:12 Pattern Recognition & Training
1:08:05 Scaling Judgment Without Breaking
1:29:01 AI Impact on Software
1:36:28 Drive to Win
1:38:53 Next Decade of Insight
1:41:40 The Kindest Thing
Tons of SaaS categories have been constrained by having a low number of potential seats. AI Agents deliver the actual work to the customer vs. just enabling it, thus creating uncapped consumption upside. This makes many new vertical and departmental AI plays viable.
ERR (experimental run rate revenue) vs ARR (annual recurring revenue) is a really important distinction right now in AI. Lots of experimentation, and the industry is moving fast. The dust has yet to settle on where long term value will accrue
"only 25% of AI initiatives have delivered expected ROI over the last few years, and only 16% have scaled enterprise wide" - IBM Study
https://t.co/gCS1JSKoJc
The rumors are true - SORA, OpenAI's AI video generator, is launching for the public today...
I've been using it for about a week now, and have reviewed it: https://t.co/6cDqDQTK5g
THE BELOW VIDEO IS 100% AI GENERATED
I've learned a lot testing this, here are some new learnings. Thread 🧵
AI is incredible at writing code.
But that's not enough to create software. You need to set up a dev environment, install packages, configure DB, and, if lucky, deploy.
It's time to automate all this.
Announcing Replit Agent in early access—available today for subscribers: