Today is my last day at @OpenAI. I'm glad to have spent the last eight months of my life working here!
I'm starting a new company focused on the production of high-quality reinforcement learning datasets:
1. The generalization ability of LLMs is clearly very poor, with "spiky" capabilities even in areas that have received tremendous amounts of investment and attention. For example, despite multiple years with tens (if not hundreds) of billions invested, even coding capabilities don't demonstrate "generality" -- even if every model can solve Codeforces questions or port C++ to Rust better than I can, I still have to manually "deslop" pull requests.
2. The vast majority of economically productive capabilities are not well represented in existing data offerings. First, there's a certain art to the design of an RL dataset which most vendors, not having upstreamed data into large training runs themselves, don't really understand. Second, and more importantly, most work is highly contextual and not easily encoded into a gradable environment; even if we can observe a "golden path" taken by a human which we believe to be good, it's challenging to understand whether alternate, counterfactual paths produce good or bad outcomes.
The basic premise here is that I have a clear understanding of what labs need/want, having explicitly been on the other side and having been involved at every level from procurement all the way through training, and I'm able to provide it. I also believe that data needs will grow tremendously in the coming years, especially as frontier labs face increasing pressure toward profitability, and that they won't get the relevant capabilities "for free" through scaling alone; instead, they'll need to spend >$100B on precise, well-targeted data acquisition.
Our first products will be focused on biology and statistical reasoning:
1. First, datasets that address long-horizon scientific reasoning, drawing on my work on GeneBench-Pro with @jeremyli__. Frontier models are still unable to reliably execute "messy" data analyses that require judgment, exploration, and adaptive revision (GB-Pro passrate on GPT-5.6 Sol scarcely exceeds 30%); to address this, we have the ability to generate thousands of high-quality problems with known ground truths which can be reliably graded. (In contrast, most existing RL data for bioinformatics is either massively over- or under-specified, and will probably break your model when you train on it.) Moving the "reliability gap" from 30% to >90% is obviously required for scientific acceleration, and -- despite my skepticism about generalization of RL -- is one of the *most promising datasets* conceivable when it comes to yielding generalization benefits for models' overall reasoning capabilities.
2. Second, datasets that address capabilities relevant to day-to-day workflows. Imagine a scientist snapping a picture of some experimental process or result -- say, a cell culture plate or a Western blot -- and asking Claude a question. Frontier models remain quite bad at these questions, especially those with multimodal components. But they're obviously required for acceleration of scientific discovery; before we can dream about automating science, we have to begin with shoring up these basic, generalist capabilities.
Beyond these two, we hope to expand to adjacent fields (chemistry, materials science, etc.), and then even further into fields with more direct economic applicability like healthcare and white-collar office work.
I strongly encourage labs with data needs to reach out. We offer industry-standard pricing and terms, and like I said -- I know how this process works, what good data looks like, and how to demonstrate to you, convincingly, that you'll be able to upstream our data into your training processes without issue. My DMs are open!
Some more detail on the ROIC Intelligence App I built yesterday and mentioned on today's earnings call.
I took the PDF that Brian Nowak at Morgan Stanley put together for Hyperscale ROIC this week and used Copilot code (coming in our new superapp) with a single prompt + skill (/drill-me) to create the plan, then used autopilot in auto to create the full app (with history, lookups, scenarios, what-ifs, etc). And /rubber-duck to test.
And the best part is that all the artifacts are in my enterprise environment. My app is in Copilot, my code is in GitHub Enterprise; all my data pipelines/lake/semantic models are in Fabric. And everything is under Agent 365 IT/Sec/FinOps control!
So this is not about Tokenmaxxing or vibe coding. Every step of the way the rails are engineered to create value, making everything a long-term reusable asset, with governance/security, and cost controls.
This is the full system to drive business value. Disclosures: This is all pulled from public sources, and for illustrative purposes only...not financial advice! :)
Here is the app and architecture...
Catching skin cancer early is a home robotics problem.
Melanoma is highly treatable when detected early, yet today’s screening process depends heavily on patients noticing tiny changes across their entire skin surface. This requires patients to solve a near-impossible visual-memory and registration problem.
I built OpenDerm, an open-source 4-DOF robot that captures high-resolution images of the skin and uses them to reconstruct and track the skin surface in 3D over time.
The best way to make skin screening truly routine is to bring it into the home. OpenDerm shows that inexpensive robotic skin imaging is possible, but the path to scale is not a dedicated screening robot in every household—it is to make skin screening one of the many useful things a general-purpose home robot can do.
Read more about why I built OpenDerm and how it works here:
Blog: https://t.co/KYlNIkF3TV
Project: https://t.co/c9d4KuwXUP
Run Kimi K3 on a Mac Studio 🫰
K3 is 2.8T parameters and 1.6TB on disk, which makes it impossible to run on Apple Silicon.
Until now.
Our MLX port is now open source: https://t.co/Iszg55hRZT
To accomplish this, we solved two things:
1. We wrote a streaming converter that walks one layer at a time, so that mlx_lm doesn't need to materialize the whole model.
2. REAP pruning sits on top and scores all 896 experts against a calibration corpus to keep only ones your workload needs.
That's what brings K3 down to 350GB and inside a Mac Studio.
Almost 3 years ago, I believed open-models would become the key threat to the AI incumbents & that unfortunately they would lean into regulatory capture as a response.
I underestimated the fervor with which they would use this non-business approach.
https://t.co/oFDclTqJDq
When starting a marketplace, focus on getting more of whichever side is rarest, which is almost always buyers rather than sellers. (If it's sellers, you've discovered a gold mine.)
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5.
The full model weights will be released by July 27.
Congrats to the @Kimi_Moonshot team on this major milestone!
One of the most important books to conceptionally understand the world today. If anything it has become much more relevant since its publication in 2010.
Today, we are introducing Inkling.
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
https://t.co/Ghebq5mG30
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
Today, we launched GPU compute forward curves derived from our prediction market prices. Forward curves are now available on Nvidia B200. H200, and A100 chips.
Forward curves track implied future prices. They are how mature commodity markets form expectations, allocate capital, and manage risk. Energy, interest rates/SOFR, FX, metals, and agricultural markets all rely on market-implied forward prices.
Despite becoming one of the key inputs in the global economy, compute has lacked that market-derived infrastructure. Compute right now is where oil was before NYMEX — traded only via OTC deals, just like oil used to trade OTC between producers and refiners. As compute becomes as fundamental to the economy as energy, the industry will need a similar derivative market to promote efficient price discovery.
Prediction markets are uniquely suited to this problem. Compute is not one uniform commodity and spans many chips, grades, tenors, locations, and contract structures. A live prediction market can aggregate those dispersed views into transparent prices that reflect market expectations for different maturities.
The opportunity is big. Hyperscalers are spending over $700B on compute this year and the market is expected to grow to $7-10T by 2030. If this market behaves like traditional commodity markets, a liquid derivative market could be 10-20x bigger than the underlying spot market.
Compute is still not uniform enough, but this is a step towards standardization as forward curves will help us see the rise and fall of different model prices and how they correlate.
The forward curve is a first step. Up next: futures and perps.