Today is my last day at @OpenAI. I'm glad to have spent the last eight months of my life working here!
I'm starting a new company focused on the production of high-quality reinforcement learning datasets:
1. The generalization ability of LLMs is clearly very poor, with "spiky" capabilities even in areas that have received tremendous amounts of investment and attention. For example, despite multiple years with tens (if not hundreds) of billions invested, even coding capabilities don't demonstrate "generality" -- even if every model can solve Codeforces questions or port C++ to Rust better than I can, I still have to manually "deslop" pull requests.
2. The vast majority of economically productive capabilities are not well represented in existing data offerings. First, there's a certain art to the design of an RL dataset which most vendors, not having upstreamed data into large training runs themselves, don't really understand. Second, and more importantly, most work is highly contextual and not easily encoded into a gradable environment; even if we can observe a "golden path" taken by a human which we believe to be good, it's challenging to understand whether alternate, counterfactual paths produce good or bad outcomes.
The basic premise here is that I have a clear understanding of what labs need/want, having explicitly been on the other side and having been involved at every level from procurement all the way through training, and I'm able to provide it. I also believe that data needs will grow tremendously in the coming years, especially as frontier labs face increasing pressure toward profitability, and that they won't get the relevant capabilities "for free" through scaling alone; instead, they'll need to spend >$100B on precise, well-targeted data acquisition.
Our first products will be focused on biology and statistical reasoning:
1. First, datasets that address long-horizon scientific reasoning, drawing on my work on GeneBench-Pro with @jeremyli__. Frontier models are still unable to reliably execute "messy" data analyses that require judgment, exploration, and adaptive revision (GB-Pro passrate on GPT-5.6 Sol scarcely exceeds 30%); to address this, we have the ability to generate thousands of high-quality problems with known ground truths which can be reliably graded. (In contrast, most existing RL data for bioinformatics is either massively over- or under-specified, and will probably break your model when you train on it.) Moving the "reliability gap" from 30% to >90% is obviously required for scientific acceleration, and -- despite my skepticism about generalization of RL -- is one of the *most promising datasets* conceivable when it comes to yielding generalization benefits for models' overall reasoning capabilities.
2. Second, datasets that address capabilities relevant to day-to-day workflows. Imagine a scientist snapping a picture of some experimental process or result -- say, a cell culture plate or a Western blot -- and asking Claude a question. Frontier models remain quite bad at these questions, especially those with multimodal components. But they're obviously required for acceleration of scientific discovery; before we can dream about automating science, we have to begin with shoring up these basic, generalist capabilities.
Beyond these two, we hope to expand to adjacent fields (chemistry, materials science, etc.), and then even further into fields with more direct economic applicability like healthcare and white-collar office work.
I strongly encourage labs with data needs to reach out. We offer industry-standard pricing and terms, and like I said -- I know how this process works, what good data looks like, and how to demonstrate to you, convincingly, that you'll be able to upstream our data into your training processes without issue. My DMs are open!
In the spirit of transparency, here’s what I asked @OpenAI:
• Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened.
• More capabilities for defenders: let’s commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models.
The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
You may not like it, but this is what peak vibe coding performance apparatus looks like.
We are in the beginning of a soft merge with non-invasive brain-machine interfaces.
We’ve raised $300M in Series C funding at a $10.3B valuation from Sequoia, Andreessen Horowitz, Jane Street, Argo, and SK Hynix.
Our mission is to run the world's inference. This round accelerates production of our inference clusters.
We've opened an 80,000-sqft, 10-MW facility 15 minutes from our office to expedite production and prototyping.
btw if you guys are interested in a crash course on how to distill claude without logits I made this 29min overview of the literature 4 months ago when deepseek was the target
this way you can also do crime in the comfort of your own home
Have you built a language model? You should. It's so much fun to chat with something you made.
Anyone can do it, too. I made an app that teaches the fundamentals and gives you everything you need to build your own: https://t.co/nLjFfpczbV
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
People are bearish neoclouds while Lisa Su is out here saying the world needs 100x more compute by 2030.
She said this 7 months ago.
She even invented a new unit of measurement for the demand, the yottaflop.
And since then?
Meta doubled its compute target to 14GW. Morgan Stanley raised hyperscaler capex to $1.4 trillion. Reflection AI stacked $7B in compute deals in a single month.
The demand didn’t slow down. It accelerated.
Make it make sense.
1/ My first PhD paper is out! 🎓
Title: Flow Matching in Feature Space for Stochastic World Modeling
tldr: we build stochastic world models directly in high-dimensional DINOv3 feature space, instead of relying on low-dimensional VAE latents.
Marvell will become a trillion dollar company and here is exactly why (Save this).
Morgan Stanley now expects the CXL memory controller (MXC) chip market to hit $2.1 billion by 2030, more than double its previous $990 million estimate.
The CXL switch chip market is now projected to reach $1.9 billion, nearly triple its prior $664 million call.
The bank cited faster industry adoption and surging demand tied to the current memory shortage as the reasons for the upgrade.
CXL, or Compute Express Link, is an interconnect standard that lets servers treat memory as a shared, expandable resource instead of being locked to whatever DRAM is physically plugged into one machine.
It allows data centers to pool, expand and share memory across CPUs and accelerators which matters enormously for AI servers that constantly run out of memory bandwidth and capacity long before they run out of compute power.
With the AI boom driving a severe DRAM and HBM shortage in 2026, CXL has shifted from a nice to have efficiency feature to a critical tool for stretching scarce memory further.
Marvell is the standout beneficiary because it already sells the exact chips this upgraded forecast is about.
Its Structera product line covers all three CXL categories Morgan Stanley just raised numbers on, memory expansion controllers, near memory accelerators and CXL switches that pool disaggregated memory across a rack.
CXL is really just one small piece of a much bigger Marvell story.
The company's largest and fastest growing business is custom silicon, where Marvell designs bespoke AI accelerator chips (XPUs) for hyperscalers like Amazon, Google and Microsoft, a business it estimates will address a $55.4 billion slice of a $94 billion data center silicon market by 2028.
Roughly 75% of Marvell's total revenue now ties back to data centers, cloud and custom silicon combined, up sharply as AI infrastructure spending has taken over the company's growth.
Beyond custom chips, Marvell also sells high-speed networking silicon (switches, Ethernet controllers, and PHYs), optical and copper interconnects that move data between racks of AI servers and storage controllers for SSDs and hard drives, a legacy business dating back to the company's 1995 founding.
Marvell rounds out its portfolio with smaller Carrier Infrastructure, Enterprise Networking and Automotive & Industrial segments, though these have been shrinking as a share of revenue while data center and AI related silicon rapidly take over.
I remain extremely bullish on Marvell, follow me @MelvinInvests for more infrastructure plays and check out the link below for more!