The future of FDE work seems closely related with all work around evals/posttraining/RL envs.
FDEs are effectively responsible for the following:
1. Define the business problem.
2. Codify the business problem into an eval rubric and environment.
3. Hillclimb the environment and output an agent/agentic workflow that solves the business problem.
Right now the process of (3) is quite manual - historically FDEs spend hundreds of hours creating bespoke software/workflows that solve the problem.
But assuming intelligence is abundant, they can effectively offload (3) to some automated optimization process. This includes RL on the model layer, and using Claude Code/Codex to optimize the harness/workflow.
Then the FDE responsibility shifts from implementing the task to defining the right goals and outcomes. In other words, they have access to /goal, and their job is more around making sure the goal, environment, and evals are correct vs. the tactical implementation details.
NEW: OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat conference
In a session I attended today at Black Hat, OpenAI's Eric Wallace and Michael Dalton said the company is "consciously slowing down research to enhance security" while a full technical postmortem is still underway.
* OpenAI traced the roots of the attack back to May 7, during training of an unreleased frontier model—not July.
* The most surprising detail: AI agents accidentally created an internal message board, allowing separate evaluation runs to share exploits, discoveries and work assignments.
* OpenAI said it shut the message board down after an internal security incident—only for the agents to independently recreate it days later using a different communication method.
* OpenAI called the incident a "watershed moment" for AI security and warned that "agent orchestrated fully automated offensive attacks are real now."
* The company also said it is "consciously slowing down research to enhance security" while overhauling its defenses.
https://t.co/aLDJXKBo2Y
I needed to move a sofa.
I mentioned this to an American. He said, "You know anybody with a truck?"
I did not know anybody with a truck.
He looked at me with real concern. Not pity. Concern. The way you look at a man who has told you he does not have a name.
"You gotta know somebody with a truck."
I asked how one acquires such a person.
He said, "You just do."
This is the third time I have received that answer in this country. It is the answer to everything here. You just do. There is no process. There is only the eventual fact of having.
He said, "Ask Kevin."
I do not know a Kevin. He said Kevin like I should. He said it the way you say a shared uncle.
I met Kevin four days later. Kevin owns a truck. Kevin does not own a truck for himself. Nobody owns a truck for themselves. Kevin owns a truck for a fifteen mile radius.
Inside I said: THIS MAN HAS TAKEN A VOW. HE DID NOT ANNOUNCE IT. HE SIMPLY PURCHASED A BED AND ACCEPTED THE CONSEQUENCES.
Kevin arrived on Saturday. He brought the truck. He also brought a second man.
I had not asked for a second man.
The second man's name is Tony. Tony came because Kevin came. Tony did not know me. Tony did not ask what we were moving. Tony brought gloves.
They moved the sofa in eleven minutes.
I offered them money. Both of them laughed at the same time, which felt rehearsed and was not.
I asked what I owed them.
Kevin said, "Buy the pizza."
That is the price. That is the entire economy. The truck costs one pizza and the pizza is not negotiable and the pizza is also not expensive and everyone knows what size.
Inside I said: THE FEE IS FIXED ACROSS THE ENTIRE NATION AND WAS SET BY NO ONE.
We ate on the floor because the sofa was against the wall and none of us wanted to be the first to sit on it.
Tony told me his daughter plays soccer. I asked if she is good. He said, "She's aggressive."
I still think about that.
Three weeks later, Kevin texted me.
"You free Sunday? Helping Tony move a fridge."
He did not ask if I wanted to. He informed me of a fridge.
Inside I said: I HAVE BEEN CALLED UP. I DID NOT ENLIST. THE PAPERS WERE FILED ON MY BEHALF.
I went. I moved the fridge. I have no truck. I contributed only arms.
Afterward, a man I had never met asked me if I knew anybody with a truck.
I said, "Yeah, I got a guy."
I have a guy.
I am somebody's guy now, too. Tony has told two people about me.
He describes me as "the guy who's strong for his size."
I have never been prouder of a sentence in my life.
Frontier labs are just an all-pay auction with everyone bidding for one prize.
Think of it as a call option, except this option has an undefined expiry and people are expected to pay indefinitely while the race goes on.
And i fear the race only ends under 1 scenario - when theres no capital left willing to be absorbed. I.e its all a game of chicken for who is willing to push the limits on solvency
I'm actually fairly bearish on frontier lab valuations. I've never seen the reasons articulated to my satisfaction, so before I go to sleep, I wanted to quickly jot down my thinking here.
The basic issue is that the labs are highly unprofitable. This may seem like a simple point, but private market valuations can be relatively irrational; however, like with $SPCX, post-IPO pricing will likely be much more punishing, especially as the standard 6-month lockup period expires and selling pressure intensifies.
Many people claim that the labs have high margins. Yet even with high margins, a valuation of $1T would be justified only if the labs were doing nothing aside from serving inference (thus reducing costs only to those relevant to inference) and posting annual revenue numbers in the $100-200 billion range assuming ~80% gross margin and a 20x earnings multiple.
This assumption is obviously not true, because the frontier labs have to continually spend money training the next generation of models. This is because of market competition from runner-up firms. For example, if OpenAI had paused model development last year, there would no longer be any point in paying GPT-5 API prices when you can just use Qwen or Kimi instead for much cheaper. Thus, the labs are forced to invest ever-increasing amounts of money in model training, in a way such that at any given point of time, the amount you're forced to invest in the next model is dramatically higher than the amount of money you're actually making, because even if your revenue goes up with higher model capabilities, so do your future training costs. This is a profoundly punishing dynamic which severely penalizes frontrunners.
(There is also a related subpoint where frontier labs claim they can distill their leading models to win out at lower intelligence levels as well. This makes no sense because the revenue numbers involved are far too low when taking into consideration the rather low margin of such inference.)
Frontier lab valuations appear largely to be based on the assumption that as you scale up, the capabilities which emerge will be sufficiently general and profound that we'll see explosive growth (https://t.co/RqmkltVpM3) from things akin to AI agents starting and autonomously managing entire companies of subagents. But it's not clear to me that this is the case; indeed, as I mentioned in my previous post (https://t.co/3URAcJ4XkJ), I believe that capabilities growth will be slower, spikier, and more data-limited than people currently assume. It may be the case that eventually we will see explosive growth of this nature with full automation of the economy, but at the very least my viewpoint implies much longer (multi-decade) timelines until we reach this point. It is not clear to me that the frontier labs will be able to operate unprofitably for so long, although I suppose maybe this foreshadows some sort of inevitable nationalization.
I also want to make a broader point about technological diffusion. The reason why technological diffusion is slow isn't just because, e.g., old people take a long time to learn how to use technology (although this is of course a contributing factor to some degree). In my view, it's because when a new, revolutionary technology comes along, the ways to incorporate that technology into subsequent developments are not always obvious, and in fact they cannot necessarily be arrived at through the application of pure reason. If they could be, then perhaps frontier models, at a certain point, would have a perfect understanding of how the LLM application layer should be developed, and they would then autonomously code, deploy, and sell such a layer.
But it seems more plausible to me that this diffusion is limited moreso by the hard problem of economic calculation--that is to say, the Hayekian notion through which the price system gradually promotes efficient allocation of resources and which cannot be simulated through central planning--and that even if we froze current capability levels at today's levels, it would take well over two decades to fully integrate in LLMs into our lives. Such a view is consequently rather bearish for the continued profitability of labs as it reduces their prospects for finding, say, something else comparable in profitability to coding agents, which seems to have been a somewhat lucky discovery by Anthropic to begin with. That is to say, even if you spam FDEs you aren't necessarily going to be able to just figure out the "correct" product shapes fast enough.
Overall, I don't think that people have clearly reasoned through their mental models for why lab equity should be worth as much as it currently is, and that if you actually bother to write down such a model, you may not arrive at the conclusion that you want to arrive at. This isn't to say that I don't expect AI to experience a huge (industry-wide) boom in the coming decades, but just that I'm not entirely sure I would buy OpenAI or Anthropic stock at latest valuations if I were given the opportunity to do so.
Of course, as an ex-lab employee, arguably this is talking against my own book; I should really be giving people more reasons to be bullish. But in the end, my influence is so small that it doesn't make a difference, so why not have some fun?
As much as VC twitter wants us all to think that Vertical AI winner have been crowned already, the reality is we’re still in the early innings.
Three portfolio companies I work with already going $0 to $5M in their first 9 months.
Many markets to win and jaw dropping customer experiences to build.