I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.
OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.
> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.
最後の一文すごいですね...
https://t.co/LdYGF43KSE
some investors have been asking me lately about the wave of inference chip startups. My answer: we need real performance data before forming a view.
So here’s ours. Across three public models on InferenceX, Jalapeño delivered:
• 1.5–1.9× more AI work per watt
• 1.7–3.6× lower end-to-end latency
• 2.1–4.1× higher performance on highly interactive workloads
When I joined OpenAI, Richard had a small team and was just getting started. very proud to be his first strategic finance partner and what the team has built.
https://t.co/QFOvCDLPZX