@thsottiaux Just sent a DM. I use ChatGPT Work extensively for a specific search workflow. Great speaking with you few mths back at the Codex hackathon. Happy to help test!
You might have noticed that the multi-agent feature was also released in beta for the Responses API with GPT-5.6 models: https://t.co/ER717WM9OF
It's now supported in our Agents SDKs as well!
https://t.co/GYlD1BFdwl
As engineering, product, design, DS, etc. melt into a new kind of role, I was reflecting on what roles might look like in the future. For example, when I look at the Claude Code team I see what I think is five archetypes:
1. Prototyper: comes up with brand new ideas; churns out many ideas, most of which don't ship
2. Builder: quickly turns a prototype/idea into production-grade product/infra
3. Sweeper: cleans up the UI, simplifies the code and system, unships, optimizes performance
4. Grower: takes a product that has been built and iterates on it to improve Product-Market Fit
5. Maintainer: owns a mature system to make it secure, reliable, fast, and efficient as it scales
Many people span across 2 roles, and sometimes 3 roles. I also notice that these roles are not really tied to job function -- eg. across Anthropic, some designers match category 1, some 2, some 3; same for engineers, PM, DS.
A healthy team needs a mix of these, depending on the product:
- A product that is new and pre-PMF needs people that are strong at 1+2+3
- A product that is growing and has found PMF needs 2+3+4 and some 5
- A product that has strong PMF needs 3+4+5 and some 2
Maybe product roles of the future will look more like this, and less like the domain-specific roles of today?
@willccbb always been a huge fan. also just pushed this on environments hub. would be great to chat sometime in SF abt applied RL in finance and legal domains.
https://t.co/H8DCejwgF0
https://t.co/NDLrThWYvS
specialization beats generality (when you own the eval). prime intellect + morgan stanley fine-tune open qwen models for the Q language (finance) and surpass frontier models on their Q test. models + code are open!
A super long overdue (3+ years?) post on scaling laws.
Compute is expensive. Scaling laws are a way to help us reason about the optimal compute allocation between data and model size before committing to a large run.
The post covers what scaling laws predict, how compute-optimal allocation works, why Kaplan et al. and Chinchilla disagree, and how data limits + fitting details make extrapolation tricky.
https://t.co/HP26eJvjHB
PPO had a second wave in the LLM era for reasons unanticipated by the original paper
- the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise
- the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO, https://t.co/sBo9DeFS5Y)
Today we're launching Exa Agent: Opus/GPT 5.5 quality web research at 2-10x lower cost.
It's our most powerful endpoint, particularly good at deep research and list-building - as cheaply as possible.
Basically you can now use a deep research API for close to the cost of a search API.
As the world wrestles with the cost of big models, we'll keep pushing the price/performance frontier 🫡
@GregKamradt@steipete would love to join. i've been building with and contributed to the oai agents-sdk, would be excited to chat about it. also keen to catch up on the ARC presentation from devday last yr. Sounds fun.
If you’re getting started with Codex Mobile, @Dimillian’s guide is worth a read!
Your phone isn’t a tiny terminal. It’s a control center for directing and reviewing work running on your laptop or in the cloud.
Thomas helped build the experience, and it shows.
Awesome to see this innovation in text diffusion. DiffusionGemma is lightning fast, 4x faster than other Gemma 4 models! Congrats to @bodonoghue85 and the team who worked so hard on this - excited to see what people build with it!
We are releasing River API, our first product, in early access. The API gives you access to the same battle-tested tools that we’re using internally at River for post-training, reinforcement learning and continual learning. Check it out and let us know what you think!
Training a generator with PROPEL roughly doubles the yield of goldilocks training tasks (not too easy, not too hard) across three domains: 2.4x on code induction, 1.7x on math, and 2.0x on a 27B SWE agent.
@WilliamBryk Speaking of Einstein, I love using Exa for scientific search and built an app with the api (bit of a mix of scholar and openalex). Think that a lot of these researcher-oriented features could be native. Happy to chat more or share user feedback if helpful.
*When you think of AI, think of humanity too*
This was the main message of the GStar Summit on AI & Humanity that I organized two days ago (as a breather after Google I/O). For a long time, I was deeply focused on advancing model capabilities toward superhuman reasoning, and I found immense joy when we accomplished our goal, e.g., our IMO-gold milestone in July 2025. I always assumed that someone, somewhere, would just take care of the social impact side for me.
It really hit me in Feb 2026 when our math research agent, Aletheia, solved FirstProof problem #7 by flawlessly utilizing heavy mathematical machinery, as noted by our surprised mathematicians. I started to wonder: what's left for humanity?
Discussing human values with my wife, Wendy Nguyen, and our dear friend Jean DeSombre, I realized I hadn't been thinking much about humanity while actively charting the course of frontier AI. But as our conversations continued over the past few months, I noticed myself becoming more "humanity-aware."
For example, during a team lunch at Google DeepMind, we were chatting about what we would do after we "solved" AI for math. I heard a number of suggestions, e.g., robotics or world models, but nothing explicitly centered on humanity. I told the team that we need to think about empowering humans: one way is to build a model capable of asking questions, from basic paraphrasing all the way to forming conjectures, which would be incredibly useful for both education and model training!
At GStar Summit 2026, aside from the deep dives into agentic AI with @edchi, @YiTayML, @preslav_nakov, Phong Nguyen, Myungsub Choi, Noriyuki Kojima, and Hung Bui, we dedicated half of the summit to talking about humanity!
@PoShenLoh, awesome as always, proposed the "Thought+Full" philosophy and reminded us how people find joy in helping others. Jay Kim shared with us new perspectives about human health in space, showing that space is not that scary to be a part of. Together with other speakers and panelist, Wendy Nguyen, Jean DeSombre, Marc Woo, Laurent El Ghaoui, Tuong Nguyen, Tuoc Huynh, Tuan Cao, and @CurtisSChin, we discussed all aspects of humanity in the age of AI.
Our hope is that everyone who walked out of the summit can now find joy in talking and working with each other on human values (either before or after AI)!
Thanks to everyone for coming from all over Asia Pacific and Silicon Valley! Many people told me that it's very rare in Vietnam for people to stay from 8am - 6pm with a fully-packed hall of 1,000+ people!
See you again soon!