sigh.. i have to say something here
imagine you hired a new developer to your company, and on day one he did some terrible work, over-engineering your codebase, speaking jargons all day without context, making all kinds of “genuine mistakes”, lost trust with everyone around him
and then he goes on and tell you - in order for me to do a better job, you need to delete your entire company’s culture and workflow, and have everyone bend over to do things his way, only then can he do reasonable work
oh - and no one tells him how to do his job, NO ONE. he’s always the smartest one in the room and despite doing a terrible job on day one, despite the only results on his resume were vibe coded demos of games that already existed, he demands that you give him a big charter and let him go dark with no communication, taking no feedback
would you have hired a teammate like this?
when humans feel frustrated after using a model, let’s figure out how to RL the model better so they become a better teammate
don’t let a bad model RL _you_
sigh.. i have to say something here
imagine you hired a new developer to your company, and on day one he did some terrible work, over-engineering your codebase, speaking jargons all day without context, making all kinds of “genuine mistakes”, lost trust with everyone around him
and then he goes on and tell you - in order for me to do a better job, you need to delete your entire company’s culture and workflow, and have everyone bend over to do things his way, only then can he do reasonable work
oh - and no one tells him how to do his job, NO ONE. he��s always the smartest one in the room and despite doing a terrible job on day one, despite the only results on his resume were vibe coded demos of games that already existed, he demands that you give him a big charter and let him go dark with no communication, taking no feedback
would you have hired a teammate like this?
when humans feel frustrated after using a model, let’s figure out how to RL the model better so they become a better teammate
don’t let a bad model RL _you_
We’re also upgrading Auto-review in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna.
Combined with Luna’s new price, we expect Auto-review to cost about 10x less, making your agentic workflows more cost-efficient.
We’re also upgrading Auto-review in the ChatGPT app and Codex CLI from GPT-5.4 to GPT-5.6 Luna.
Combined with Luna’s new price, we expect Auto-review to cost about 10x less, making your agentic workflows more cost-efficient.
Been seeing my @ChatGPT Codex quota reduce way too fast lately. After looking at analytics, ~95% of my turns are for `codex-auto-review` (aka Approve for Me)
Just got Codex to look at my transcripts and it's used 1.74B tokens across 23.5k tool reviews over the past 6 weeks 🫨
📣 @Kimi_Moonshot's Kimi K3, an open-weight model, is now generally available and rolling out in GitHub Copilot.
The model shows frontier-level abilities on agentic coding with highly cost-effective pricing. It is hosted by @FireworksAI_HQ. Learn more. 👇
https://t.co/lI3GgiZJ5M
📣 @Kimi_Moonshot's Kimi K3, an open-weight model, is now generally available and rolling out in GitHub Copilot.
The model shows frontier-level abilities on agentic coding with highly cost-effective pricing. It is hosted by @FireworksAI_HQ. Learn more. 👇
https://t.co/lI3GgiZJ5M
Build a plugin once and use it across compatible agent clients.
Introducing Agent Plugins, an open standard developed with @awsdevelopers, @cursor_ai, @github, @code, and @vercel that packages Agent Skills and supports MCP server configurations in a shared format.
Introducing Alpamayo 2 Super: the frontier open reasoning VLA for autonomous vehicles and robotaxis.
34B parameters. Full-surround awareness. Available for commercial use under under OpenMDW-1.1.
Download now on @huggingface → https://t.co/pyyW2NP4BW