OK, GPT-5.6 Luna is a bit of a beast. Given the 80% price drop today I decided to try it in Datasette Agent, and it's furiously quick and generates all the SQL, HTML and JavaScript (for Datasette Apps) I could possibly want
Turns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3.
Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of our canonical compaction implementation.
https://t.co/wHjaNsvIv8
Very excited to welcome @NoamShazeer to OpenAI as our new lead for architecture research! His work on transformers, MoE, and efficient decoding have shaped modern AI.
He’s extremely AGI-pilled and is super thoughtful about making it all go well. Welcome, Noam!
I’m excited to share that I’ll be joining OpenAI and look forward to working with the exceptional team there.
It was a difficult decision to move on. I’m incredibly proud of the amazing team at Google and everything we’ve built together. It has been an honor and a pleasure to work with all of you.
haven't commented on this until now but this sounds genuinely misanthropic.
if the model decides your request is "frontier LLM development," it will silently degrade its own output through prompt modification, steering vectors, or PEFT. no refusal. no notification. no fallback to another model. you just get worse work and never know why. sounds like even being bio-adjacent is enough to get limited.
a refusal is honest. it tells you where the line is and lets you route around it. silent degradation is something else entirely. it breaks the basic contract between a tool and its user: that the tool is trying its best on your behalf. this is also terrible precedent for alignment and scalable oversight. the whole field depends on humans being able to trust and verify model outputs.
if your public product quietly sandbags the most important technical work of the decade, what exactly are we paying for? 0.03% of traffic sounds small until you realize who that 0.03% is: the researchers and builders pushing the frontier. precision targeting of the people the tool exists to serve.
refuse if you must. but degrading work silently is wrong, full stop.
BREAKING NEWS: Anthropic's latest model will NOT help you if it thinks your ML research/ML engineering is interesting, and/or will secretly degrade its IQ so that the average engineer won't notice. We are already seeing Anthropic's latest model's moderation filters our GPU inference research and programming 😭
OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new way to build on Amazon Bedrock with OpenAI through the security, compliance, and governance workflows they already use.
This is also the beginning of a broader expansion of OpenAI capabilities on AWS, including future availability for cybersecurity capabilities like Daybreak.
https://t.co/vMws0YU6Q3