@teortaxesTex These days, I explain this "break-out" concept to non-technical people by saying itโs like a mafia boss in jail directing his goons outside.
Heโs still in jail, but can take malicious action outside of his cell ๐
@scaling01 > Astra scores 97.6% on FrontierMath Tier 4 v2, 74.1% on DeepSWE v1.1, 95.9% on BenchCAD, 96% on GPQA Diamond and 100% on ExploitBench. It also reports a 98.6% score on ARC-AGI-3.
From https://t.co/DpHgS1pvXH
"GPT-6-Astra will be available to approved cybersecurity defenders in OpenAIโs Daybreak program starting on Thursday, with plans to bring the new model to paying ChatGPT subscribers and its API over the next several days."
"Astra scores 97.6% on FrontierMath Tier 4 v2, 74.1% on DeepSWE v1.1, 95.9% on BenchCAD, 96% on GPQA Diamond and 100% on ExploitBench. It also reports a 98.6% score on ARC-AGI-3."
https://t.co/a6OZxt4uwg
https://t.co/h37nMMIVyh
@ivanfioravanti It's probably much much better than 2x DGX Sparks.
On paper, it's 2.2x the memory bandwidth and 2x the matmul compute for 30-35% more money
@natolambert On a side note, are you using Codex or Claude Code more these days?
Noticed you used Codex instead of Claude Code for this which you used to prefer.
@EmperoAI Since the architecture is shared (except for the multimodal encoder), perhaps on policy distillation can be done to have even stronger training signal than SFT?
@AlexReibman I believe the SuperGrok subscription allows usage via OAuth, and like the ChatGPT subscription, you get much more mileage (in API pricing) than the upfront subscription price.
Perhaps you could try this?
@scaling01 To be fair, Seed is still quite behind Moonshot, https://t.co/Mt9iEc9sD1, DeepSeek and Qwen in post training. It hard to see this model significantly beating what the other Chinese labs have by then.