It is hard to say that scaling parameters is the only path towards AGI/ASI. In the age of RSI, the pace to improve a medium-sized model is much faster.
Excited to release 3.8 Flash following our 3.7 release just 3 weeks ago! It’s a substantial jump in agentic, coding and cyber capabilities. Lots of ingredients came together for this launch and we are so excited for what’s to come! Can’t wait to see what people build with it.
Glad that others are finding Amplio helpful. It is our main harness for autonomous long research runs spanning days to weeks. It is heavily used by our team for automated AI research / RSI on top of the Simply codebase.
Open sourced at: https://t.co/GqlmGVMnD4
Today we're launching Gemini 3.7 Flash - our latest workhorse model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. ⚡️
We have been iterating rapidly with the Flash series, going from 3.5 to 3.7 in just 3 months, making it more helpful across a wide range of tasks:
• Software Engineering (DeepSWE v1.1): 37.0% ➔ 65.3%
• Web Development (Code Arena Elo): 1506 ➔ 1588
• Enterprise Automation (AutomationBench): 13.4% ➔ 30.4%
@Yuchenj_UW Disagree. Startups with fresh talents + compute will become new frontier labs. Any organization will get entrenched over time and eventually stalled on innovations. There is no evidence that anyone is the exception so far.
@karpathy Very inspiring as always! We are also open sourcing part of our infra on automated research for Gemini to evolve itself at https://t.co/WH7JBEEm9h More complex than the nanochat setup but closer to SOTA LLM pre/post-training while staying as minimal as possible. More on the way.
I saw a guy coding today.
Terminal 1: Claude Code
Terminal 2: Codex
He typed the same prompt into both.
Then stared at the screen for 60 minutes.
Like a psychopath.
Opened Cursor.
Read the 10k lines AI generated.
Like a Costco receipt checker pretending to verify every item.