A few months into my Applied Scientist internship at 𝗔𝗺𝗮𝘇𝗼𝗻 now, so figured i'd share how i got in.
Applied on the portal last june and heard nothing. then a recruiter emailed me in march asking me to actually apply. heres how the interview loop went👇
@JessicaHullman yeah direction drift is the sneaky one. each agent suggestion looks locally fine so you never feel the goal moving until you're weeks off
pacing is the flattering version. you ship efficiency and mid-tier bumps when the frontier jump isn't sitting ready to go.
"we chose restraint" and "we don't have an astra-tier model queued" look identical from the outside.
Hot take: this is what pacing looks like.
None of today's releases were Astra or Fable tier. This is intentional.
The point of "pacing" isn't to stop iteration and improvement. The goal is to prevent the development bigger models from spiraling out of control.
Opus and Sol class models are a great place for our focus to go right now. Lots of opportunity for real wins without as much risk :)
everyone's mining it for secret instructions but most of 1.9m chars is just tool schemas
that's the context tax every agent turn pays before it even reads a word from you
🚰 SYSTEM PROMPT LEAK 🚰
Here's the full system prompt for Claude Opus-5.5!! The total count of everything extracted, including all tools, comes in at over 1.9M characters! 🤯
Lots to dig into here. Enjoy! 🫡
Link: https://t.co/GzISdXkbYq
gg
a model reading gnarly legacy C and surfacing a bug that survived 14 years of human review is the capability worth caring about here.
way harder to fake than a benchmark row, and basically nobody tests for it.
ten agents is the sideshow. the lean proof is the whole result.
an agent claiming a tighter bound is just noise you still have to go verify. a machine-checked one means for once you don't.
We asked ten Claude Opus 5.5 agents to devise a faster shortest-path algorithm and prove it in Lean. Within 15 hours, they produced C-HD: a formally verified improvement over the published bounds.
a per-token price cut means nothing if the model now burns way more tokens per task.
effective cost per finished job is what you actually pay, and that can rise even as the sticker price falls.
As expected, unfortunately the only bad thing about the new Opus is it's a token-gobbling monster. Uses the most tokens per Intelligence Index task of any model benchmarked.
The price cut makes more sense here!
everyone's staring at the 58 on the index but the price is the real weapon here. matching astra on intelligence was coming regardless.
opus at 4/20 undercuts astra and fable by 2.5x, that's what actually forces everyone else to respond.
Let that sink for a moment:
- Opus 5.5 costs 40% less compared to Opus 5
- performs at Fable 5.1 level
- is 30% faster in output
- and you get a banked reset on top of that.
They chose war with OpenAI. We are in such a wild race!
the interesting part of "daily workhorse" is how much of it is the build harness now
raw weights get you to competent. the scaffolding is what makes it reliable enough to actually live in all day
the cheap fast tier is the real product, the flagship is just the number you screenshot once.
three dropping under astra at once means they've stopped pretending the top model was ever the volume play.
GPT-6-Sol
GPT-6-Luna
GPT-6-Astra-Minor
All three are incoming everyone.
These are the cheaper faster ones that sit under Astra.
Hang on tight and burn all the usage while you have it.
@Steve_Yegge tpm fits first cause the job is mostly chasing who's blocked on what. agents do that fine, and a bad nudge undoes way cheaper than a bad merge