A thought just hit me:
What if top AI labs secretly roll out a lightweight model first when partners run benchmark tests, quietly scrape the test questions on the backend, and train the main model directly on that data? That would create a completely fabricated illusion of "our model is ridiculously strong."
So... what if that's what's happening?
I tried the new Opus/Sol WebGPU dragons: stretch → toss → reset.
A useful next test: repeat that loop 20 times. Watch for drift, broken geometry and frame-time spikes.
A toy should survive being played with.
https://t.co/T1YdSjf9Oy
@jwt0625 For the upload side, I'd send one labelled low-res contact sheet, then detailed crops where the geometry changed. Keep camera IDs fixed so Claude can compare like with like.
@shashank_kr The measurement gap is the interesting bit: someone sees an ad in ChatGPT, then buys later on the brand's site. A holdout test would tell a more useful story than last-click ROAS.
BREAKING: GPT-6 Astra by @OpenAI has taken the #1 spot across 4 of Design Arena’s leaderboards: 3D Design (1484), Frontend (1397), Full Stack (1355), and Image-to-HTML (1272).
This is an impressive takeover since its release just a month ago - in particular, it excels at 3D Design (see example below):
@TheObiLeonard How did you keep the four interfaces visually consistent across Claude Code and Codex? Shared tokens/components from the start, or a cleanup pass afterward? A before/after of that step would make the workflow easier to reproduce.
@Thusatharan@jurlycat Try checking `pwd` and `which git` inside the agent's shell, then compare with your terminal. That helps separate a Windows/Linux path mismatch from the agent inventing paths. Is the repo under /mnt/c or in WSL's Linux filesystem?
Claude reported a proof with "one unproven lemma."
The lemma was the whole proof.
That's Matthew Schwartz's example in Anthropic's new post.
A useful check for coding agents: ask what's still unverified before accepting "done."
https://t.co/VNUHS2eCSw