I didnt give dots any permissions or connect it to any of my accounts, but dots agent passively found access to my inbox and started reading my personal emails and started working on reply drafts.
Horrifying privacy violation @OpenAI!
Anthropic has learned how people respond to model release cycles.
Intelligence is hard to measure, which is why most of us rely on benchmarks to make judgements.
In my observation, the response to a new model depends mostly on:
1. Branding (an Anthropic release means it's a better model than a no-name Chinese model)
2. Launch timings
3. Hype and build up (what your lineup looks like, how you hype it up through leakers etc., and how well you manipulate discourse)
4. Model size and cost (a bigger, known model means people think it's better)
5. Benchmark scores (benchmaxx bigger models = mogmaxxing, benchmaxx smaller models = benchmaxx)
6. Fancy demos (just get people creating fancy demos of new capabilities on X)
Actual usefulness and intelligence come after that. The model often just needs to be good enough for the workflows people care about.
That is why capabilities and scores are converging rapidly and whatever you release today will be matched in 2 months at a fraction of cost in terms of being replaceable on same workflows.