@bitcloud It happens every single time. It could also have to do with RL data. They get enough power AI responses to improve their system using RL and so dumb it down to save costs and optimise the next model.
Most people hate the fact they use their phones so much, yet they still do.
Mark saying people probably don’t want the physical world to be cluttered when wearing his glasses is beside the point.
"Maybe in the future, my AI girlfriend is on the other side of the screen or something."
Mark Zuckerberg responds to Dwarkesh's fears of getting reward-hacked by AR.
"Maybe in the future, my AI girlfriend is on the other side of the screen or something."
Mark Zuckerberg responds to Dwarkesh's fears of getting reward-hacked by AR.
The weirdest thing about AI assistants is that they consistently get worse over time, not better. I watched Claude 3.5 go from impressively thoughtful to increasingly lazy in just a few months.
At first it would provide detailed, nuanced responses with careful reasoning. Now it defaults to minimal bullet points even when the task clearly requires depth. And nobody seems to acknowledge this regression.
What's happening behind the scenes? My suspicion is that it's a combination of cost-cutting measures, training drift, and optimisation for the wrong metrics. The models get "aligned" toward minimal viable outputs rather than maintaining their initial quality.
The most frustrating part is the inconsistency.
Sometimes you get the brilliant assistant you remember, other times you get a glorified list-maker that seems determined to do as little work as possible.
Meanwhile, Claude 3.7 coding in Cursor has quietly evolved in a fascinating way. It's now much better at respecting task boundaries than before.
A few weeks back there was this weird phenomenon where it would eagerly solve problems you explicitly asked it not to tackle. Either Anthropic made adjustments or Cursor modified something, but the improvement is noticeable.
It now behaves more like Claude 3.5 in terms of following instructions correctly, while maintaining its expanded capabilities.
When will they solve these quirks as I want consistency.