LMArena is a cancer on AI.
I was hoping it would die out 6 months ago, after Maverick showed what it gets you.
But it keeps rearing its head. The WSJ talks about how important it is. A new VP asks what their team is doing to climb it. The cycle continues.
It’s fundamentally broken, optimized for the wrong incentives:
Users spend 2 seconds skimming responses before clicking their favorite. They're not reading carefully.
They're not fact-checking. They're just picking whichever model response catches their eye.
This means that the easiest way to win on LMArena is by...
Being verbose - longer responses look more authoritative!
Formatting aggressively - bold headers look like polished writing!
Vibing - wild, colorful emojis grab your attention!
It doesn't matter if a model completely hallucinates. If it looks impressive, LMSYS users will vote for it over a correct answer.
After all, remember this and all the sycophancy issues we've seen this year?
@dvargas92495 Just getting started, but seems to me that gpt4 function calling struggles with planning prompts, e.g. think() in autogpt. Using functions literally causes worse thinking.
Curious if you've seen that too?
@razibkhan Even if they scaled enrollment and set it to perfect racial proportions...the university model would still make no sense. IMO that’s what causes the fear of change.