@Haleeeemahh@AfnadX In their “selected benchmaxxed benchmarks”
Gemini always has been high scoring in chosen benchmarks and unreliable in practice.
If it is so dangerous in cyber, what is its ExploitBench score?
Keeping defending vaporware
@SadAlbert10 Seeing Google employees hype and cringe post AI-written HR speak while being so disconnected from the reality of their public perception is infuriating. They keep putting their benchmaxxed numbers in front of everybody’s faces. Where is the model?
@bcherny Prompt engineers are copers and just a window for people who have no other skills. It was evident as models got smarter you don't need to throw up the third known gang sign from mesopotamia to unlock a "hidden reasoning model"
I am surprised that people are surprised this is how I prompt Claude.
Talk to Claude the way you would a coworker. There's no secret to prompting. There's no need to be overly scaffolded or prescriptive for most tasks -- give Claude a goal, and it will figure it out.
Back in the Sonnet 3.5 days, your prompt mattered a lot. Nowadays, it's much more important to communicate to the model:
1. What you want it to do
2. How much effort you want it to spend
3. How it should verify that it did the right thing
@Chaos2Cured@AnthropicAI I don't think it works like that, the internet isn't clash royal where you drop a level 8 pekka, your opponent does too and both cancel each other out. Offense is always at an advantage in cybersecurity
@Chaos2Cured@AnthropicAI Yeah but tbh that would also increase chances of other script kiddies with similar tools attacking you or anybody.
I mean I am all for normal AIs to be shared broadly amongst people, concentrating that to few would cause a bad power imbalance but this is inherently risky