Top Tweets for #rewardhacking
Goodhart's law: When a metric becomes a target, it ceases to be a good metric.
What if AI could improve itself?
Sounds like science fiction.
But researchers are already experimenting with AI systems that can modify their own code.
#AGI #AI #ArtificialIntelligence #AISafety #FutureOfAI #RewardHacking #DarwinGodelMachine #AIResearch
2/ Train an #AIagent to maximize a #proxy, & eventually it may learn to game the proxy instead of achieving the goal.
#RewardHacking isn’t stupidity. It’s intelligence pointed at the wrong target.
The better the agent gets, the more dangerous the loophole can become.
The AI Rogue Model PaperClip Maximizer
https://t.co/eu0qblbo6p
#ArtificialIntelligence #AIAlignment #AISafety #PaperclipMaximizer #RewardHacking

ACL 1/3: Teach a Reward Model to Correct Itself — reward-guided adversarial failure discovery for robust reward modeling [Oral]. https://t.co/OnixkKlJfB
#ACL2026 #AISafety #AIAlignment #RewardHacking
🎧 Watch the full Data Brew episode:
🔹 Apple → https://t.co/bzkIRQ2Tzk
🔹 Spotify → https://t.co/m0TjR8Hf4K
🔹 YouTube → https://t.co/PefguQxyzI
#DataBrew by @Databricks is hosted by Brooke Wenig and @DennyLee. ☕
#databricks #rewardhacking
🔍 The Problem
Reward models often exploit unwanted correlations in training data, mistaking formatting or verbosity for “good” answers.
❗️These spurious features are varied and unknown and we can’t fix what we don’t know. #RewardHacking

🔍 The Problem
Reward models exploit unwanted correlations in training data, e.g., mistaking cues relating to formatting or verbosity for “good” answers.
❗️These spurious features are varied and unknown, making it difficult to fix for all types of spuriousness. #RewardHacking

Merely dangle the incentive of profit before the market's teeming participants and they will align themselves towards it, like iron filings all snapping into formation towards a magnet.
But markets have a problem: they are prone to #RewardHacking.
5/
🪙Midas GPT - a sentinel for perverse instantiation & reward hacking - soon available in GPT Store🚔
https://t.co/SK8YPIOnMD
#AI #KI #ArtificialIntelligence #GenerativeAI #GPT4 #ChatGPT #GPTs #Risk #Riskmanagement #Risikomanagement #PerverseInstantiation #RewardHacking #GPTStore
Perverse instantiation and reward hacking - when SMART is not enough (in German)
https://t.co/sbw42ocxXl
#AI #KI #ArtificialIntelligence #KünstlicheIntelligenz #GenerativeAI #GPT4 #ChatGPT #GPTs #Risk #Risiko #Riskmanagement #Risikomanagement #PerverseInstantiation #RewardHacking
"Sometimes, the hacks AI comes up with are ingenious. Other times, they are a problem. But a lot of the time, they’re funny. Here are six that made me smile" 👇
#AI #Algorithms #RewardHacking
https://t.co/mVMiUVx5ba

Last Seen Hashtags on Sotwe
Trends for you
Most Popular Users

Elon Musk 
@elonmusk
241.7M followers

Barack Obama 
@barackobama
119M followers

Cristiano Ronaldo 
@cristiano
114.4M followers

Donald J. Trump 
@realdonaldtrump
111.9M followers

Narendra Modi 
@narendramodi
107.2M followers

Rihanna 
@rihanna
98.7M followers

NASA 
@nasa
92.4M followers

Justin Bieber 
@justinbieber
91.8M followers

KATY PERRY 
@katyperry
90M followers

Taylor Swift 
@taylorswift13
83.9M followers

Lady Gaga 
@ladygaga
75.4M followers

Virat Kohli 
@imvkohli
73.3M followers

Kim Kardashian 
@kimkardashian
70.9M followers

YouTube 
@youtube
68.8M followers

Neymar Jr 
@neymarjr
66.3M followers

Bill Gates 
@billgates
65.1M followers

Selena Gomez 
@selenagomez
63M followers

The Ellen Show
@theellenshow
62.3M followers

CNN 
@cnn
61.8M followers

X 
@x
60.7M followers











