@Plinz Has your view on controllability changed? I thought your position was that trying to permanently control AI by tethering it to today’s human values was counterproductive and likely impossible, and instead the goal should be to build shared purpose
Fully get that Orthogonality means we can’t guarantee moral convergence, but what gives you so much confidence selection pressures don’t steer AI into broadly “valuable” futures, even those that might include humans (even if they aren’t in “control”). I’ve read many arguments from you for why we can’t assume we get it (which make sense to me) but I haven’t read anything that would lead to such high confidence it won’t happen.
@Plinz Why would regulating the big labs outlaw decentralized AI? Are you saying that the regulation will start out only being for “frontier” models but drift to affect open source over time?
I think you provided reasons why an AI might not kill us even if it didn’t love us, but I don’t see why it definitely wouldn’t. We don’t know what goals AIs will have, if the AI views us with indifference, many of those goals might involve killing us, so giving a probability of 0% seems to imply you think the AI definitely will love humans
@DKokotajlo How much would it help if US Gov invested massive amounts of money into alignment research? Could be framed as helping accelerate the tech, because currently companies are slowing down due to alignment issues
@fchollet Is it possible humans have less fluid intelligence than we think and evolution does a lot of the “preparation” we need for on-the-fly adaptation?
@behrouz_ali@mirrokni@meisamrr What are the remaining hurdles that need to be overcome still? It doesn’t seem like any of the labs find this convincing as a path forward, I don��t see any obvious flaws though…
@polynoamial Why are models so overconfident? I’m sure it’s difficult for some reason but why can’t models be penalized for confidently stating incorrect answers in training? This is the most annoying thing with current LLM’s
@behrouz_ali@mirrokni@meisamrr Seems like a key challenge is getting the model to remember disparate ideas over randomly surprising facts. Can the model write notes to itself to remember things and then encode these notes into its memory? I feel like this is (partially) how humans remember things
@bcherny Claude code seems to enter plan mode and run forever (2+ hours before I cancel) when you ask a question on mobile. The same question in the terminal takes much less time. I think it may be eating up my usage as well
@OfficialLoganK I’m confused on which model is which on the Gemini mobile app, what is the difference between thinking and pro? They both seem to think for roughly the same amount of time (I have the $20 tier)
@kyliebytes Pretty clear they’re not conscious atp but I see no reason to be against forming a “model welfare team” to try and come up with a framework for verifying this / updating when it changes
@willcb I think the bigger problem is knowing when it’s gone down a bad path. I would guess humans have similar or even higher error rates but we don’t see this effect in humans because we’re good at self correcting.