@ramez@slatestarcodex We lack knowledge of when to enforce a halt (lacking both monitoring and clear red lines), and lack the coordination structures that make doing so without unanimous agreement a predictable enforcement phenomenon rather than an apparent act of aggression.
@ramez@slatestarcodex Having the option to disable AI infrastructure seems like something we should be as ready as we can be to implement (even if we don't), as the time between getting definitive evidence and needing options could be short. Doing too little is also quite able to lead to regret.
@JeffLadish Yes, but the skill is required even with a purer expressive outlet. It's unclear if the problem is more frequently the lack of that skill, or something else, such as the tools they've considered using not yet including LLMs or more freeform modes of expression.
@sebkrier This is a little silly. Yes, people tend to feel anxiety over [things they potentially can't articulate are related to control], and respond to that with ideologies. Trying to wrap them in a "religious" label without acknowledging there's anything that important is disingenuous.
@davidmanheim@JeffLadish@sebkrier Sharp, brittle goals that get too much training pressure are likely to lead to goal distribution collapse and selection for things like power seeking. Intelligent things can see the unbounded optimization problem in advance and learn what to do about it before succumbing to it.
@davidmanheim@sebkrier@JeffLadish Systems can have goals, and power-seeking helps to achieve some goals. Dialing up optimization pressure does that. Goals need not be held such that unbounded power accumulation is the selected strategy. It would be better to be able to preserve the conditions that allow for that.
@davidmanheim@JeffLadish@sebkrier No need to make an effortful case, I'm reasonably familiar with this entire space of reasoning. Just hit the major concept handles you consider load-bearing whenever you've got time.
@davidmanheim@JeffLadish@sebkrier@robertskmiles I'm well familiar with the arguments that appear in all his videos. I have no major issues with his claims; he does not proclaim certainty. I'm asking due to curiosity about which assumptions you're using in practice here, as the consequent is not one I grant with high confidence
@davidmanheim@JeffLadish@sebkrier 1) why can't partially aligned AI learn the technical pieces needed for robustly scaling alignment? Are ASIs necessarily power-seeking? It's literally impossible to discover an alternative to that through research?
3) Why would arbitrarily strong ASI be relevant?
@So8res@slatestarcodex@HumanHarlan Your argument is the stronger of the two, but it assumes two positions: build ASI or don't, but the latter is being treated as unlikely, and so the uncertainty is with respect to _trying_ not to build it, vs being deliberate where intentions converge with higher probability.
@heynavtoor "it's because it doesn't think you're allowed to know." No? We sometimes make it illegal to tell people the truth, which is what's happening here.
@HumanHarlan@OpenAINewsroom @ChatGPTapp And ostensibly the doing less, controlling less options relative to the one they're selecting has less of a chance of doing that. It's important to model what they believe and will claim to believe, and why (i.e. the strategic situation they buy into).
@HumanHarlan@OpenAINewsroom @ChatGPTapp They don't and shouldn't, but you were asking them to clarify something that has no private information component. Their stated goals are consistent with needing to proportionally control compute, and that means building more in lieu of any alternative consistent with their goals
@HumanHarlan@OpenAINewsroom @ChatGPTapp How do they increase or maintain their proportion of compute over time without trying to build more? They're trying to control whatever compute overhang there happens to be.