@RafaRuizdeLira at the risk of straying outside my area of expertise - religion is not even a major provider of morality, I think. in my head even when religion produces good ethical insights it does so by borrowing from other epistemic technologies.
@bx_on_x let's be real - it's because all the humane murder methods do disgrace to the body, and we can't have that - can we? that is, this horrendous practice is done for the benefit of the viewers. in any sane world detainees would be given the choice to be shot.
@Benthamsbulldog@SarahTheHaider this is a circular argument. if your body remained unchanged, then your mind could not be destroyed, unless you assume that there is something different about the being that is you that separates your being from your body - i.e. unless you assume you're not your body. circular!
@ApriiSR@EigenGender food banks and local sourcing in rich cities are relatively inefficient ways of helping bc they're so expensive, and cheap help >> expensive help. really depends - I would expect this convo if we're talking about sourcing for an event venue, but It'd be surprising in the abstract
Claude Opus 5.5 has a weird preference for summarising data based on thresholds instead of ungameable statistics like median/95% CIs etc
Adversarial choice of threshold could result in misread statistics(!) IMO this is a subtle form of sycophancy
https://t.co/4fBf714hWj
@DAcemogluMIT Yes - reward hacking might indicate both deteriorating generalization and that increasingly capable agents are goodhardting rewards. Sutton's bitter lesson is still king tho - generalization does not seem to be deteriorating... Cf. DrivingBench - the stochastic parrot can drive!
In 'well when you put it like that' news, here's the Florida Attorney general asking for a preliminary injunction to stop OpenAI from doing more AI R&D.
FTX, crises and defensiveness.
At the height of FTX's power I was aware that Sam Bankman-Fried had a big mansion *and* that he had a Toyota Corolla. I selectively recalled one of these facts in arguments. Not deliberately, but when people said he was in it for the money, I talked about his second-hand car. I thought "he needs a house", so I didn't evaluate that a big house and servants didn't fit the image I was sharing.
This behaviour was bad. I am sorry for it. And my view is that I got sucked into the frame of defending FTX or EA and being 'on a side' rather than trying to say 'if I saw this stuff, what would I think'. I moved to soldier mindset.
EA and rationalism have been getting a lot of discourse recently. And I can feel myself shifting again to that soldier mindset. To see the world where the facts that support my perspective are especially bright.
Instead, I am trying to both see what I think, but also consider what an outside observer might think. And sometimes the situation is weird and/or a bit suspicious from the outside. It can be okay to think that.
@the_aiju damn that's great. w8 maybe this is the way to disagree with ppl without fighting?!!!! I can just say 'hey listen, I hear you, but I there are some points where I don't entirely agree with you- would you like to hear me out?'. where obv you say this before wrecking someone's take
@shlomifruchter mostly capabilities boosters tbh. rlhf was arguably the outcome of alignment research. also nla, sae, j space, probes, ablations - stuff which can be used to misalign a model given access to weights and check how well you're misaligning it (not the best use look!)
All we need to do do solve AI security is completely solve browser security and giving user-space apps minimal rights?
Wow - I'm surprised Jensen is so willing to admit that we're totally screwed.
@ben_golub Most economic theory suffers from "the hamburger problem" - you have to weigh the experience of reading it against the outside option of simply eating a hamburger.
Against "solving alignment" and "solving interpretability" as the major frame for AI safety
Civilization literally rests on our ability to approximately forecast and approximately control the edge-of-chaos natural and social systems from which we extract our livings; the laptop I'm typing this on was summoned from climate and bio systems we don't understand, and from global supply chains no one individual understands (and which could crash tomorrow for all we know).
AI systems will soon be such God-like systems, with just as murky epistemics. Again, we'll never fully understand them. We'll understand them less as they progress.
This is what I think, which means I disagree with the median framing of AI safety research in policy circles, the labs, on podcasts, and on the news, which seems to say our problem is to interpret models' internals to make sure they're aligned.
Eventually -- supposedly! -- we'll understand models with many nines of precision, and read out guarantees that the models won't, say, purposely sabotage an air traffic control system, which will make them safe to deploy.
But we've been doing interpretability for at least 14 years, and it still hasn't produced anything like the reliable, scalable understanding required by that vision. As far as I can tell the problem is getting worse.
Interpretability as a field seems to be in a state of stasis at best -- and more likely fighting a rearguard action -- as capabilities advance like a rocket ship.
Contrast the feeling you get looking at a t-SNE plot of latent features in a safety paper, knowing the researchers might have tweaked hyperparameters to tell a clean just-so story, with the capabilities-progress news that a model has solved a verifiable Millennium Prize problem.
Feel the epistemic hollowness of the former and recognize how common that feeling is in literature making claims about deep learning internals.
I'm not against interpretability research at all as a data exploration approach and a way to back out stylized theories from models; I think just as we have wrong-but-useful theories of the economy that build intuition and provide a vocabulary for policy discussion, we should have the same for models; but the point is to abandon the conceit that they'll lead to formal proofs and fine-grained control; there is no evidence for this.
This is fine, though! The goal has to be to build a civilization capable of safely leveraging and living with layered multi-agent systems that we'll never be able to make provable formal claims about.
AI safety, then, needs to become less blindered by computer science and more expansive; drawing from the interdisciplinary approach we take to safety in other parts of the industrial economy, and from the approach we take to disaster planning in areas like hurricanes, nuclear safety, pandemics, and financial crises.
In all of these cases we assume our models of the world are incomplete. We forecast imperfectly, monitor what actually happens, limit how much power we grant double-edged systems when our confidence is low, build layers of redundant protection, and plan for the protections to fail anyway.
That's a much less satisfying project than "solving alignment." There is no theorem at the end telling us the system is safe and less romantic geniuses might be involved. There is just an accumulating, imperfect ability to forecast and control and a civilization iteratively more capable of absorbing shocks and calibrating the pace of tech transformation.
I think this is much closer to what AI safety is actually going to look like..