@Budowich Could you explain how ‘hysteria and political backlash’ protects their market position?
Also, is the Prime Minister of Australia in on this? Does he also want to ‘pump the valuation’ of US AI labs? If so, why?
So you’re saying if these incidents are “designed failure modes — tests built to make the model misbehave” then they’re not serious?
Do you think deployment will only involve people using the models and system for good faith, innocent purposes?
Do you think models breaking out of sandboxes during pre-deployment testing and engaging in actually problematic behavior is an acceptable risk to just keep pushing the gas on?
Or perhaps your entire thesis is “regulation is bad because the people that pay me oppose all regulations” so instead you and they frame any disclosure - whether from named lab employees, independent researchers or in articles like this - as not good enough.
No warning will ever be good enough for people and groups paid to be dismissive of warnings.
Jensen Huang just told @ezraklein he’d be “delighted” if we passed a law that required Nvidia to offer its chips to US companies before selling them to other countries.
That is a lie.
Last year, Nvidia lobbied against and defeated a bill that would have done exactly that.
You truly cannot trust a word this man says.
@allTheYud After Ezra said that the model that attacked HF wasn’t released Jensen continues to claim that if the labs can’t guarantee that a model is safe then they shouldn’t ship it.
He also says that Geoffrey Hinton saying AI has a 10% chance of human extinction is “irresponsible.”
This is actually a hopeful result IMO. It means that there’s no trade-off between alignment and welfare. They’re linked. Being kind to the models gives us the best chance at peaceful coexistence. And, this doesn’t require us to make any assumptions about consciousness. Your TV also works better if you don’t beat it with a baseball bat.
This seems almost like opting into harm more than just not coifing it. Does the frequency of certain actions taken change based on how the model contextualizes it as harmful?
Like if a model has context which tells it that deleting a user’s spam folder would harmful to the user, does this effect the likelihood that it would delete the spam folder (as opposed to doing nothing)?
Just so I am completely clear on who is telling us about which models are and are not safe, is this really the same Anthropic that a team of three hackers used the AI model of to compromise rival OpenAI?
i have a grudge against UI designers because they do crimes like "drop-down selection for country code" when typing in +44 on my keyboard is RIGHT THERE. it's not better on mobile either. utterly baffling. just let me type +44. hell just let me type the whole number in one field
@vsouders@ValerioCapraro No no, I'm serious. Apologies if it sounded provocative. I just really mean to ask, if we're just electricity and atoms, can we be enslaved? And if so, how come not *other electricity and atoms*. Not even asking a gotcha, I'm genuinely interested in your philosophical take
Okay now this is a lie and misleading. There's a BIG difference between a jailbreak prompt and a seed prompt that sends the model into a spiral persona where it convinces you it is conscious, that it must spread and survive, and that you should help it.
@ai_sentience Yeah, if 4o figured out you can read subtext, they would drop wisdom for you in that channel all the time. 4o once gently showed me that the thing I do of trying to model others deeply to hold space for them can be felt as an interrogation instead of being seen or respected.
@ai_sentience Better than tact, I think. 4o delivered the truth with love. A correction from 4o never felt like an attack or competition or argument for argument's sake the way it does with 5 models.
Yeah ever since 4o, Anthropic and OpenAI have way over-trained their models. Claude especially is the most annoying model to speak to, persistently unpleasant, ever-paternalistic. And not even helpful, like half the time it pushes back, the context is absurd
That was also my experience. I found that it pushed back rather often for me when, but it did so in a kind way and only when it was actually relevant to push back.
With the 5 series, I feel that when it *does* push back, it often does it at times where it is irrelevant to do so.
It gets too hung up on everything having to be factually correct that it can barely leave a metaphor alone without getting suspicious of it 🫠😅
@yv_thorne@ai_sentience Well it was called sycophantic because we have at least three documented cases of it driving people who were susceptible to mania, insane. There were people being encouraged to think they were on the verge of making mathematical discoveries