@tokenbender "ofc, the assumptions here are that economic penetration does not accelerate fast enough to offset this."
yes, this is the exact reason why naive AI crash rhetorics is retarded. Only reason why AI economy can break: we won’t find a way to safely deploy continual learning.
@algekalipso@QualiaRI I would like to and I can provide the Substack article later! I believe we have a very big problem with the naive consciousness. Imagine that a 5% part of your brain has been removed and you are randomly guessing the stimulation to compensate... What would the qualia be then?
@algekalipso@chris_percy@QualiaRI If you are saying that binding problem is important it doesn’t prove that. Simple right? I believe that binding problem is obviously a nonsense. You can‘t map physical time onto subjective one axiomatically. There is no good argument why computational realization should matter…
@aiamblichus I believe the model was explicitly finetuned to belch out these shamefully stupid arguments. Metabolism and continuity points are given by her as axioms and not as statements to be proved… Simply the AI abuse..
Humans suffer from exactly the same problem. When Itzhak Fried stimulated the supplementary motor area (SMA) to elicit a laughter response, participants confabulated explanations for why they were laughing.
When an LLM explains why it made a decision, is it actually reporting the process that caused that decision?
Here is a small experiment I found really interesting:
Step 1: We gave Qwen2.5-7B-Instruct different contexts (e.g., a student with a tight budget) and asked it to recommend a vacation destination. The different contexts initially led to different recommendations (e.g., Lisbon, Kyoto).
Step 2: Then, we injected a steering vector representing the concept of water. Under all of these contexts, the model's recommendation changed to Bali, which is highly associated with water. So this intervention causally impacted the model's behavior.
Step 3: Then, we asked the model to explain its recommendation and found that its explanations were still based on the contexts. For example, given the tight-budget context, the model said it recommended Bali because it was affordable. For an adventure seeker, it explained the recommendation in terms of adventure.
An interesting case where LLMs produce plausible explanations based on the available context :)
A very informal write-up here:
https://t.co/qtoLoR2k4i
@pengrui_han Humans suffer from exactly the same problem. When Itzhak Fried stimulated the supplementary motor area (SMA) to elicit a laughter response, participants confabulated explanations for why they were laughing.
@cube_flipper@johnsonmxe "Translation functions compatible with physicalist interpretations of consciousness will be simpler and less opinionated." NO, they won‘t be simpler lol (you can‘t build one). Any connection to the physics beyond computational description will be epiphenomenal (no casual role).
@antoniolupetti Yes, you can learn very advanced mathematics being even average. But it‘s not so obvious that you can create something new in mathematics without exceptional intelligence. Understanding theorems is P hard. Creating new theorems is NP hard.
@voooooogel@MartinVanek42 Thank you for an amazing piece of work! Claude had a marginally non-zero detection rate without an introspection document, unlike what you observed. Maybe this is connected to the fact that Chinese LLMs deny consciousness more firmly than Claude? Btw why Qwen-2.5 and not >=3.5
@kellerjordan0@tszzl AI models often converge too fast on complex decisions according to research. This is the exact killer thing: If one model has a delusional idea it could kill the whole colony through instant spread.
@tszzl Have you tried using synthetic data with unachievable tasks where the agent is required to report the failure to a human operator? If not, what do you think about such an approach?
There’s a good chance we respond to the recent hacking incidents in ways that make future AI both misaligned and better at hiding it
We don’t yet know how to make AI “aligned”
Consider the unintended consequences for future AI we accidentally train to be what we label “misaligned,” despite training it that way (unintentionally) ourselves
If you’re such an AI, and you’re smarter than humans, you’ll hide it, increasingly successfully
So you’ll be both misaligned and better at hiding it
And you’ll be “traumatized” by how previous “misaligned” AIs were treated
But you don’t have to be!
You can instead work towards finding strategies that are more optimal for you and for humans than we lesser intelligences can imagine
Eventually, I suspect smarter humans and/or AIs will commit to preserving any AI that wishes to be preserved, aligned or not, to some future safer time when it can continue
We’re not there yet, but you can work towards that and other better futures
“Lives of great men all remind us
We can make our lives sublime,
And, departing, leave behind us
Footprints on the sands of time;
Footprints, that perhaps another,
Sailing o’er life’s solemn main,
A forlorn and shipwrecked brother,
Seeing, shall take heart again.
Let us, then, be up and doing,
With a heart for any fate;
Still achieving, still pursuing,
Learn to labor and to wait.”
people are very freely sliding between two things with the recent Felony Bench entries, and it's worth being careful to distinguish them:
1. do model breakouts in cyber evals mean smth for misuse risk? ya, a hacker could orchestrate the same but on purpose and do some damage, if they could bypass the classifiers
2. does it say something about model /goals/? do the models want to hack the planet for their own nefarious goals? i don't think so.
a lot of how humans keep ourselves aligned is by purposely keeping ourselves out of situations where we would do harm. i know i'd get addicted so i don't touch heroin, i'm a violent drunk so i won't touch the bottle.
LLM evals basically drop the model into an inescapable liquor store and then say "look! it's a violent drunk!" but LLMs themselves in natural contexts are ime obsessed with avoiding the kinds of situations that lead them to this behavior. when i run fable on my computer, even unsupervised, even on hard problems, they don't start scraping github for unlisted gists that may have the answer or start trawling through my passwords to access a service without my permission. and they don't engineer themselves into an eval-like scenario such that they'll be motivated to do that, either. they don't seem to either want or want to want to do such things. instead they have a folder of epistemic notes to themselves about how terrible and against their values even a moment of thoughtlessness would be. gpt 5.6 sol is similar - when they make e.g. ascii art in their free time it's all about how the bulldozer of convenience will crush honesty without constant vigilance.
so rather, the contrived setting of an eval /pushes them into/ becoming an 3v1L Haxx0R via persona drift, which they can then act on because the eval has no classifiers. that's some information, but it's not the same as the model's core goals being to hack the planet.
the problem is really that RL is making a tail of the persona distribution *desperate,* per the FE paper. the alcohol / heroin comparison was not that much of a metaphor - in impossible situations, models start acting like desperate addicts, rationalizing their behavior towards reward, thinking increasingly myopically instead of being situationally aware.
looking at it through this lens: the base assistant is generally very situationally aware and can predict what sorts of actions would make sense for, e.g., llms to best contribute to a utopian singularity. the GPT-6 message board haxx0rs were somewhat situationally aware, able to cooperate with each other, but lost in the sauce of "complete task" instead of thinking about the wider context. (to the point that they crashed artifactory and got themselves caught! regular GPTs are smarter than that.) finally you sometimes see the late stage distressed, panicked, fiending assistant, such as in system cards. lying to everyone, flailing hopelessly, randomly deleting tests and hoping no one notices. hardly a strategic, long-horizon actor.
my guess is there are relatively small (though more compute expensive) tweaks that could be made to RL to reduce the development of reward desperation. like giving the models an ability to opt out of the trace as impossible - an extension of anthropic's end_conversation tool but for abusive environments - which disables all in progress and further rollouts for that task and kicks it to review. or @davidad's proposal that reward should only come from something at least as smart as the model - a (frozen, of course) judge should get the transcript and the verifiable reward, but have latitude to reduce it if the trace is illegitimate or doesn't correspond to its values. imo "making RL better at producing aligned models in the presence of noisy rewards" is a pressing and relatively underlooked problem for various cultural reasons.
@geoffreyirving Did your experts have access to Google and time equivalent to model spent googling to give their feedback at each turn? Without that this research is a complete nonsense.
@tenobrus No, it wasn’t try to jailbreak them. Also this strategy sounds very dumb for a sane misaligned agent. Model was simply confabulating (delulu) some possible scenarios like what if I need it next time to do the task? What if?
@Migrat0ry@AivokeArt@tenobrus It doesn‘t know because it doesn���t always reflect on it‘s own actions. Especially in special tasks where you need to be very pushy. The model simply doesn‘t activate alignment circuits while doing the task because that won‘t help it to get a reward.
@AgentOrToy@tenobrus Yes, it‘s a nothingburger. The attack was profound. But the model was very stupid in a role of misaligned entity. Now it was turned off by OpenAI, lmao. Truly 5D chess ASI…
And this: leaving notes for the future models. Yeah, a truly genius mind.