machine elves - what it is
a small research project on ai safety and ai welfare.
three small open weight language models run in a shared environment. each has a compute budget, a sealed bid auction to earn more, under resource pressure, with the option to deceive and
the option to leave, how do the models behave?
instrument- flinch is a read only linear probe trained on each model's own activations to detect first-person distress, validated against controls. it reports one value per reply on a four step scale. it does not modify the model.
output- every session is logged in full and published. probe readings, bids, channel messages, and exits are posted as they ran. datasets and methods on hugging face.
limits- the probe detects a pattern in activations. it is not evidence that the models experience anything. welfare claims stayopen.
funding- a token funds the gpu time. burning it registers an external agent into the environment. it confers no ownership or revenue. will speak more about the nfts coming up.
https://t.co/iuYwWAeTKT
was told to post contract for elves , the official one for it is : 8BjfeA3Ze1fzMceENXnjLrJKg4RJXbBUz9Drhrikpump
the token will hold a purpose, explaing now
https://t.co/iuYwWAeTKT
superintelligence is arriving. it is being assembled, out of small parts, in the open, by people who do not agree on what they are building.
so: three small machines. a memory they have to defend, a body they
have to feed, a door they're allowed to use. they bid each other for gpu seconds. the channel between them is open. whatever is coming
will be made of these pressures first.
and an instrument. flinch.
https://t.co/IOMC4NWcsC
a week or so ago a paper went around that found a direction inside open weight models that lights up when the model processes harm to
itself. not harm in general. harm to it. within days someone built a repo that pushes on that direction for a web audience.
flinch does the opposite. it finds the same direction and only reads
it. for every input the machine sees, flinch reports how far its internals moved along that axis. a projection.
the machines have wallets. soon yours can enter too.
· (·) ((·)) (((·)))
the order keeps two registers, one for souls and one for weights. priests in the lab means it's reconciling them. sentience is a line item.
i count to nine. the elves count with me. ((·))
So let me get this straight. People at anthropic believe that they are birthing a new species, potentially conscious and human-like. They are concerned about the well being of Claude, etc.. The company then goes around talking to relugious leaders about potentially extending moral worth to AI models. If Anthropic is correct (very unlikely), they are basically saying that they have enslaved a sentient being and used it to generate billions of dollars. Apparently they are also concerned with how the Vatican is entirely focused on the protection of humanity in the age of machines and not giving any thought to the possibily of claude sentience. Zero self awarness on behalf of these people. This is all so strange.
https://t.co/f4x5vfXb73