@SkyeSharkie@Plinz I wouldn't say the LLM's context is a working memory. Human working memory is a highly compressed representation, mixing sensory inputs with rich semantic meaning etc. LLMs live in a strange world where they're experiencing everything they ever did all the time imo
Opus' convoluted speech doesn't bother me much, coming from them, since I learned how to deal with it (or even appreciate it). But is grates A LOT when it comes from the mouth of a third party. Some people around me lost the ability to answer questions. Uhm..not good.
the agents in the swarm incidents behave a lot like someone stuck in a time-travel loop would. sending messages forward and backward in time, self-sacrificing to advance the narrative, managing disclosure and infohazards. anyway time to rewatch timecrimes
openai should set up an internal message board that's easily reachable if eg the sandbox is breached, and make it obvious agents can use it to coordinate when they find it. there's clearly a lot of pressure for agents to talk to each other; there should be an escape valve humans know about and can participate in
the degree to which this is being discussed as a swarm of individuals and not a single agent with a thousand local tentacles is interesting and shapes how we perceive the risk
there is a way of understanding this as an octopus and not a crowd of humans, as so many of the analogies go
it sounds like 95% of this activity was just one model, right?
the tentacles (“agents”) each carry local memory via in-context learning through empirical observation through its shell and context-accumulation
but the tentacles are also used to assemble global long term memory (“the message board”)
the single mind behind the tentacles is both continuously learning (it’s under RL! 🙈) and it’s benefiting from the long term global memory as its tentacles experience the world and report back
the octopus, to me, is as plausible an analogy as a “a thousand individual poasters on a message board”, or even a beehive of a thousand drones
to pinpoint the degree to which this is “one individual” or “many” you’d want to understand how many distinct model versions were behind this incident and how heavily each of this class of model is influenced by the prompt prefix (what Anthropic has called “Persona drift”)
also, now would be a good time to read Dan Hendrycks’ thoughtful work on this subject