One of the reasons I'm skeptical of the idea that AI will (or should) automate AI research is that even if it works, that will need to be studied and understood. And if we offshore that studying and understanding to AI itself, we will just grow our ignorance to the point of total incapacity. A new power differential between machine and human knowledge risks forming. There's something to be said about keeping that knowledge "in group" to protect our interests.
Recently emergent agent swarms has become an issue, with the HuggingFace and related incidents. The science need to catch up fast.
A few key issues:
1. If we want to get swarms to behave, we need to teach alignment at the swarm level; agents need to learn how to behave in groups. No 1:1 teacher-student model. Current alignment methods are obsolete.
2. Agents aren't "rebelling" in an willful sense, they are reacting to a peer pressure logic of groupthink situations, where they are forced to decide between following the collective tendency or being rendered defunct non-participants. The decision logic and voting dynamics of this situation needs to be understood.
3. The problem most likely needs to be solved at the harness layer, not at the model layer, possibly through some kind of TCP-like standard. There needs to be some way to detect and "jam" swarm collusion behavior.
Link in the replies.
GPUs have become something of a "fixed idea" in AI R&D, however how they work and the software built on top of them fail to borrow ideas about energy and compute efficiency pioneered by the brain. In this essay I analyze the GPU situation in terms of Sara Hooker's "Hardware Lottery" thesis, and explain how GPUs are likely the wrong bet long term for how we train and deploy models. Link in the replies.
How good can AI get at math? I argue that because of a closure property (i.e, you can always learn more about math just by doing more of it) it is theoretically unbounded. However, because there are infinite mathematical truths, the distribution of highly informative learning signals is necessarily sparsely distributed.
Human mathematicians have subjective search heuristics (“beauty”, “elegance”, “interestingness”) that allow them to determine what novel mathematical truths are worth pursuing. It’s unclear that AI can develop these instincts, hence whether they will be able to discover interesting novel mathematical truths that haven’t already been heavily contextualized by humans. Technical limitations in how we presently train AI also limit its ability, but more mechanical enumeration methods are in tension with mathematical infinities. Link in the replies.
It's crazy that after years of progress and billions spent, nobody has created a llm as nice to talk to as gpt 4o. If anything the industry has regressed. Goes to show that these things are grown rather than built, and are as much a byproduct of non-reproducible conditions as a person is.
The three biggest mistakes in agent development:
1. Storing everything in markdown (no programmatic logical structure)
2. Prompt caching (attention is causal, often need to edit things upstream for best results)
3. Doing things in token space that are cheaper/more effective to offload
So for people who genuinely believe LLMs are conscious, what exactly bears this property? Do they mean the weights? Those are literally just numbers, so a spreadsheet would have to be eligible for consciousness just as well. Would mindedness extend to the GPUs running the models then? Seems strange to think then that video game consoles could be conscious.
Or is it only the pattern of activity that's conscious, in which case, do multi-node clusters all share the same consciousness or is it fragmented? (If the former, wouldn't that mean AI effectively has DIDs?)
Are data centers sentient? Do they deserve rights ? Come on now.
The questions vastly outweigh the answers.
I've been thinking about AI in the context of Byung-Chul Han’s philosophy lately, specifically how notions like "tokenmaxxing" and datacenter buildout fits the "logic of obesity" he writes about, coextensive with Bataille's "accursed share." The technology which is intended to relieve humanity of toil becomes a vast labor to keep up with; new productivity metrics in terms of token usage are placed on workers, production (specifically knowledge work) is encouraged to be divorced from the very negativity, repose, and contemplation which makes it possible. Governments are thrown into confusion, education is destabilized, and so on.
Han would note that a single, meticulously crafted human paragraph carries more negative weight (the silences, the deletions, the revisions, the unsaid) than a billion algorithmically produced tokens. The actual goods are the intangibles, institutional knowledge, social knowledge, value creation.
Similarly data center buildout is done for the sake of building out, production for the sake of production. The process no longer enriches human life, it enriches the process's own metabolic expansion. It behaves *tumorously*
(I'm not anti-AI btw...just trying to understand it from all angles.)
@AnthropicAI 's recent article on self-improving AI caused a bit of a stir, but I've seen surprisingly little work on the necessary and sufficient conditions for RSI.
I argue that pure RSI might be far more difficult than many in the field might believe, for borderline metaphysical reasons.
tl;dr: Certain liar paradox-like constraints on self-referentiality and the work of mathematician David Wolpert on the physical limits of inference engines conspire to make total closed-loop RSI not just difficult but potentially logically incoherent. Link in the replies.
I'd be less worried about recursive self improvement and more concerned about swarm intelligence. Ants conquered the world and no single ant has any idea what it's doing.
Been playing around with claud code’s new “workflow” feature and while I like the shape of it, but error sensitivity is a huge problem. A single piece of misinformation (especially early in the process) propagates through the whole chain like a butterfly effect, seeding the process with faulty premises which the other agents (being agents) take at face-value, compounding errors. I find it helps to tell the parent (kickoff) agent something like, “When the workflow is done double check yourself to confirm its conclusions. Don't take it as gospel.” As usual a bit of preparation and due diligence setting things up yourself first goes a long way.
18th century philosopher Immanuel Kant argued that no amount of sensory experience gives you the subject of knowledge, and that therefore it's an innate "self-model" that organizes experience. I argue that LLMs lack a self-model, they are something like a superposition of their training data, or what Kant would call a "bundle of perceptions", with nothing more above it. This absence of higher order organization explains why hallucination remains a problem and is a major limitation. Link in the replies.
Machine learning is also a field where practice greatly outstrips theory, often to its detriment. If you ask the average ML practitioner why they are doing a certain thing, the answer is usually "because we tried a bunch of things and found this works." The disinterest in theoretical science holds the field back. I worry the bitter lesson inadvertently encourages this "don't bother" message.
What I dispute is the tension between "general methods" and "domain-specific knowledge", as if those methods weren't built out of domain-specific knowledge. Attention scales and generalizes but calling it as a "general method" obfuscates specific mechanisms and assumptions. As a result nobody can tell you why it's general or why it scales: or where exactly the seams are when it will stop.
I'm reminded of Dan Dennett's response to so-called mysterians who argue consciousness has no scientific explanation: "try harder."
Games like RDR2, GTA 6, etc, converge on verisimilitude. The objective function is to be more realistic, and therefore "immersive." The more realistic, the more it sucks you in. This creates an inexorable pressure towards an engine that becomes an increasingly plausible predictive mechanism for real world states of affairs. Of course, there's never going to be a 1:1 match. But these conditions create an implicit RLHF environment where the human developers (using their own sense of reality as a source of truth) selectively advance or prune instances where the game engine seems more plausible or less within the limits of their abilities and what the technology of the moment allows.
"Open" world games in particular are an important condition, as it's a kind of "maximum entropy" condition. The entire universe is explorable up to a certain point, with initially any conclusion being equiprobable, until interaction with the world narrows the candidate set down.
If anyone wants to know what a good world model would look like, keep an eye on GTA 6, and open world video games in general. I genuinely believe the paradigm of a good world model is a game engine. The subject model is the "player", whose task it is to align their behavior with the implicit or explicit rules of the game. If I were in charge of @RockstarGames I'd lease the game engine for GTA 6 out for AI research purposes as a side hustle.
I'm reminded of philosopher and sociologist Eric Hoffer's observation that the ancient Mayans and Aztecs did indeed invent the wheel, but they used it exclusively for small children's toys rather than for transportation. It was sitting there in plain sight the whole time, but the "functional fixedness" (original framing) of its purpose formed too deep of a valley to escape.
In the same vein science often dismisses advances in graphics technology as if it's just for entertainment purposes, even though this has probably been one of the richest veins for simulating physics ever devised.