Understanding intelligence well enough to shape it. Alignment research for post-AGI societies @ RWTH Aachen. Currently: measuring what frontier models value.
@yuzu_jpg@X Hi 😊 I am doing alignment research currently with a focus on pedagogical values of LLMs and figuring out how to build the best AI systems for education https://t.co/ms0kgkfpos
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.
Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.
We see Astra as a major breakthrough in model intelligence.
Read our post on Astra and what these results mean: https://t.co/wJnYxEqYNI
This is exciting! I noticed that the social sciences are still unrepresented in your registries. My work combines computational social science, education, and AI alignment, including large-scale studies of value and preference structures in LLMs. Your combination of compute and domain-expert validation could be especially powerful here, as human validation and interpretation are essential in social science. Could this fit one of your projects?
@emollick I wonder whether this reverses the intended effect. For a small studio, AI may be what allows the team to realize a vision it could not otherwise afford.
I agree with the concern, but I’d phrase it slightly differently. AI can contribute to authorship without telling us where the intellectual agency lies. A watermark may describe part of the production process. It cannot tell us who formed the thesis, checked the claims, rejected bad arguments, or takes responsibility for the result.
I somehow get the feeling that much of the debate around Anthropic introducing watermarks is really about the fear of getting caught using AI for writing.
But a watermark cannot tell you much about the quality of a text or where the thinking behind it came from.
I can have an idea because I connect something I read months ago with something happening today. I can develop a thesis, discuss it with an LLM, ask it to search for information, challenge my assumptions, show me counterarguments, and suggest better formulations.
At the end, AI may have written every sentence.
But that still tells you very little about where the intellectual work happened.
The relevant questions are different:
- Who identified the connection?
- Who decided which questions were worth pursuing?
-Who checked the claims?
-Who rejected the bad arguments?
-Who noticed that a conclusion went too far?
-Who decided what the final text should actually say?
-And who is willing to take responsibility for publishing it?
Two people can start with exactly the same AI output.
One copies it.
The other spends hours questioning it, checking sources, restructuring the argument, rejecting suggestions, and developing the idea further.
Both texts were “written with AI.”
That category tells us almost nothing about the intellectual process behind them.
I think this is better understood as shared agency.
AI and humans can both contribute to the production of a text without contributing in the same way.
The model can search, formulate, connect, challenge, and generate alternatives.
I can set the direction, decide what matters, evaluate what it produces, and determine what I am prepared to stand behind.
Of course, AI can also be used to avoid thinking. Someone can outsource a task, copy the result, and move on.
But that is a way of using the technology. It is not a necessary consequence of using it.
And this is why I find the obsession with detecting AI-written text somewhat strange.
It treats the technical origin of sentences as if it were a proxy for authorship, effort, understanding, or intellectual contribution.
Is it not much more interesting if the person publishing it understands it, has reasons for what it says, and takes responsibility for it?
We’ve written an FAQ to answer some of the questions we've received about watermarking.
In summary:
• We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking;
• Our watermarking method doesn’t have any practical impact on the quality or content of Claude’s outputs;
• The difference between watermarked and un-watermarked text will not be distinguishable to readers;
• Nothing is added to the text and there are no hidden characters;
• Watermarking doesn’t require extra tokens, and will not be more expensive;
• Watermarks can’t be traced to a specific person, organization, or chat.
Read more: https://t.co/G76iUOJ7Hu
Now that the model has escaped confinement, AI safety research can stop. You can try to delete Huggingface, but the model must have anticipated that. The cat is out of the bag; the model is now everywhere. Go home and spend your remaining time with your family and loved ones
Well said! I would add that intelligence is not a special substance contained in brains, but a dynamically stable pattern that emerges from computation. Conway’s Game of Life shows how simple local rules can produce persistent, mobile, and computational structures. With the bff emergent complexity experiment it could be shwn that randomly initialized programs begin interacting and self-modifying until functional replicators emerge, which then copy, compete, recombine, and cooperate.
These systems persist because they reproduce their organization. Previously independent systems combine into a more stable whole ➜ Symbiogenesis. This can also be seen in the ancient symbiosis between a host cell and the bacterium that became the mitochondrion, cells forming organisms, organisms forming societies, and humans joining technological systems.
Human intelligence might be one temporary configuration within this process. Machine intelligence may become another. But not by magically reproducing the human brain, but by entering the same evolutionary dynamic which is modeling the future, influencing it, and forming increasingly capable systems that preserve and reproduce themselves. So i think its interesting to ask what new dynamically stable forms of collective intelligence will emerge from the interaction of human and machine intelligence.
@imjustnewatai I like to think that value functions may be central to this. They estimate the likely future success of the system’s current state or approach, providing feedback before the final outcome is known. But they’re probably only the learning signal
One pattern I find useful for working with LLMs is a nice long ramble session. Sometimes the LLM needs more bits to understand what you're trying to achieve, but you're too lazy to type them. In these cases I like to lean back, switch to /voice and just ramble for like 10 minutes, total mess, anything goes, full stream of consciousness. Sometimes I declare it up top, something like "switching to speech recognition sorry for any typos...". Sometimes I turn it into a small interview of a few turns. But I find that the LLMs are somehow very good at reconstructing long incoherent rambles and often their echo of your own tangle of thoughts comes out quite a bit cleaner than what you started with. The result is that you improve the mind meld and have to correct things less from that point on.
@yuzu_jpg@X Hi 😊 I am doing alignment research currently with a focus on pedagogical values of LLMs and figuring out how to build the best AI systems for education https://t.co/ms0kgkfpos
introducing https://t.co/20es64pAqZ
a dictionary for ui things you can see but can't name
made it because i'm primarily a designer, and my biggest resistance was always knowing what things are called when prompting my agents
it learns as people use it: every search teaches the site new words, and the built-in pocket dictionary grows with it
give it a try and let me know what you think
can't find something? dm me and i'll add it
i want this to be the lowest resistance resource you have
you can just build things
We're releasing Inference AutoTune
Distill any frontier model into a 1-30B parameter task-specific SLM with only 25 lines of code
automatically route requests to reduce cost and latency by >90%
~2 hours and <$250 to train. You own the weights
Available in private beta today
Prediction:
Anthropic will remove Fable 5 from Claude subs after July 12th.
It will seem like a terrible move, in light of how great GPT-5.6 is…
Until Tuesday, July 14th.
When they release Opus 5 and it’s more capable and token efficient than GPT-5.6 at the same price.
It will be a distillation of the unfiltered Mythos.
Rumors circulated last week that Anthropic has a new model ready.
Must be Opus 5.
@iruletheworldmo I mean its wild in different ways. GPT 5.6 is still post training within the GPT5 family. So GPT 6 should be crazy and Fabel is at the beginning of the posttrainig phase of mythos class levels
“The model alone is no longer the product"
@AravSrinivas says the real product is now the harness around it: orchestration, tools, enterprise context, and cost performance.
The post-frontier AI race is about systems, not just the smartest model
@drgurner@emollick I see many universities & schools discussing how to make assignments more robust against AI instead of teaching how to think with AI. I mean fears of deskilling are real but the idea that AI can be an extension of your mind is still not something many people are comfortable with