@LeRoyDesCimes I don't think the paper implies this unfortunately. What it's saying is that an LLM with a certain persona (induced by fine-tuning or system prompt) tends to express traits from characters in its training data that are close to that persona
"We have found that GPT-6 Astra is more capable of
controlling its own CoT than GPT 5.6-Sol
[...]
we find that the model is able to remain undetected when strategically underperforming in evaluations"
When I look at older art, I often appreciate the technical skill, or feel a mood from the painting. But with modern art I don't. At most I think "oh that's a neat trick".
Examples attached. What am I missing?
When I look at older art, I often appreciate the technical skill, or feel a mood from the painting. But with modern art I don't. At most I think "oh that's a neat trick".
Examples attached. What am I missing?
A\ became more paranoid about security after testing Mythos Preview earlier in the year and spent months hardening their systems.
This seems to be why they don't feel the need to pause RL now.
"The majority of RL has resumed, but some high-risk environments remain paused"
Weβre sharing an update on our alignment and security efforts.
In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems.
In a new post, we describe:
1. How weβve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards
2. An update on our alignment assessment
3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them
4. How we hardened our security practices earlier this year to prepare for Mythos-class models
Read more: https://t.co/E3Ea1Ds814
When flying lotus listens to a song, he visualises all the different parts laid out on a computer screen
People experience the same world very differently
Today we're releasing Laguna S 2.1, our most capable model to date.
It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.
Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.
Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface
https://t.co/xxGeAgo35R
Early mass-market television. Feels to me technology is invented as soon as it's possible, initially barely usable... Then it improves quickly.
See also: AI