I'm obsessed. Took one of my favorite books, The Charisma Myth, and turned it into a tool you can actually practice with. Now presence, warmth, and power are things you can train. Building with AI is too fun rn ✨
Thank you @claudeai@AnthropicAI !! Comment if you want to try :)
@EmperoAI Seems like the dataset is closed (sad, but understandable). To get a bit of an intuition, would you mind sharing the tokens count (for the 70k traces)
I like MATS (Machine Learning Alignment & Theory Scholars). It has produced excellent research and made valuable contributions to AI safety.
Still, I see a worrying trend: some MATS research appears designed to portray Chinese open-weight models as “evil.”
Anthropic helps shape part of MATS’s research agenda. Anthropic also helps define the evaluation standards through its researchers, concepts, and model-based judges. This influence deserves scrutiny, especially when Anthropic is a closed American AI company competing with Chinese open models.
Terms such as “evil persona” are anthropomorphic and normatively loaded. Anthropic’s persona-vector research first defines “evil,” generates opposing examples, and then extracts a direction corresponding to that definition. The result inevitably reflects assumptions embedded in the experimental design.
There is also a structural asymmetry: Chinese open models are altered and dissected, while closed American models rarely face equivalent scrutiny. American models then often judge the resulting behavior. Technically valid findings can still receive geopolitically biased interpretations.
Stronger research should apply identical tests across model families, use independent human annotation, include culturally diverse evaluators and multiple judge models, and clearly distinguish an original model from a deliberately corrupted derivative.
MATS does excellent work. Greater methodological symmetry and cultural neutrality would make its research more credible.
As you might guess, this suggests that distilling reasoning traces may have been possible for a long time without ever breaking the cryptography.
An anecdote: we find that prefilling Kimi-K3 reasoning with a few tokens of Opus reasoning measurably shifts its response toward Opus’s🤷♀️
A small memorization analysis showed that specific Claude and GPT reasoning spans are up to ~6 orders of magnitude easier to extract from Kimi-K3 than from the next-closest model.
@chris_j_paxton@luk3hans3n Could be both - SFT cold-start + rewarding lingual consistency (English consistently/other language consistently).
IIRC that was what R1 did
@TheStalwart 1. Grab a sample of 1K reviews done by Flash 2.0
2. Run GEPA with your existing prompt, using Flash 2.5, to modify your prompt enough so that 2.5 reviews agree with the 2.0 eqv