Today, we are launching Belief Updates, the Simplex research blog, with two posts.
Read the welcome post here: https://t.co/CWSKmCaxlB
And see our post about nonergodicy in the quote.
We are starting Belief Updates for a few reasons: 🧵👇
Training data for LLMs is made up of many sources. Given that, what structure should we expect in the activations?
New work from Simplex shows how the belief geometry over this type of data forms telescoping cones, and transformers represent them! 🧵👇
https://t.co/hBi2r1e32j
A longstanding dream of interp is to decompose activations into distinct, interpretable parts.
But when should we expect that to work, and what even are such parts?
New from Simplex: transformers factor their world into orthogonal subspaces, even when it costs accuracy.🧵👇
Happy to share what @yaoliucs, myself and @EmmaBrunskill have been working on over the last year: "Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance Sampling" https://t.co/CSifgqFIel