Activation patching confirms it: ablating the manifold directions tanks probe accuracy far more than random-direction controls, showing these subspaces really do carry the task-relevant info.
π https://t.co/EwC8CfSHC2
Tasks where the ordinal variable is locally computable from token identity β clean 1D manifolds with place-cell-like tiling.
Tasks needing cross-position integration or semantic extraction β higher-dimensional, less coherent representations.
Architecture matters too: Qwen3-4B shows much stronger geometric "twisting" than Gemma for indentation and unlike its numeric twisters, these twists preserve ordinal order.
Do LLMs represent all ordinal concepts the same way?
We tested 4 tasks (bracket depth, indentation, table position, numeric magnitude) across Gemma-2-2B/9B & Qwen3-4B and found the answer is no, the geometry depends on how the model has to compute the value.
This is the craziest thing I have heard this week, crazier than MosiacML news.
This reminds me of the famous dialogue from The Social Network. A Billion Dollars.
Excited to announce that weβve raised $1.3B to build one of the largest clusters in the world and turbocharge the creation of Pi, your personal AI.
https://t.co/p5AfRXGPan
Today Iβm excited to announce the first version of our new personal AI, Pi... https://t.co/wYpgcXdB1t
Pi is smart, kind and supportive. Itβs designed to be better at natural, flowing conversation than lists, plans, or code.
just read code written by someone who I really looked up to and from a leading research lab.
have really mixed feelings after reading through functions of >400 lines.