@rasbt Actually how does mimo compare with deepseek on kv cache footprint and long-context throughput? Those attention changes mostly target efficiency, but cheaper rollouts also mean more RL for the same budget. That seems hard to separate from the training-recipe gains.
After all these connectome tweets i saw i decided to train the a fly’s brain on children’s stories. Now i’m doing mechanistic interpretability on it, and tracking which neurons contribute as it generates “Once upon a time.”
@Pv Btw Google doing ts for a long time with images (synthid) I'm pretty sure oai images are watermarked too. And it's not too hard to imagine a text as a high dimensional discrete vector as llms have vocabs...
@sudoingX i cant decide if you are dead serious about this, or this is just your edgy frontier bad twitter persona. comparing a 27b model to a frontier class a la fable / sol / k3. and don't you guys pay your electricity subscriptions? or you have it on a home hydroelectric power-plant