‼️huge ssi news.
ilya is about to take his first tentative steps out of the age of research and back into the age of scale. it’s time to smell what ssi is cooking.
ssi have built a small reasoning engine that can compete with much larger training runs because his data is better curated to meta learning. but, more importantly.
we��re about to step into the era of TTT (test time training. gradient descent happening in real time to solve your problems). so instead of a context window you get actual learning.
and because it’s so sample efficient it can be trained on hard to verify tasks that other paradigms can’t touch. everyone else’s weights are frozen, they struggle out of distribution. ssi have created something that has a bundle of knowledge but can truly learn in real time and use that to your advantage.
current approaches are trying to hack their way to ‘learn’ with memory tricks, this thing will updates its weights, remember key lessons, and finally feel like a human level reasoner. this is a huge paradigm shift from the king. early results are very impressive. we can stop watching memento on repeat.
it’s learning all the way down, the descent is real.
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.
It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
https://t.co/aZDvqqhHRi
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.
It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
https://t.co/aZDvqqhHRi
maybe the nueralese thing is too strongly formulated
but I think the incentives in multi-agent RL are stronger than in single agent RL
and as we scale time horizons and become more safety concerned this seems like something that could emerge
it's still not super likely imo without the right (wrong) environments and pressures
@scaling01 Another example is bytedance, they have the best resource and the size of talents in china.
They are pretty good at training image/video gen model. Why they are lagging so far behind other labs? I think there might be some legal issue that drive them to avoid distillation
@scaling01 Nothing to do with distillation. So the real gap can be found by compare deepseek v4 pro (they might release an official version soon) and the OpenAI/anthropic ones.
@scaling01 Deepseek is the Chinese lab with the least amount of distillation.
Although one of my friends told me a renowned professor in llm extraction attack had a strong evidence that shows ds r1 distilled o1, I do believe they are the best Chinese lab and their feats almost have
@scaling01 I believe google might have at least a k3 level model. But it is not good enough for them to release and name it as Gemini 3.5 pro. For lab like deepmind, people have a high expectation. If Gemini 3.5 pro could not pair with fable 5, it would break their narrative and reputation.
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final
((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)