“What Trump is doing is he's instituting a worldwide trade extortion scheme. The tariffs are not about any particular outcome. They're about going to each country and saying, ‘Hey, nice country you got there. Would be a shame if something happened to it.’ And what he's seeking from each of these countries is submission. He wants them to submit their economic Independence and sovereignty to him.”
“And of course, the effect that that has in those countries is the opposite of what he wants. It puts the incentive to all of the politicians in Canada to oppose him, because if they do submit, they'd be thrown out of office right away. So what he's done is he's ensured the failure of his own policy and in doing so, ensured that the retaliation will come.”
“Now, while it's true that Canada will suffer more than the Americans, it's also true that the Canadians are so pissed off at us right now - 80% of Canadians are against the United States, for the first time in living memory - that they're willing to suffer more for the point of maintaining their sovereignty and standing up to Trump. And that's not the case in Michigan. That's not the case in Maine. People in America are not a fan of these trade wars. They’re not on board with Trump’s strategy.”
“So there will be political consequences, not for Trump, because he is not running, but for Republicans. But it's pretty clear that Trump doesn't care about that.”
@RolandMemisevic@akarshkumar0101@chrmanning I think that's a good bet too. Actually let me clarify that I don't really think of SMT as "distill transformers." It's joint optimization of a memory-bottlenecked transformer and an RNN. So it may be closer to your bet than it seems? See fig 10 for example.
Excited to welcome Andrew Gordon Wilson to our research team. He will be reporting to Denis and lead new research efforts on continual learning, synthetic data, long horizon RL environments and architectures. We’re hiring! Please reach out to Andrew!
Pretraining Recurrent Networks without Recurrence by @akarshkumar0101 & @phillip_isola is a great paper!
It makes an end-run around the problems of RNNs via a transformer teacher to learn good predictive state representations & supervised learning of a memory transition function
When teaching CS224N, I thought it important to introduce students to a broad neural toolbox – not just transformers but FFNs, CNNs, LSTMs, tree-recursive NNs, BiDAF QA nets, highway nets, …. I think the resurgence of work using recurrence shows the importance of this approach.
Pretraining Recurrent Networks without Recurrence by @akarshkumar0101 & @phillip_isola is a great paper!
It makes an end-run around the problems of RNNs via a transformer teacher to learn good predictive state representations & supervised learning of a memory transition function
Transformers are Inherently Succinct by Pascal Bergsträßer, Ryan Cotterell & Anthony Widjaja Lin is an interesting contribution.
Various work (Hahn 2020, Li & Cotterell 2025 i.a.) has shown that transformers are less powerful than RNNs … yet somehow they do so well in practice.
I’ve been looking forward to today for 3 years
Today we’re announcing Foundation - Chroma’s solution to memory
Our research preview of this technology builds self-improving memory from your agent sessions. Try it out at https://t.co/DpynS0vnSd
I remember my first day as an intern at a big company.
They said probably 5 times during orientation that everyone's calendar was open and visible. Setting up meetings was encouraged.
So I looked up the CEO's calendar...it was set to private. I asked and they looked at me like I was a monster, and said "why would you want to see his calendar?"
They couldn't have taught me better about corporate politics if they had tried. Honestly admirable.
A few people have recently claimed my tweets are actually AI. Apparently I go to great lengths to bypass AI detectors, but these people have seen right through it, as they are connoisseurs of fine human writing.
I find this hilarious and don't plan to change my writing style, which has had some AI characteristics long before AI writing existed. As I've said before, it's not AI's style itself that makes it nauseating to read, but the disconnect between the punchy turns of phrase and the shallow substance behind those words. I'm confident enough in the substance of my writing that if my style is AI-like, I take it entirely as a compliment.
It's sad what the prevalence of slop has done to us, making us constantly second-guess everything we read. I find it best to stop worrying about whether something is AI and instead read it a second time to see if it reveals more depth than the first time or starts to feel hollow.
p.s. Not a brand new paper, and, indeed, an Outstanding Paper at ICLR 2026, but it’s hard to keep up these days….
p.p.s. I didn’t actually read the proofs in the appendix. I am on vacation.
p.p.p.s. Ryan Cotterell appears to have deleted his Twitter account….
These succinctness results suggest that people working on mechanistic interpretability might have a very hard row to hoe.
But otoh the circuits transformers actually learn from data are very unlikely to include the weird counting transformers underlying the proofs in this paper.