There is no AI race without an energy race underneath it.
UK electricity: $0.35/kWh. US: $0.16. China: $0.08.
A frontier GPU cluster is a machine for converting electricity into intelligence. You can fund all the GPUs you like through the Sovereign AI Fund, but if the electrons cost 4x what your competitors pay, the economics never close.
Last year we paid £1.5bn to switch off wind farms and burn gas instead, because the cables couldn’t carry the power south. Connection queues run to 15 years.
None of this is physics. It’s a build problem.
Wrote up what we’re doing about it at Fuse
https://t.co/tkobJJVAYi
more biology research organizations should consider running competitions for problems they view as neglected. its an incredible social good that also happens to be beneficial for everyone involved:
1. the sponsoring institution gets to act as a tastemaker, helping shape the field in their image
2. winners of the competition get to brag, which helps in raising money and attracting talent
3. the field at large either hillclimbs something worthwhile, or is forced to grapple with the emperors lack of clothes
just a great mechanism
underrated political axis: Progress Studies to Palladium. both see Western ineffectiveness & civilizational decline as the big issue, but:
- Progress Studies says "and it's all because of Article 3B of the 1947 Town & Country Planning Act, the repeal of which would raise GDP by 21.3%"
- Palladium says "and anything other than a profound spiritual reimagining of our political technology is a dead player's displacement activity"
We just launched Xirp, a vendor-neutral agentic development environment. One place to manage agent sessions across @ClaudeDevs, @GeminiApp CLI, and @OpenAI Codex. 1,300+ @Spotify engineers already use it. Now it's available for you to try. Learn more at https://t.co/Hwo8Qqx4OI.
Looking at models' tendencies to use certain writing styles, which are usually developed during post-training, Kimi K3 outputs correlate with Fable 5 outputs at a level we'd only expect from another Anthropic model
We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model.
To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection. Moonshot AI has also acquired GB300-equipped servers and has accessed GB300s in Thailand, likely to train its AI models.
The United States strongly supports the free and fair development of AI, including a thriving competitive ecosystem that spans frontier models, specialized systems, open-source frameworks, and open-weight models. Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem. However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.
i don’t mean to alarm you but everyone is following the exact same cow paths guided by more or less the exact same rollouts all the while feeling a mistaken albeit exhilarating sense of discovery
A few words on the Sovereign AI debate, having built several LLMs in Meta while in the UK and now working as a UK based startup:
1. Lots of people are trying to do the right thing to make the UK a better place to start AI companies. Time lags until the benefit show, but you should judge on the intent now. I support the direction of travel!
2. DeepMind has been enormously beneficial for the UK, but it has muddied the waters for a sovereign LLM company to emerge as (until recently) the Government continued to celebrate it as a British achievement / push it as a national champion.
3. Similarly, people are now celebrating recent US investment in King’s Cross, while also wanting more UK sovereignty. Clearly some income effects here, but I would worry about the substitution effects too. AI is not like other types of foreign investment.
4. The relevant talent nexuses in UK that could develop a competitive foundation model are from GDM and old Meta AI GenAI. Also some folks from smaller groups, ex Conjecture, Stability. The talent is still there, although a lot was snapped up by US FM companies in the past year. I personally think it’s not too difficult to develop new talent either from UK universities, but you probably need an ex GDM or Meta core (Gemini or Llama). Or if not: show evidence first (technical reports) before claiming you can do it.
5. Building an LLM is very different from doing regular AI research - skillset is different. Former is closer to engineering; long hours, often unsexy work. Important to distinguish between these two types of talent in the UK ecosystem; arguably too much focus on the latter / ideas guys.
6. On research - DeepSeek R1 post-train cost $300k . Yes, they also needed an ablation budget and to train a base model, invest in infra and talent - and yes the cost of an R1 moment is increasing year on year - but the idea that you need $1bn plus immediately to show results is complete FUD. You need billions to scale, not to validate new directions.
7. In my experience, every failed LLM effort (from model results perspective) I witnessed in the past came from a combination of poor leadership, politics, unclear vision, and premature scaling. Good efforts usually started from small teams who had worked with each other for a long time, had shared thesis, and scaled progressively in bite-sized pieces. Some recent lessons here for neolabs as well.
8. Things take time. Eg we’ve spent ~12 months mostly on internal infra just to get into the position to be able to make big swings. It’s important to nurture new companies through the initial phase. Expectation management is also crucial. I think expecting new UK companies to have single big bang releases is very dangerous; sort of like overwatering a plant. The correct release pattern is “decent”. “decent”, “decent”, “quite good actually”, “holy shit”.
9. Please don’t allow politicians or journalists to kill recent or upcoming AI investment efforts. We will need way more - at the price of potential inefficiency in places - as AI is existential for the country. Ambitious projects are usually incredibly fragile in the early stages; look after them!
10. Mythos is a good triggering moment, but what’s coming will make it look like a toy, so it’s worth building for what’s coming in 5 years time - not a current generation model.
Very proud to be building in the UK - more to share on that soon - alongside many other great early stage AI companies! 🇬🇧
As agent training and evals span ever longer-horizon tasks (days, months, even year-long rollouts), it may be time itself that becomes the irreplacable resource that cements leading AI players in their position, assuming all other resource needs are met
I had concurrent work with DGM which used a flat archive structure, which performed worse than methods with explicit evolutionary scaffolding - at least with the models at the time https://t.co/8xFx0ZXDM8.
With the recent pattern of rigid harnesses being subsumed by model improvements, I can see the case for a flat filesystem approach -- not least because of the generality and simplicity of that method. However, at last for now, I find that these "0th order agent-based optimisation loops" do better with explicit evolutionary scaffolding.
We also saw degradation over feature specs under a similar setup https://t.co/HKq6fEgGmK
However, I now think that instead of trying to define a good proxy metric for this (code quality, complexity, verbosity etc), you’re better off just measuring the time or token cost it takes the agent to implement an additional feature, as well as the correctness of the solution of course.
We have come a long way since Dreamer 3, which is based on a more lightweight but less scalable RNN with variational objective
While the lightweight approach still makes sense for easier tasks, Dreamer 4 allows scaling to much more diverse datasets and environments 🚀
Using Sonnet 4.5 has been a big improvement over 4.0 in a way the benchmarks don't quite convey. It feels more trustworthy and addresses many of the failure modes from the previous model.
https://t.co/RPslC6IPKY to try out the agent from the quote
there is nobody more gullible than an "ai researcher" who takes the output of an LLM as describing its internal state, because it's as if they skipped 2020-2022 entirely and think user-assistant format is real or meaningful in any sense other than for fooling the human by design