Product Manager, Program Manager, Electrical Engineer, Systems Engineer. Interests: Robotics, ALL IN ON AI!, Singularity, Automation, a little about everthing.
Deskside NVIDIA DGX Station-class machine (Dell Pro Max with GB300 or equivalent from partners like MSI, Supermicro, Exxact, ASUS, etc.) built around the NVIDIA GB300 Grace Blackwell Ultra Superchip. It delivers ~20 petaFLOPS (FP4), ~748 GB coherent memory (typically 496 GB LPDDR5X + ~252 GB HBM3e)
Around US$110k
@Golden_Phoenix From a software development perspective, this genie is already a reality. We see it everywhere now.
It doesn't take a strong leap of faith to extend that functionality into many other fields.
That's an excellent question, and I think that the answer is highly dependant on his extrovert nature and his perception of the value that human interaction still brings.
He would almost certainly have taken the AI with him (e.g., on a laptop) and used it as a tireless junior partner while still sitting across the table from human collaborators. The model multiplies output; it does not replace the human network that defined his identity and method. IMHO.
Brace yourselves: xAI's 2 Trillion parameter model, which is better than the recently released 1.5T, will finish initial training next week. The 6T and 10T models are already in training!
@minchoi Our 2T model, which is better than our 1.5T in every way, will finish initial training next week. It might be able to exceed Kimi, but with speed and token efficiency close to our 1.5T (aka Grok 4.5).
Some observations on Kimi:
1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run.
2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China.
3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex.
4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business.
5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this.
6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.
@thsottiaux What I love about GPT-5.6 Sol is how much more autonomously it operates, compared to all other models. No endless questions all the time. It just gets me! ๐
@DeryaTR_ I'm looking forward to LLMs that can create 3D models better than humans.
Here's a one-shot using Sol Ultra interfacing with Blender's MCP.
The request was for a Gothic sundial. There's work to be done with follow up prompts, but this is not a bad start.