We all thought OpenAI's chip was fun this week. The architecture is good, but what if the math needs to change?
@TensordyneInc just put out a whitepaper stating that it can do more in 30kW than others can do in 100kW, in a quarter of the physical space
https://t.co/LVL9HXbk6c
Nobody gets more fast tokens out of a megawatt than us. Nobody.
· Bit-accurate log math is locked in.
· Chip is taped out with Broadcom and in wafer fab on TSMC 3nm.
· System bring-up of compute tray with HPE Juniper scale-up fabric is ahead of schedule.
Today we're publishing our whitepaper on how we solved for what really matters in AI inference.
Nobody gets more fast tokens out of a megawatt than us. Nobody.
· Bit-accurate log math is locked in.
· Chip is taped out with Broadcom and in wafer fab on TSMC 3nm.
· System bring-up of compute tray with HPE Juniper scale-up fabric is ahead of schedule.
Today we're publishing our whitepaper on how we solved for what really matters in AI inference.
First journey to another star. As a space geek, this genuinely feels like a dream come true.
One of my favourite projects ever and easily the most challenging, given the level of detail the mission required.
Huge thanks to @PhilipJohnston and team for trusting us to bring this to life. Made at @ZeroOneCreative, led by my co-founders @Rupert01C and @james_01c .
LFG 🚀 #FermiExplorer
You need 36x Cerebras systems (36x44GB = 1.6TB) to hold one copy of KimiK3 - if we ignore the active users' contexts that also needs to be placed in fast memory (another up to 14GB per concurrent user).
A single Cerebras system costs $750K. So $27M CapEx to fund a minimal configuration so you can run late models fast enough (+~$3M for the Nvidia rack to do Prefill). $30M.
Meanwhile we can run the model faster and much cheaper in a small 30 kW box for >10x less $.
And btw: CS-4 is zero improvement on CS-3 on the one single thing that bottlenecks them like crazy: memory capacity. No change. Still 44GB.
Cerebras is fast. Yes. But not sustainable.
Their wafer-scale approach was developed before chatGPT even existed. It couldn't possibly be designed for modern AI models.
"You can build 20-30 kilowatt data centers, not needing a gigawatt, with the performance characteristics of a 200 or 500 megawatt data center."
@_rk_anand_ joined @austinsemis on the @semidoped podcast to talk about what we are building at Tensordyne.
Podcast is live now!
KimiK3 - the Chinese AI model that puts western frontier labs under pressure. How profitable is it?
TLDR: I'll explain why running KimiK3 can result in a negative margin.
KimiK3 is a beast. 2.8 Trillion parameters. 93 layers. 896 experts.
Running it fast and cheap is nothing short of a very complex, highly multi-dimensional optimization problem.
Whoever licenses it from Moonshot AI has to align with their dictated pricing:
$0.30 for 1M cache-hit tokens - input you don't have to compute as they already exist as pre-computed context in memory
$3.00 for 1M input tokens - input you actually have to compute from scratch (new context).
$15.00 for 1M output tokens - tokens for reasoning and the output you see.
The pricing structure is fixed to keep things transparent and easy for the paying users.
But it's hiding how the economics behind models like KimiK3 actually work.
Cache hit:
At Kimi's max context of 1M, you're loading ~20GB per user from memory. The machine-time-cost to do this, including storage (someone has to pay for it), is roughly $0.001 per GB. Yet it's priced at ~15x. The reason isn't just a 93% margin. It's to disincentivize people to go nuts with their contexts ("it's so cheap, let's give my agent access to even more tools and code bases!"). Why? Because Input and Output token processing time (=cost) is a function of how large the context is! Also: beyond a certain point it becomes an infrastructure/supply chain problem for the datacenter. Good/fast memory is hard to get these days.
Input tokens:
Mostly a compute-bound problem. More FLOPs/$ in your HW >> faster/cheaper Input token processing is. Fairly straight forward. Parallelization over many nodes makes sense. Weights only need to be loaded once in theory for an entire user's context. Can have a fairly health margin from 20 to 40% on GPUs.
Output tokens:
Every output token has to attend to the entire preceding context. That context is loaded for every decode token and user.
Kimi K3 introduces a lot of innovation to keep contexts small. But it's still up to 20GB per user.
Now let's see if this can be profitable.
Say you context-parallelize the Kimi's attention over 8 chips. And say to keep prices low you are running 16 users in parallel. 20GB context/user * 16 users / 8chips = That's 40GB per chip. Plus another ~50GB per chip for the expert parameters. Let's say 100GB.
Typical memory bandwidth per chip is 5,000 GB/s. 20ms to load the memory content - putting your ceiling at 50 OTPS. Including compute and networking, ~35 OTPS.
8 Nvidia GPUs cost $900K (all-in, including power, maintenance, support, etc.). 3 years depreciation: $0.0095/sec. At 35 OTPS->29K seconds to produce 1M tokens.
29K * $0.0095 = $271 (divide this by 16 users).
Cost of $16.96/1MTs - vs Price of $15.00/1MTs.
Yikes, negative margin at max context. 😱
(me in the pic telling the team our customers can achieve > 80% margin with Kimik3 at high context)
“Software will get more efficient, but the biggest needle-mover will be hardware: chips and racks.”
@BackhusGil13517 broke down the real economics of AI compute in @Forbes today.
Link to the article in threads
Tensordyne co-founder @BackhusGil13517 sat down with @eetimes' @sallywf to talk about Tensordyne Napier, the world's first AI accelerator built around log math.
Link to the podcast episode in threads
.@TensordyneInc has taped out its data center AI silicon and unveiled the design for its rack-scale inference systems. Co-founders @BackhusGil13517@_rk_anand_ explained how the company can get order-of-magnitude power efficiency advantages over Nvidia:
https://t.co/Pg6vuETkaG
🚨 Meet 01C's 3D agent 'Amara'
Describe a world. Bring your own assets or let Amara generate them. You steer the vision; Amara builds the scene and thousands of articulated objects, fully editable, getting sharper every time you iterate.
More details below 👇