I spent the last six months trying to deconstruct Taalas's patents. We think it can be better: 100x fewer memfetches than their bitROM, better software by better quantization than their hardware team could.
So we wrote a compiler that:
> takes a huggingface checkpoint,
> quantizes the model
> descends the weights down metal layers to RTL and GDS, do your DRC, Yosys, PEX, make the electrical waveform execute a mat-vec from your huggingface checkpoint (thanks Cambricon tech papers)
in ~7000 lines of human-readable code.
we're looking for someone with contacts with a foundry/access to 7nm PDKs or contacts at a cheap EuroPTW shuttle? I'm broke and unemployed.
@itsclivetime this is what i think is the future of perplexity per picojoule. @zerohedge@zephyr_z9 you guys wanna see 40,000 tokens/sec?
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
Grok 4.5 is not yet using our internally developed C/C++ inference software that exact maps to the GB300 hardware. Doubling or more of the current speed is probably achievable.
LI Auto's MindVLA-o1 model addresses VLA limitations with multi-modal MoE Transformer architecture, 3D ViT encoding, enhanced multi-modal reasoning, predictive modeling, and integrated hardware-software design for improved efficiency and adaptability.