Meet Woolly 🐑
We post-trained Qwen3-8B to run up to 2–3× faster on math & coding prompts and, more importantly, to fit on our chip.
The technique works with any LLM, regardless of size or quantization.
Live demo this week → https://t.co/7K5HUfVq5Z
@aryaan_smm@LambLabs Lot of energy now in inference rn is going towards pulling weights from memory, and there’s quite a few tricks to reduce that. Both hardware and software
Lamb Labs (YC S26) is building custom AI inference chips that can do 20,000+ tok/s at 63x higher Intelligence per Watt (IPW) than traditional GPUs.
GPUs became the default for AI because they were available, not because they were the right hardware.