UC Berkeley open-sourced FreeToken. Wild results:
A single RTX PRO 6000 runs the 753B GLM-5.2 at 14.9 tok/s!
An 8GB RTX 4060 laptop (~$1,000) runs Qwen3.6-35B at 39.3 tok/s!
FreeToken is 2��4x faster than Ollama across consumer GPUs. Local AI inference is getting very real. Great work by @Andy_ShuoYang and UC Berkeley Sky Lab!
Unitree just posted a humanoid doing a 2m standing jump, beautiful show of its explosive power.
And also its top running speed is 12.66 m/s. That is about 28.3 mph, slightly above the ~12.4 m/s peak measured for Usain Bolt.
Those are ofcourse peak-performance demos, so there is still a lot we don't know about endurance, acceleration, energy consumption or repeatability.
But hardware iteration speed compounds.
New tools are launching faster than most of us can keep up with.
And sometimes, the hard part isn’t learning the tool. It’s figuring out where to start.
Guidy reminded me of my first job, when having a mentor meant someone could quickly point you in the right direction.
Now, you can get that kind of guidance directly on your screen, step by step.
Especially useful if you’re not a tech person and constantly finding yourself learning new software.
Honestly, this feels like something we’ll need more and more.
Worth checking out: https://t.co/tzIY2G02AO