@badlogicgames You can check for early access of the thinky machines micro sequence model and plug it into the TTY without any tools. One letter for a letter. If the user wants to interfere, they need to send messages to the agent from another terminal session …
@ivanfioravanti@balamuralimenon On the M5 Pro low power pre-fill and token generation goes down to 2/3 of the full power performance using about 40% of the energy in MLX (based on Mactop).
@deliprao Comparing SVG-output to SVG-output: this is Gemini Pro 3.1. It took more than a minute. Generally vector generation especially for diagrams or animations can be more efficient than generating millions of pixels … What’s the name of the startup?
@simonw I totally agree with the sweet spot of the 3B model, but at least the reasoning version does not produce any usable German with Q8. Based on the great multilingual performance of the Magistral and Mixtral models I was hoping for better results.
@osanseviero a) Gemma3n is doing a good job in following instructions for thinking. b) Allowing to set a thinking budget between 0 and X,000 would be interesting. c) Add thinking as an adapter. d) Optimize for token efficiency during thinking.