Ollama v0.33 is here!
You can now easily configure Claude Desktop to seamlessly work with Ollama as a third-party gateway provider.
One toggle. Cloud & local models just work👇
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: https://t.co/tzOmB7gdZP
Available now across all official platforms:
Weights: https://t.co/9LRMahY9Wa
API: https://t.co/VcaQnzYmS9
Coding Plan: https://t.co/Nk8Y98HNhU
ZCode: https://t.co/Peepqv4XSx
Chat: https://t.co/WCqWT0qCQb
AutoClaw: https://t.co/aGEG5HqTTb
Just FYI. If you are switching models in a session all the time - you are doing it wrong.
Every time you switch, your entire prompt cache is invalidated on the new model you switch to, and you have to repay the full input tokens price for all of it.
Stop doing this unless those models are free.
This is not a hermes thing - this is a fundamentals of inference thing.