GLM 5.2 is a great model but it's not multimodal. Not being able to send reference images or errors was sort of a deal breaker.
On ZCode it has a vision MCP that lets the model work with images. I wanted something similar for me in OpenCode, thus I've made a multimodal plugin that lets all models work with images, audios and other files with the help of another (preferably cheap) model.
Try it out if you are looking for something similar!
https://t.co/1yOLbqGTQO