@MiaAI_lab@huggingface Alright. What's your tps? I get approx 40-50 right now with your latest recipe. Seems weird why... Drafter is 66 tps and 40-50 acceptance rate.
@MiaAI_lab@pradeepXkapoor This is something that I will believe only when I am running it on my machine. If it's true, I will delete all the other models. ๐
@MiaAI_lab Actually it's not just smarter but handles long context better and can solve problems with less thinking, especially when it's using web search.
@MiaAI_lab@thatcofffeeguy Soketimes I have. Probably the use case is different. (3dsmax sdk and other low level macro heavy stuffs)
For that it's not always that reliable unfortunately. Mimo handle those cases better I think.
@MiaAI_lab What I noticed that it's weirdly slow especially for an A3B model.
I am using the NVFP4 quant with DFlash. Reducing the mtp to 8 and also the temperature make it a bit faster, but it's still somewhere close to 30 tps.
@MiaAI_lab@aleksandar_xyz Yeah but the quality difference is significant. Actually I don't what language / framework / sdk / api. For me it's 3ds max mostly and it has a pretty big and macro heavy code base. In that case Mimo gives noticable better results.
@aleksandar_xyz@MiaAI_lab It's much better in planning. The code quality seems more structured. It keeps the context much better and degrades much less over time as the context gets bigger.
@MiaAI_lab Mimo is doing the same thing. Actually it's a good thing because it's not going to make BS and the chancen of having an actial working solution is much higher.