The inference debate often comes down to cost per token, and I agree it's important
But the cheapest model isn't always the right model for the task, because domain expertise and task specific performance matter too, it matters a lot
My answer is @OpenAlmondHQ, an open market for specialized inference