@patloeber flash-lite should be fine tuned on tasks where we are worried about latency. Specifically, document retrieval, OCR, summarization. For production, I care less about instant code and more about those kinds of tasks. So if flash-lite were best in class at that, game changing.
@pashov You should not be charging much higher than the direct token price. If you do, you’ll get out priced by other people pricing more fairly. E.g 2X the token price of opus is reasonable, giving you a 100% margin