Groq has set a world record in LLM inference API speed by serving Llama 3.2 1B at >3k output tokens/s π
Meta's Llama 3.2 3B and 1B models are well positioned for two categories of use-cases. Firstly, applications running on edge devices or on-device, where compute resources are limited. Secondly, use-cases which require very fast response times and/or very cheap token generation.
@GroqInc with their custom LPU chips are taking fast and cheap token generation to the extreme by serving the models at >3k tokens per second and pricing at $0.04/1M input/output tokens. To put this in context, this is ~25X faster than GPT-4o's API and ~110X cheaper.
While intelligence of these models is not comparable to the much larger frontier models, not all use-cases require frontier intelligence. Consumer apps which require real-time interaction and cheap token generation, live monitoring and classification are both example use-cases which suit these smaller models.
Link below for our analysis of how Llama 3.2 3B & 1B compare to other smaller models, and of the providers serving them π