Google has released Gemini 3.1 Flash-Lite Preview! This upgrades the fastest, lowest-cost Gemini model series, scoring 34 on the Artificial Analysis Intelligence Index while served at over 360 output tokens/sec, significantly faster than other first-party API endpoints
Key takeaways:
➤ Improved intelligence over Gemini 2.5 Flash-Lite: @GoogleDeepMind's Gemini 3.1 Flash-Lite Preview scores 34 on the Artificial Analysis Intelligence Index, up 12 points from Gemini 2.5 Flash-Lite (09-25). However, Gemini 3.1 Flash-Lite Preview had limited gains in tool use capabilities, matching Gemini 2.5 Flash-Lite (09-25) on Tau2-Telecom with 31%, and scoring 958 on GDPval-AA, 12 points behind gpt-oss-120b (high)
➤ Leading speed and latency: Gemini 3.1 Flash-Lite Preview maintains the same high speeds and low latency as Gemini 2.5 Flash-Lite (09-25), measuring at over 360 output tokens/s, with an average answer latency of 5.1s. To measure latency for reasoning models, we use time to first answer token, which accounts for both prefill processing, and thinking time
➤ Continued differentiation in multimodal capabilities: Gemini 3.1 Flash-Lite Preview scores 78% on our multi-model reasoning evaluation, MMMU-Pro. This places it above frontier models such as Claude Opus 4.6 (max) and Kimi K2.5 (reasoning). This adds to Google’s broader strength in multimodal reasoning, with the other Gemini 3 models taking out the top three spots on MMMU-Pro.
➤ Increased pricing: Gemini 3.1 Flash-Lite Preview is priced at $0.25/$1.5 per 1M input/output tokens, an over 3x increase from its predecessor based on a blended 3:1 input/output token price. Its cost to run the Artificial Analysis Intelligence Index is $94, an increase over Gemini 2.5 Flash Lite’s $38
➤ Other model details: Gemini 3.1 Flash-Lite Preview retains the same 1 million token context window as its predecessor, and includes support for tool calling, structured outputs, and JSON mode