[Gemini 2.5 Flash Lite](/models/google-gemini-2-5-flash-lite) delivered the fastest median response time among all measured models over the last 24 hours, recording 576ms with a 100% check success rate. In addition to leading latency metrics, it remains the lowest-cost model observed in the current price catalog, billed at $0.100 per million input tokens and $0.400 per million output tokens.

Across all monitored endpoints, operational availability remained high over the 24-hour evaluation window, with every checked model logging a 100% uptime rate across hourly checks. However, latency and pricing continue to span a wide range across different model architectures and capability tiers.

Prices

Pricing observed across Google's Gemini model family demonstrates clear differentiation based on tier. Gemini 2.5 Flash Lite provides the lowest cost baseline, while Gemini 3.1 Pro (preview) sits at the top of the price range at $2.00 per million input tokens and $12.00 per million output tokens.

ModelInput Price ($/M)Output Price ($/M)
Gemini 2.5 Flash Lite$0.100$0.400
Gemini 3.1 Flash Lite$0.250$1.50
Gemini 2.5 Flash$0.300$2.50
Gemini 2.5 Pro$1.25$10.00
Gemini 3.5 Flash$1.50$9.00
Gemini 3.1 Pro (preview)$2.00$12.00

Sub-dollar input rates are confined to the Flash Lite and Flash variants of the 2.5 and 3.1 generations. Higher-capacity variants like Gemini 2.5 Pro ($1.25 input / $10.00 output) and Gemini 3.5 Flash ($1.50 input / $9.00 output) establish a distinct middle-to-high pricing tier.

Speed

Speed measurements collected from hourly checks throughout the past 24 hours show that lightweight models reliably sustain sub-second response times. Gemini 2.5 Flash Lite led all systems at 576ms median response time, followed by GPT-5.4 Mini at 760ms and Gemini 3.1 Flash Lite at 818ms.

Mid-tier latency results included GPT-5.6 Luna at 1.06s, GPT-5.6 Terra at 1.17s, Gemini 2.5 Flash at 1.25s, and GPT-5.6 Sol at 1.45s. Longer response durations were logged for Gemini 3.6 Flash at 2.14s, GPT-5 Mini at 2.19s, Gemini 3.5 Flash at 2.43s, and GPT-5 Nano at 2.52s.

The highest median latencies in the 24-hour test window came from larger pro offerings. Gemini 3.1 Pro (preview) recorded a median response time of 3.63s, while Gemini 2.5 Pro registered 3.93s. Every model evaluated achieved a 100% success rate across all hourly automated checks.

Benchmark

Results from standing benchmark evaluations over the last 7 days show two models sharing the top position. Gemini 3.6 Flash and GPT-5 Mini each captured 100% of available points, with both incurring $0 in total benchmark spend.

Behind the leaders, Gemini 3.1 Flash Lite earned 85.2% of available points ($0 spent), closely followed by GPT-5.6 Terra at 84.9% of points ($0 spent). Further down the 7-day standings, GPT-5.6 Luna achieved 76% ($0 spent), while Gemini 2.5 Flash reached 74% ($0 spent).

These measurements reflect automated performance, baseline pricing, and benchmark scores collected under specific API test conditions. They do not predict performance under heavy concurrency, long-term provider reliability, or behavior on specialized enterprise workloads.