In our latest round of direct evaluation conducted on October 6, 2026, Gemini 3.6 Flash recorded a top-tier result by capturing 100% of available points on our standing seven-day benchmark while incurring zero billed expenditure. This matching total puts Gemini 3.6 Flash at the apex of our accuracy testing alongside GPT-5 Mini, which also achieved 100% accuracy at zero recorded spend during the same seven-day window.

Speed

Over the past 24 hours of continuous hourly testing, every monitored endpoint delivered a 100% success rate without a single dropped request. Response latencies across the suite revealed sharp performance divisions between compact lightweight architectures and larger reasoning models.

[Gemini 2.5 Flash Lite](/models/google-gemini-2-5-flash-lite) registered the lowest median response time across all tested models at 585ms. OpenAI's GPT-5.4 Mini followed as the second-fastest option, recording a median latency of 706ms. Gemini 3.1 Flash Lite held third place at 834ms, completing the set of models operating under the one-second threshold.

ModelMedian Response TimeUptime
Gemini 2.5 Flash Lite585ms100%
GPT-5.4 Mini706ms100%
Gemini 3.1 Flash Lite834ms100%
Gemini 2.5 Flash1.0s100%
GPT-5.6 Terra1.1s100%
GPT-5.6 Sol1.2s100%
GPT-5.6 Luna1.3s100%
Gemini 3.6 Flash2.0s100%
GPT-5 Nano2.3s100%
GPT-5 Mini2.3s100%
Gemini 3.5 Flash2.3s100%
Gemini 3.1 Pro (preview)3.6s100%
Gemini 2.5 Pro3.9s100%

Models in the mid-range band clustered between 1.0s and 2.3s median latency. Gemini 2.5 Flash checked in at 1.0s, while OpenAI's GPT-5.6 family demonstrated consistent scaling across its variants: GPT-5.6 Terra posted 1.1s, GPT-5.6 Sol logged 1.2s, and GPT-5.6 Luna registered 1.3s. Further back, Gemini 3.6 Flash reached 2.0s, followed closely by GPT-5 Nano, GPT-5 Mini, and Gemini 3.5 Flash, which all registered 2.3s median response times.

The slowest tier comprised the flagship developer previews and production Pro endpoints. Gemini 3.1 Pro (preview) recorded a median response time of 3.6s, while Gemini 2.5 Pro came in at 3.9s.

Prices

Our price tracking for Google's Gemini lineup shows distinct tiers based on model generation and target workload size. Gemini 2.5 Flash Lite remains the cheapest entry point in the catalog, billed at $0.1 per million input tokens and $0.4 per million output tokens. Gemini 3.1 Flash Lite costs $0.3 per million input tokens and $1.5 per million output tokens, placing it slightly above Gemini 2.5 Flash at $0.3 per million input and $2.5 per million output tokens.

ModelInput Cost (/M)Output Cost (/M)
Gemini 2.5 Flash Lite$0.1$0.4
Gemini 3.1 Flash Lite$0.3$1.5
Gemini 2.5 Flash$0.3$2.5
Gemini 2.5 Pro$1.3$10.0
Gemini 3.5 Flash$1.5$9.0
Gemini 3.1 Pro (preview)$2.0$12.0

At the high end of the pricing spectrum, Gemini 2.5 Pro is billed at $1.3 per million input tokens and $10.0 per million output tokens. Gemini 3.5 Flash costs $1.5 per million input tokens and $9.0 per million output tokens. The most expensive option in the measured set is Gemini 3.1 Pro (preview), which commands $2.0 per million input tokens and $12.0 per million output tokens.

Benchmark

Our standing benchmark results over the last seven days measure task completion accuracy against total execution expenditure. Gemini 3.6 Flash and GPT-5 Mini shared the top position by securing 100% of available points at $0 total spend.

GPT-5.6 Terra captured third place with 86.7% of points at $0 spend, followed by Gemini 3.1 Flash Lite at 84.9% of points. Gemini 2.5 Flash scored 75.5% of points, outperforming GPT-5.6 Luna, which recorded 72.9% of points at $0 spend.

These measurements reflect observed network responses, billed rates, and task completion metrics collected directly during the specified sampling windows. They do not demonstrate unmeasured internal architectural changes, underlying hardware provisioning adjustments, or long-term operational stability beyond the reported timeframes.