Gemini 3.6 Flash and GPT-5 Mini recorded perfect 100% results in our 7-day standing benchmark evaluation. Both models logged $0 in spend over the evaluation period while maintaining full accuracy across tested points. GPT-5.6 Terra followed with an 88.5% score, while Gemini 3.1 Flash Lite earned 85.9%.
Benchmark
Our standing benchmark measures performance as a percentage of available points accumulated over a 7-day period. Every model evaluated during this window incurred $0 in total spend.
| Model | Benchmark Score | Total Spend |
|---|---|---|
| Gemini 3.6 Flash | 100% | $0 |
| GPT-5 Mini | 100% | $0 |
| GPT-5.6 Terra | 88.5% | $0 |
| Gemini 3.1 Flash Lite | 85.9% | $0 |
| GPT-5.6 Luna | 78.1% | $0 |
| Gemini 2.5 Flash | 70.8% | $0 |
Gemini 3.6 Flash and GPT-5 Mini achieved identical 100% accuracy ratings, outpacing third-place GPT-5.6 Terra by 11.5 percentage points. Gemini 3.1 Flash Lite placed fourth with 85.9%, demonstrating strong performance relative to lower-tier options like GPT-5.6 Luna at 78.1% and Gemini 2.5 Flash at 70.8%.
Speed
Response speed across all monitored systems remained stable over the last 24 hours, with every model completing 100% of hourly uptime checks without failure.
Gemini 2.5 Flash Lite recorded the fastest median response time at 558ms, followed closely by GPT-5.4 Mini at 666ms and Gemini 2.5 Flash at 860ms. OpenAI models occupied much of the mid-tier latency band, with GPT-5.6 Luna logging a median response of 1.01s, GPT-5.6 Sol at 1.06s, and GPT-5.6 Terra at 1.16s. Gemini 3.1 Flash Lite returned a 1.22s median response time.
On the slower end of the spectrum, Gemini 3.6 Flash reached 1.84s median latency despite its 100% benchmark score. GPT-5 Nano followed at 1.89s, while GPT-5 Mini registered 2.09s and Gemini 3.5 Flash reached 2.17s. The slowest response times belonged to preview and flagship tier models: Gemini 3.1 Pro (preview) recorded a median response time of 3.52s, and Gemini 2.5 Pro recorded 3.58s. All 13 monitored endpoints maintained 100% success rates across all hourly checks during the 24-hour observation window.
Prices
Pricing models for Google endpoints show clear distinctions across generations and sub-tiers per million tokens processed.
| Model | Input Price / M | Output Price / M |
|---|---|---|
| Gemini 2.5 Flash Lite | $0.100 | $0.400 |
| Gemini 3.1 Flash Lite | $0.250 | $1.50 |
| Gemini 2.5 Flash | $0.300 | $2.50 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |
| Gemini 3.1 Pro (preview) | $2.00 | $12.00 |
The lowest pricing tier observed is Gemini 2.5 Flash Lite at $0.100 per million input tokens and $0.400 per million output tokens. Stepping up to Gemini 3.1 Flash Lite increases input costs to $0.250/M and output costs to $1.50/M. Gemini 2.5 Flash is priced at $0.300/M for input and $2.50/M for output.
At the higher end of the pricing spectrum, Gemini 2.5 Pro costs $1.25/M input and $10.00/M output. Gemini 3.5 Flash comes in at $1.50/M input and $9.00/M output, featuring higher input costs but lower output costs than Gemini 2.5 Pro. Gemini 3.1 Pro (preview) represents the highest listed rate across measured endpoints, charging $2.00/M input and $12.00/M output.
These measurements reflect direct synthetic tests gathered across controlled hourly API checks. They demonstrate latency, uptime, benchmark points, and pricing figures as captured during this specific evaluation window, but they do not prove long-term model reliability or behavior under varying real-world enterprise workloads.

