Gemini 3.6 Flash achieved a 100% accuracy score on the seven-day standing benchmark without incurring execution spend, setting the top mark across all evaluated models. During the same 24-hour observation window, six OpenAI models—GPT-5 Nano, GPT-5 Mini, GPT-5.4 Mini, GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna—returned to responding following prior failed checks. However, availability metrics across the entire test suite recorded a uniform 25% success rate for all models.

Benchmark

The seven-day standing benchmark evaluates model accuracy relative to total token spend. Gemini 3.6 Flash captured all available points, followed by GPT-5.6 Terra at 94.5%.

ModelAccuracyTotal Spend
Gemini 3.6 Flash100%$0
GPT-5.6 Terra94.5%$0
Gemini 3.1 Flash Lite87.5%$0
Gemini 2.5 Flash80.2%$0
GPT-5.6 Luna71.9%$0

Gemini 3.1 Flash Lite recorded an 87.5% score, placing third overall, while Gemini 2.5 Flash reached 80.2%. GPT-5.6 Luna rounded out the benchmark group at 71.9%. Every tracked model completed its evaluation runs with $0 in total spend recorded over the seven-day window.

Speed

Across 24 hours of hourly automated checks, median response times ranged from 606ms to 3.69s. Gemini 2.5 Flash Lite posted the fastest median response time at 606ms, followed by GPT-5.4 Mini at 661ms and Gemini 2.5 Flash at 830ms.

Gemini 3.1 Flash Lite reached a median speed of 1.08s. GPT-5.6 Terra logged a 1.12s median response time, while GPT-5.6 Luna measured 1.16s and GPT-5.6 Sol recorded 1.63s. Higher latency models included Gemini 3.5 Flash at 2.10s, GPT-5 Mini at 2.31s, Gemini 3.6 Flash at 2.47s, and GPT-5 Nano at 2.56s. Gemini 3.1 Pro (preview) and Gemini 2.5 Pro registered the slowest median speeds at 3.47s and 3.69s, respectively.

Every monitored endpoint sustained a 25% success rate across all hourly automated checks throughout the 24-hour cycle. Six OpenAI models returned to responding status after recording prior failed checks: GPT-5 Nano, GPT-5 Mini, GPT-5.4 Mini, GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna.

Prices

Billed API costs per million tokens vary significantly across model tiers and generations. Google's Gemini 3.1 Pro (preview) remains the highest-priced model in the observed dataset at $2.00 per million input tokens and $12.00 per million output tokens.

Gemini 3.5 Flash is billed at $1.50 per million input tokens and $9.00 per million output tokens. Gemini 2.5 Pro costs $1.25 per million input tokens and $10.00 per million output tokens. Among the lightweight variants, Gemini 2.5 Flash costs $0.300 per million input tokens and $2.50 per million output tokens, while Gemini 3.1 Flash Lite runs $0.250 per million input tokens and $1.50 per million output tokens. Gemini 2.5 Flash Lite represents the lowest entry price point in the catalog at $0.100 per million input tokens and $0.400 per million output tokens.

These empirical metrics report observed response speeds, success rates, prices, and accuracy under standardized test conditions. They do not demonstrate internal provider infrastructure conditions, future operational uptime, or unmeasured workload behavior.