Comparison

Gemini 3.5 Flash vs GPT-4o

Per input token, Gemini 3.5 Flash runs 40% cheaper than GPT-4o. Output pricing favors Gemini 3.5 Flash, 10% less per token than GPT-4o. Gemini 3.5 Flash has the larger context window than GPT-4o.

FieldGemini 3.5 Flash
GPT-4o
Legacy
ProviderGoogleOpenAI
Context window1M128K
Max output65.5K
Input / 1M$1.5040% cheaper$2.50
Output / 1M$9.0010% cheaper$10.00
Cached input / 1M$0.1588% cheaper$1.25
Batch discount50%
Release date

Worked cost examples

Estimated monthly cost on three common usage shapes

WorkloadGemini 3.5 FlashGPT-4o

Chatbot

800 input / 200 output tokens per request, 100,000 requests/month

$300.00cheaper$400.00

RAG pipeline

6,000 input / 700 output tokens per request, 30,000 requests/month

$459.00cheaper$660.00

Batch summarization

20,000 input / 1,500 output tokens per request, 5,000 requests/month

$217.50cheaper$325.00

Gemini 3.5 Flash vs GPT-4o FAQ

Is Gemini 3.5 Flash cheaper than GPT-4o?
Gemini 3.5 Flash costs $1.50 per 1M input tokens versus $2.50 for GPT-4o — Gemini 3.5 Flash is 40% cheaper per input token.
Which is cheaper for output tokens, Gemini 3.5 Flash or GPT-4o?
Gemini 3.5 Flash charges $9.00 per 1M output tokens; GPT-4o charges $10.00 per 1M output tokens.
Which has the larger context window, Gemini 3.5 Flash or GPT-4o?
Gemini 3.5 Flash does, with a 1,048,576-token context window.
Do Gemini 3.5 Flash and GPT-4o offer cache pricing?
Gemini 3.5 Flash discounts cached input to $0.15 per 1M tokens. GPT-4o discounts cached input to $1.25 per 1M tokens.

Looking for another matchup? See all comparisons.