Comparison
Gemini 2.5 Flash-Lite vs GPT-4o
Gemini 2.5 Flash-Lite costs 96% less per input token than GPT-4o. On generated (output) tokens, GPT-4o costs 96% more than Gemini 2.5 Flash-Lite.
| Field | Gemini 2.5 Flash-Lite | GPT-4o Legacy |
|---|---|---|
| Provider | OpenAI | |
| Context window | — | 128K |
| Max output | — | — |
| Input / 1M | $0.1096% cheaper | $2.50 |
| Output / 1M | $0.4096% cheaper | $10.00 |
| Cached input / 1M | $0.0199% cheaper | $1.25 |
| Batch discount | 50% | — |
| Release date | — | — |
Worked cost examples
Estimated monthly cost on three common usage shapes
| Workload | Gemini 2.5 Flash-Lite | GPT-4o |
|---|---|---|
Chatbot 800 input / 200 output tokens per request, 100,000 requests/month | $16.00cheaper | $400.00 |
RAG pipeline 6,000 input / 700 output tokens per request, 30,000 requests/month | $26.40cheaper | $660.00 |
Batch summarization 20,000 input / 1,500 output tokens per request, 5,000 requests/month | $13.00cheaper | $325.00 |
Gemini 2.5 Flash-Lite vs GPT-4o FAQ
- Is Gemini 2.5 Flash-Lite cheaper than GPT-4o?
- Gemini 2.5 Flash-Lite costs $0.10 per 1M input tokens versus $2.50 for GPT-4o — Gemini 2.5 Flash-Lite is 96% cheaper per input token.
- Which is cheaper for output tokens, Gemini 2.5 Flash-Lite or GPT-4o?
- Gemini 2.5 Flash-Lite charges $0.40 per 1M output tokens; GPT-4o charges $10.00 per 1M output tokens.
- Do Gemini 2.5 Flash-Lite and GPT-4o offer cache pricing?
- Gemini 2.5 Flash-Lite discounts cached input to $0.01 per 1M tokens. GPT-4o discounts cached input to $1.25 per 1M tokens.
Looking for another matchup? See all comparisons.