Gemini 2.5 Flash

google/gemini-2.5-flash
VisionAudioToolsJSONReasoning
by Google · 2025-06-17

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

ctx1.05M tokens
Max output65.5K
Inputfile + image + text + audio + video
Outputtext
p50 TTFT10.00 s
INPUT$0.30/ 1M tokens
OUTPUT$2.50/ 1M tokens
p50 TTFT10.00 s7d
p95 TTFT10.00 s7d
TRAFFIC433.8Ktokens / 7d

Google Gemini 2.5 Flash is a multimodal language model developed by Google. It belongs to the Gemini 2.5 series, designed for efficient processing of text, images, audio, and video inputs. The model…

What is Google Gemini 2.5 Flash?

Who should use Gemini 2.5 Flash?

What are the key specifications of Gemini 2.5 Flash?

How does Gemini 2.5 Flash differ from other Gemini models?

Code samples

Call from any SDK

OpenAI-compatible — keep the SDK you already use

  • OpenAI SDKhttps://api.orcarouter.ai/v1
  • Gemini SDKhttps://api.orcarouter.ai
from openai import OpenAI

client = OpenAI(
    base_url="https://api.orcarouter.ai/v1",
    api_key="$ORCAROUTER_API_KEY",
)

response = client.chat.completions.create(
    model="google/gemini-2.5-flash",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

Supported parameters

  • include_reasoning
  • max_tokens
  • reasoning
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_p

Pricing

Input / 1M tokens$0.300
Output / 1M tokens$2.50
Cache read / 1M$0.030
CurrencyUSD

Cost calculator

Tokens / month10MM
Input share70%%
Estimated / month $9.60 · With prompt caching $8.66

Estimate based on list price

Token & cost estimator

Input tokens: 20Cost per request: $0.001256

Estimate only — actual token counts depend on the provider's tokenizer.

Performance

p50 TTFT
10.00 s
Output speed
1247 tok/s
p95 TTFT
10.00 s
Error rate
2.3%

Public benchmarks

17.8
AA Coding
Better than 15% of models compared
#105 of 123
14.1
AA Intelligence
Better than 14% of models compared
#107 of 125
60.3
AA Math
Better than 44% of models compared
#45 of 81
AIME
50.0
AIME 2025
60.3
GPQA Diamond
68.3
Humanity's Last Exam
5.1
IFBench
39.0
LiveCodeBench
49.5
Long-Context Recall
45.9
MATH-500
93.2
MMLU-Pro
80.9
SciCode
29.1
TerminalBench Hard
12.1
τ²-Bench
14.9
Source: artificialanalysis.ai

How it compares

Gemini 2.5 FlashGemini 3.1 Pro PreviewGemini 3.1 Pro Preview Custom ToolsGemini 3 Flash Preview
Input $/M$0.30$2.00$4.00$0.50
Output $/M$2.50$12.00$18.00$3.00
Context1.0M1.0M1.0M1.0M
Quality3/1010/1010/109/10
Compare side-by-sideCompare side-by-sideCompare side-by-sideCompare side-by-side

FAQ

How much does Gemini 2.5 Flash cost per token?
Input is $0.30 per 1 million tokens, output is $2.50 per 1 million tokens. These are Google’s provider rates with zero markup through OrcaRouter.
What is the context window size of Gemini 2.5 Flash?
The context window is 1,048,576 tokens (1M tokens). Maximum output is 65,536 tokens.
What are the main strengths of Gemini 2.5 Flash?
It achieves a MATH-500 score of 93.2, demonstrating strong mathematical reasoning. It also supports multiple input modalities (file, image, text, audio, video) and has a very large context window at a competitive price.
How does Gemini 2.5 Flash compare to Gemini 2.0 Flash?
Gemini 2.5 Flash is more expensive but offers improved math performance (93.2 vs. lower score on MATH-500). Both have 1M context windows. 2.0 Flash costs $0.15/$0.60 per 1M tokens, while 2.5 Flash costs $0.30/$2.50.
Does OrcaRouter add any markup to the model price?
No, OrcaRouter bills at the exact provider rate with zero markup. You pay $0.30 per 1M input and $2.50 per 1M output tokens.
Can I use Gemini 2.5 Flash via an OpenAI-compatible API?
Yes. Use base URL https://api.orcarouter.ai/v1 and model ID "google/gemini-2.5-flash". All OpenAI SDKs and HTTP clients work with standard chat completion parameters.
What input modalities does Gemini 2.5 Flash support?
It supports file, image, text, audio, and video inputs. You can mix these in a single request.
How is my data handled when using Gemini 2.5 Flash through OrcaRouter?
Data handling follows Google’s API privacy policies and OrcaRouter’s terms. Google may process data for model improvement unless you opt out via your Google Cloud account. Check Google’s data processing documentation for details.

Embed this badge

Google: Gemini 2.5 Flash$0.30/M in10000ms p50via OrcaRouter
HTML <a href="https://www.orcarouter.ai/models/google/gemini-2.5-flash" target="_blank"> <img src="https://www.orcarouter.ai/embed/google/gemini-2.5-flash.svg" alt="Google: Gemini 2.5 Flash on OrcaRouter" /> </a>
Markdown [![Google: Gemini 2.5 Flash](https://www.orcarouter.ai/embed/google/gemini-2.5-flash.svg)](https://www.orcarouter.ai/models/google/gemini-2.5-flash)

Model card as data

GET /api/public/models/google/gemini-2.5-flashOpen
Machine-readable:/llms.txt/llms-full.txt