Gemini

Gemini 3.7 Flash

GeminiBalanced
ThinkingTool UseVisionStructured Output

About this model

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Performance Tier

Balanced

Gemini 3.7 Flash is a balanced model from Gemini : strong performance at a reasonable price.

Strong cost-performance ratio. Reliable for most professional use cases without premium pricing.

Pricing

This model is included in Elosia plans
Affordable

Low cost. Suitable for sustained use and high-volume interactions.

Typeper 1M tokens
Input (prompt)$0.375
Output (completion)$1.88
Image$0.375
Internal reasoning$1.88
Cache read$0.037
Cache write$0.021

Capabilities

Context Length1.0M
Max Output Tokens66K
TokenizerGemini
Inputtext, image, video, file, audio
Outputtext
Release DateAugust 13, 2026

Benchmarks

General Intelligence
MMLU
Not reported
GPQA Diamond
94.5%
Mathematics
MATH-500
Not reported
Programming
HumanEval
Not reported
Reasoning
Humanity's Last Exam
47.9%
Multimodal
MMMU-Pro
85.5%
Agentic
Terminal-Bench 2.1
85.8%

Where does Gemini 3.7 Flash stand?

Compare its performance index against every other model.

View the performance leaderboard

Recommended Use Cases

CodingAnalysisGeneral ChatResearchData Extraction

Strengths

  • Strongest agentic terminal work in the Flash line (Terminal-Bench 2.1 85.8%, up from 78.0% on Gemini 3.6 Flash)
  • Large software-engineering gain over its predecessor (DeepSWE v1.1 65.3% vs 48.6%, FrontierCode 1.1 Main 43.6% vs 34.4%)
  • Leading web development among Flash models (Code Arena 1588 Elo, up from 1538)
  • Near-perfect long-context retrieval (GDM-MRCR v2 8-needle 97.0% at 128k average); separately, the model supports a 1M-token context window with 64K output
  • Broad multimodal coverage: video (LVBench 85.4%), computer use (OSWorld-2.0 47.9%)
  • Best relative standing against rival models on multidisciplinary reasoning (HLE-Verified 53.6%)

Limitations

  • The only regression versus Gemini 3.6 Flash in Google's table is CharXiv Reasoning (84.5% vs 85.2% without tools, 88.7% vs 89.4% with tools), a step back for chart-heavy analysis and data extraction
  • Last of the four models Google compares on GDPVal-AA v2 knowledge work (1525 Elo versus 1578 to 1628 for the rival models)
  • Behind GPT-5.6 Terra on DeepSWE v1.1 (65.3% vs 69.6%), Terminal-Bench 2.1 and 3.0, and OSWorld-2.0 (47.9% vs 50.2%), and behind Claude Sonnet 5 on Agent's Last Exam (26.3% vs 33.3%)
  • Terminal-Bench 3.0 at 14.9% shows the harder agentic frontier remains largely unsolved
  • A refinement of Gemini 3.6 Flash rather than a new pretraining run
  • Google publishes no MMLU, GPQA Diamond or MMMU-Pro figures; the GPQA Diamond, MMMU-Pro and Humanity's Last Exam scores here come from third-party testing (Artificial Analysis)

Frequently asked questions

Resources

This model may use your data for training

Similar Models