GPT

GPT-5.6 Sol Pro

GPTFlagship
ThinkingTool UseVisionStructured Output

About this model

GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

Performance Tier

Flagship

GPT-5.6 Sol Pro is a flagship model from GPT : the most capable in their lineup.

Best-in-class model from this provider. Highest performance across benchmarks, ideal for demanding tasks.

Pricing

This model is included in Elosia plans
Premium

Highest cost level. A long conversation can quickly consume your monthly cap.

Typeper 1M tokens
Input (prompt)$5.00
Output (completion)$30.00
Cache read$0.500
Cache write$6.25

Capabilities

Context Length1.1M
Max Output Tokens128K
TokenizerGPT
Inputfile, image, text
Outputtext
Release DateJuly 9, 2026

Benchmarks

General Intelligence
MMLU
Not reported
GPQA Diamond
Not reported
Mathematics
MATH-500
Not reported
Programming
HumanEval
Not reported
SWE-bench Verified
Not reported
Reasoning
IFEval
Not reported
ARC-AGI-2
85.4%
Agentic
SWE-bench Pro
64.6%
Terminal-Bench 2.1
88.8%

Where does GPT-5.6 Sol Pro stand?

Compare its performance index against every other model.

View the performance leaderboard

Recommended Use Cases

CodingAnalysisResearchMathematicsData Extraction

Strengths

  • Flagship of the GPT-5.6 family, served in pro reasoning mode for maximum quality on complex tasks
  • Leading abstract reasoning: ARC-AGI-2 85.4 at high effort, reaching 90.0 (xhigh) and 92.5 (max) with more reasoning
  • Frontier agentic coding: 88.8 on Terminal-Bench 2.1 and 64.6 on SWE-bench Pro (single-agent), strong on command-line and long-horizon work
  • 1M-token context window with programmatic tool calling and multimodal (text, image, file) input

Limitations

  • Pro reasoning mode consumes more tokens per query than base Sol at the same per-token price ($5 / $30), so it uses up quota faster
  • Trails the current SWE-bench Pro leaders on in-codebase resolution (64.6 vs about 80)
  • OpenAI published no MMLU, GPQA or math benchmarks for GPT-5.6, so direct comparison with older flagships is harder

Frequently asked questions

Resources

This model may use your data for training

Similar Models