Grok

Grok 4.5

GrokFlagship
ThinkingTool UseVisionStructured Output

About this model

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Performance Tier

Flagship

Grok 4.5 is a flagship model from Grok : the most capable in their lineup.

Best-in-class model from this provider. Highest performance across benchmarks, ideal for demanding tasks.

Pricing

This model is included in Elosia plans
Moderate

Moderate cost. A balanced choice for regular use without constant cap watching.

Typeper 1M tokens
Input (prompt)$2.00
Output (completion)$6.00
Cache read$0.300

Capabilities

Context Length500K
Max Output Tokens
TokenizerGrok
Inputtext, image, file
Outputtext
Release DateJuly 8, 2026

Benchmarks

General Intelligence
MMLU
Not reported
GPQA Diamond
Not reported
Mathematics
MATH-500
Not reported
Programming
HumanEval
Not reported
SWE-bench Verified
Not reported
Reasoning
IFEval
Not reported
Agentic
SWE-bench Pro
64.7%
Terminal-Bench 2.1
83.3%

Where does Grok 4.5 stand?

Compare its performance index against every other model.

View the performance leaderboard

Recommended Use Cases

CodingAnalysisResearchGeneral ChatMathematics

Strengths

  • xAI flagship, strong on coding, agentic tasks and STEM knowledge work
  • Competitive agentic coding: 83.3 on Terminal-Bench 2.1 and 64.7 on SWE-bench Pro (single-model)
  • Notably token-efficient on agentic coding versus comparable flagships, at about 80 tokens/second (per xAI)
  • 500K-token context window with multimodal (text, image) input, function calling, web and X search, and code execution

Limitations

  • xAI published only coding and agentic benchmarks; no MMLU, GPQA, math or reasoning scores to gauge broader capability
  • Trails the current SWE-bench Pro coding leaders (64.7 vs about 80)
  • Smaller third-party tooling ecosystem than OpenAI or Anthropic

Frequently asked questions

Resources

This model may use your data for training

Similar Models