MoonshotAi

Kimi K3

MoonshotAiFlagship
ThinkingTool UseVisionStructured Output

About this model

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Performance Tier

Flagship

Kimi K3 is a flagship model from MoonshotAi : the most capable in their lineup.

Best-in-class model from this provider. Highest performance across benchmarks, ideal for demanding tasks.

Pricing

This model is included in Elosia plans
High

High pricing. Reserve this level for tasks that demand maximum quality.

Typeper 1M tokens
Input (prompt)$3.00
Output (completion)$15.00
Cache read$0.300

Capabilities

Context Length1.0M
Max Output Tokens
TokenizerOther
Inputtext, image
Outputtext
Release DateJuly 16, 2026

Benchmarks

General Intelligence
MMLU
Not reported
GPQA Diamond
93.5%
Mathematics
MATH-500
Not reported
Programming
HumanEval
Not reported
SWE-bench Verified
Not reported
Reasoning
IFEval
Not reported
Humanity's Last Exam
56%
Multimodal
MMMU-Pro
81.6%
Agentic
Terminal-Bench 2.1
88.3%

Where does Kimi K3 stand?

Compare its performance index against every other model.

View the performance leaderboard

Recommended Use Cases

CodingAnalysisResearchGeneral Chat

Strengths

  • Moonshot flagship at 2.8T parameters, activating 16 of 896 experts per token (LatentMoE)
  • Graduate-level reasoning at the top of the field (GPQA Diamond 93.5%), independently confirmed by Artificial Analysis
  • 1M-token context, four times the 256K window of the Kimi K2 family
  • Native vision capabilities, with strong multimodal reasoning (MMMU-Pro 81.6%)
  • Kimi Delta Attention with attention residuals, for roughly 2.5x the scaling efficiency of Kimi K2 per Moonshot

Limitations

  • Moonshot official evaluation focused on agentic, reasoning and vision benchmarks: MMLU, MATH-500 and HumanEval are not part of the reported suite
  • Thinking is always on and reasoning_effort accepts only "max", so there is no instant path for simple calls
  • Most expensive model in the Kimi line by a wide margin ($3/M input, $15/M output)
  • Terminal-Bench 2.1 is self-reported at 88.3%, while independent runs land lower (Artificial Analysis 85.0%, Vals AI 80.9%)

Frequently asked questions

Resources

This model may use your data for training

Similar Models