Together AI · open model benchmarks

Open models
caught the frontier.

Kimi K3's weights are out. The open frontier now holds 93% of the top score, 4 points off the frontier, on weights you can take anywhere.

26 models · 24 providers · no pay-for-placement

Frontier indexopen vs closed · AA v4.1
204060CLOSED61OPEN572024now
Gap in 2024
0 pts
Gap today
0 pts
Open leader
0
Kimi K3 · weights out
0Models tracked15 open source
0Open leaderKimi K3
0.0×Intelligence / $best open vs best closed
0Providers comparedincl. first-party APIs

One index hides the fit. Pick your use case.

Every model on one chart

The whole field on one composite index: filter, sort, and inspect any model.

26 models shown · click a point or row to inspect

Intelligence vs cost to run the index

◤ most attractive quadrant$50$100$200$500$1K$2K$5Kcost to run the full index (USD, log)102030405060Intelligence IndexZClaude Opus 5GPT-5.6 SolGPT-5.6 TerraGrok 4.5GPT-5.6 LunaKimi K3DeepSeek V4 Flash 0731gpt-oss-20b
↑ higher quality · ← cheaper is better · dashed = Pareto frontier26 of 26 modelsAttributedArtificial Analysis v4.1

Leaderboard

Ranked by Intel / $: index points per $100 of index-run cost. Click # for the Pareto frontier order.
Model
1
DeepSeek V4 Flash 0731open
DeepSeek
5069.4$72
2
gpt-oss-20bopen
OpenAI
1541.2$36
3
GPT-5.6 Lunaclosed
OpenAI
5129.3$174
4
DeepSeek V4 Proopen
DeepSeek
4425.0$176
5
gpt-oss-120bopen
OpenAI
2424.9$96
6
MiniMax-M3open
MiniMax
4421.6$204
7
Llama 3.3 70Bopen
Meta
911.2$81
8
Muse Spark 1.1closed
Meta
519.3$548

Kimi K3

open
Moonshot (Kimi) · Open source · 2800B
Intelligence Index57
Intelligence / $1002.3index pts per $100 run cost
Index run cost$2.4K
Context1M
Price in / out$3.00 / $15.0per 1M tokens
Cached input$0.3010× cheaper
Capability breakdown0 to 100
Agents
50
Coding
76
GPQA
94
HLE
44
Long context
75

Scores attributed to Artificial Analysis & underlying public evals.

Open is catching up, fast

The open frontier gained 48 index points in under two years. Kimi K3's weights shipped July 27 and cut the gap to 4 points, the closest open has ever been.

Best open vs best closed intelligence

DeepSeek V3DeepSeek R1gpt-oss-120bK2 ThinkingDeepSeek V4 ProGLM-5.2Kimi K3Closed 61Open 57Oct ’24Jan ’25Apr ’25Jul ’25Oct ’25Jan ’26Apr ’26Jul ’260204060
↑ higher is better · shaded band = the gapAttributedArtificial Analysis (historical)

Run the open frontier yourself.

Same OpenAI-compatible API, open weights you keep, at a fraction of closed-API cost.

Copied
# uv add together
from together import Together
client = Together()
response = client.chat.completions.create(
model="moonshotai/Kimi-K3",
messages=[{"role": "user", "content": "What are the top 3 things to do in New York?"}],
)
print(response.choices[0].message.content)