How the numbers are made

Quality scores come from Artificial Analysis and public leaderboards. We never grade our own models. Prices are Together's published rates for the models we serve, and the lab's API rate otherwise. Speed for every host, Together included, comes from Artificial Analysis provider pages, with the live run widget as a spot check.

Where each number comes from

Quality

Attributed

A 0-100 intelligence index from Artificial Analysis and the public evals behind it. We publish the weights, but we don't run the evals or assign the scores.

  • Source-linked on every figure
  • Versioned + capture-dated
  • Reasoning is a sub-view of the same board

Price

First-party

Per-1M-token rates from Together's published pricing and OpenAI-compatible API, split into input / output / cached, blended at a 7:2:1 cache:input:output ratio, normalized to tiktoken tokens.

  • Cached-input discount surfaced
  • Index-run cost ties price to quality
  • Competitor rates attributed

Speed

Attributed

Output tokens/sec and time to first token come from Artificial Analysis provider-page medians, so every host is measured the same way. The live run widget streams a real completion as a spot check.

  • Quantization disclosed where the source labels it
  • Same harness for every host
  • Live widget runs a real request

Open Index v1, sourced from Artificial Analysis v4.1

A weighted average of named, contamination-resistant public evals. Scores are recalibrated so top models land in the 50s-60s, leaving headroom as models improve.

CategoryEvaluationWeight
Agents34%GDPval-AA v2Real-world knowledge work across 44 US occupations, scored as Elo against human experts (human = 1000).20%
τ³-Bench BankingAgentic banking tool-use over a large policy knowledge base.14%
Coding24%Terminal-Bench 2.1Terminal, SWE and sysadmin tasks executed in a real shell.16%
SciCode288 scientific-computing subproblems drawn from research code.8%
Scientific Reasoning24%Humanity's Last ExamFrontier academic questions across disciplines, run without tools.12%
GPQA Diamond198 graduate-level, Google-proof science MCQs.6%
CritPtResearch-level physics problems.6%
General18%AA-Omniscience Accuracy6,000 knowledge questions scored for accuracy.8%
AA-Omniscience Non-HallucinationThe same suite, scored for declining to guess when unsure.4%
AA-LCR100 long-context reasoning questions over ~100K tokens.6%

Weights mirror Artificial Analysis Intelligence Index v4.1. This site re-presents AA's published per-model scores; it does not run these evaluations.

How the per-category frontiers are built

The use-case explorer rolls public leaderboards into one composite score per category. Here are the rules, all of them checkable.

Benchmark grouping

Capability categories and slugs mirror whichLLM so both tools share one dataset. Design and the domain frontiers (legal, healthcare, math, back-office support, finance, customer support) are curated additions. Security and multilingual are excluded: no public leaderboard covers them credibly yet, and we don't invent scores.

Vetted sources only

Benchmarks come from Artificial Analysis, Epoch AI, Arena, Scale AI or Vals AI leaderboards; an official project page is used only when no aggregator covers the suite. Every chip and table column links to its source.

Field normalization

Each benchmark is min-max scaled against the tracked field with small margins (floor = min - 0.12 x range, ceiling = max + 0.05 x range), so hard suites, Elo ratings and dollar-scored runs contribute comparably.

Composite averaging

A category score is the mean of the normalized benchmarks a model has results for. A missing result stays a gap (no vision support, or not run on a frontier-only suite) rather than counting as a zero.

Value pick

The cheapest open model that holds at least 65% of the open leader's category score. The quality floor keeps ultra-cheap models from winning every category on price alone.

Percent presentation

Quality deltas are shown as a share of the closed leader's composite score. The composite is field-relative, which makes this conservative: it never flatters open, and raw per-benchmark gaps in the tables are typically smaller.

Cross-listing

One published result can count in every category it speaks to: GDPval-AA v2 informs both world knowledge and agents; the Tau benchmarks inform agents, finance and customer support.

Same rules for everyone

Open and closed models run through identical grouping, normalization and averaging. No category exists because it flatters an open model.