How the numbers are made
Quality scores come from Artificial Analysis and public leaderboards. We never grade our own models. Prices are Together's published rates for the models we serve, and the lab's API rate otherwise. Speed for every host, Together included, comes from Artificial Analysis provider pages, with the live run widget as a spot check.
Where each number comes from
Quality
AttributedA 0-100 intelligence index from Artificial Analysis and the public evals behind it. We publish the weights, but we don't run the evals or assign the scores.
- Source-linked on every figure
- Versioned + capture-dated
- Reasoning is a sub-view of the same board
Price
First-partyPer-1M-token rates from Together's published pricing and OpenAI-compatible API, split into input / output / cached, blended at a 7:2:1 cache:input:output ratio, normalized to tiktoken tokens.
- Cached-input discount surfaced
- Index-run cost ties price to quality
- Competitor rates attributed
Speed
AttributedOutput tokens/sec and time to first token come from Artificial Analysis provider-page medians, so every host is measured the same way. The live run widget streams a real completion as a spot check.
- Quantization disclosed where the source labels it
- Same harness for every host
- Live widget runs a real request
Open Index v1, sourced from Artificial Analysis v4.1
A weighted average of named, contamination-resistant public evals. Scores are recalibrated so top models land in the 50s-60s, leaving headroom as models improve.
| Category | Evaluation | Weight |
|---|---|---|
| Agents34% | GDPval-AA v2Real-world knowledge work across 44 US occupations, scored as Elo against human experts (human = 1000). | 20% |
| τ³-Bench BankingAgentic banking tool-use over a large policy knowledge base. | 14% | |
| Coding24% | Terminal-Bench 2.1Terminal, SWE and sysadmin tasks executed in a real shell. | 16% |
| SciCode288 scientific-computing subproblems drawn from research code. | 8% | |
| Scientific Reasoning24% | Humanity's Last ExamFrontier academic questions across disciplines, run without tools. | 12% |
| GPQA Diamond198 graduate-level, Google-proof science MCQs. | 6% | |
| CritPtResearch-level physics problems. | 6% | |
| General18% | AA-Omniscience Accuracy6,000 knowledge questions scored for accuracy. | 8% |
| AA-Omniscience Non-HallucinationThe same suite, scored for declining to guess when unsure. | 4% | |
| AA-LCR100 long-context reasoning questions over ~100K tokens. | 6% |
How the per-category frontiers are built
The use-case explorer rolls public leaderboards into one composite score per category. Here are the rules, all of them checkable.
Benchmark grouping
Capability categories and slugs mirror whichLLM so both tools share one dataset. Design and the domain frontiers (legal, healthcare, math, back-office support, finance, customer support) are curated additions. Security and multilingual are excluded: no public leaderboard covers them credibly yet, and we don't invent scores.
Vetted sources only
Benchmarks come from Artificial Analysis, Epoch AI, Arena, Scale AI or Vals AI leaderboards; an official project page is used only when no aggregator covers the suite. Every chip and table column links to its source.
Field normalization
Each benchmark is min-max scaled against the tracked field with small margins (floor = min - 0.12 x range, ceiling = max + 0.05 x range), so hard suites, Elo ratings and dollar-scored runs contribute comparably.
Composite averaging
A category score is the mean of the normalized benchmarks a model has results for. A missing result stays a gap (no vision support, or not run on a frontier-only suite) rather than counting as a zero.
Value pick
The cheapest open model that holds at least 65% of the open leader's category score. The quality floor keeps ultra-cheap models from winning every category on price alone.
Percent presentation
Quality deltas are shown as a share of the closed leader's composite score. The composite is field-relative, which makes this conservative: it never flatters open, and raw per-benchmark gaps in the tables are typically smaller.
Cross-listing
One published result can count in every category it speaks to: GDPval-AA v2 informs both world knowledge and agents; the Tau benchmarks inform agents, finance and customer support.
Same rules for everyone
Open and closed models run through identical grouping, normalization and averaging. No category exists because it flatters an open model.
What comes from where
- Artificial Analysis Intelligence Index v4.1, per-eval scores + capability indexes
- Per-provider speed/latency/price (AA provider pages)
- Epoch AI benchmarking hub (SWE-Bench Verified, FrontierMath)
- Arena leaderboards (LMArena, WebDev Arena)
- Scale AI leaderboards (MCP Atlas)
- Vals AI boards (Legal Research, HLAB, MedScribe, ProofBench, CorpFin, EMB, Finance Agent, TaxEval)
- Official boards (EQBench, Design Arena, OSWorld, MMMU-Pro)
- Historical frontier trajectory (v4.1 rescored)
Freshness & limitations
Every figure here was read from its cited source on Aug 3, 2026. Rankings shift with every model release, and Artificial Analysis refreshes about 8× a day, so expect these numbers to drift until the next snapshot.