Where the numbers come from

Quality is Artificial Analysis's, never ours. Prices are published rates, Together's for the models Together serves. A script re-reads both every day, so what you see is what the sources said on Sep 16, 2026.

The short version

1We do not grade models

Every score on this site is Artificial Analysis's, taken from their public model pages and capability boards. One harness, every model, same conditions. We never run an eval or assign a score.

2Cost is what it cost

Alongside quality, AA publishes what it cost them to run the index against each model. That captures verbosity and reasoning burn, which a list price hides. Prices are the provider's published serverless rates.

3Nothing is hand-typed

A daily script reads the sources and rewrites the data file. No one edits a number by hand, so nothing drifts quietly out of date and there is no room to flatter an open model.

What the Intelligence Index measures

Ten public evaluations, weighted to 100 across four categories. This is AA's composition of Intelligence Index v4.3, reproduced here so you can see what a score is made of.

Agents 30%Coding 20%Scientific Reasoning 20%General 30%
  • 15%
    AA-Briefcase Agents Long-horizon business deliverables, scored as Elo against expert work and normalised for the index.
  • 15%
    AA-Omniscience General 6,000 knowledge questions, scored 10% for accuracy and 5% for declining to guess.
  • 10%
    GDPval-AA v2 Agents Real-world knowledge work across 44 US occupations, scored as Elo against human experts.
  • 10%
    Terminal-Bench 4.0 Coding 66 terminal, SWE and sysadmin tasks executed in a real shell.
  • 10%
    SciCode Coding 288 scientific-computing subproblems drawn from research code.
  • 10%
    Humanity's Last Exam Scientific Reasoning Frontier academic questions across disciplines, run without tools.
  • 10%
    CritPt Scientific Reasoning Research-level physics problems, the hardest suite in the index.
  • 10%
    GDP.pdf General Document understanding over long real-world PDFs, scored all-or-nothing per document.
  • 5%
    AutomationBench-AA Agents Multi-step workplace automations run end to end, scored for partial completion.
  • 5%
    AA-LCR v1.1 General 100 long-context reasoning questions over roughly 100K tokens.

Artificial Analysis Intelligence Index v4.3, with AA's capability indexes at v1.1. The use-case boards use the same source: either one of AA's published capability indexes, or the plain average of the AA evals named on the board. No re-weighting, no normalizing against the field. Recomputing this table against AA's raw eval results returns AA's own published score for every model carried here, to within rounding.

The rules that are ours

Which models appear

Every model Artificial Analysis has scored on the full index and published a price for: 71 of the 215 models AA tracks. The field is not trimmed to a top slice, because a Pareto frontier is defined by the models on it - cut one cheap model and the line moves. Charts cap what they show; the data does not.

The value pick

The cheapest open model that still keeps 70% of the open leader's score on that board. Without a quality floor, "value" is always won by whatever is cheapest and weakest.

Missing results

A model with no result on a measure keeps a gap, shown as a dot. It is never counted as a zero, and the category average is taken over what the model does have.

Which speed is shown

AA's leaderboard speed is measured against the model creator's own API, which for an open model is rarely where you run it. Where Together serves the model and AA has measured that endpoint, the site shows Together's number instead. Prices already worked this way.

Evaluations AA has dropped

v4.3 removed GPQA Diamond and τ³-Bench Banking from the index and replaced Terminal-Bench 2.1 with 4.0. AA still runs them, so they stay visible in the score breakdowns, labelled and never averaged into a board score.

Scores that go negative

AA-Omniscience rewards accuracy and penalises hallucination, so a model that guesses wrong more often than it declines to answer scores below zero. That figure is shown as AA publishes it; it is not floored at zero, and a bar simply does not draw.

Where Together loses

Together runs inference and appears in these comparisons. Competitors are always selectable and are shown winning where they win. Speed is Together's own endpoint, which is slower than the field on some models, and those read slower here.

Sources

Read on Sep 16, 2026 by scripts/refresh-data.mjs. Rankings move with every release, and AA refreshes many times a day, so expect small drift between runs.