The Placement Index

What a workload costs, on every chip, right now

Every other leaderboard ranks models. This ranks placements: the same workload on different silicon, priced per million tokens delivered inside its latency target.

The premier chip is the wrong chip. An H100 PCIe ranks sixth of nine here on cost per delivered token, and the gap is not small. Price alone gives a fairly stable ordering. Feasibility does not: tighten the first-token budget and placements begin failing it outright, and the cheapest chip that can still do the job is a different chip. A placement that misses the deadline has no cost per compliant token, so ranking on price alone answers a question nobody is asking.

Workload controls

Placement ranking

# Placement $/Mtok Premium TTFT TPOT tok/s Provenance

Premium is the multiple over the cheapest placement that still meets the latency budget. A chip that cannot meet the budget has no cost per compliant token: not a high one, an undefined one, and it is marked infeasible rather than ranked.

Provenance is the whole point. MEASURED means the estimate was checked against traces from that hardware and landed inside a 15 percent gate published before the run. prior means spec sheet arithmetic, unverified. Measured cells have run at 4.2 to 10.6 percent error. A prior cell is unbounded, and one has been observed above 40 percent. We would rather show you which is which than average them into a single confident number.

List prices, on-demand, single tenant, everything assumed available. Real placement is also constrained by quota, region, contract terms, reserved capacity and what a provider will actually hand you this week. Those move the ranking and none of them are modelled here. Every figure is produced by the same estimator the CLI ships, and you can reproduce any row with berth estimate.

Why the ordering looks stable. Seven of the nine placements above carry spec-sheet priors, and spec sheets are smooth: compute and bandwidth are quoted as though they scale together, so a fleet of priors produces a tidy ranking. Measurement is what breaks it. An L40S fitted a bandwidth above its own microbenchmarked ceiling. The fixed prefill floor differs by 24 percent between two serving stacks on the same card. Mixture-of-experts decode moves roughly 1.6 times the bytes active parameters predict. None of that is on any datasheet, and every one of those was found by running the hardware rather than reading about it. The ranking above will get less tidy as the corpus fills, and that is the point of filling it.

The corpus behind it

3 of 72 cells measured

A cell is one accelerator running one model under one serving stack. All 72 are below, measured or not, because a table showing only what has been done tells a reader nothing about what to trust.

Validated is measured and inside the 15 percent gate published before the run. Measured means traces exist and a question is open, published regardless. Running is dispatched to an operator now. Queued is committed next. Everything else is open, and that is most of the table.

bf16, vLLM unless noted
Accelerator Llama-3-8B
dense
Llama-3-70B
dense
Qwen3-30B-A3B
MoE
Mixtral-8x7B
MoE
DeepSeek-V2-Lite
MLA
DeepSeek-V3
MoE
L40SNVIDIAVALIDATED60 traces, 10.6% TPOT
SGLang cell running
openqueuedqueuedopenopen
H100 PCIeNVIDIAVALIDATED60 traces, 4.2% TPOTopenopenopenqueuedopen
H100 SXMNVIDIAopenqueuedopenopenopenopen
H200 SXMNVIDIAopenopenopenopenopenopen
A100 80GBNVIDIArunningdense control, runningopenMEASURED60 traces published
control run dispatched
openopenopen
B200NVIDIAopenopenopenopenopenopen
GB200 NVLNVIDIAopenopenopenopenopenopen
MI300XAMDqueuedopenopenopenopenopen
MI325XAMDopenopenopenopenopenopen
TPU v5eGoogleopenopenopenopenopenopen
TPU v6eGooglequeuedopenopenopenopenopen
Trainium2AWSopenopenopenopenopenopen

Measured cells have predicted decode latency to between 4.2 and 10.6 percent using untuned priors, against a 15 percent gate published before the first run. Every trace behind them is downloadable, and every open question has a pre-registered run against it.

If your workload is a blank square, that is the ask. One run is ninety cells: roughly an hour on a rented box, a few dollars of GPU, and nothing leaves your cluster unless you choose to send the trace.

What has to hold for any of this to generalise

Four axes, and the honest status of each. The roofline is one functional form with constants recovered from a microbenchmark that costs minutes, so the expensive question is whether the form transfers, not whether the constants do.

AxisStatusEvidence
Memory technology holds An untuned prior predicted decode within 1 and 6 percent on GDDR6 and HBM2e, on cards it had never seen.
Serving stack holds Decode 0.854 under SGLang against 0.850 under vLLM, same card, same model. Serial prefill 1.01 against 1.09.
Architecture family open The first mixture-of-experts cell missed the gate. Decode appears to move about 1.6 times the bytes active parameters predict; a dense control on the same box is pre-registered.
Vendor untested Every measured cell is NVIDIA. AMD and TPU carry priors, and one of them currently tops the ranking above.
Quantisation untested fp8 and fp4 paths exist in the code and in the trace schema. No cell has been run at anything but bf16.
Locality and availability not modelled Region, quota, reserved capacity, contract terms and interconnect topology. All of these move a real placement decision and none are in the estimate.

Where the closed form stops being the right object is a question with an answer, and finding it is part of the work rather than an embarrassment: disaggregated prefill and decode, chunked prefill, speculative decoding, multi-node interconnect. At some boundary arithmetic gives way to a learned surrogate. Knowing precisely where that boundary sits is itself the contribution, because a surrogate trained without it is a black box and one trained with it tells you which regions need learning and which are simply arithmetic.

Unmeasured silicon

Rows marked prior are spec-sheet arithmetic on hardware we have not yet run. They are shown because a two-entry fleet answers nobody's question, and marked because a number you cannot check is not a measurement. Measured cells have landed between 4.2 and 10.6 percent error against a gate published before the run. Untick to see only those.

By workload class

The same fleet, four questions

Each block below is the same eight placements asked a different question. The orderings are not the same, and that difference is the reason this page exists.

Where each chip ranks

One row per accelerator, one column per workload

Rank by cost per million tokens delivered inside that workload's latency budget. A dash means the placement cannot meet the budget at all, in which case its cost per compliant token is not high, it is undefined.