Two workloads. Same model, best hardware for each. 40x apart.
To measure what compute produces, and move every workload to where it produces most.
pip install berth-placement
Two services · identical weights · each already on its cheapest feasible placement
Cost per million served tokens. The difference is workload shape, not hardware. Neither team could see it, because no metric they had was denominated in delivered work.
AI is the first trillion-dollar infrastructure that cannot yet state, in a standard unit, what its output costs.
That spending buys hours. What anyone actually needs is work delivered before a deadline. No standard unit connects the two, so nobody can say what a workload costs, compare two places to run it, or prove a change made anything better.
Every efficiency gain in AI today is unmeasurable. It is why capital cannot tell a good decision from a lucky one, and why two teams running the same model, each on the cheapest hardware that meets their target, can be 40x apart on what a delivered token costs them.
We built the missing meter, and the system that acts on what it reads.
One output token delivered inside its service-level objective and above its quality floor.
price notation $/MSVT · dollars per million served tokens
A token that arrives after its deadline is not a cheap token. It is not a token.
Every existing measure drops something.
The served token is the first measure of AI compute that carries all of it at once: the work, the deadline, the quality, and the invoice.
Deciding which chip, which provider, and at which price a model should run under a stated p99, then moving it as prices, capacity, and traffic shift.
The problem: that answer changes. A model version ships. A price moves. Traffic shifts. Most teams decide once and never check.
What makes it adaptive: it never stops deciding. Prices move, models ship, capacity changes. So does the right placement.
Three pieces. The unit is measured, the measurement is checked, and the decision is acted on.
Predicts what a workload costs and how fast it runs on a given accelerator, before you rent it.
Measures the real thing and checks the prediction.
Watches for change, adapts the placement, and proves what the change saved.
$ pip install berth-placement
All open source. All reproducible from published traces.
The model is a page of physics: bytes moved, work done, time under load. No learned components and no training data, so every number is auditable by hand.
That matters because an answer is only worth having if you can check it. Every prediction ships with the term that produced it and a label saying whether that term rests on a measurement or a specification sheet.
Where the physics stops describing the machine, we will learn the difference. Knowing exactly where it stops is what makes that possible, and mapping that boundary is the research.
An untuned model, fitted on one accelerator, predicts hardware it has never seen to within ten percent. The bar was set and published before the first measurement was taken.
Every trace is downloadable. Every number is reproducible. When the instrument itself is wrong, we publish that too, with the mechanism and the test that catches it if it returns.
The instrument is free and stays free. What we sell is holding a placement correct over time, and proving it.
We hold a class against its service level, re-measure when the corpus moves, and issue a conforming receipt each quarter. A class is one model, one serving configuration, one service level.
Above a threshold, we are paid on the difference we prove under the published Holdout Protocol. Below it the holdout costs more than the arrangement returns, assurance is the right instrument, and we will say so.
The measurement record, for allocators, lenders and operators underwriting compute. Including series nobody else publishes.