Research
One paper, a set of open problems, and the rules we work under. All of it published, including the parts that did not go our way.
Paper · 2026
Why AI inference has no unit of account, what the missing one costs, and how the served token closes it.
Electricity had lamp-months before it had the kilowatt-hour. Freight had negotiated rates before it had the ton-mile. In both cases the unit arrived after the industry was already enormous, and everything that made a market followed it rather than preceding it. The paper argues that AI inference is at the same point, specifies the served token, and replays production traces from two independent operators to show what the absence is costing.
Stating what we cannot yet model is rarer than publishing what we can, and it is the more useful half. These are open, and we would rather they were solved by someone else than not at all.
A page of physics predicts most of inference well. It does not describe everything, and the interesting question is the boundary: disaggregated serving, speculative decoding, and the regimes where a scheduler rather than the silicon sets the answer. Mapping that boundary is what makes a learned residual tractable rather than a replacement for understanding.
A constant fitted on one accelerator transfers to another it has never seen, within ten percent, across different memory technology. Whether that survives a different compiler and a different memory system is the measurement that turns a useful result into a general one.
Workloads with different deadlines share capacity. Solving each in isolation leaves the headroom that a tight-deadline class reserves unused by the loose one. Whether solving jointly beats solving independently by more than the uncertainty is an open question with a pre-registered kill criterion.
Once output is priced in a common unit, allocation becomes a market design problem rather than a scheduling one. What contracts, what settlement, and what happens to depreciation schedules when an external unit prices what the asset produces.
A measurement organisation is only worth what its method is worth. These are the rules, and they are the reason to believe anything above.
Essays and technical notes, on the unit, on placement, and on what the corpus turns up.
If you work on performance modelling, systems learning, or market design, the problems above are genuine, the verifier is deterministic, and the environment is live.