Topic
Inference
3 stories on this topic, newest first.
NewsAnalysis
Inference prices keep falling. Here is who actually benefits
Per-token prices for frontier-class models have dropped by roughly an order of magnitude every 18 months. The savings are real, but they land unevenly across the stack.
ResearchAnalysis
Test-time compute changed the scaling roadmap. Here is what it costs
For a decade, progress meant bigger training runs. Now labs can trade inference compute for capability instead. That shifts the economics from capex at the lab to opex at the user, and it changes what "a better model" means.
IndustryExplainer
The real cost of running an AI product, line by line
Token spend is the line everyone watches and rarely the largest. A working breakdown of where the money goes in a production generative AI product, from inference and evaluation to the humans in the loop.


