Writing from the team
Notes on running
your own models.
Benchmarks we ran, decisions we made about pricing and hardware, and the occasional argument about where European inference should live.
Archive
-
01
Per-token pricing is the right way to buy inference until a certain volume. This is the arithmetic for finding that volume, and what to do on either side of it.