Skip to content
← notes

note · 9 October 2026 · 5 min readdraft

Unmeasured is not zero

What I learned building a per-user ledger for every token, page and image an AI service pays for, and why “we don't know” has to be a value.


For a while, the honest answer to “what did this turn cost?” was “some of it”. Chat calls were logged against a user named after the service itself. Embeddings, OCR pages, vision pages and generated images weren't metered at all. One model was costed at a fifth of its invoice.

None of that raised an error. That's the trouble with spend: when it goes missing, nothing breaks. Nothing 500s, no test goes red, nobody complains. The money just stops being attributable.

One event for everything we pay for

The ledger has one event shape for everything that's billed: a chat call, an embedding, a vision page, an OCR page, a generated image. The question it answers isn't “how many tokens” but “what did this turn cost”. So every event says who it was for: the user, the turn, the stage, the document.

That identity is set once, at the request boundary, in a context variable, so a meter can't forget who it bills. Python doesn't carry context variables into a thread pool, so the one place that fans work out to threads re-binds it explicitly. And a test fails on any raw submit() that would leak spend to nobody.

A token is never estimated

Numbers come only from the provider's own usage payload. Providers report usage in at least four shapes, so one reader normalises them all, including cached-input and reasoning tokens, which are the easiest to drop. A call that failed after generating was still billed, so usage is read off the exception too.

Three states, not one

fig. · a zero is a claim. These are three different facts.

A zero means “this cost nothing”. A payload nobody could read isn't zero, it's unknown. So it's recorded as usage_missing. A model with no rate isn't free, it's unpriced: pricing_missing. At startup the service names every model it can spend on that has no price, so the invoice is never the first to tell us.

One rate table

There used to be two pieces of arithmetic for cost: the ledger's, and an older diagnostics path with its own private rate table. They disagreed for months. The dashboard and the invoice told different stories. I folded them into one function and checked it against 308 cases captured before the change. All 308 came out identical, which is the only way I'd ship a change to money.

What it caught

  • a model under-costed five times against its invoice;
  • OCR under-reported 6.7×: an environment variable changed the price, but not the model actually called;
  • generated images recorded as free, because one provider bills per megapixel;
  • cancelled image generations that were billed and invisible, because cancellation is a BaseException and slipped past the meter.

Every one of them had been silent.

If you're building one

Make “we don't know” a value. Put identity in the request, not in the call. Keep rates as data with dates on them, so a promotional price can't bill past its end. And when two numbers must agree, delete one of them.


written from my own design docs, commit messages and tests; the numbers are theirs.

next note →Spreadsheets are not documents