The first engagement
We build a vocabulary on your corpus and measure what it does to your token counts — on held-out text, against the tokenizer you run today. You receive a reproducible figure for your own traffic, the cost it removes at your own rate, and the context headroom it returns.
Four figures, each specific to your corpus and each reconstructible from the measurement that produced it.
A lexicon trained on your corpus and scored on a held-out partition it never saw, against your current tokenizer. A deterministic count over fixed text — the same inputs return the same number on re-run, which is what separates a measurement from a projection.
Your ratio applied to your throughput at your own price per million tokens. Across published sectors that lands between $226k and $466k per terabyte at $3.00/M; your figure is computed from your volume, with the arithmetic shown so you can substitute your own rate.
A window is counted in tokens, not in meaning. At the measured agentic ratio a 128K window carries 373K tokens-worth of your material, and 1M becomes 2.92M. This is the half of the return that is never billed and rarely priced.
A ranked decomposition across the traffic classes in your corpus, so deployment begins where the return per unit of work is highest. Agentic and structured traffic consistently lead; document-style prose follows.
Established empirically rather than assumed, by training at increasing corpus sizes against one fixed holdout.
Sample size is the question every engagement opens with, so we measured it. Lexicons were trained on disjoint pools from 1 MB to 1.6 GB and scored against a single held-out partition, making corpus size the only variable. The answer splits cleanly: two traffic classes reach a ceiling and a modest export is enough, while source code rewards everything you can send.
| Traffic class | 25 MB | 50 MB | 100 MB | Ceiling reached at | We ask for |
|---|---|---|---|---|---|
| Technical & research prose | 97.8% | 99.4% | 100% | 100 MB | 50 MB |
| Agentic & tool traffic | 94.7% | 97.4% | 99.1% | 200 MB | 100 MB |
Both curves above were carried past the point where they stopped moving, which is what makes the percentages meaningful: they are shares of a measured saturation point, not of the largest sample we happened to run. Prose settles early — beyond 50 MB an eightfold increase in corpus moved the ratio by 0.04%. Agentic traffic carries more structural variety and settles at 200 MB. For these two, breadth across your traffic classes is what moves the number — a representative slice, not volume for its own sake.
Code never stopped improving. We carried it to 1.6 GB — eight times the largest sample the table above needed — and every doubling of corpus was still adding compression. So we quote no ceiling for code and we do not ask for a token sample: more code keeps paying, and here is the schedule it pays on.
| Code you send | Fewer tokens than o200k | Gain over previous step |
|---|---|---|
| 100 MB | 1.64× | — |
| 400 MB | 1.71× | +4.2% |
| 800 MB | 1.77× | +3.5% |
| 1.5 GB | 1.82× | +2.6% |
Read the right-hand column as the reason to send more: at 1.5 GB an extra doubling of your corpus was still worth 2.6% on every token you will ever spend. We ask code-heavy customers for 1.5 GB where the corpus supports it, and 400 MB as a working floor. A repository export reaches those numbers without a data program — source code is the one traffic class most organisations already have in volume.
The results page quotes 1.68× for developer tooling from a standard engagement. The figures above are what a code corpus returns when it is built for scale instead — which is why we ask for the larger sample.
Your ratio is computed against your own tokenizer, not a published figure borrowed from another vendor's stack.
We measure against six production tokenizers spanning four model families. Across all of them the spread is only 7%, which means tokenizer choice moves the result far less than domain composition does — and our published figures are quoted against o200k_base, the most token-efficient of the six. Whatever you run, the comparison you receive is computed on your stack.
Scope, timeline and fee are set against the size and composition of your corpus. Tell us what you run and we will come back with both.
A lexicon is built and measured on your own corpus before any figure exists — engineering against real data, held to the same standard as every number we publish. What you acquire is a quantified account of what your traffic costs you today, on your infrastructure, at your rate.
The deliverable is yours outright. It stands as the baseline every subsequent decision is measured against, and it is reproducible — the same corpus and the same method return the same number, by anyone, at any time.
Next step
Domain, approximate volume, and whether you are self-hosted or on an API. That is enough for us to scope the audit.