Oculi·Dei
Oculi Dei LLC — Michigan

The first engagement

The Token Audit

We build a vocabulary on your corpus and measure what it does to your token counts — on held-out text, against the tokenizer you run today. You receive a reproducible figure for your own traffic, the cost it removes at your own rate, and the context headroom it returns.

50 MB–1.5 GB
Representative sample, sized by traffic class — prose reaches its ceiling at 50 MB, source code keeps paying all the way up
1.82×
Fewer tokens than o200k on source code at 1.5 GB — every doubling of corpus we have measured has added compression
6
Production tokenizers benchmarked against, spanning four model families
Yours outright
A quantified baseline of what your traffic costs you today

What the audit establishes

Four figures, each specific to your corpus and each reconstructible from the measurement that produced it.

01

Your compression ratio

A lexicon trained on your corpus and scored on a held-out partition it never saw, against your current tokenizer. A deterministic count over fixed text — the same inputs return the same number on re-run, which is what separates a measurement from a projection.

02

The cost it removes

Your ratio applied to your throughput at your own price per million tokens. Across published sectors that lands between $226k and $466k per terabyte at $3.00/M; your figure is computed from your volume, with the arithmetic shown so you can substitute your own rate.

03

Your context headroom

A window is counted in tokens, not in meaning. At the measured agentic ratio a 128K window carries 373K tokens-worth of your material, and 1M becomes 2.92M. This is the half of the return that is never billed and rarely priced.

04

Where the gain concentrates

A ranked decomposition across the traffic classes in your corpus, so deployment begins where the return per unit of work is highest. Agentic and structured traffic consistently lead; document-style prose follows.

What we need from you

Established empirically rather than assumed, by training at increasing corpus sizes against one fixed holdout.

Sample size is the question every engagement opens with, so we measured it. Lexicons were trained on disjoint pools from 1 MB to 1.6 GB and scored against a single held-out partition, making corpus size the only variable. The answer splits cleanly: two traffic classes reach a ceiling and a modest export is enough, while source code rewards everything you can send.

Share of achievable compression captured at each sample size, measured against a ceiling each of these two classes actually reaches.
Traffic class25 MB50 MB100 MBCeiling reached atWe ask for
Technical & research prose97.8%99.4%100%100 MB50 MB
Agentic & tool traffic94.7%97.4%99.1%200 MB100 MB

Both curves above were carried past the point where they stopped moving, which is what makes the percentages meaningful: they are shares of a measured saturation point, not of the largest sample we happened to run. Prose settles early — beyond 50 MB an eightfold increase in corpus moved the ratio by 0.04%. Agentic traffic carries more structural variety and settles at 200 MB. For these two, breadth across your traffic classes is what moves the number — a representative slice, not volume for its own sake.

Source code is different, and it is worth knowing why

Code never stopped improving. We carried it to 1.6 GB — eight times the largest sample the table above needed — and every doubling of corpus was still adding compression. So we quote no ceiling for code and we do not ask for a token sample: more code keeps paying, and here is the schedule it pays on.

Source code, measured against o200k on held-out traffic your build never saw.
Code you sendFewer tokens than o200kGain over previous step
100 MB1.64×
400 MB1.71×+4.2%
800 MB1.77×+3.5%
1.5 GB1.82×+2.6%

Read the right-hand column as the reason to send more: at 1.5 GB an extra doubling of your corpus was still worth 2.6% on every token you will ever spend. We ask code-heavy customers for 1.5 GB where the corpus supports it, and 400 MB as a working floor. A repository export reaches those numbers without a data program — source code is the one traffic class most organisations already have in volume.

The results page quotes 1.68× for developer tooling from a standard engagement. The figures above are what a code corpus returns when it is built for scale instead — which is why we ask for the larger sample.

Benchmarked against what you actually run

Your ratio is computed against your own tokenizer, not a published figure borrowed from another vendor's stack.

We measure against six production tokenizers spanning four model families. Across all of them the spread is only 7%, which means tokenizer choice moves the result far less than domain composition does — and our published figures are quoted against o200k_base, the most token-efficient of the six. Whatever you run, the comparison you receive is computed on your stack.

How the engagement runs

1 — SampleYou provide a representative slice of the text your models process, sized per the table above. Where data handling is constrained, the transfer is scoped to your requirements before anything moves.
2 — PartitionA portion is reserved out of training entirely. Every figure you receive is measured on text the lexicon has never encountered, split by content hash so the two sets are disjoint by construction.
3 — Build and measureA lexicon is trained on your corpus and scored against your current tokenizer, with the round trip verified byte-exact on your own documents — including control bytes, mixed encodings and unusual scripts.
4 — DeliverYour ratio, its value at your rate, your context headroom, and the ranked decomposition of where the return concentrates — with the method behind each figure.

Scope, timeline and fee are set against the size and composition of your corpus. Tell us what you run and we will come back with both.

What you are paying for

A lexicon is built and measured on your own corpus before any figure exists — engineering against real data, held to the same standard as every number we publish. What you acquire is a quantified account of what your traffic costs you today, on your infrastructure, at your rate.

The deliverable is yours outright. It stands as the baseline every subsequent decision is measured against, and it is reproducible — the same corpus and the same method return the same number, by anyone, at any time.

Next step

Tell us what you run

Domain, approximate volume, and whether you are self-hosted or on an API. That is enough for us to scope the audit.

Request an audit →