Oculi·Dei
Oculi Dei LLC — Michigan

Every figure, checked

Verification

Every number published here is measured on text the lexicon never trained on, recomputed independently from its source measurement, and held to thresholds the tooling enforces automatically. The figures are reproducible on request, and we are glad to walk through the methodology in detail.

Lossless at scale

807,663
Documents round-tripped
1.36 GB
Total text verified
0
Round-trip failures
0
Adversarial failures

Whole documents, not truncated samples. The adversarial set covers the cases that break tokenizers quietly: zero-width-joiner emoji, right-to-left scripts, combining marks, raw control bytes, mixed line endings, and private-use codepoints — the places where a collision corrupts text without raising an error anywhere downstream.

Measured on held-out traffic

Split by content hashEvery sector figure is measured on text the lexicon never trained on. Training and evaluation sets are disjoint by construction.
Matched-train, all elevenEach sector's lexicon is trained on 50 MB or more of that sector's traffic, so the sectors are directly comparable to one another.
Independently recomputedEach published figure is recalculated from its source measurement by a separate validator — ratio, tokens per terabyte, percentage and dollars — and both must agree before it ships.
Quoted against the hardest baselineFigures are published against the most token-efficient tokenizer we test. Measured across six production tokenizers from four families, the same lexicon performs better against every alternative.
Coverage figures carry their scopeVocabulary activation rises with corpus breadth, so the scope travels with the number: 92.04% of 186,002 slots active across 33 domains and 1.36 GB, measured against o200k_base on the identical text at 72.60%. The figure that carries the cost is the remainder — 14,807 slots never fire, against 54,807 for o200k. A dead slot still occupies a row of the embedding matrix and a column of the output projection on every forward pass.
Document length is controlled forSector figures are scored on a fixed document unit, so a result cannot be produced by sampling convenient fragments. Rescoring the same holdouts as whole, uncut documents moves the published ratios by at most 0.15%. Where a corpus runs to long files the ratio is length-dependent, which is one more reason the figure that governs an engagement is the one measured on your own traffic.
Thresholds enforced by toolingMinimum corpus size and maximum plausible ratio are applied mechanically when the record is generated, not left to judgment.

The gain survives tools the lexicon has never seen

The question worth asking of any tokenizer result, answered with a measurement rather than a reassurance.

Holding out by content hash still draws both halves from one pool of tools, so we went further: the agentic corpora were re-partitioned by tool. A portion of the tool vocabulary was reserved out of training entirely, and the lexicon was then scored only on records using tools it had never encountered.

55–81% of the gain is structural and carries across to an entirely new tool set, with the remainder specific to tools already seen. The range widens with training scale — a larger corpus captures more tool-specific structure — so we quote the floor rather than the ceiling, measured on 7,021 evaluation records using 446 tool schemas reserved entirely out of training.

What we quote, and what we don't

THE CONSERVATIVE CASE

We lead with the floor, not the ceiling

Where a lexicon meets traffic in an unfamiliar format the ratio is lower, and that lower figure is the one we put in front of you. A tuned result is never quoted for untuned traffic.

YOUR OWN CORPUS

The number you get is measured on your traffic

Gains vary by domain, so every engagement opens with a measurement of the customer's own corpus. Nothing is committed before that figure exists.

PROVENANCE

Each figure names the corpus behind it

Training size and source travel with every sector result, so the basis for a number is always visible alongside it rather than a page away.

SCOPE

Compression and privacy, stated as such

A tuned vocabulary changes what a request costs and who can read it. We describe it as exactly that, and size every claim to what has been measured.