Oculi·Dei
Oculi Dei LLC — Michigan

Private vocabulary for trained models

Your model,
speaking your language.

Inference is billed by the token, and most organizations have never measured what their own traffic actually costs them. We measure it, then reduce it — losslessly, on the models you already run, verified on your corpus before anything changes.

2.92×
Fewer tokens on agentic tool traffic, against GPT-4o's tokenizer — the hardest baseline we test
807,663
Documents round-tripped losslessly on the shipping lexicon — 1.36 GB, zero failures
92.04%
Vocabulary in active use across 33 domains and 1.36 GB — against o200k's 72.60% on the same text, where 54,807 slots never fire at all
3 / 3
Model families passing initialization gates on their real weight matrices, including a sparse mixture-of-experts

The part nobody prices

Fewer tokens is also a bigger context window

The same window, holding more of your content. You are not billed for the difference and you do not change models to get it.

A context window is measured in tokens, not in meaning. Encode the same material in 2.92× fewer tokens and the window you already pay for carries 2.92× more of it — the whole codebase instead of a slice, the full trace instead of the tail. Cost falls and capability rises together, which is why this lands on both sides of a buying conversation.

Shown at the measured 2.916× agentic ratio. A held-out blend of prose, code and agentic traffic in one stream measures 1.97×; tuned to one customer's own traffic it reaches 3.14×.

The line item nobody audits

Tokens are the unit of both billing and compute

Which makes the count a cost centre that compounds with every request, for the life of the deployment.

Organizations audit cloud spend, storage and egress. Almost nobody audits the token count itself — the quantity every invoice and every GPU-second is denominated in. It is treated as a fixed property of the text rather than something measurable, and therefore something reducible.

It compoundsA reduction lands on every request from the day it ships, and scales with volume rather than against it.
It is domain-specificConcentrated, structured traffic carries far more reducible overhead than diffuse text. Measured across eleven sectors, the spread runs 1.53× to 2.92×.
It is invisible without measurementThere is nothing to compare a token count against until someone measures the same text a second way. That measurement is the engagement.
It cuts cost and raises capability togetherThe same window carries proportionally more of your material, so the return lands on the bill and on what the model can see.

Where to look

The detail, in full

TOKEN AUDIT — START HERE

Measure your own traffic first

The first engagement: a lexicon tuned to your corpus, scored on held-out text, so you get your reduction, its dollar value at your rate, and your context headroom before committing to anything.

Request an audit →

RESULTS

Eleven sectors, measured and priced

Token reduction per sector on held-out traffic, the value that represents per terabyte, and the context-window expansion that comes with it.

See the results →

VERIFICATION

How every figure is checked

Lossless round-trip at scale, held-out by content hash, and how much of the gain carries across to tools the lexicon has never seen.

See the checks →

CONTACT

Operators, researchers, investors

Tell us the domain, the rough volume, and whether you are self-hosted or on an API. That is enough to say what is measurable about it.

Get in touch →