Swap the base URL
Keep your SDK, your model names, your streaming. Only the endpoint changes.
client = OpenAI( base_url="https://gw.tok-net.de/v1", api_key=TOKNET_KEY, )
All your favorite AIs for about 71% cheaper. Tok Net AI is your drop-in API. It does much more than just reduce costs. See all features down below.
There are already teams shipping with us today, and these are the numbers we achieved:
Median token reduction across production accounts
More usable context inside the same window
Added latency at p50, including routing
Gateway uptime over the last 12 months
No rewrites, no prompt migration, no proxy you have to babysit. Point your existing client at the gateway and keep shipping.
Keep your SDK, your model names, your streaming. Only the endpoint changes.
client = OpenAI( base_url="https://gw.tok-net.de/v1", api_key=TOKNET_KEY, )
We do our magic.
Per-feature cost, every compression diff, every routing decision.
week 1 $18,402 → $5,338 evals 97.4% → 97.9% alerts 0 floor breaches
Every request passes through our systems before it reaches a model. Quality floors and budgets sit on top, so cheaper never means worse.
A minimum eval score per route. Nothing is ever traded below it.
What the provider would have charged you directly is a hard ceiling. We only bill a share of what we save you, so every request comes out below the original price. Never above it!
Savings are the reason teams switch. This is the reason they stay: one gateway that already does the boring, load-bearing work. Open a category to see inside it.
Two ways to use Tok Net AI. Pick one to see how it is priced.
We bill 15% of the spend we take off your bill. Nothing else. You get the whole platform for free.
Self-hosted, private VPC and custom compressors: talk to us about Enterprise.
A chat app like ChatGPT or Claude, except you are not tied to one lab.
Pricing lands with the beta; the savings-share idea stays.
Then you pay us nothing. We bill 15% of the saving we measure, so no saving means no invoice, and the fee can never exceed the saving itself. That is the promise: your bill with us is never higher than the same traffic sent straight to your provider.
Per request. We price the call as it would have run unchanged at your provider's list rate, subtract what it actually cost after compression, caching and routing, and bill 15% of the difference. Every line sits in the dashboard with the prompt diff and the model that answered it, and the whole ledger exports to your warehouse, so you can check our arithmetic against your provider invoice.
No. We have a quality floor of 97.2%. There will never be a request where we deliver a response that is below this threshold.
One line: the base URL. Your SDK, model names, streaming and tool calls all stay as they are, and you keep using your own provider keys. Compression and caching need no configuration, so the first reduction usually shows up within the hour; routing gains arrive after the first eval run, typically two or three days in.
No. We do not train models on your traffic.
Anthropic, OpenAI, Google, Mistral, and any OpenAI-compatible endpoint, including vLLM and Ollama behind your own network. Routing pools are yours to define — you choose which models are eligible, and we never route outside that set.
And not just that. You will get all the features from our platform for free.