Enterprise AI spend control

Your company pays top prices for routine work.

Most companies send every AI request to the most expensive model available, because deciding case by case is hard. Krites makes that decision automatically — and keeps the expensive model for the work that actually needs it.

30-minute walkthrough on your own prompts. No data or API key needed to evaluate.

What it is

A single decision point between your people and your model bill.

Every request passes through Krites on its way to a provider. It is read, checked against your policy, sent to a model, and recorded — so the same request is handled the same way every time, and you can see why.

01

A request comes in

A summary, a draft, a code review, a contract check — from the chat workspace or from your own application.

02 · Krites

Krites assesses it

How difficult the task is, how much is riding on it, how sensitive the data is, and which models this team is permitted to use.

03

The right model answers

Routine work is handled cheaply. High-stakes work stays on the strongest model, determined by policy rather than by budget.

04

The cost is recorded

Each request is attributed to a team, a person and a use case as it passes through, so reporting requires no reconciliation later.

One real decision

What happens to a single prompt

Request
Krites
DeepSeek V4 Flash
GLM 5.3
Claude Opus 5
THE REQUEST

“Summarize these 12 support tickets into three bullets for the standup.”

WHAT KRITES SEES
Kind of work
Summarizing
Difficulty
Low · 0.22
Risk if wrong
Low · 0.10
Sensitivity
Internal
WHERE IT GOES

DeepSeek V4 Flash

Allowed by policy, capable of the task. Falls back to GLM 5.3 if it fails.

94× cheaper than premium

RULE

Legal, security, finance and executive work — anything scored high-risk — is never quietly moved to a cheaper model to save money. A team that is over budget gets slowed down or flagged, not given a worse answer.

For the engineers in the room

Open any of these for the mechanics.

Under the decision

Choosing a cheaper model is only the first saving.

Four more sit underneath it. Every request goes through all five, in this order, and they compound: a cheaper model thinking less about a shorter prompt whose repeated half was served from cache. Each layer is measured on its own line, so a saving that turns out not to be real cannot hide inside one that is.

  1. 01

    Choose a cheaper model

    The routing decision itself: the cheapest model on the ladder that the task's difficulty and risk allow, picked from the set your policy permits rather than the one the budget would prefer.

  2. 02

    Choose a cheaper reasoning level

    The same model can think hard or barely at all, and thinking is billed like any other output. The effort is proposed from the classifier signals routing already has, capped by your policy, then clamped to what that endpoint actually accepts — three layers, each of which can only lower it.

  3. 03

    Send fewer tokens

    An agent resends its whole history every turn. A re-read of a file already in that history becomes a diff, an identical repeat becomes a one-line note, and old turns fold into a deterministic summary — without rewriting bytes the provider has already cached.

  4. 04

    Reuse cached tokens

    The stable part of a request — tool definitions, the system prompt, the conversation so far — is marked so the provider serves it from cache at a fraction of the input rate, with breakpoints placed to extend the cached prefix rather than rebuild it.

  5. 05

    Retry intelligently, and count the real cost

    Failures retry on the decision's fallback and escalate a tier when an answer fails its quality check, all inside one usage record, so a retry can never bill twice. Cache savings, the premium paid to write those cache entries, and tokens compaction never sent are three different kinds of claim and are reported as three numbers.

Each layer has its own switch and its own measurement, so one can be turned off without disturbing the others — and none of them is allowed to report a saving the provider's own token counts do not support.

The bill

Routing the same work costs 83% less than sending it all to a premium model.

Here is the same request priced four ways — 2,000 tokens in, 600 out, at today's list prices.

Claude Opus 5$0.0250
GLM 5.3$0.0054
DeepSeek V4 Flash$0.0003
Routed by Krites$0.0043

A 60/30/10 mix across the three tiers — 83% below all-premium, with the hard work still on the best model.

Illustrative, at current registry list prices — not a customer result. Your real mix comes out of the eval harness on your own traffic before you commit to anything.

Spend by team, person, use case

Attributed as it happens, so there is nothing to reconcile at month end.

A forecast you can question

Run rates projected to month end, broken into what is driving them such as more volume, longer answers, more premium use.

Waste, named

Premium models used on easy work, teams over budget, repeated prompts — each with a monthly cost and a recommendation.

Security & data

Krites works without keeping what your people wrote.

Routing and reporting run on details about each request — the kind of task, the model, the tokens, the cost. Storing the text itself is a separate setting, and it is off unless you turn it on.

Fingerprints, not prompts

Each prompt is stored as a one-way SHA-256 fingerprint. It identifies repeats and reuses cached answers, and it cannot be turned back into the original text.

No training on your data, ever

Every request goes upstream with a no-training flag that cannot be switched off. You can additionally require zero-retention endpoints. Policy tightens, never loosens.

Fails closed

If no model meets your privacy rules, the request is refused with a clear error. It is never retried somewhere weaker to get an answer.

Permissions enforced on our servers

Who can see which workspace is decided on the server for every request, not in the browser — so access cannot be gained by tampering with the page.

Book a demo

See it route your own prompts

Bring three or four prompts your team actually sends. We classify them live, show which model each one lands on, and put your current spend through the maths. The invite arrives immediately with a Meet link.

  • Screen-shared, no slides
  • Live routing on your prompts, nothing stored
  • Policy modelling for your teams
  • A concrete before-and-after on your bill

Calendar not loading? Open it in a new tab.

Krites — the enterprise AI spend control plane