Skip to content

Pricing

PRICING

AI pricing is a black box. This one isn't.

Most AI pricing hides the two things a buyer needs most: what a month will actually cost, and who gets served when demand spikes. You can’t forecast the bill, you can’t see what drove it, and when the service is busy a throwaway request and a mission-critical one compete for the same capacity — with no way to put yours first.

Sentinel answers both. You pay only for what you use, on a meter you can read line by line — or, when you run the engine yourself, one flat annual licence and no meter at all. And capacity is never a free-for-all: your place in line is something you buy, and it holds defence › paid › free, every time.

Reseller SaaS

Pay-as-you-go

  • Input — prompt ≤ 2,048 tokens — $0.05 / 1M
  • Input — prompt > 2,048 tokens — $0.40 / 1M
  • Output$2.00 / 1M
  • Or reserve a lane: Chat Lane $99 / month · Context Lane $699 / month
Learn more

Qila

Flat annual licenceper deployment

  • Dedicated slice, one tenant — $4,500 / month
  • Qila licence: annual, per deployment — priced per engagement
  • Not usage-metered — charged for the slice, never for what runs on it
Learn more

Why your bill is predictable

On our cloud you pay by the lane — one live request the sovereign node is handling. That’s the honest unit of cost: the engine only runs a fixed number of requests at once and queues the rest, so pricing the lane means your bill tracks real capacity you can point at, not a number you have to trust.

There are two lane classes, because a long prompt is not a short one. The node keeps a fixed pool of working memory, handed out in small blocks. A short chat request needs one block; a long agent prompt of around 30,000 tokens needs about fifteen — so one Context Lane uses about eight Chat Lanes. That’s the whole reason there are two plans instead of one, and it’s measured on the node, not guessed.

Lane classWhat it carriesWeight
Chat LaneShort conversational and API requests — assistants, support, classification, extraction1
Context LaneLong-context and agent work — large documents, codebases, tool-using agents8

The names describe the shape of the request, not a screen you log into. Every request is billed by what it actually uses, so a Chat Lane plan sending 30,000-token prompts is charged as the heavier work it really is — fair in both directions, and no surprises at the end of the month.

The two classes are kept apart because they crowd each other out. Left in one pool, a few agent sessions grab all the working memory and every short request behind them waits. Splitting them is what stops one customer’s heavy workload quietly slowing everyone else’s.

Why your priority is protected

Capacity is finite, so we don’t pretend otherwise. The node’s measured capacity is 128 chat-equivalents, and we sell at most 96 of it. The rest is held back for bursts, the free tier, and defence priority — so the capacity we publish is capacity that really exists. We turn away new orders past that line rather than oversell and leave you queuing.

When demand does spike, admission runs defence › paid › free across everything, and every key carries a hard rate cap. No tier gets an uncapped key — which is what keeps any one account from crowding out the rest.

Your price buys your place in line: defence, then paid, then free When the shared slice is at capacity, higher-priced traffic is served first. Defence is admitted first and never bumped; paid is served ahead of free and gains priority as spend grows; free is daily-capped and yields to both. Every key of every tier carries a hard per-key rate cap — there are no uncapped keys. WHEN THE SLICE IS FULL, YOUR PRICE BUYS YOUR PLACE SERVED FIRST 1 · Defence — served first, always never bumped · usage never counted · buys a place nothing can take 2 · Paid — ahead of every free request the more you spend, the higher you sit · pre-empts free at peak 3 · Free — try it at no cost daily-capped · pre-emptible · yields to paid and defence Every key in every tier is rate-capped, so no one account can crowd out the rest.
What each tier buys you when the shared slice is busy: your price sets your place in line — and no tier, including defence, gets an uncapped key.

Billing follows where you run it — pay-as-you-go on our in-country cloud, one flat fee when you host it yourself — but it never changes what you can see. Your own usage and Sentinel Watch look identical under a flat licence and on the meter.

FOR TELCOS & BUILDERS

You buy lanes, or you pay per token, and either way you can resell it under your own brand. See the telco door.

Lane subscriptions — buy a guaranteed place

A lane is reserved capacity: it’s yours for the month whether or not the node is busy, and it’s the only thing that guarantees you a place when it is. Buy it when you can’t afford to wait in line.

SKUWeightAvailablePrice / month
Chat Lane1pool of 128, sold to 96$99
Context Lane8up to 12$699

A Context Lane costs about eight Chat Lanes because that’s what it occupies. There’s no extra charge for long context beyond the capacity it uses.

Each lane comes with a generous monthly token allowance. Normal use stays inside it; anything beyond meters at the rates below, at the same prices everyone else pays. So a lane is a floor under your capacity, not a ceiling on your bill — you’re never cut off for outgrowing it, and never charged twice for the same request.

Developer API — /app

Pay-as-you-go, on prepaid credits, drawing on the shared pool rather than a reserved lane. You can buy any workload this way, including long-context work. Every request runs the full Mesh pipeline — guardrails, credit check, sovereign completion, pricing, metering, settlement — recorded in your currency, PKR by default, so nothing lands on your bill unseen.

Price / 1M
Input — prompt ≤ 2,048 tokens$0.05
Input — prompt > 2,048 tokens$0.40
Output$2.00

One boundary, one multiplier. The input price steps up at exactly the block size that separates the two lane classes, and it steps by exactly the factor that prices a Context Lane: . A long prompt costs eight times as much because it occupies eight times the capacity — the same measured fact, whether you buy that capacity as a lane or by the token. Nothing about the model changes at 2,048 tokens; what changes is how much of the node your request is holding, and that’s all the price is tracking.

Output carries the highest rate for the same reason, not a different one. Reading a prompt is a single fast pass; writing a reply holds the lane open word by word for as long as it runs. So the price sheet is really one idea said three times: you pay for how much you occupy. Shorter prompts and tighter answers genuinely cost less, and a long-winded answer is never something you pay for without seeing it.

Both ways of buying come from the same measurement, so their prices stay in the same range: a lane costs roughly what a busy month of that lane’s traffic would cost on the meter, and its allowance is sized to what the lane can realistically serve. What you pay extra for is the reservation — metered traffic runs in the shared pool and gives way at peak; a lane does not.

Rate limits grow with your total spend and account age (usage tiers, at /app/limits). Paid keys go ahead of free.

Tier caps and lane grants are set by the tier, not by hand — move an account between tiers and its rate limit, daily cap and lane allocation move with it, so the numbers we publish are the numbers in force.

Playground — /playground

Free to try, mobile-first, running in the reserve; higher tiers are a subscription paid through local rails.

PlanPrice (indicative)Includes
Free$0Daily-capped, in-country, pre-emptible
PlusmonthlyHigher daily cap · priority over free · sovereignty badge
PromonthlyHighest cap · priority · power features

Reseller — wholesale

Wholesale lanes and wholesale tokens to you; you set your own retail markup and bill your customers under your own brand, with a per-tenant white-label theme across the Playground and Platform your subscribers see. Operator console at /app/admin.

Wholesale, retail, and the margin between them are recorded on every request and totalled per tenant, per period — now broken out by lane class, so you can see which of your customers is buying cheap conversational traffic and which is buying capacity. See Reseller.

Payment rails: JazzCash / Easypaisa / bank transfer for top-ups, with Stripe as a config-gated card option. Amounts display in PKR; the API keeps a USD reference.

RUN IT ON YOUR OWN INFRASTRUCTURE

Mesh on-premise — flat licence

Not everyone wants their traffic on our cloud. The same full Mesh gateway — the whole routing engine, not a cut-down copy — can be self-hosted on your own servers for a single organisation’s internal use. You point it at your own chosen providers, and it routes your traffic the way our cloud routes ours.

There’s no meter to run on hardware we don’t operate, so on-premise isn’t priced by the token or the lane. You pay one flat annual licence per deployment, offline, and none of your traffic is billed online — a fixed, forecastable line item instead of a variable bill.

Mesh on-premiseFlat annual licence, per deployment — priced per engagement

Your own usage and Sentinel Watch look exactly as they do on the cloud — a licence changes how you’re billed, not what you can see. What it doesn’t change is where your traffic goes: when you self-host, whether anything leaves your network is your setting to make, with your own provider keys. That’s why on-premise is not sovereign and not air-gapped. If you need it to be impossible for anything to leave — no outbound traffic allowed — rather than a setting you own, that’s a different product — see the defence door.

FOR DEFENCE

Sentinel Sovereign — dedicated slice

A slice of your own, a flat licence, and no meter. See the defence door.

Defence doesn’t buy lanes out of a shared pool. It buys the whole slice: a dedicated slice of in-country hardware running one tenant’s traffic, with no cross-access between slices. Nothing else runs alongside it, so there’s no neighbour whose long-context job can push your work into a queue.

Dedicated sliceWhole slice, one tenant — $4,500 / month
Qila licenceAnnual, per deployment — priced per engagement

Not usage-metered. That’s a design decision, not a billing convenience: counting classified traffic is itself a leak, so the enclave doesn’t count it. Volume, timing and frequency are exactly the metadata an adversary wants, and a record of usage is a record of all three. You’re charged for the slice, never for what runs on it.

What the licence covers: air-gapped deployment with no outbound traffic allowed, an internal certificate authority, accreditation-grade audit through Sentinel Watch, and custom deployment terms.

See Sentinel Sovereign.