Pricing

You pay for what you use. That is the whole model.

Add credits, starting at $5. Each request takes a little off that balance at the published price of whichever model answered - usually less than the provider would charge you directly, because Zumik keeps a slice of what it saves you rather than adding a fee on top. No subscription, no seats, no plan to choose.

Pay-as-you-go
From $5 in prepaid credits

drawn down per request - whatever you do not use stays on your balance

Get started
  • Every model, on one key - no per-provider accounts
  • Automatic model choice, or name one yourself
  • Failover to another model when a provider goes down
  • Per-request cost, broken down by model and API key
  • Prompt caching set up for you, and measured
  • Spending caps and alerts at 50%, 80% and 100%
  • A free check of how much of your bill is repeated work
  • Delete your data and get a receipt for it
  • Use your own provider keys instead, if you prefer
Larger teams

When you need your own hardware.

Guaranteed response times, private networking, a specific region, confirmed deletion, and models running on infrastructure you control. We only switch this on once a replay of your real traffic shows it beats the managed providers, so you never pay for hardware you did not need.

  • Replay-backed BYOC runtime profiles (Dynamo, SGLang, LMCache, Mooncake)
  • Kubernetes profiles on llm-d, KServe, and AIBrix
  • Custom retention, region, and procurement policy
  • Volume credit pricing and committed-use discounts
Talk to us

Calculator

Estimate your monthly spend

Drag the sliders to match roughly what you send. It compares paying full price for every request against the same traffic with caching working - the whole bill is credits, there is no platform fee on top.

600,000 requests/mo · 7,200,000,000 input tokens · 420,000,000 output tokens.

$93,000
List price, no reuse / mo
$57,360
With 55% reuse / mo
Estimated monthly saving
$35,640
38%
Monthly subscription$0.00
Estimated credit spend$57,360/mo

Estimate only, at provider list prices; Zumik's published per-model rates land at or under list on managed routing. Actual reuse depends on prompt ordering and retention locality. Output is excluded from cache savings.

Note

This is an estimate. How much you actually save depends on how your prompts are arranged and how often you repeat yourself - run a free check to measure yours, or read what each provider charges for cached content.

Compare

Check the numbers yourself.

Zumik only makes money when routing genuinely costs you less. See which models are cheapest once caching is working, or how Zumik stacks up against the other gateways.

Frequently asked

Pricing, answered plainly.

How am I charged?

You add credits up front, and each request takes a little off that balance - based on how much text went in, how much came back, and the published price of the model that answered. Nothing is charged monthly, and you are never billed for a month you did not use.

What is the smallest I can spend?

$5. That is a real top-up, not a trial, and whatever you do not use stays on your balance. There is no seat fee, no monthly minimum, and no plan to pick.

Can I stop it running away with my money?

Yes, and you should. Set a hard cap and Zumik refuses requests past it rather than charging you. You also get alerts at 50%, 80% and 100% of that cap, and you can give individual API keys their own budget.

Is there a free tier?

Not for running requests - every one of those costs us real money at the provider. You can create an account, explore the console, and run a free check on your workload without a payment method.

What if I already pay OpenAI or Anthropic directly?

You can keep those accounts and use their keys through Zumik. The provider bills you for the inference as usual, and Zumik charges a small per-request fee for the routing and the reporting. Same if you run models on your own hardware.

How does caching actually change my bill?

Content the provider has already seen is billed at their cache-read rate instead of full price - often around a tenth. On the agent workloads we have measured, roughly half of a typical request qualifies. The calculator below shows the difference on your numbers.

Why is this cheaper than going direct?

Because a lot of requests do not need the expensive model you sent them to, and a lot of tokens are repeats you should not be paying full price for twice. Zumik takes a slice of what it saves you, so on managed routing the price you pay lands under provider list.

Start with $5

Create an account, add credits, change one setting in your app or coding tool. You only ever pay for what you actually run.