Pricing

Pay for tokens, not a plan.

No subscription, no seats, no tiered feature gates. Load prepaid credits from $5 and spend them per token at published per-model rates - on managed routing the price lands under provider list, because the savings Zumik captures are its margin.

Pay-as-you-go
From $5 in prepaid credits

billed per token at published per-model rates - unused credits stay on your balance

Sign up
  • Exact OpenAI-compatible /v1 and native /v2 APIs
  • Reuse opportunity and realized-capture metrics on every request
  • Reproducible model-alias releases and routing records
  • Workload diagnostics with Workload Reuse Score
  • Metadata-first tracing with opaque, tenant-scoped handles
  • Provider-native caching, batch, and service-tier routing
  • Purge jobs with profile-specific receipts
  • BYOK across all five first-class providers
Enterprise & BYOC

When the workload earns it.

Dedicated SLOs, private networking, regional isolation, runtime-confirmed purge evidence, and explicit KV orchestration. Activated only after a replay run proves it beats managed providers, so you never pay for infrastructure you do not need.

  • Replay-backed BYOC runtime profiles (Dynamo, SGLang, LMCache, Mooncake)
  • Kubernetes profiles on llm-d, KServe, and AIBrix
  • Custom retention, region, and procurement policy
  • Volume credit pricing and committed-use discounts
Talk to us

Calculator

Estimate your monthly spend

Drag the assumptions to match your workload. The calculator compares list-price inference against the same traffic with a cached stable prefix - the whole bill is credits, there is no platform fee on top.

600,000 requests/mo · 7,200,000,000 input tokens · 420,000,000 output tokens.

$93,000
List price, no reuse / mo
$57,360
With 55% reuse / mo
Estimated monthly saving
$35,640
38%
Monthly subscription$0.00
Estimated credit spend$57,360/mo

Estimate only, at provider list prices; Zumik's published per-model rates land at or under list on managed routing. Actual reuse depends on prompt ordering and retention locality. Output is excluded from cache savings.

Note

This is an estimate. Realized reuse depends on prompt ordering and retention locality, run a workload diagnostic to measure yours, or read how providers cache on the prompt caching pages.

Compare

Cheaper than recomputing, by design.

Zumik only earns when routing genuinely costs you less. See which models are cheapest once caching is working, or how Zumik compares to gateways and routers.

Frequently asked

Pricing, answered plainly.

How much does Zumik cost?

Zumik is pay-as-you-go on prepaid credits. Load any amount from $5; credits are drawn down per processed input and generated output token at published per-model rates. There is no subscription and no free inference tier.

Do credits expire or carry a minimum?

Unused credits stay on your balance; the smallest top-up is $5. There are no seats and no monthly fee. Hard budget caps, 50/80/100% alerts, per-key budgets, and opt-in overage keep spend predictable.

Is there a free tier?

No free inference tier. The workload diagnostic is free and needs no payment method, so you can measure reuse opportunity before loading credits.

Is BYOK or BYOC extra?

On BYOK you bring provider keys and the provider bills your inference directly; on BYOC you run the model on your own infrastructure. On both paths Zumik charges a small control-plane fee per request for routing, state, and diagnostics. The fee applies on every path.

How does reuse change my bill?

Reused prefix tokens are billed at the provider cache-read rate instead of full input price. The calculator shows the difference; in our corpus, median realized reuse is around 53% on agent workloads.

Pay-as-you-go credits from $5

Sign up, point an OpenAI client at Zumik, and only pay for the inference and diagnostics you run.