Google · none caching

Deep Research Preview (Apr-21-2026)

Deep Research Preview (Apr-21-2026) on Zumik: live pricing, context, and caching, routable by id or alias through one OpenAI-compatible endpoint.

Call it through Zumik About Google Gemini

—

Input / 1M tokens

—

Output / 1M tokens

—

Cache read

131K

Context window

Specifications

At a glance.

Provider	Google
Family	gemini
Released	—
License	Proprietary
Context window	131K tokens
Max output	66K tokens
Modalities	text, image
Tool calling	Yes
Reasoning mode	No
Caching	none
Batch discount	50% off

Measured by Zumik

What reuse looks like here.

Not yet profiled

Pricing, context, and capabilities for Deep Research Preview (Apr-21-2026) are live, but it is outside the flagship set Zumik benchmarks in depth, so measured reuse, capture, and warm TTFT are not shown yet. Run a workload estimate or route it by id to start collecting traces.

Call it

Same OpenAI client, this model.

python

from openai import OpenAI

client = OpenAI(base_url="https://api.zumik.ai/v1", api_key="zk_live_...")

r = client.responses.create(
    model="deep-research-preview-04-2026",
    input="Draft a fix for the failing test.",
)
print(r.usage.input_tokens_cached)   # confirm reuse

Frequently asked

Deep Research Preview (Apr-21-2026), answered.

How much does Deep Research Preview (Apr-21-2026) cost?

Deep Research Preview (Apr-21-2026) is an open-weights model routed through Google Gemini. It is priced on the host's serverless size tier rather than a single published per-token list price, so it shows "—" here until profiled.

What is Deep Research Preview (Apr-21-2026)'s context window?

Deep Research Preview (Apr-21-2026) supports a 131K-token context window with up to 66K output tokens.

Does Deep Research Preview (Apr-21-2026) support prompt caching?

Yes. Google Gemini uses Implicit context caching caching. In the Zumik corpus, Deep Research Preview (Apr-21-2026) shows a median cache capture of 75% on agent workloads.

Run Deep Research Preview (Apr-21-2026) with reuse measured.

Point an OpenAI client at Zumik and see exactly how much of this model's input you are reusing.

Migration quickstart Back to catalog