OpenAI · none caching
GPT-3.5-turbo-16k
GPT-3.5-turbo-16k on Zumik: live pricing, context, and caching, routable by id or alias through one OpenAI-compatible endpoint.
Specifications
At a glance.
| Provider | OpenAI |
| Family | gpt-3.5 |
| Released | — |
| License | Proprietary |
| Context window | 16K tokens |
| Max output | — |
| Modalities | text |
| Tool calling | Yes |
| Reasoning mode | No |
| Caching | none |
| Batch discount | 50% off |
Measured by Zumik
What reuse looks like here.
Pricing, context, and capabilities for GPT-3.5-turbo-16k are live, but it is outside the flagship set Zumik benchmarks in depth, so measured reuse, capture, and warm TTFT are not shown yet. Run a workload estimate or route it by id to start collecting traces.
Call it
Same OpenAI client, this model.
from openai import OpenAI
client = OpenAI(base_url="https://api.zumik.ai/v1", api_key="zk_live_...")
r = client.responses.create(
model="gpt-3-5-turbo-16k",
input="Draft a fix for the failing test.",
)
print(r.usage.input_tokens_cached) # confirm reuseFrequently asked
GPT-3.5-turbo-16k, answered.
How much does GPT-3.5-turbo-16k cost?
GPT-3.5-turbo-16k is an open-weights model routed through OpenAI. It is priced on the host's serverless size tier rather than a single published per-token list price, so it shows "—" here until profiled.
What is GPT-3.5-turbo-16k's context window?
GPT-3.5-turbo-16k supports a 16K-token context window.
Does GPT-3.5-turbo-16k support prompt caching?
Yes. OpenAI uses Automatic prefix caching caching. In the Zumik corpus, GPT-3.5-turbo-16k shows a median cache capture of 79% on agent workloads.
Run GPT-3.5-turbo-16k with reuse measured.
Point an OpenAI client at Zumik and see exactly how much of this model's input you are reusing.
