OpenAI · none caching

GPT-3.5-turbo-instruct-0914

GPT-3.5-turbo-instruct-0914 on Zumik: live pricing, context, and caching, routable by id or alias through one OpenAI-compatible endpoint.

Input / 1M tokens
Output / 1M tokens
Cache read
4K
Context window

Specifications

At a glance.

ProviderOpenAI
Familygpt-3.5
Released
LicenseProprietary
Context window4K tokens
Max output
Modalitiestext
Tool callingYes
Reasoning modeNo
Cachingnone
Batch discount50% off

Measured by Zumik

What reuse looks like here.

Not yet profiled

Pricing, context, and capabilities for GPT-3.5-turbo-instruct-0914 are live, but it is outside the flagship set Zumik benchmarks in depth, so measured reuse, capture, and warm TTFT are not shown yet. Run a workload estimate or route it by id to start collecting traces.

Call it

Same OpenAI client, this model.

python
from openai import OpenAI

client = OpenAI(base_url="https://api.zumik.ai/v1", api_key="zk_live_...")

r = client.responses.create(
    model="gpt-3-5-turbo-instruct-0914",
    input="Draft a fix for the failing test.",
)
print(r.usage.input_tokens_cached)   # confirm reuse

Frequently asked

GPT-3.5-turbo-instruct-0914, answered.

How much does GPT-3.5-turbo-instruct-0914 cost?

GPT-3.5-turbo-instruct-0914 is an open-weights model routed through OpenAI. It is priced on the host's serverless size tier rather than a single published per-token list price, so it shows "—" here until profiled.

What is GPT-3.5-turbo-instruct-0914's context window?

GPT-3.5-turbo-instruct-0914 supports a 4K-token context window.

Does GPT-3.5-turbo-instruct-0914 support prompt caching?

Yes. OpenAI uses Automatic prefix caching caching. In the Zumik corpus, GPT-3.5-turbo-instruct-0914 shows a median cache capture of 79% on agent workloads.

Run GPT-3.5-turbo-instruct-0914 with reuse measured.

Point an OpenAI client at Zumik and see exactly how much of this model's input you are reusing.