Leaderboard

Highest measured cache capture

Median capture rate from the Zumik corpus: realized reused tokens over candidate reusable tokens. High capture means the savings actually land.

#ModelProviderCapture rateList blendedCache disc.
1Claude Fable 5Anthropic94%$20.0090%
2Claude Opus 4.7Anthropic94%$10.0090%
3Claude Opus 4.8Anthropic94%$10.0090%
4Claude Sonnet 4.6Anthropic94%$6.0090%
5Claude Haiku 4.5Anthropic91%$2.0090%
6GPT-5.5OpenAI88%$11.2590%
7GPT-5 MiniOpenAI86%$0.6990%
8Gemini 3.5 FlashGoogle85%$3.3890%
9GLM 5.1Fireworks83%$2.1581%
10Gemini 3.1 Pro PreviewGoogle83%$4.5090%
11Kimi K2.6Fireworks82%$1.7183%
12DeepSeek-V4-ProFireworks81%$2.1792%
13OpenAI gpt-oss-120bFireworks81%$0.2690%
14GPT-5.5 ProOpenAI80%$67.500%
15Grok 4.3xAI79%$1.5684%
Method

Median captureRatePct across eligible requests, with provider-reported evidence where available.

Rank models for your workload.

A diagnostic measures your real reuse and re-ranks the catalog for the way you actually call models.