This is the English edition. 한국어판 and 日本語版 are also available.

GPT-6 Prompt Caching Explained: Costs, Diagnostics and the Cache-Miss Traps

2026-09-26 · AI · United States · Zoogom Editorial

#OpenAI#GPT-6#prompt caching#API cost#AI agents#developers

A stable prompt prefix reused across several long-running AI agent requests

The listed token price is no longer enough to estimate the cost of a GPT-6 agent. Two applications can send the same number of input tokens and produce very different bills. One recomputes a long system prompt, tool catalog and conversation history on every turn. The other writes the stable prefix once and pays a much lower read rate when later requests reuse it.

Caching can also backfire. If an application pays to write context that is never reused, or changes an early tool definition on every request, it absorbs the write premium without receiving the read discount. OpenAI’s new Prompt Cache Diagnostics is designed to make that failure visible.

Three takeaways

Normalized cache economics

Normalized cache economics: Request pattern, Cost with caching, Cost without caching, Practical meaning

These figures normalize the ordinary uncached input rate to one. The actual dollar bill depends on the selected model, processing tier, long-context rules and any other model or tool charges.

What a prompt cache actually stores

OpenAI Developers’ prompt-caching guide says the cache preserves computed key-value states, or KV tensors, rather than storing the tokens as the cache object itself. If a later request has an eligible matching prefix, the model can reuse those states instead of recalculating the entire shared segment.

The compared prefix is broader than the text an application labels as a prompt. It includes the rendered context seen by the model: OpenAI-provided instructions, tool names and schemas, developer messages, conversation history, text, images, documents and supported audio. A seemingly minor change to a tool schema can therefore prevent everything after that point from matching.

The key GPT-5.6-and-later numbers

The minimum cacheable prefix for GPT-5.6 and later is 1,024 visible input tokens. Hidden OpenAI-provided system content does not count toward that minimum.

Writing an eligible prefix costs 1.25 times the ordinary uncached input rate. Reading a matching cached prefix costs 0.1 times that rate. One write plus one complete reuse therefore costs 1.35 times the base input cost, compared with 2 times for processing the same prefix twice without caching. One write and nine complete reads cost 2.15 times, compared with 10 times without caching.

This is an intentionally simple example. New user input and new tool results after the shared prefix still incur input processing, and a partial match saves less than complete reuse.

Implicit versus explicit caching

In implicit mode, OpenAI places a cache breakpoint at the end of the latest eligible message, such as a user message or the final tool response in a consecutive group. It is the convenient default for multi-turn conversations that keep appending to stable history.

Explicit-only mode lets the developer place prompt_cache_breakpoint markers on selected content blocks. If no explicit breakpoint is present, that request neither uses prompt caching nor creates a cache write. This makes it possible to cache a long policy, shared reference set and stable tool definitions while leaving a customer-specific or rapidly changing suffix uncached.

A request can create up to four cache writes. Multiple boundaries can serve content that changes at different rates, but every additional write needs a credible reuse case. More breakpoints do not automatically mean lower cost.

A matching stable prefix on one side and a tool or context change breaking reuse on the other

The settings most likely to cause a miss

The first is the model. Different models use different weights and caching behavior, so an application should not assume that a prefix written for one model can be read by another.

The second is the tool catalog. Names, descriptions, schemas, ordering and parallel-call settings can change the rendered prefix. If a request should not call tools, keeping the catalog stable and using tool_choice: "none" is often more cache-friendly than deleting the tools from the request.

The third category includes output and reasoning settings. A Structured Outputs schema, reasoning.effort and response verbosity can alter model-side instructions. For supported GPT-6 conversations, OpenAI recommends appending a configuration_update when changing reasoning effort instead of rewriting the earlier request-level setting and invalidating a reusable prefix.

The fourth is compaction. Context management can replace earlier conversation content with a compacted representation. That saves context space, but it may stop cache reuse from the first changed token onward. Context length and cache reuse are related operational costs, not the same optimization.

Finally, dynamic values placed too early cause avoidable misses. A timestamp, account-specific permission or session identifier at the beginning of a developer message can make all later shared material appear different. Stable instructions should precede the breakpoint; changing values should follow it whenever the application’s safety and behavior allow.

What Prompt Cache Diagnostics adds

OpenAI made Prompt Cache Diagnostics generally available for GPT-5.6 and later supported Responses API models on September 8, 2026. It can compare reuse against a previous response, identify why a cache missed and point to troubleshooting guidance. OpenAI API changelog

The Prompt Caching Dashboard answers a different question. It shows cache hit rates over time, reads per write and the distribution of cache-read, cache-write and uncached tokens. Diagnostics explains an individual mismatch; the dashboard reveals whether the production system is improving.

The role of prompt_cache_key also changed. Earlier models benefit from a stable key for routing related requests. GPT-5.6 and later handle cache routing automatically, so the key is optional for optimization. It remains useful for separating accounting by customer or workspace and reducing the chance that one user could infer another user’s matching prefix through cache-hit behavior.

A production design checklist

Start by dividing every request into stable, periodically changing and per-turn content. Safety policy, permanent product documentation and fixed tool schemas belong in the stable layer. Account policy, timestamps and session state change more often. The current question and newest tool output belong at the end.

Measure reuse instead of assuming it. A large shared policy used by thousands of requests is an obvious candidate. A one-off document uploaded for a single answer may not justify a deliberate write. Long-running agents benefit only if earlier history and tools remain sufficiently stable.

Do not sacrifice behavior for a better hit rate. Moving a crucial user-specific rule after the wrong breakpoint could change model behavior. Security, correctness and tenant isolation take precedence over token savings.

Track cost per completed task. Input, cache writes, cache reads, output tokens, retries, container or tool charges and human review all contribute to the economic outcome. A high cache-hit rate can coexist with an expensive or unreliable workflow.

For U.S. SaaS teams, allocate cache usage by customer or workspace when explaining bills. The cache key can support separate accounting, but it is not a substitute for authorization or data-isolation controls.

Frequently asked questions

Is prompt caching enabled automatically?

Supported models can use prompt caching and implicit mode is the normal starting point. If an application selects explicit-only mode and provides no breakpoint, no cache write occurs.

Is the cache write always a bad deal?

It is more expensive than ordinary input on the first request. A full first reuse changes the total to 1.35× rather than the 2× required to process the same prefix twice without caching.

Does continuing the same conversation guarantee a hit?

No. A durable session can preserve history, but a changed model, tool catalog, setting or compacted prefix can still prevent reuse. Session continuity and cache matching are separate concepts.

Does the prompt cache store the prompt text as its cache object?

OpenAI’s documentation describes the cache as preserving computed KV tensors, not the tokens themselves. Data-retention and privacy decisions should still rely on the account’s full data-control terms, not that technical description alone.

Conclusion

The most effective GPT-6 cost optimization may be architectural rather than a model downgrade. Keep reusable instructions and tools stable, place changing content after intentional boundaries, and inspect real misses with diagnostics.

Prompt caching is not a free discount. It is a trade: a higher write price in exchange for very inexpensive reuse. The correct metric is total cost per successful task, including misses, output, retries and review—not the cached-token percentage in isolation.

Sources and use notice

OpenAI, GPT and related marks belong to their respective owners. This independent article is not sponsored or approved by OpenAI. It analyzes official OpenAI documentation reviewed on September 26, 2026. Model support and pricing can change, so production decisions should be checked against the latest official documentation. The images are independent editorial illustrations rather than platform screenshots.

Source: OpenAI Developers · Includes original screenshots or graphics