GPT-6 Prompt Caching: Why 90% Cheaper Reads Do Not Mean a 90% Smaller Bill

Verwandte Modelle/Anbieter: GPT OpenAI OpenAI Anbieter
GPT-6 Prompt Caching: Why 90% Cheaper Reads Do Not Mean a 90% Smaller Bill
Article image
Article image

Asking an AI to turn a collection of documents into a report and presentation rarely ends with the first draft. A request to move the conclusion forward, followed by another to add a competitor comparison, may look like two small edits. Behind the interface, however, the application may send the original documents, instructions and conversation history back to the model each time. Without a cache hit, that material must be processed again, adding cost and delay.

OpenAI announced a GPT-6 Prompt Caching upgrade on September 22, 2026, aimed at reducing this repeated work. The update improves default cache-hit rates and introduces more detailed monitoring, diagnostics and controls.

Article image

Under the published pricing described in the announcement, using the same material 10 times could reduce the cost of that repeated input by 78.5%, provided the first use writes it to the cache and all nine subsequent uses hit it. Cached reads cost one-tenth of ordinary input processing. That does not mean the entire bill falls by 90%.

Article image

What prompt caching actually preserves

When a model processes text, it calculates intermediate representations known as key-value, or KV, states. These states support the generation of subsequent responses. Prompt caching retains those computed results so that a later request with an unchanged beginning can reuse them rather than repeat the initial processing.

Article image

The documentation illustrates how previously computed KV states can support later steps. This is not permanent memory, nor does it mean copying a previous answer. The model still generates a fresh response to the new request; it simply avoids some repeated input processing.

Reuse requires an exact match at the beginning of the requests, known as a shared prefix. That prefix can contain system instructions, tool definitions, chat text, images, documents and audio context.

For GPT-5.6 and later models, the shared prefix must contain at least 1,024 visible input tokens to qualify. Cached content remains on the server for at least 30 minutes after its latest write or use. Reusing it during that period refreshes retention without another write charge. Staying in the same chat does not guarantee a hit: routing to another server or region can still prevent reuse.

Article image

The difference between cheaper reads and a cheaper bill

Cache creation has an upfront cost. For GPT-5.6 and later models, the stated cache-write price is 1.25 times the ordinary input rate, while a successful cached read costs 0.1 times that rate.

Article image

Suppose processing a document normally costs one unit. Sending it 10 times without caching costs 10 units. With one cache write and nine successful reads, the calculation becomes 1.25 + (9 × 0.1) = 2.15 units—a reduction of 78.5%.

That saving applies only to the reused document. New questions, generated output and other billable computation still incur their usual charges. If material is written to the cache but never reused, its initial processing instead costs more than ordinary input. Actual savings therefore depend on reuse, not just the discounted read price.

Article image

Why small changes can break reuse

Semantic similarity is insufficient for caching: the shared prefix must match exactly. An application that places a changing timestamp at the start of every prompt can prevent reuse of otherwise identical documents that follow it. The practical design rule is to place stable material first and append changing content afterward.

GPT-6缓存输入最高省九成,你的账单为啥没打一折?

The upgrade adds explicit breakpoints that identify the boundary of a reusable prefix. These mark cache boundaries; they do not instruct the model to pause its reasoning.

Article image

An interactive example distinguishes a reused prefix in green from separately processed content in red. Tool definitions also need to remain stable: renaming tools or changing parameter order can disrupt matching. OpenAI recommends preserving tool definitions and their order, using control parameters to make an unnecessary tool temporarily unavailable rather than removing its definition.

For GPT-6, changing the global reasoning-effort setting can likewise disrupt prefix reuse. Adjusting reasoning effort through incremental updates within the conversation can preserve the cache while changing how much reasoning the model performs.

These are primarily application-development decisions, but their consequences reach users through response times and costs.

Article image

Making cache losses visible

The update includes a cache dashboard that separates reused input from input requiring fresh processing and shows when hit rates decline.

Article image

The September 22 dashboard example displays both cache-hit rates and input composition. Diagnostic tools compare a current request with previous responses to identify changes in tools or settings and estimate how many potentially reusable tokens were lost. In one official example, a small tool-name change prevented 5,629 tokens from hitting the cache.

Another feature, Warmup, allows an application to process known system instructions, tool descriptions or background material before the user submits a request. Moving that work ahead of the interaction can reduce time to the first output token. It does not eliminate the computation or its price: cache writes still cost 1.25 times the ordinary input rate.

Manus reports that its cache-hit rate for OpenAI models increased from approximately 85% to consistently above 90%. Across trillions of requests, GitHub Copilot reports that the share of prompt tokens requiring repeated processing fell by more than 50% relative to its previous baseline, with a noticeable improvement in time to first token.

Article image

Who receives the savings?

For developers paying directly for API usage, successful caching can appear directly in usage reports and bills. For customers of ChatGPT, Manus or other subscription and bundled products, the connection is less direct because retail pricing sits between infrastructure costs and the user.

This update is not a ChatGPT price-cut announcement. Lower processing costs could support cheaper subscriptions, larger usage allowances or more demanding features. They do not automatically produce any particular customer discount.

As AI applications move beyond short answers toward tasks lasting several hours, request design and resource management become more important competitive factors. Two applications may pay the same model rates yet deliver different costs and response times because they organize prompts and reuse context differently. The remaining question is how much of that efficiency will reach users as tangible value.

Article image
Article image
Article image

Diesen Artikel teilen