The 272K cliff: GPT-6.1 Sol long-context pricing

One token over 272,000 and the entire request re-prices — not just the tokens past the line. Here is the band table and the arithmetic.

Short version: below the threshold GPT-6.1 Sol bills at $2 input / $10 output per million tokens. Above it, the same request bills at $4 / $15. The premium is not prorated.

The two bands

Input lengthInput /1MCached input /1MCache write /1MOutput /1M
≤ 272,000 tokens$2.00$0.10$2.50$10.00
> 272,000 tokens$4.00$0.20$5.00$15.00

Input, cached input and cache writes double; output rises by 1.5×. The same 2× / 1.5× rule applies across the GPT-6 family, and to GPT-5.6 Sol and Terra.

Why it is a cliff and not a slope

Most tiered pricing charges a premium rate only on the tokens above the threshold. OpenAI does not. Crossing the line re-prices every token in the request, which means the cost curve is discontinuous — a step, not a ramp.

Take a request with 271,000 input tokens and 1,000 output tokens on Standard:

RequestInput costOutput costTotal
271,000 in / 1,000 out$0.5420$0.0100$0.5520
272,001 in / 1,000 out$1.0880$0.0150$1.1030

Adding 1,001 tokens — about 0.4% more input — takes the request from $0.55 to $1.10. The input line exactly doubles. If that request runs 2,000 times a day, the difference is roughly $33,000 a month for a prompt that grew by a thousandth.

The one place caching narrows the gap

Cached input is the only line where the absolute increase is small: $0.10 to $0.20 per million. It is still a doubling, but a cache-heavy workload with a modest uncached tail feels the cliff far less than a workload that ships raw context.

Concretely, 100,000 cached tokens plus 10,000 ordinary input tokens cost $0.03 below the line and $0.06 above it — the ratio holds, but the absolute number stays small.

What to do about it

The threshold is per request. Splitting one 400K-token job into two 200K-token requests keeps both below the line — same tokens, half the input rate.

Work out your own number

The composer on the home page flags the cliff automatically: set input tokens above 272,000 and it shows the monthly cost before and after trimming, alongside the tier and cache levers. It runs entirely in your browser — nothing is uploaded.

Related