What does your OpenAI request actually cost?
Almost everyone leaves service_tier on the default. It is the single
biggest lever on an OpenAI bill that nobody touches. Compose your real cost here —
tier, caching, reasoning effort and the 272K long-context cliff — across
GPT-6.1 Sol, GPT-6 Astra and GPT-6 Luna.
1. What are you doing?
This sets a starting configuration. Everything stays editable.
2. Tune the request
Defaults are typical for an API-backed app.
Cost
Every line below is billed separately on the usage dashboard.
Cache writes: $0 per day (set the count above; each write bills at 1.25× the input rate).
Same workload, every tier
Cheapest first.
| Tier | Rate | Speed | Per call | Per 30 days |
|---|
Same workload, every model
At the tier you selected. Sorted on your inputs, not on a benchmark score.
| Model | Price / 1M | Per call | Per 30 days |
|---|
Monthly cost at a glance
Same numbers, drawn to scale.
Two levers, and almost everyone only pulls one
An OpenAI request has a headline rate — $2 / $10 per million tokens for
GPT-6.1 Sol — and almost every cost calculator stops there. But the headline rate is
only what you pay on the Standard tier, at short context, with no caching.
Three multipliers sit on top of it, and they are where the money actually is.
The rule of thumb: set the tier to the slowest one your workload tolerates, then cache the stable part of the prompt, then keep effort at the lowest level that still produces a correct answer. In that order.
Lever 1 — the service tier
service_tier is a per-request parameter, not a plan you buy. Batch and Flex
run at half the standard rate. Fast runs at double.
Ultrafast runs at six times — and on GPT-6 Astra that is six times the
speed for exactly six times the price, published as $300 per million output
tokens against a standard $50.
The failure mode is not picking a bad tier. It is never picking one at all: overnight enrichment, backfills, eval runs and embeddings refreshes sit on Standard at full price for no reason, because the default is Standard and nobody revisits it.
Lever 2 — the 272K cliff
Above 272,000 input tokens the whole request is re-priced: input and cached input double, output rises 1.5×. The premium applies to every token in the request, not just the tokens past the line. A 271,999-token prompt and a 272,001-token prompt are billed at rates that differ by 100% on the input side.
That makes prompt trimming a first-class cost lever, not a housekeeping chore. If an agent loop drifts past the line, the fix is to shorten the history — not to buy a faster tier.
Lever 3 — reasoning effort
Reasoning tokens are billed as output tokens, and the output rate is five times the input rate on every current model. Raising effort therefore multiplies the expensive half of your bill while leaving the cheap half untouched.
The multipliers in the slider above are planning assumptions, not vendor-published figures — OpenAI does not publish a thinking-token multiplier per effort level. Measure your own traffic, then set them to match.
Where each tier belongs
| Tier | Rate | Reach for it when |
|---|---|---|
batch | 0.5× | Overnight enrichment, backfills, eval runs, embeddings refresh |
flex | 0.5× | Non-production traffic, unpredictable spikes, retryable jobs |
standard | 1× | Interactive traffic you cannot predict |
fast | 2× | Latency-critical requests — bought per request, never as a default |
ultrafast | 6× | Bulk generation where wall-clock time is the constraint. Astra only, no EU endpoint |
Two mistakes that cost real money
- Leaving the tier on the default forever. Standard is the correct choice for interactive traffic and the wrong one for everything else. Sort your workloads by whether a human is waiting, then move the ones that are not onto Batch or Flex.
- Paying for a faster tier when the prompt is the problem. If a request is slow because it carries 300K tokens of history, Fast mode makes an expensive request expensive and quick. Trim it below the cliff instead.
Frequently asked
Is Ultrafast worth six times the price?
Only when wall-clock time is the actual constraint. OpenAI publishes up to 8× faster token generation in Codex and up to 6× in the API — which is exactly the price multiple, so you are buying speed at par, not at a premium or a discount. It is also Astra-only today, runs at deliberately low default rate limits, and has no EU endpoint.
Do these prices include prompt caching?
They include it as a lever you control. A cache read on GPT-6.1 Sol costs $0.10 per million against $2.00 uncached — a 95% discount — while a cache write costs $2.50, or 1.25× the input rate. Set the hit rate and write count above and the calculator composes both.
How current are the numbers?
The data file behind this page records a verification date, shown in the footer. Rates in this market move monthly, and service tiers were still changing through the second half of 2026 — always confirm against the vendor's own pricing page before you commit a budget.