OpenAI service tiers, explained
batch, flex, standard, fast,
ultrafast — five tiers spanning a 12× spread in price,
and a decision most teams never make at all.
Short version: the tier is a per-request parameter. Pick the slowest tier your workload tolerates — Batch and Flex are half price, Fast is double, Ultrafast is six times. Standard is the correct default only for interactive traffic you cannot predict.
The five tiers
| Tier | service_tier | Rate | Speed | Residency |
|---|---|---|---|---|
| Batch | batch | 0.5× | ≤24 h | EU + global |
| Flex | flex | 0.5× | queued | EU + global |
| Standard | default | 1× | 1× | EU + global |
| Fast | fast / priority | 2× | up to 2.5× | global only |
| Ultrafast | ultrafast | 6× | up to 6× | US + global only |
Rate is the published multiplier on the Standard per-token price. Speed is the vendor's published ceiling, not a guarantee for your workload.
What the multiplier actually does to a bill
The multiplier applies to all four lines of the usage dashboard — uncached input, cache
reads, cache writes and output — so it scales the whole request, not one part of it. On
GPT-6 Astra, whose standard rates are $10 input and $50 output per
million tokens, Ultrafast publishes $60 input and $300 output —
six times on every figure.
That last number is worth sitting with. Astra on Ultrafast costs $300 per million output tokens. GPT-6.1 Sol on Standard costs $10. The same workload can be thirty times more expensive purely because of a parameter nobody on the team remembers setting.
Fast and Ultrafast are not available everywhere
Both premium tiers run on global processing and US data residency. Neither has an EU endpoint, and OpenAI's Ultrafast guide states it directly: the mode does not support EU or other non-US regional processing endpoints. A European customer can still send Ultrafast traffic through global processing — but cannot keep the inference inside the EU while doing it.
If your contracts specify where inference happens, Fast and Ultrafast may simply be off the table, regardless of price. Check that before you model the savings.
Ultrafast also starts throttled
Default Ultrafast token rate limits for GPT-6 Astra are deliberately low — 500,000 tokens per minute for API tiers 1 to 3, rising to 5M at tier 5. OpenAI describes the mode as available to all API users "at low rate limits" and points organisations to their account team for more. Budget for the conversation, not just the tokens.
Picking a tier, by workload shape
- Nobody is waiting. Nightly enrichment, backfills, eval runs,
embeddings refresh →
batch. Half price, and the largest safe saving available to most teams. - Traffic is spiky and retryable. Non-production endpoints, batch-ish
interactive work →
flex. Also half price, but it may queue. - A human is waiting, volume unpredictable. →
standard. Cache the system prompt instead of buying speed. - A human is waiting and the latency is the product. →
fast, set per request. - Wall-clock time is the constraint and volume is large. →
ultrafast, and only on Astra today.
A five-day order of operations
Tier changes are cheap and reversible, which makes them the right first move — before any prompt engineering or model swap:
- Sort every workload by one question: is a human waiting for this response?
- Move the "no" pile to
batch. Watch the async failure rate. - Move the "maybe" pile to
flexbehind a feature flag, and watch retries. - Leave the "yes" pile on
standard, and cache its stable prefix. - Only then look at
fast— for the specific requests where latency is actually worth 2×, not as an account-wide default.
Work out your own number
The composer on the home page takes your token counts, cache hit rate and daily volume, and shows the per-call and monthly cost on every tier at once. It runs entirely in your browser — nothing is uploaded.