OpenAI logoOpenAIopenai.com

GPT-5.4

GPT-5.4 is OpenAI's value pick in the frontier tier: most of GPT-5.5's reasoning and coding strength, at a much lower price. It's built for hard professional work like deep reasoning, long research, and real software engineering. OpenAI calls it 'a more affordable model for coding and professional work.' It's also their first mainline reasoning model to fold in frontier coding. The context window is large: 1,050,000 tokens, with up to 128,000 tokens of output. That gives you room to reason over whole contracts, full filings, research sets, or an entire codebase in one prompt. So when do you move off it? Step up to GPT-5.5 when a task sits at the very edge of hard, like the trickiest reasoning or your highest-stakes code. Step down to a mini model for simple, high-volume jobs like short replies or basic extraction. For most frontier work, though, GPT-5.4 is the one that fits both the task and the budget. You get frontier-level answers without paying the frontier-level rate. For a lot of teams, that trade is the whole reason to pick it.

Context window
1,050,000 tokens
Max output
128,000 tokens
Knowledge cutoff
Aug 2025
API identifier
gpt-5.4

Pricing · per 1M tokens

Input$2.50
Output$15.00
Cached input$0.25

Cached input is ~10× cheaper — it pays off when you send the same context again and again.

Real numbers

What GPT-5.4 costs in practice

Real enterprise workloads, with the token assumptions shown openly. Your mileage varies — these are honest starting points, not guesses.

FinanceInformation Technology

Modernizing a legacy trading system

$0.145per request
$290per month

Your engineers feed old service code into GPT-5.4 and ask it to refactor and move it to a modern stack. The large context holds several files at once, so the logic stays consistent across the change. You get frontier coding on a big migration without paying GPT-5.5 rates for every file.

28,000 input tokens5,000 output tokens2,000 requests / month

Input and output roughly balance here. If you re-run the same files while iterating, cached input at a tenth of the input rate takes a real bite out of the bill.

HealthcareResearch & Innovation

Answering questions across long clinical documents

$0.147per request
$176.4per month

Your research team drops trial protocols and study papers into one prompt and asks grounded questions. The context window is large enough to hold the full set, so answers stay tied to the source. It reasons across the documents instead of guessing from a short summary.

48,000 input tokens1,800 output tokens1,200 requests / month

This one is input-heavy, with long documents in and short answers out, so most of your spend sits on the input side. Reusing the same document set makes cached input well worth setting up.

ManufacturingFinance & Accounting

Pulling key terms from supplier contracts

$0.059per request
$354per month

Your finance team feeds supplier contracts to GPT-5.4 and asks for pricing, penalties, and renewal dates in a clean structure. It reads the messy legal language, pulls the terms, and flags anything that looks risky. That saves hours of manual review on every deal.

14,000 input tokens1,600 output tokens6,000 requests / month

Middleweight on both sides. GPT-5.4 earns its price when contracts are messy and need judgment; for clean, simple forms, a mini model is the better call.

How we work these out: cost = (input ÷ 1M × $2.50) + (output ÷ 1M × $15.00), then × monthly volume. List prices only, no cached-input discount applied — so these are the ceiling, not the floor.

FAQ

Common questions about GPT-5.4

How much does GPT-5.4 cost?+

GPT-5.4 costs $2.50 per million input tokens and $15.00 per million output tokens. Cached input costs far less, at $0.25 per million, a tenth of the standard input rate, so reusing the same context saves real money. That pricing sits well below GPT-5.5, which is the point of the model. Your actual bill depends on how many tokens you send in and get back.

GPT-5.4 vs GPT-5.5?+

GPT-5.4 gives you most of GPT-5.5's capability at a lower price. Both are frontier reasoning models, but GPT-5.5 is stronger on the very hardest tasks. Choose GPT-5.5 when a job sits at the edge of what's possible. Choose GPT-5.4 when you want frontier-level work without the frontier-level bill.

What is GPT-5.4's context window?+

GPT-5.4 has a 1,050,000-token context window and can return up to 128,000 tokens in one response. That is room for long contracts, full financial filings, or an entire codebase in a single prompt. It makes GPT-5.4 a strong fit for research and retrieval work across large document sets.

Is GPT-5.4 good for coding?+

Yes, coding is one of GPT-5.4's core strengths. It's OpenAI's first mainline reasoning model to fold in frontier coding, so it reasons through a problem instead of just autocompleting. That makes it a strong pick for refactoring, code review, and building features across a large codebase. For the most complex, highest-stakes code, GPT-5.5 is the step up.

GPT-5.4 vs GPT-4o?+

GPT-5.4 is a newer, reasoning-first model built for harder coding and professional work than GPT-4o. It carries a 1,050,000-token context window and a knowledge cutoff of August 31, 2025, so it handles long documents and more recent context well. If your work needs deep reasoning or large inputs, GPT-5.4 is the better fit. For lighter, general-purpose chat, an earlier model like GPT-4o may be all you need.

Pricing as of 8 Jul 2026 · reviewed monthlySources: OpenAI API pricing · OpenAI model docs

Not sure GPT-5.4 fits your workload?

We'll size the right model for your use case and give you a real cost estimate. One free call, no slides.

Talk to us