For most business use cases, Claude Sonnet 5.5 and GPT-6 Sol are the two models to test first. Both cost $2 per million input tokens and $10 per million output tokens, five times less than Claude Fable 5.1 or GPT-6 Astra. Claude Opus 5.5 is the pick for long, open-ended work where the model has to decide the approach on its own. For volume work such as email triage or simple extraction, GPT-6 Luna and Gemini 3.8 Flash do the job for a fraction of the price. In four weeks, Anthropic, OpenAI and Google all refreshed their lineups, so this guide gives you the current price table as of September 28, 2026, a model choice per business task, a worked cost per document, and a method to decide on your own data.
The current generation at a glance
Seven models matter for a business at the end of September 2026, plus Gemini 3.1 Pro, which still comes up in most comparisons. Prices are public API list prices in dollars per million tokens (input / output), excluding caching, batch discounts and paid tools.
| Model | Released | Input / output | Positioning | Good for |
|---|---|---|---|---|
| Claude Fable 5.1 | Sept. 1 | $10 / $50 | Anthropic's most capable model | Very long agents, hard research and analysis, as an escalation tier |
| GPT-6 Astra | Sept. 3 | $10 / $50 | OpenAI's most advanced model | Computer use, browsing, security work, the hardest cases |
| Claude Opus 5.5 | Sept. 22 | $4 / $20 | Replaces Opus 5, announced at Fable 5.1 level | Agents that plan on their own, long analyses, large refactors |
| Claude Sonnet 5.5 | Sept. 28 | $2 / $10 | Mid tier, very close to Opus 5.5 according to Anthropic | Well-scoped everyday tasks, documents, writing, bug fixing |
| GPT-6 Sol | Sept. 22 | $2 / $10 | Demanding everyday tasks, including code | Assistants, writing, code, agents with a defined path |
| Gemini 3.8 Flash | Sept. 2 | $0.75 / $3.75 then $1.50 / $7.50 from Jan. 1, 2027 |
Fast, multimodal, one million token context | PDFs and images at volume, low-cost agents |
| GPT-6 Luna | Sept. 22 | $0.10 / $0.50 | High volume, tasks with a clear goal | Email triage, simple extraction, summaries, quick answers |
| Gemini 3.1 Pro | Preview | $2 / $12 (up to 200k tokens, as of Sept. 1) | Google's high-end model, still in preview at that date | Multimodal work, teams already on Google Cloud |
Three price tiers stand out. At the top, Fable 5.1 and GPT-6 Astra at $10 / $50. In the middle, Opus 5.5 at $4 / $20, then Sonnet 5.5 and GPT-6 Sol at $2 / $10. At the bottom, Gemini 3.8 Flash and GPT-6 Luna, 13 to 100 times cheaper than the top tier on output tokens.
What moved the most is the price of the mid tier. In early September, a sensible default model cost $5 / $25 (Claude Opus 5) or $4 / $20 (GPT-5.6 Sol). Today it costs $2 / $10 at both Anthropic and OpenAI, which says the GPT-6 series costs half as much as the 5.6 series. On the Anthropic side, Opus 5.5 is 20% cheaper per token than Opus 5 and, according to Anthropic, about 40% cheaper in real use for a level comparable to Fable 5.1. It replaces Opus 5 two months after that model's release.
Outside this table
Anthropic says Claude Haiku 5.5 is coming "in the coming weeks". DeepSeek V4.1-Flash, released around September 10 with open weights under the MIT license, mostly matters if you want to host the model yourself. If your data cannot leave infrastructure you control, none of the proprietary models above solves that on its own, and an open-weight model deployed in your environment is worth evaluating. Our guides on self-hosted RAG with open models and Mistral vs OpenAI vs Anthropic cover that path and its GDPR implications.
Sonnet or Opus, GPT-6 or Claude: which one wins?
No model wins everywhere. Each head-to-head is decided by the nature of the task and by the ecosystem you already work in.
Sonnet 5.5 or Opus 5.5?
Anthropic draws the line itself. In its Sonnet 5.5 announcement, it places Sonnet 5.5 two points below Opus 5.5 on GDPval-AA, a professional work evaluation, while stating that Opus 5.5 remains clearly stronger at complex, open-ended work that requires sustained judgment. Sonnet 5.5 is described as strongest at well-scoped everyday tasks, bug fixing and producing polished documents.
So the question is simple. Can you describe the expected result? If yes, pick Sonnet 5.5, at half the price. If the model has to choose its own approach, over a long session or dozens of steps, test Opus 5.5.
GPT-6 Sol or Claude?
GPT-6 Sol and Sonnet 5.5 have exactly the same price. OpenAI says Sol makes about half as many mistakes as its predecessor on its internal factuality evaluation, reaching reliability close to Astra. OpenAI also claims that Sol and Luna outperform Fable and Opus. That is a vendor claim, reported by TechCrunch without detailed comparative figures, so verify it on your own cases.
At equal price, tooling is often the tie-breaker. If your team works in ChatGPT Work and Codex, or your integrations already run on the OpenAI API, Sol is the natural candidate. If you are on Claude, or on AWS, Google Cloud or Azure with Claude, Sonnet 5.5 and Opus 5.5 plug in without changing your stack. Against Opus 5.5, Sol costs half as much. Compare it with Sonnet 5.5 first, and bring in Opus 5.5 only if Sol fails on your hardest cases.
For day-to-day use in ChatGPT or Claude, the gap between these models is smaller than the gap in how two colleagues use them. A team that frames its requests well gets more out of Sonnet 5.5 than a team asking vague questions to Opus 5.5.
Gemini 3.1 Pro vs Opus 5.5?
The question keeps coming up because Gemini 3.1 Pro is cheaper than Opus 5.5 per token ($2 / $12 versus $4 / $20). It has lost most of its relevance, though. On September 1, Gemini 3.1 Pro was clearly behind Opus 5 on the Artificial Analysis Intelligence Index (48 versus 61 at high effort), it was still in preview, and it now costs more on output than Sonnet 5.5 or GPT-6 Sol.
For a long, complex agent, test Opus 5.5 first. If you are on Google Cloud, the useful comparison is Gemini 3.8 Flash, which natively reads PDFs, images, audio and video, against Sonnet 5.5 or GPT-6 Sol.
Which model for which business task
The right model is the cheapest one that reaches your quality bar on your task. Below are the starting points we test first, and the signal that justifies stepping up a tier.
| Task | Start with | Step up to | When to step up |
|---|---|---|---|
| Document extraction | Gemini 3.8 Flash or GPT-6 Luna | Sonnet 5.5 or GPT-6 Sol | Highly varied layouts, fields that need interpretation |
| Email triage | GPT-6 Luna | Sonnet 5.5 or GPT-6 Sol | Several requests per email, long history, reply to draft |
| Internal assistant on documents | Sonnet 5.5 or GPT-6 Sol | Opus 5.5 | Questions that span several documents |
| Writing | Sonnet 5.5 | Opus 5.5 | Long documents where substance matters as much as style |
| Multi-step agent connected to your tools | Sonnet 5.5 or GPT-6 Sol for a defined path, Opus 5.5 otherwise | Fable 5.1 or GPT-6 Astra | Repeated, measured failures on long cases |
| Code | Sonnet 5.5 or GPT-6 Sol | Opus 5.5, then Fable 5.1 or Astra | Large refactors, migrations, hard incidents |
Document extraction. Invoices, purchase orders, certificates, statements. The output format is fixed and there is little room for interpretation. Start with Gemini 3.8 Flash if your documents arrive as scanned PDFs or photos, and with GPT-6 Luna if the text is already cleanly extracted. Move up to Sonnet 5.5 or Sol when layouts vary a lot (ten suppliers, ten invoice templates) or when a field requires interpretation, such as a price revision clause in a contract. Our breakdown of PDF data extraction architecture and our guide to AI invoice OCR and validation go deeper on this case.
Email triage. Classify an incoming email, identify the customer, extract the intent, route it to the right team. GPT-6 Luna is enough in most cases, and OpenAI positions it precisely on these clear-goal tasks. Step up for emails that mix several requests or depend on a long thread, or when the next step is drafting the reply. On this use case, most of the work lies in the CRM connection and the escalation rule to a human, as shown in our article on AI ticket classification.
Internal assistant on your documents. An assistant that answers from your procedures, contracts or technical sheets. Sonnet 5.5 is our first pick, for its cost and its clearer writing, one of the improvements Anthropic highlights. GPT-6 Sol is the OpenAI equivalent at the same price. Opus 5.5 is only justified when questions require cross-referencing several documents and reasoning over them, such as comparing two versions of a contract. Answer quality depends first on retrieving the right passages, as the five RAG failure modes we keep seeing in production show.
Writing. Customer replies, meeting notes, sales proposals, internal memos. Sonnet 5.5 by default. Luna can be enough for short, tightly framed texts such as an acknowledgment or a follow-up. Opus 5.5 earns its price on long documents where substance matters, such as a technical proposal or a synthesis of dozens of pages.
Multi-step agent connected to your tools. An agent that reads a case file, queries the ERP, updates the CRM and prepares a deliverable. Everything depends on the path. If it is well defined, with known steps and clear tools, Sonnet 5.5 or GPT-6 Sol are often enough. If the agent has to plan for itself and hold up over a long session, start with Opus 5.5. Fable 5.1 and Astra are escalation tiers for cases where Opus fails repeatedly and measurably. Our comparison of workflows vs AI agents helps decide whether you need an agent at all, and our AI agent engineering work follows that logic.
Code. For bug fixing and everyday changes, Sonnet 5.5, which Anthropic presents as strong on exactly that, or GPT-6 Sol if your team works in Codex. Opus 5.5 for large refactors and migrations. Fable 5.1 and Astra for the hardest investigations, when a senior engineer's hour costs more than the API bill.
Torn between two models for your use case?
We run Sonnet 5.5, GPT-6 Sol and a lighter model on 30 to 50 of your real documents, emails or questions, with the same prompt, and show you quality, cost per task and latency side by side.
What one task actually costs, model by model
A price per million tokens means little until you turn it into a price per task. Take a common case, extracting data from a document, with deliberately simple assumptions.
- One document with 3,000 input tokens, instructions included.
- A structured answer of 500 output tokens.
- 10,000 documents per month.
- No caching, no batch discount, no reasoning tokens, no retries.
The math for Sonnet 5.5
- Input, 3,000 × $2 / 1,000,000 = $0.006
- Output, 500 × $10 / 1,000,000 = $0.005
- Total, $0.011 per document, so $0.011 × 10,000 = $110 per month
The same calculation for each model gives the table below.
| Model | Input per document | Output per document | Cost per document | 10,000 documents per month |
|---|---|---|---|---|
| Fable 5.1 or GPT-6 Astra | $0.030 | $0.025 | $0.055 | $550 |
| Opus 5.5 | $0.012 | $0.010 | $0.022 | $220 |
| Sonnet 5.5 or GPT-6 Sol | $0.006 | $0.005 | $0.011 | $110 |
| Gemini 3.8 Flash (until Dec. 31, 2026) | $0.00225 | $0.001875 | $0.004125 | $41.25 |
| Gemini 3.8 Flash (from Jan. 1, 2027) | $0.0045 | $0.00375 | $0.00825 | $82.50 |
| GPT-6 Luna | $0.0003 | $0.00025 | $0.00055 | $5.50 |
In this scenario, moving from Fable 5.1 to Sonnet 5.5 cuts the bill by five, and moving from Sonnet 5.5 to Luna cuts it by another twenty. In production, three effects push this calculation off course.
Reasoning. Thinking tokens are generally billed at the output rate. If the model thinks for 2,000 tokens before answering, output goes from 500 to 2,500 tokens. Sonnet 5.5 then costs $0.006 + $0.025 = $0.031 per document, or $310 per month, almost three times more. For simple extraction, a low reasoning effort is often enough. Gemini 3.8 Flash also consumes more tokens than version 3.7 at the same unit price, so its real bill can drift above the table.
Caching. If 2,000 of the 3,000 input tokens are fixed instructions read from cache at $0.20 per million, Sonnet 5.5 input drops to $0.0004 + $0.002 = $0.0024. The document then costs $0.0074, or $74 per month instead of $110. Cache writes, billed at $2.50 per million for Sonnet 5.5, are paid again each time the cache is refreshed and weigh little on a continuous flow.
Errors. This is often the heaviest line. Assume, for illustration, that a light model gets 10% of documents wrong and a mid-tier model 3%. On 10,000 documents, that is 700 extra corrections every month. At two minutes per correction, about 23 hours of human work, to save roughly $105 in API fees. If the gap is only 0.5 points, that is 50 corrections and under two hours a month, and Luna wins again. The metric that matters is cost per validated task, not price per token. A router that sends only the hard cases to the expensive model, which we build as part of our LLM integration work, is often the cheapest setup overall.
What launch announcements do not tell you
Each vendor publishes results measured on its own protocols. They help you pick which models to test, not which one to put in production.
OpenAI claims GPT-6 Sol and Luna do better than Fable and Opus. Anthropic places Opus 5.5 at the level of Fable 5.1. These claims can be true on the tests each vendor chose and false on your task. Harnesses, prompts and effort levels differ from one lab to the next, and none of them tests your invoices, your emails or your domain vocabulary.
The most overlooked criterion is format compliance. A more capable model that adds one sentence before the expected JSON breaks an integration the previous model handled fine. That gap shows up in no benchmark table, which is why we treat structured outputs in production as a test criterion of its own.
How to compare models on your own data
- Collect 30 to 50 real cases. Easy ones, hard ones, and errors already seen in production, each with the expected answer validated by someone who knows the business.
- Set the pass bar before testing. For example 95% of fields correct, no invented amounts, valid JSON on every call. A bar set afterwards tends to bend toward the model you already prefer.
- Keep the prompt identical. Same instructions, same documents, same tools, effort level recorded. Adapt prompts per model only in a second round, once the raw ranking is known.
- Measure four things. Quality (field-level accuracy, or a correct and sourced answer), real cost per task from tokens consumed, latency (median and worst case), and output format compliance.
- Keep the cheapest model that passes the bar. The edge cases it misses can be routed to the tier above.
The slowest part is building the cases and their expected answers. Running the models can be automated and replayed at every new release. Our guide on how to choose an AI vendor covers the evaluation criteria in more detail.
Should you migrate now?
Yes if a test on your cases shows a gain, no as a reflex. The answer mostly depends on the model you run today.
| You run on | Test | Why now |
|---|---|---|
| Claude Opus 5 | Opus 5.5, with Sonnet 5.5 in parallel | Opus 5.5 replaces Opus 5 at 20% less per token, and Sonnet 5.5 costs 2.5 times less than Opus 5 |
| Claude Sonnet 5 | Sonnet 5.5 | Same price, over 30% faster and up to 30% cheaper per task according to Anthropic |
| GPT-5.6 Sol or Terra | GPT-6 Sol | Sol goes from $4 / $20 to $2 / $10, below Terra's launch price, and the GPT-6 series has no Terra tier |
| GPT-5.6 Luna | GPT-6 Luna | Roughly half the price, $12 versus $5.50 per month in our example |
| Fable 5.1 or GPT-6 Astra | Opus 5.5 and GPT-6 Sol | $550 per month versus $220 or $110 in our example, if quality holds |
| Gemini 3.8 Flash | GPT-6 Luna, before December | The price doubles on January 1, 2027, from $41.25 to $82.50 per month in our example |
The test on your evaluation set always comes first. A prompt tuned for Opus 5 or GPT-5.6 can behave differently on the new version, even one that is better on average, especially on output format. Then run the new model in parallel on real traffic for a few days, without using its answers, and compare. Finally, keep the previous model configured as a fallback for as long as the vendor supports it.
The release pace argues for an architecture where the model is just a parameter. Opus 5 was replaced two months after launch. If switching models means reworking your application, fix that before any migration, as covered in our guide to deploying LLMs to production.
Nothing in production yet? Do not wait for the next model. Between two models of the same tier, the gap is small compared with the gap between a well-chosen use case and a poorly chosen one. Start by identifying one to three tasks that suit AI and check their feasibility on your data, which is what our AI audit is for. The model choice comes last, on your own cases.
Already running on Opus 5, GPT-5.6 or Gemini 3.8 Flash?
We replay your real cases on Opus 5.5, Sonnet 5.5 or GPT-6 with the same prompt, and you migrate only if the quality or cost gain is measured. Your production stays in place during the test.
Frequently asked questions on Sonnet 5.5, Opus 5.5 and GPT-6
Further reading
- Opus 4.8 vs GPT-5.5 vs Gemini 3.1 Pro, our earlier frontier model comparison and its decision grid.
- How to choose the right AI vendor, why a public benchmark is not enough and how to build your own evaluation.
- Structured outputs in production, keeping JSON answers valid when you switch models.
- PDF data extraction with AI, the architecture behind reliable document processing.
- Workflow vs AI agent, when a fixed workflow beats an autonomous agent.
- Mistral vs OpenAI vs Anthropic, the sovereignty and GDPR angle of model choice.
- Deploying LLMs to production, self-hosting versus API with real cost numbers.