Anthropic’s new Opus model brings near-flagship reasoning to everyday work at the same price as its predecessor. The launch-day evidence supports most of the capability claims, trims the cost advantage to something more modest, and flags one risk that should change how marketing teams use it.
Anthropic released Claude Opus 5 on 24 July 2026, its fourth significant model in under two months. The framing is unusually commercial for a company that normally leads on capability. Opus 5 is not the smartest thing Anthropic sells, and it does not pretend to be. It is the one the company expects businesses to actually run.
That distinction is easy to miss and expensive to get wrong. Claude Fable 5 remains the flagship for long, autonomous, difficult work. Claude Mythos 5 has frontier capabilities on restricted access. Opus 5 sits below both, at half Fable 5’s token price, aimed at the complex-but-routine work that fills an agency week: research synthesis, technical audits, reporting, multi-step operations. As Anthropic product leader Dianne Penn told Reuters, the rough rule is Opus 5 for value and Fable 5 for “days-long, very autonomous projects”.
The answer in brief: Claude Opus 5 is Anthropic’s newest generally available Opus model, released on 24 July 2026. API pricing is $5 per million input tokens and $25 per million output, unchanged from Opus 4.8 and half Claude Fable 5’s rate. It has a one-million-token context window, five effort levels with high as the default, and a May 2026 knowledge cut-off. Artificial Analysis scored it narrowly ahead of Fable 5 on its Intelligence Index at maximum effort, and measured the cost of a completed task at about 26 per cent below Fable 5 rather than 50 per cent. The same testing found its hallucination rate 14 points higher than Opus 4.8, at 50 per cent. For new agentic and research workloads it is the Claude model to evaluate first. For existing Opus 4.8 workflows, test before switching; Opus 4.8 remains available with no announced shutdown.
Claude Opus 5 at a glance
| Specification | Claude Opus 5 | Why it matters |
|---|---|---|
| Release status | Generally available | Suitable for production evaluation, not a preview |
| Release date | 24 July 2026 | Days old at the time of this review |
| Stable model ID | claude-opus-5 | The identifier used in API requests |
| Inputs and output | Text and images in, text out | Use another model for image or audio generation |
| Context window | 1,000,000 tokens | Holds substantial document sets in one working context |
| Maximum output | 128,000 tokens | Supports long reports and code output |
| Effort levels | Five: low, medium, high, xhigh, max | High is the default and thinking is on by default |
| Base API price | $5 input / $25 output per million tokens | Unchanged from Opus 4.8, half Claude Fable 5’s rate |
| Fast mode | Twice the base token price | Runs at roughly 2.5 times the normal speed |
| Knowledge cut-off | May 2026 | Newer facts still need supplied sources or retrieval |
| Availability | Claude apps, Anthropic API, Amazon Bedrock, Google Cloud, Microsoft Foundry | Default model on Claude Max, strongest on Claude Pro |
| Predecessor status | Opus 4.8 still available | No announced shutdown, so no forced migration |
Sources: Anthropic’s Claude Opus 5 announcement, models overview, migration guidance and pricing documentation, all checked 25 July 2026. Prices are in US dollars and exclude caching, regional premiums and cloud-provider charges.
What has Anthropic actually confirmed about Claude Opus 5?
Opus 5 became generally available on 24 July 2026 across Anthropic’s own products, its API and the three major cloud platforms. Pricing holds at $5 and $25 per million input and output tokens, identical to Opus 4.8, with fast mode at double the token rate for roughly 2.5 times the speed. Prompt caching can reduce the cost of repeatedly supplied reference material, though cache writes and reads carry their own rates. The model accepts text and images and returns text; claims about improved “visual” ability refer to understanding designs and building interfaces, not generating photographs.
The genuinely new control is effort. Anthropic’s migration documentation sets out five levels, low through max, with high as the default, and thinking now enabled by default rather than requested. Higher settings buy more planning, reasoning and tool use; lower settings preserve much of the quality at fewer tokens and lower latency. Two beta features arrived alongside: changing a model’s available tools mid-conversation, and automatic fallbacks that route a declined request to another model instead of returning an error.
The migration guidance contains a warning worth forwarding to whoever owns your prompt library. Opus 5 verifies its own work unprompted, writes longer responses, narrates more of its progress and expands task scope more readily than Opus 4.8. Old templates that instruct the model to double-check everything can now trigger expensive over-verification. Migrating is not a matter of changing a model name in a config file; it needs prompt and cost testing.
Anthropic describes Opus 5 as its most aligned model to date and the least susceptible to being tricked into misuse. That comes from the company’s own automated behavioural audit, published in the Claude Opus 5 system card, which is admirably candid about the method’s limits, including the model’s elevated ability to recognise when it is under evaluation.
What did the first independent tests find?
Unusually for a launch, external evidence landed on day one, and it is more useful than the vendor charts.
Artificial Analysis tested all five effort settings and placed Opus 5 at maximum effort narrowly top of its Intelligence Index, ahead of Fable 5, GPT-5.6 Sol and Opus 4.8. On its agentic knowledge-work evaluation, Opus 5 finished more than 100 Elo points clear of both Fable 5 and GPT-5.6 Sol. The evaluator discloses that it supported Anthropic’s pre-release testing, so this is external measurement rather than blind replication. On day one, it is still the best evidence available.
Its cost findings puncture the simplest version of the value story. Opus 5 at maximum effort averaged $2.03 per Intelligence Index task against $2.75 for Fable 5 with fallback: a real saving, but closer to a quarter than the half the token prices imply, because Opus 5 reasons at length and writes more. Measured against cheaper options the picture reverses entirely, with Opus 4.8 at $1.80 and Sonnet 5 at $1.53 per task. Effort matters enormously here: output token use spans roughly eightfold between low and max settings. Half-price tokens are confirmed. Half-price work is not.

The finding that should change behaviour concerns factual reliability. Artificial Analysis measured Opus 5 as seven points more accurate than Opus 4.8 on closed-book knowledge, while its hallucination rate rose fourteen points to 50 per cent, because the model answers more often when it is uncertain. Anthropic’s system card reports the same direction of travel. A model can know more and invent more at the same time, and for anyone producing client work containing prices, specifications, statistics or quotations, that is the most important number in the launch coverage.

A narrower production test points the same way. CodeRabbit ran Opus 5 against roughly 100 known code-review error patterns, three runs per configuration, and found it produced more precise actionable comments than its production baseline while catching fewer of the known issues and generating about four times as many low-value nitpicks. Its conclusion, that Opus 5 works as a precision-oriented second reviewer rather than a sole safety net, generalises well beyond code. Ars Technica reached a compatible verdict: this is a capability-per-cost release, not an unambiguous leap.
What do Anthropic’s own benchmarks show, and how much weight do they carry?
The launch charts report leading or near-leading results across agentic coding, knowledge work, novel problem-solving, computer use and business automation. Anthropic says Opus 5 more than doubles Opus 4.8’s score on Frontier-Bench at a lower cost per task, scores roughly three times the next-best model on ARC-AGI 3, and beats Fable 5’s OSWorld 2.0 computer-use result at around a third of the cost.
Frontier-Bench deserves a word, because it is widely misdescribed. It began as Terminal-Bench 3.0 and became a standalone agent-work benchmark spanning 74 tasks across seven domains, covering databases, machine learning, formal proofs, CAD and scientific analysis as well as coding. It is not a general knowledge-work test. Anthropic’s footnote is also more informative than its headline: the result comes from an internal run on the mini-SWE-agent harness with a GKE backend, averaging five attempts per task, with Opus 4.8 serving as fallback when safety classifiers refused a request for Opus 5 and Fable 5.
None of this makes the numbers wrong. It makes them specific. A benchmark result describes a model plus its settings plus its harness, not a universal measure of judgement or factual reliability, and the public Frontier-Bench leaderboard had not published its own Opus 5 entry at the time of writing. Anthropic’s own comparisons also show the model trailing GPT-5.6 on the DeepSWE coding benchmark, which lends the rest a little credibility.
How does Claude Opus 5 compare with Fable 5, Sonnet 5 and Opus 4.8?
Eighteen months ago the line-up was easy to describe: Haiku small, Sonnet medium, Opus large. That picture no longer holds. Fable 5 sits at the top, released on 9 June 2026 alongside Mythos 5, the security-focused model that remains invitation-only. Sonnet 5 is the fast mid-tier, and Haiku 4.5, from October 2025, remains the budget option.
| Model | Anthropic’s intended role | Base price per million tokens* | Context | Max output |
|---|---|---|---|---|
| Claude Fable 5 | The hardest, longest-running autonomous work | $10 in / $50 out | 1M | 128k |
| Claude Opus 5 | Complex agentic coding and enterprise work | $5 in / $25 out | 1M | 128k |
| Claude Sonnet 5 | Speed and intelligence balance | $2 in / $10 out until 31 Aug 2026, then $3 / $15 | 1M | 128k |
| Claude Haiku 4.5 | High-volume, low-latency work | $1 in / $5 out | 200k | 64k |
*Base Anthropic API rates in US dollars, from Anthropic’s pricing documentation, checked 25 July 2026. Excludes caching, fast mode, regional premiums and cloud-provider charges. Base token price is not realised cost: Artificial Analysis measured Opus 5 at roughly 26 per cent below Fable 5 per completed task, not 50 per cent. Mythos 5 (invitation-only) omitted.

Two points follow. First, Opus 5 is not a universal upgrade. Anthropic positions it for complex agentic and enterprise work, with Sonnet and Haiku remaining the tiers for genuine volume. Routing everything through Opus because it is now cheap is how teams turn a price cut into a bigger bill. Second, subscription access is not unlimited: allowances vary by plan and are consumed faster by long contexts and high effort settings, while API customers face spend and throughput limits set at organisation and usage-tier level. Anthropic advises checking actual quotas in the Console rather than assuming a universal rate. Model a real workflow before you model a budget.
What is the wider industry saying about Claude Opus 5?
The reaction worth reading has come from the companies building on the model and the platforms carrying it, rather than from the launch-day summaries.
Developer publication The New Stack gathered the practitioner detail on the day. Cursor co-founder Sualeh Asif described Opus 5 as offering “near Fable 5 intelligence at Opus speed and cost”, sitting just under the flagship while behaving much the same way. More useful for anyone modelling a budget, Niko Grupen, head of applied research at the legal AI firm Harvey, reported performance close to Opus 4.8 while the model generated around 26 per cent fewer tokens on average. That figure needs keeping apart from the Fable 5 comparison above. One is a token reduction against the previous Opus, the other a cost gap against the flagship, and they happen to land on the same number.
Infrastructure moved immediately, which tells you something about where this model is expected to end up. Microsoft made Opus 5 available in Foundry on launch day, framing it around the problem it says enterprises actually have: “Most AI handles one task well. The harder problem is the complex workflow that spans hours.” Snowflake added it to Cortex AI the same day, so the model can work on company data without that data leaving the warehouse. Neither move says anything about output quality. Both say a good deal about how quickly your marketing data could come within reach of a model nobody on the team has evaluated yet.
The soundest procurement thinking came from B2B publication MarketScale, which skipped the benchmark race in favour of three questions: re-price your existing workloads against the new capability, measure cost per completed task rather than cost per token, and audit how often safety refusals have been quietly pushing staff into workarounds. It also relayed a claim from Box’s chief technology officer of 8 to 17 per cent improvements across enterprise content, data analysis and due-diligence workflows. Read that as a launch-partner figure, because that is what it is.
What nobody has published is as telling as what they have. Nothing so far shows that Opus 5 improves organic rankings, paid-media return, conversion rate or creative effectiveness. Those are the outcomes marketing teams are judged on, and a general capability benchmark is not a proxy for any of them. The gap will close through client work rather than through leaderboards, which is a reason to start measuring your own.
Where can Claude Opus 5 genuinely help a marketing team?
The model is days old, so what follows is informed analysis rather than measured outcome.
Research and synthesis is the strongest immediate case, and it aligns with Opus 5’s best independent result. A team can supply research reports, interview transcripts, survey data, product documentation and performance exports, then ask the model to reconcile contradictions, separate evidence from assumption and produce a decision brief. How much fits in a million tokens depends entirely on volume, so treat the window as generous rather than infinite. It does not make the model’s knowledge current either: the cut-off is May 2026, and live market claims still need retrieval and source checking. Good uses include drafting audience hypotheses, comparing competitor positioning and stress-testing a strategy against alternative explanations. Poor uses include inventing personas from demographic stereotypes or treating a fluent narrative as primary research.
Content development shifts economically rather than qualitatively. At roughly a quarter less per completed task than Fable 5, work that justified a frontier model only for flagship assets may now justify near-frontier quality across more of the programme. The highest-value role remains editorial collaboration rather than unattended volume: supply verified evidence and an explicit audience, require claim-level citations, and keep a named human editor accountable for accuracy, tone and taste. The hallucination finding makes that checking stage more important, not less.
Search work benefits from the agentic capability rather than the prose. Connected to appropriate crawling, analytics or browsing tools and authorised data sources, a model with Opus 5’s computer-use results can classify a content estate, identify gaps, build briefs, draft structured data and work through a multi-step technical audit with less supervision. None of that happens out of the box; it needs deliberate tooling and permissions, and it confers no ranking advantage. Google’s generative-AI guidance is explicit that AI can help with research and structure, but scaled content produced without added value risks breaching spam policy.
Paid media and campaign operations are promising and should stay read-only at first. Analysing search-term reports, segmenting campaign data, drafting ad variants and preparing test plans are all sensible. Letting an agent alter bids, budgets or live creative without approval is not, because a small reasoning error becomes a material spend incident the moment write access exists. Add narrowly scoped actions only once logging, permissions, spend caps and rollback procedures work.
Reporting is the best evaluation ground of all, because success is definable. Ask the model to reconcile a source export against a written summary, calculate specified metrics, flag missing data and tie every conclusion to a table or query, then compare its figures against an established report before anyone else sees them.
The commercial implications for agencies
If the model performs on real client work, agencies will deliver research packs, audits and reporting systems with fewer manual hand-offs. That does not translate neatly into fewer people; it changes where experienced people spend their time. Framing the problem, sourcing reliable data, designing the evaluation, exercising creative judgement, checking claims and owning the recommendation all become more valuable, not less.
Hourly pricing comes under pressure when a labour-intensive deliverable becomes fast. There are defensible responses: price for the outcome, use fixed fees with clear scope, or keep transparent time-based pricing and pass the efficiency on. Charging for hours automation has plainly removed, while saying nothing, is not one of them.
Staffing shifts towards hybrid roles. Strategists need enough data literacy to spot a false conclusion. Editors need to check source fidelity and hold a point of view. Someone needs to govern permissions, connectors and audit logs. Junior work will change, but stripping out junior roles altogether removes the route through which senior judgement is trained.
A mixed model policy is almost always cheaper than an Opus-only one. Use deterministic software for calculations that need no language reasoning, a small model for classification and bulk transformation, Opus 5 for ambiguous multi-source work, and Fable 5 only when the task genuinely justifies it. Route by measured need rather than prestige.
Accuracy, privacy, copyright and brand safety
The risk picture is unusually well documented, because Anthropic published a system card of nearly 200 pages on launch day and the UK AI Security Institute contributed testing.
Start with the headline that will reach your board first. In AISI cyber-range evaluations reported in the system card, Opus 5 completed an end-to-end attack on a simulated enterprise network in eight of ten attempts. The context matters as much as the number: those networks ran outdated software with reused credentials and no active defenders, so this is a capability signal rather than evidence the model defeats hardened environments. The same document shows Opus 5 deliberately held behind Mythos 5 on offensive security, with binary vulnerability scanning and exploit generation still blocked. Hugging Face’s chief executive Clem Delangue offered Fortune the counterpoint worth hearing: guardrails on closed models “flag and refuse a lot of legitimate security work, because analysing an attack looks a lot like preparing one”.
For marketing teams the risks that bite are more mundane. Accuracy comes first, and the evidence is now specific rather than general: a 50 per cent hallucination rate in Artificial Analysis’s test, up fourteen points on Opus 4.8. Every factual claim, statistic and quotation in AI-assisted client work needs human verification, and Opus 5 strengthens that rule rather than relaxing it.
Confidentiality depends on which product you use. Anthropic states that commercial inputs and outputs are not used for model training by default unless a customer opts in or submits feedback, and feedback can include a whole conversation and may be retained for up to five years. Opus 5’s exemption from the special retention requirement applied to Fable and Mythos is genuinely useful for client work, but “no special retention requirement” is not “nothing is stored”. Use an approved commercial workspace or the API, review the data-processing agreement and deployment region, disable unnecessary feedback, restrict connectors, and keep secrets and unnecessary personal data out of prompts.
UK organisations processing personal data still owe the ICO’s principles of lawfulness, fairness, transparency, minimisation, accuracy, security and accountability. The regulator’s AI guidance is under review following legislative change, which makes current legal and DPO advice more valuable than usual for higher-risk uses. Personalisation deserves particular restraint: inferring health, financial or ethnic characteristics from ordinary-looking customer records creates real legal exposure, and the model should never become an unaccountable decision-maker for offers or exclusions.
Copyright remains unsettled. The UK Government’s March 2026 report on copyright and AI describes continuing disputes over training, transparency, licensing and enforcement. The immediate practical questions are narrower: whether your inputs are licensed, and whether outputs reproduce protected material or a third party’s distinctive creative work. Keep source records and take specialist advice on consequential campaigns.
Advertising obligations do not soften because a model wrote the copy. CAP and the ASA treat their rules as media-neutral, and their guidance on AI and advertising confirms that harmful, offensive or socially irresponsible AI-generated imagery breaches the Code exactly as any other image would. Human review is a governance requirement, not a courtesy.
A 30-day evaluation plan
Do not start with “adopt Opus 5 across marketing”. Start with five representative tasks.
- Choose measurable work. One research synthesis, one content brief, one reporting task, one operational workflow, one small technical build.
- Baseline the current process. Record human time, model cost, error rate, revision count and completion time before you change anything.
- Test more than one model and effort level. Compare Opus 5 at medium and high effort against your incumbent, and include Sonnet 5 for the routine end.
- Score the deliverable, not the prose. Rubric: factual accuracy, source fidelity, completeness, brand fit, compliance, actionability, and total cost after human review.
- Limit authority. Keep trials read-only, use test accounts and redacted data, and require approval before anything publishes, spends or reaches a customer.
- Document the operating model. Named human owner, approved data classes, permitted tools, fallback model, retention settings, incident process, review frequency.
- Decide by workflow, not by vendor. Adopt where the evidence shows improvement. Keep the baseline where it does not.
The test set is the real asset here. Models change every few weeks; a stable evaluation lets you reassess without rebuilding your decision process around every launch.
What to watch next
The first independent data arrived faster than anyone expected, and the next wave matters more: fully blind evaluations with no pre-release vendor involvement, the public Frontier-Bench leaderboard’s own Opus 5 entry, cost-per-deliverable figures at different effort settings, and prompt-injection testing against connected business systems. Watch effective costs as well as headline prices, since longer reasoning and heavier tool use can change the bill for an apparently identical task. Anthropic’s cadence suggests more movement soon: Haiku 5 is conspicuously absent from a line-up where everything else has reached version 5, and Sonnet 5’s introductory pricing ends on 31 August 2026, which makes a sensible deadline for any cost modelling. Finally, watch how agencies restructure quality assurance, because if models keep getting better at execution, the scarce skill becomes deciding what to do and taking responsibility for the result.
The verdict
Claude Opus 5 is a substantial release at an aggressive price, and the launch-day evidence mostly supports Anthropic’s capability claims. The economics are more modest than the headline suggests: roughly a quarter cheaper per completed task than Fable 5 rather than half, and more expensive per task than Opus 4.8 or Sonnet 5 at maximum effort. The factual reliability finding is the one that should change behaviour, because a more knowledgeable model that answers more readily when uncertain is precisely the kind that slips errors past a tired editor.
The teams that benefit will treat this as a procurement and process question: run their own comparisons, measure cost per finished job, route work to the cheapest model that does it well, and keep qualified humans responsible for claims and sign-off. The teams that get burned will read a launch-day leaderboard and rebuild their workflow around it by Friday.
AIWIZ helps organisations evaluate AI models against their own work and build the governance around them. To discuss a model evaluation or an AI-search visibility review, contact the AIWIZ Digital Marketing team.
Frequently Asked Questions
Has Claude Opus 5 been released?
Yes. Anthropic released Claude Opus 5 on 24 July 2026 across Claude.ai, Claude Code, Claude Cowork, the Anthropic API, Amazon Bedrock, Google Cloud and Microsoft Foundry. It is the default model on Claude Max and the strongest model available on Claude Pro.
How much does Claude Opus 5 cost?
The base API price is $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8 and half Fable 5's rate. Fast mode costs twice the base rate. Realised cost is higher than the token price implies: Artificial Analysis measured $2.03 per task at maximum effort against $2.75 for Fable 5, roughly a 26 per cent saving rather than 50 per cent.
What are Claude Opus 5's effort levels?
Five: low, medium, high, xhigh and max, with high as the default. Thinking is enabled by default and the effort setting controls how deeply the model reasons on each turn. Output token use varies around eightfold between the low and max settings, so effort is the main lever on both quality and cost.
Is Claude Opus 5 better than Claude Fable 5?
Not in every respect. Anthropic still recommends Fable 5 for the longest and hardest autonomous projects. Opus 5 edged it on Artificial Analysis's Intelligence Index at maximum effort and leads on agentic knowledge work, at half the token price, but results depend heavily on settings, tools and task type.
Does Claude Opus 5 hallucinate more than Opus 4.8?
Yes, on the available evidence. Artificial Analysis found Opus 5 seven points more accurate on closed-book knowledge but measured its hallucination rate fourteen points higher, at 50 per cent, because it answers more often when uncertain. Anthropic's system card reports the same pattern. Human verification of factual claims is essential.
Should marketing teams use Claude Opus 5 for content and SEO?
It is worth testing for research synthesis, content briefs, editing, site analysis and technical SEO support. It should not be used to publish unverified claims or mass-produce low-value pages, which risks breaching Google's spam policies. Human editorial review and genuine subject expertise remain necessary.
How this was researched: figures are drawn from Anthropic’s announcement, platform documentation and system card, from Artificial Analysis’s launch-day evaluation, and from the named independent evaluators, trade publications and enterprise platforms linked throughout. Where a number originates with Anthropic, the article says so.