Claude Opus 5: What Anthropic's New Model Changes
Anthropic launched Claude Opus 5 on 24 July 2026 with new per-token pricing, a Fast Mode option and benchmark claims aimed at long-running AI agent work.
Anthropic released Claude Opus 5 on 24 July 2026, describing it as "a step change improvement for the Opus tier powering long-running agents" [1]. The company made it available the same day "on all platforms," priced at $5 per million input tokens and $25 per million output tokens, with an optional Fast Mode running at roughly 2.5 times the default speed for twice the base price [1]. Separately, Anthropic's Claude Code documentation confirms Opus 5 ships with a 1-million-token context window, and lists Fast Mode pricing there as $10/$50 per million tokens — consistent with the "twice the base price" figure from the main announcement [2].
Unlike OpenAI's GPT-5.6 launch two weeks earlier, which split into three separate models, Anthropic's update is a single new flagship replacing Opus 4.8 in the Opus tier.
What Anthropic announced on 24 July 2026
The headline claim is that Opus 5 is built specifically for agents that run for extended periods across many steps, rather than for single-turn chat quality. Anthropic frames this as a continuation of the direction it set with Claude Sonnet 5 a month earlier, which the company's Claude Code changelog describes as having "top-tier coding and tool use at Sonnet pricing" and a native 1-million-token context window [2].
Pricing and Fast Mode
At standard speed, Opus 5 costs $5 per million input tokens and $25 per million output tokens [1]. Fast Mode — used for latency-sensitive agent workflows — doubles both figures to roughly $10 and $50 respectively, matching the numbers Anthropic separately quotes in its Claude Code documentation [1][2]. Anthropic also states that Opus 5 "does not have data retention requirements for general access," language aimed at enterprise buyers concerned about prompt and output storage [1].
The benchmark claims — and their source
Anthropic's own figures, published in its Opus 5 announcement, report that the model "surpasses all other models" on an internal Frontier-Bench v0.1 benchmark and "more than doubles Opus 4.8's performance" on it at a lower cost [1]. On a coding benchmark called CursorBench 3.2, Anthropic says Opus 5 comes within 0.5% of the peak score set by a model it names Fable 5, at roughly half the cost [1]. On ARC-AGI 3, the company reports a score "three times as high" as the next-best model, and on what it calls Zapier AutomationBench, a pass rate 1.5 times higher than the next-best model [1]. On OSWorld 2.0, Anthropic says Opus 5 surpasses Fable 5's best result at just over a third of the cost [1]. In two narrower science benchmarks, Anthropic reports gains of 10.2 percentage points on an organic chemistry task and 7.7 points on protein-function prediction, though it did not name the comparison model for these two figures in the material we reviewed [1].
As with any vendor's own benchmark disclosures, these figures are Anthropic's claims about its own model, using tests the company selected. Next AI Compare has not independently reproduced them, and readers should treat repeated references to "Fable 5" and similar named comparison models as Anthropic's internal points of reference rather than independently confirmed rankings.
Safety claims alongside the capability claims
Anthropic also published safety-specific figures alongside performance ones: it describes Opus 5 as its "most aligned model to date" according to an internal automated behavioural audit, with the "lowest rates of deceptive behaviour" of any Claude model tested [1]. The same announcement states Opus 5 remains "substantially behind" a model Anthropic calls Mythos 5 specifically on cybersecurity exploit development, and that its safeguards are similar to Opus 4.8's, with "stronger guardrails on a narrow range of cyber tasks" [1]. These are self-reported audit results rather than findings from an external evaluator, and Anthropic did not publish the audit methodology in the material we reviewed.
Why this matters
Opus 5 arrives as Anthropic's clearest statement yet that it sees agentic, multi-step workflows — not conversational chat — as the main battleground for its most expensive model tier. The emphasis on Fast Mode pricing and a 1-million-token context window both point toward agents that read large amounts of context and act over long sessions, rather than answering discrete questions. Coming two weeks after GPT-5.6 and three weeks before Grok 4.6, it also confirms 2026's frontier-model race is now running on a roughly monthly cadence between the three labs.
Who should care
Teams building autonomous or semi-autonomous coding and operations agents are the most direct audience, given Anthropic's own framing. Enterprises already using Claude for high-volume workflows should note the "no data retention requirements for general access" language specifically, since data handling is often the deciding factor in procurement. Developers deciding between GPT-5.6 Sol, Grok 4.6 and Opus 5 for a new agent build now have three vendor-reported benchmark sets to weigh against each other — none of them independently verified — rather than one.
Practical implications for buyers and users
On paper, Opus 5's standard output price ($25 per million tokens) undercuts GPT-5.6 Sol's launch price ($30), but that comparison shifted once OpenAI cut Sol's pricing again on 21 August 2026 — a reminder that headline per-token prices go stale quickly across all three vendors right now, and any comparison should be checked against current published rates rather than launch-day figures. For latency-sensitive agent products, the Fast Mode surcharge (roughly double) needs to be modelled into unit economics before committing at scale. Buyers evaluating cybersecurity-adjacent use cases should note Anthropic's own admission that Opus 5 lags a named internal model on exploit development — a rare instance of a vendor flagging a capability gap rather than only claiming strengths.
Limitations, availability and unresolved questions
Anthropic's Opus 5 announcement does not state a general context window figure; the 1-million-token figure we cite here comes from Claude Code's own documentation rather than the primary Opus 5 announcement, so it may describe that product's configuration specifically rather than every access point [1][2]. Anthropic has not published its Frontier-Bench v0.1 methodology, nor detailed how "Fable 5" and "Mythos 5" — the two comparison models named repeatedly in the announcement — relate to Anthropic's public Claude lineup. No fixed retirement date for Opus 4.8 has been published.
Verdict
Claude Opus 5 is a substantial, clearly-scoped update rather than an incremental refresh, with real pricing and a stated focus on long-running agent work rather than general chat. The performance figures are extensive but entirely self-reported, and Anthropic's own safety disclosure — that the model still trails an internal comparison model on cyber-exploit development — is worth taking as seriously as the capability claims. Buyers should treat this as one of three broadly comparable July–August 2026 frontier releases (alongside GPT-5.6 and Grok 4.6) rather than an outright winner until independent benchmarking catches up.
The current AI Writing & Research shortlist
Where this sits in the wider market: our current shortlist for AI Writing & Research, what each tool is best at and the main caution to check before committing.
| Tool | Best for | Current position | Important caution |
|---|---|---|---|
| ChatGPT Best all-rounder | General writing, analysis and multimodal work | GPT-5.6 combines strong reasoning with files, images, tools and broad workflow support. It is the safest starting point when one assistant must cover many jobs. | Teams should define data-handling rules and verify important claims. |
| Claude Long-form pick | Editorial work, complex documents and careful reasoning | Claude’s current Opus and Sonnet 5 family is built for sustained professional and agentic work, with a strong reputation for readable long-form output. | The highest-capability tiers can be unnecessary for routine copy. |
| Gemini Google ecosystem | Workspace users and multimodal source material | Gemini 3.7 Flash, documented in August 2026, is the current Flash release, connecting reasoning, multimodal inputs and Google’s productivity ecosystem. | Feature availability varies by Workspace plan and region. |
| Perplexity Research pick | Fast web research and cited discovery | Perplexity is useful when the first requirement is finding and comparing live web sources rather than drafting from memory. | A citation does not guarantee that the source supports every sentence; open the evidence. |
| Jasper Brand governance | Marketing teams with repeatable brand workflows | Jasper focuses on governed marketing content, brand context and campaign production rather than being a general-purpose chatbot. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Copy.ai GTM workflows | Sales and marketing process automation | Copy.ai has evolved from a copy generator into a go-to-market workflow platform for repeatable content and sales operations. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Writesonic AI visibility | SEO content and answer-engine monitoring | Writesonic combines assisted content production with tooling aimed at search and AI-answer visibility. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Grammarly Editing layer | Everyday rewriting, tone and quality control | Grammarly works best as an editing and communication layer across existing applications rather than as the only writing system. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Notion AI Knowledge workspace | Teams whose documents and projects already live in Notion | Notion AI is strongest when it can work inside an existing team knowledge base instead of requiring constant copying between tools. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| KoalaWriter SEO drafts | Structured long-form drafts and niche publishing | KoalaWriter remains a focused option for producing structured, search-aware drafts quickly. | Human research, original experience and fact-checking are still required before publishing. |
Related reading
Sources and verification notes
Primary product documentation checked for this update: