AI Writing & Research

Claude Opus 5: What Anthropic's New Model Changes

Anthropic launched Claude Opus 5 on 24 July 2026 with new per-token pricing, a Fast Mode option and benchmark claims aimed at long-running AI agent work.

Editorial noteThis news analysis is based on the linked primary sources. Performance and product claims are attributed to the announcing vendor unless the article explicitly says they were independently tested.

Anthropic released Claude Opus 5 on 24 July 2026, describing it as "a step change improvement for the Opus tier powering long-running agents" [1]. The company made it available the same day "on all platforms," priced at $5 per million input tokens and $25 per million output tokens, with an optional Fast Mode running at roughly 2.5 times the default speed for twice the base price [1]. Separately, Anthropic's Claude Code documentation confirms Opus 5 ships with a 1-million-token context window, and lists Fast Mode pricing there as $10/$50 per million tokens — consistent with the "twice the base price" figure from the main announcement [2].

Unlike OpenAI's GPT-5.6 launch two weeks earlier, which split into three separate models, Anthropic's update is a single new flagship replacing Opus 4.8 in the Opus tier.

What Anthropic announced on 24 July 2026

The headline claim is that Opus 5 is built specifically for agents that run for extended periods across many steps, rather than for single-turn chat quality. Anthropic frames this as a continuation of the direction it set with Claude Sonnet 5 a month earlier, which the company's Claude Code changelog describes as having "top-tier coding and tool use at Sonnet pricing" and a native 1-million-token context window [2].

Pricing and Fast Mode

At standard speed, Opus 5 costs $5 per million input tokens and $25 per million output tokens [1]. Fast Mode — used for latency-sensitive agent workflows — doubles both figures to roughly $10 and $50 respectively, matching the numbers Anthropic separately quotes in its Claude Code documentation [1][2]. Anthropic also states that Opus 5 "does not have data retention requirements for general access," language aimed at enterprise buyers concerned about prompt and output storage [1].

The benchmark claims — and their source

Anthropic's own figures, published in its Opus 5 announcement, report that the model "surpasses all other models" on an internal Frontier-Bench v0.1 benchmark and "more than doubles Opus 4.8's performance" on it at a lower cost [1]. On a coding benchmark called CursorBench 3.2, Anthropic says Opus 5 comes within 0.5% of the peak score set by a model it names Fable 5, at roughly half the cost [1]. On ARC-AGI 3, the company reports a score "three times as high" as the next-best model, and on what it calls Zapier AutomationBench, a pass rate 1.5 times higher than the next-best model [1]. On OSWorld 2.0, Anthropic says Opus 5 surpasses Fable 5's best result at just over a third of the cost [1]. In two narrower science benchmarks, Anthropic reports gains of 10.2 percentage points on an organic chemistry task and 7.7 points on protein-function prediction, though it did not name the comparison model for these two figures in the material we reviewed [1].

As with any vendor's own benchmark disclosures, these figures are Anthropic's claims about its own model, using tests the company selected. Next AI Compare has not independently reproduced them, and readers should treat repeated references to "Fable 5" and similar named comparison models as Anthropic's internal points of reference rather than independently confirmed rankings.

Safety claims alongside the capability claims

Anthropic also published safety-specific figures alongside performance ones: it describes Opus 5 as its "most aligned model to date" according to an internal automated behavioural audit, with the "lowest rates of deceptive behaviour" of any Claude model tested [1]. The same announcement states Opus 5 remains "substantially behind" a model Anthropic calls Mythos 5 specifically on cybersecurity exploit development, and that its safeguards are similar to Opus 4.8's, with "stronger guardrails on a narrow range of cyber tasks" [1]. These are self-reported audit results rather than findings from an external evaluator, and Anthropic did not publish the audit methodology in the material we reviewed.

Why this matters

Opus 5 arrives as Anthropic's clearest statement yet that it sees agentic, multi-step workflows — not conversational chat — as the main battleground for its most expensive model tier. The emphasis on Fast Mode pricing and a 1-million-token context window both point toward agents that read large amounts of context and act over long sessions, rather than answering discrete questions. Coming two weeks after GPT-5.6 and three weeks before Grok 4.6, it also confirms 2026's frontier-model race is now running on a roughly monthly cadence between the three labs.

Who should care

Teams building autonomous or semi-autonomous coding and operations agents are the most direct audience, given Anthropic's own framing. Enterprises already using Claude for high-volume workflows should note the "no data retention requirements for general access" language specifically, since data handling is often the deciding factor in procurement. Developers deciding between GPT-5.6 Sol, Grok 4.6 and Opus 5 for a new agent build now have three vendor-reported benchmark sets to weigh against each other — none of them independently verified — rather than one.

Practical implications for buyers and users

On paper, Opus 5's standard output price ($25 per million tokens) undercuts GPT-5.6 Sol's launch price ($30), but that comparison shifted once OpenAI cut Sol's pricing again on 21 August 2026 — a reminder that headline per-token prices go stale quickly across all three vendors right now, and any comparison should be checked against current published rates rather than launch-day figures. For latency-sensitive agent products, the Fast Mode surcharge (roughly double) needs to be modelled into unit economics before committing at scale. Buyers evaluating cybersecurity-adjacent use cases should note Anthropic's own admission that Opus 5 lags a named internal model on exploit development — a rare instance of a vendor flagging a capability gap rather than only claiming strengths.

Limitations, availability and unresolved questions

Anthropic's Opus 5 announcement does not state a general context window figure; the 1-million-token figure we cite here comes from Claude Code's own documentation rather than the primary Opus 5 announcement, so it may describe that product's configuration specifically rather than every access point [1][2]. Anthropic has not published its Frontier-Bench v0.1 methodology, nor detailed how "Fable 5" and "Mythos 5" — the two comparison models named repeatedly in the announcement — relate to Anthropic's public Claude lineup. No fixed retirement date for Opus 4.8 has been published.

Verdict

Claude Opus 5 is a substantial, clearly-scoped update rather than an incremental refresh, with real pricing and a stated focus on long-running agent work rather than general chat. The performance figures are extensive but entirely self-reported, and Anthropic's own safety disclosure — that the model still trails an internal comparison model on cyber-exploit development — is worth taking as seriously as the capability claims. Buyers should treat this as one of three broadly comparable July–August 2026 frontier releases (alongside GPT-5.6 and Grok 4.6) rather than an outright winner until independent benchmarking catches up.

The current AI Writing & Research shortlist

Where this sits in the wider market: our current shortlist for AI Writing & Research, what each tool is best at and the main caution to check before committing.

ToolBest forCurrent positionImportant caution
ChatGPT
Best all-rounder
General writing, analysis and multimodal workGPT-5.6 combines strong reasoning with files, images, tools and broad workflow support. It is the safest starting point when one assistant must cover many jobs.Teams should define data-handling rules and verify important claims.
Claude
Long-form pick
Editorial work, complex documents and careful reasoningClaude’s current Opus and Sonnet 5 family is built for sustained professional and agentic work, with a strong reputation for readable long-form output.The highest-capability tiers can be unnecessary for routine copy.
Gemini
Google ecosystem
Workspace users and multimodal source materialGemini 3.7 Flash, documented in August 2026, is the current Flash release, connecting reasoning, multimodal inputs and Google’s productivity ecosystem.Feature availability varies by Workspace plan and region.
Perplexity
Research pick
Fast web research and cited discoveryPerplexity is useful when the first requirement is finding and comparing live web sources rather than drafting from memory.A citation does not guarantee that the source supports every sentence; open the evidence.
Jasper
Brand governance
Marketing teams with repeatable brand workflowsJasper focuses on governed marketing content, brand context and campaign production rather than being a general-purpose chatbot.Plans, limits and model availability change frequently; confirm the current vendor page before purchasing.
Copy.ai
GTM workflows
Sales and marketing process automationCopy.ai has evolved from a copy generator into a go-to-market workflow platform for repeatable content and sales operations.Plans, limits and model availability change frequently; confirm the current vendor page before purchasing.
Writesonic
AI visibility
SEO content and answer-engine monitoringWritesonic combines assisted content production with tooling aimed at search and AI-answer visibility.Plans, limits and model availability change frequently; confirm the current vendor page before purchasing.
Grammarly
Editing layer
Everyday rewriting, tone and quality controlGrammarly works best as an editing and communication layer across existing applications rather than as the only writing system.Plans, limits and model availability change frequently; confirm the current vendor page before purchasing.
Notion AI
Knowledge workspace
Teams whose documents and projects already live in NotionNotion AI is strongest when it can work inside an existing team knowledge base instead of requiring constant copying between tools.Plans, limits and model availability change frequently; confirm the current vendor page before purchasing.
KoalaWriter
SEO drafts
Structured long-form drafts and niche publishingKoalaWriter remains a focused option for producing structured, search-aware drafts quickly.Human research, original experience and fact-checking are still required before publishing.

Sources and verification notes

Primary product documentation checked for this update: