Grok 4.6 Explained: Benchmarks, Pricing, Availability
xAI launched Grok 4.6 on 12 August 2026, targeting long-running agents and matching GPT-5.6 Sol on a key index. Here's what it costs and where it runs.
xAI released Grok 4.6 on 12 August 2026, describing it as building "on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work" [1]. The model launched simultaneously in Cursor, xAI's own Grok Build tool, the API, OpenRouter, Vercel and Cloudflare, priced at $2 per million input tokens and $6 per million output tokens, with a faster variant costing double [1]. xAI has since extended Grok 4.6's reach further: the model became available in GitHub Copilot on 14 August 2026, on Amazon Bedrock on 19 August 2026, and on Google's Gemini Enterprise Agent Platform on 21 August 2026, according to xAI's own news page [2].
This is xAI's third major model release referenced in this briefing period, arriving five weeks after OpenAI's GPT-5.6 (9 July 2026) and three weeks after Anthropic's Claude Opus 5 (24 July 2026).
What xAI announced on 12 August 2026
xAI positions Grok 4.6 around sustained, multi-step work rather than single-response quality: the company says the model "stays with complex tasks across many steps," is "especially strong at turning a broad product idea into a working first version," and shows "more self-testing and verification" on longer tasks, along with "stronger first passes on visual and interactive projects" [1]. These are xAI's own characterisations rather than findings from an independent evaluator.
Benchmark scores, as reported by xAI
xAI's announcement lists a series of company-reported scores, including 61 on what it calls the AA Intelligence Index — a figure xAI says matches GPT-5.6 Sol exactly — plus 1753 on GDPVal-AA v2, 69.9% on CursorBench v3.2, 65.9% on DeepSWE v1.1, 61.3% on FrontierCode v1.1, 57.5% on APEX-Agents, 26% on Terminal-Bench v3.0, 1577 on AA-Briefcase, and 15.8% on Harvey LAB [1]. Notably, xAI's own announcement discloses that "competitor figures are drawn from the respective developers' published system cards or benchmark leaderboards" — an explicit acknowledgement that cross-vendor comparisons in the release are not independently re-tested by xAI itself, only compiled from what OpenAI, Anthropic and others have separately published [1].
Pricing and where it's available
At launch, Grok 4.6 cost $2 per million input tokens and $6 per million output tokens via the API, with a faster response variant priced at double that rate [1]. xAI also offered "2x included usage inside Grok Build and Cursor for the first week" as a launch promotion [1]. In the nine days after launch, xAI rolled the model out to GitHub Copilot (14 August), Amazon Bedrock (19 August) and Google's Gemini Enterprise Agent Platform (21 August) [2] — a distribution push across three of the largest developer and enterprise platforms outside xAI's own products.
Why this matters
The benchmark parity claim against GPT-5.6 Sol — both scoring 61 on the same index, by xAI's own account — puts Grok in direct competition with OpenAI's flagship on paper, at two-fifths of Sol's launch input price and a fifth of its output price. Just as significant is the distribution strategy: shipping into Cursor, GitHub Copilot, Amazon Bedrock and Google's own agent platform within nine days means xAI is competing for developer attention inside tools its rivals also occupy, rather than trying to pull users into a separate xAI ecosystem.
Who should care
Developers already working inside Cursor or GitHub Copilot can now select Grok 4.6 as a model option without changing tools, which lowers the switching cost of trying it against an incumbent OpenAI or Anthropic model on the same task. Teams running cost-sensitive, high-volume agent workloads should note the price difference is substantial on paper — xAI's launch pricing undercuts both GPT-5.6 Sol and Claude Opus 5's standard rates on a per-token basis, before accounting for any of the three vendors' subsequent price changes. Enterprise buyers already committed to AWS or Google Cloud now have Grok 4.6 available inside infrastructure they already use, without a direct API relationship with xAI.
Practical implications for buyers and users
Because Grok 4.6 is priced and marketed against the same benchmarks and platforms as GPT-5.6 and Opus 5, the most useful comparison for a buyer isn't the headline scores in any one company's announcement, but a task-specific test run across all three using the buyer's own workload. xAI's transparency about using competitors' self-published figures is a reasonable disclosure, but it also means none of the three-way comparisons circulating since mid-August have been produced by a neutral party. Teams already integrated with Cursor, Copilot, Bedrock or Gemini Enterprise face close to zero switching cost to trial Grok 4.6 specifically, which is not true for a net-new vendor relationship.
Limitations, availability and unresolved questions
xAI's announcement does not disclose training data sources, model size, or an independent safety evaluation alongside the performance claims. The promotional "2x included usage" offer in Grok Build and Cursor was time-limited to the first week after launch, and xAI has not published what standard included usage looks like afterwards. It is also not yet clear how long xAI intends to keep expanding Grok 4.6's platform footprint, or whether further integrations are planned beyond the four confirmed as of 21 August 2026.
Verdict
Grok 4.6 is a credible, aggressively priced entrant into the same benchmark conversation as GPT-5.6 and Claude Opus 5, and xAI's rapid multi-platform rollout is arguably a bigger story than the model itself — it puts Grok directly inside the tools developers already use rather than asking them to switch. The benchmark parity claim against GPT-5.6 Sol is worth noting but not treating as settled, since it rests on xAI's own testing and openly acknowledged use of rivals' self-reported figures for comparison. As with GPT-5.6 and Opus 5, the sensible next step for a buyer is a same-task, same-day comparison rather than trusting any one vendor's chart.
The current AI Writing & Research shortlist
Where this sits in the wider market: our current shortlist for AI Writing & Research, what each tool is best at and the main caution to check before committing.
| Tool | Best for | Current position | Important caution |
|---|---|---|---|
| ChatGPT Best all-rounder | General writing, analysis and multimodal work | GPT-5.6 combines strong reasoning with files, images, tools and broad workflow support. It is the safest starting point when one assistant must cover many jobs. | Teams should define data-handling rules and verify important claims. |
| Claude Long-form pick | Editorial work, complex documents and careful reasoning | Claude’s current Opus and Sonnet 5 family is built for sustained professional and agentic work, with a strong reputation for readable long-form output. | The highest-capability tiers can be unnecessary for routine copy. |
| Gemini Google ecosystem | Workspace users and multimodal source material | Gemini 3.7 Flash, documented in August 2026, is the current Flash release, connecting reasoning, multimodal inputs and Google’s productivity ecosystem. | Feature availability varies by Workspace plan and region. |
| Perplexity Research pick | Fast web research and cited discovery | Perplexity is useful when the first requirement is finding and comparing live web sources rather than drafting from memory. | A citation does not guarantee that the source supports every sentence; open the evidence. |
| Jasper Brand governance | Marketing teams with repeatable brand workflows | Jasper focuses on governed marketing content, brand context and campaign production rather than being a general-purpose chatbot. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Copy.ai GTM workflows | Sales and marketing process automation | Copy.ai has evolved from a copy generator into a go-to-market workflow platform for repeatable content and sales operations. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Writesonic AI visibility | SEO content and answer-engine monitoring | Writesonic combines assisted content production with tooling aimed at search and AI-answer visibility. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Grammarly Editing layer | Everyday rewriting, tone and quality control | Grammarly works best as an editing and communication layer across existing applications rather than as the only writing system. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Notion AI Knowledge workspace | Teams whose documents and projects already live in Notion | Notion AI is strongest when it can work inside an existing team knowledge base instead of requiring constant copying between tools. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| KoalaWriter SEO drafts | Structured long-form drafts and niche publishing | KoalaWriter remains a focused option for producing structured, search-aware drafts quickly. | Human research, original experience and fact-checking are still required before publishing. |
Related reading
Sources and verification notes
Primary product documentation checked for this update: