Grok 4.6 Enters the AI Frontier at 61 on the Intelligence Index, Matching GPT-5.6 at a Fraction of the Cost
SpaceXAI's latest model gains five points in just one month, putting it alongside OpenAI's most recent release while pricing input tokens at one-third the rate of its main competitors.
Grok 4.6, released to xAI's developer API over the weekend, has climbed onto the Artificial Analysis Intelligence Index frontier. The model now scores 61, in line with GPT-5.6 Sol at its maximum configuration and just behind Anthropic's Claude Opus 5 series at 63 and Claude Fable 5 at 62 with fallback enabled. The jump of five points over Grok 4.5 came less than a month after the previous version shipped, marking one of the fastest iteration cycles at the current generation's peak.
A Faster Pace at the Top
The acceleration is not limited to raw benchmark numbers. Grok 4.6 also improves its standing against the Qwen3.8 Max family and Kimi K3, both of which landed on the frontier earlier this year. For an organization that was still answering questions about its training data transparency at this time last year, the climb has been swift.
xAI did not publish a detailed methodology paper alongside the release, instead directing developers to its existing technical documentation for context window behavior, safety steering mechanisms, and tokenization changes introduced since Grok 4. Per sources briefed on the launch who spoke with Artificial Analysis, the performance gains trace to a combination of an expanded synthetic-data pipeline and a reworked sparse-attention kernel that lets the 500,000-token context window run more efficiently on batches under 32,000 tokens.
The Intelligence Index itself aggregates thirteen separate evaluation suites, ranging from multilingual coding tasks and mathematical reasoning to long-form narrative coherence and structured knowledge retrieval. Frontier status means a model lands within one point of the single best performer on the leaderboard, regardless of which organization built that model.
The synthetic-data pipeline, according to the briefing, now incorporates distilled reasoning traces from earlier Grok generations alongside newly commissioned domain-specific corpora in finance, legal reasoning, and systems programming. The sparse-attention rewrite reduces the quadratic memory overhead of full attention by approximately forty percent on sequences between eight thousand and thirty-two thousand tokens, the range where most agentic workloads operate.
Agentic Performance Leads the Pack
Where Grok 4.6 distinguishes itself most clearly is in agentic benchmarks, the category that tests whether a model can complete multi-step tasks by calling external tools. On the GDPval-AA v2 suite, a private evaluation of real-world knowledge work, Grok 4.6 posts an Elo of 1753. That slots it behind only Claude Opus 5, while the confidence interval overlaps with Claude Fable 5 and Qwen3.8 Max.
The pattern repeats across different task types. On the banking-customer-service benchmark known as tau-cubed Banking, Grok 4.6 scores 50.7 percent, placing it in the top two alongside Qwen3.8 Max at 51.3 percent. On Terminal-Bench v2.1, which evaluates models on terminal-based software-use challenges, it reaches 88.4 percent, right alongside the leaders.
According to the Artificial Analysis report, few models currently combine competitiveness across knowledge work, customer-service tool use, and terminal operations in a single scorecard. Grok 4.6 does, and at a price point that undercuts its closest rivals by more than half.

Pricing Stays Flat at $2 per Million Tokens
Perhaps the most consequential detail of the release is that headline pricing did not change. Grok 4.5 charged $2 per million input tokens and $6 per million output tokens. Grok 4.6 keeps that exact rate, even as it delivers a five-point Intelligence Index improvement.
The comparison that matters for buyers, the report notes, is against the models scoring within two points: Claude Opus 5 at $5 and $25 per million input and output tokens respectively, and GPT-5.6 Sol at $5 and $30. Grok 4.6's output-token pricing sits at roughly one-fifth of GPT-5.6 Sol's rate on that dimension where cost typically accumulates fastest.
Measured cost per task across a representative sample of agentic workloads comes to $0.84 for Grok 4.6, the same figure that Anthropic's comparable model posted three months ago. The price-to-intelligence ratio places Grok 4.6 on the Pareto frontier for every agentic evaluation in the index.
Cache-hit pricing, which applies when a model can reuse previously processed tokens, drops to $0.50 per million tokens. That is a modest increase from Grok 4.5's $0.30 rate but still lands below every major competitor's cached-input pricing.
Context Window Holds at 500,000 Tokens
The context window remains unchanged at 500,000 tokens, the same span that Grok 4.5 introduced. xAI's rationale, according to the briefing documents, is that the current frontier of practical workloads — agentic coding assistants, multi-hour research sessions, long-document summarization — rarely exceeds two hundred thousand tokens in production. Maintaining the wider window without expanding it lets the engineering team focus compute cycles on model capacity rather than memory allocation.
The model's efficiency profile bears out that choice. Across long-horizon knowledge-work tasks measured by Artificial Analysis's private AA-Briefcase benchmark, Grok 4.6 reaches an Elo of 1577. That places it in the Fable 5-tier, behind the Claude Opus 5 family but ahead of most other frontier contenders.
More telling is the resource consumption gap. Grok 4.6 resolves identical tasks in roughly fifty-three turns and uses about five hundred million input tokens on average. Claude Opus 5 (max configuration) takes roughly one hundred three turns and two billion input tokens over the same workload set. Since long-horizon agentic work accumulates context continuously, a model that reaches comparable answers in half the turns and a quarter of the input tokens carries a cost advantage that extends well beyond per-token pricing.

What This Means for Developers
For teams evaluating model choices in August 2026, Grok 4.6 presents a clear trade-off. The model matches GPT-5.6 Sol on raw intelligence scores while charging roughly one-third the input-token price and one-fifth the output-token price. On agentic workloads specifically — terminal use, customer-service automation, multi-step research — it lands statistically even with Claude's best tier.
The catch is availability. xAI is rolling out API access through a waitlist system that currently favors enterprise customers with existing cloud commitments. Self-serve developers can request access through the xAI developer console but are seeing approval windows of two to four weeks, according to three independent developer reports collected by The Decoder this week.
At the same time, Anthropic is expected to respond to the pricing pressure with a wider context variant of Claude Opus 5, internally dubbed "Project Extend," before the end of August. OpenAI's GPT-5.6 roadmap reportedly includes a "Mini" variant priced for high-volume agentic workloads, though neither company has confirmed a timeline.
A Competitive Response Takes Shape
The Grok 4.6 release arrives as the broader AI industry consolidates around two questions that dominated the second quarter of 2026: how quickly can frontier labs ship new model generations, and how aggressively can they cut costs without sacrificing performance.
xAI's answer on both fronts is "faster and harder" than expected. The five-point Intelligence Index gain in thirty days suggests that the synthetic-data and sparse-attention techniques introduced last month have reached a point of diminishing returns on smaller workloads but continue to scale on the largest training runs.
For now, the model leads on cost-adjusted performance for agentic tasks while staying competitive on static reasoning benchmarks. Whether that advantage holds depends largely on how quickly the competition responds with its own pricing adjustments, which most industry observers expect before the calendar turns to September.
The full breakdown of individual evaluation scores across all thirteen suites is available on the Artificial Analysis leaderboard, which updated its metrics for Grok 4.6 moments after the model went live on the xAI API.
Grok 4.6 model overview and benchmark results — Artificial Analysis
Explore our ongoing coverage of AI model releases and benchmark performance in the /category/ai archive.