Grok 4.6 Reality Check: AA v4.3 Scores 44; Senior SWE-Bench Shows 65.3 Pass@3, 38.9 Tasteful
Grok 4.6 now scores 44 on Artificial Analysis v4.3, but the index changed materially. Snorkel’s dated Senior SWE-Bench snapshot shows 65.3% pass@3 versus 38.9% tasteful, exposing the difference between getting a patch to work and producing merge-worthy engineering.
Why Grok 4.6 deserves a second look now
SpaceXAI released Grok 4.6 on August 12, 2026 for long-running agents, coding and knowledge work. The release itself is no longer new, but two later developments materially change how its headline scores should be read: Snorkel published a dedicated Grok 4.6 Senior SWE-Bench analysis on September 3, and Artificial Analysis moved its Intelligence Index to v4.3 on September 7.
That makes this a benchmark-reconciliation story rather than another launch recap.
Primary and independent sources:
- SpaceXAI: Introducing Grok 4.6
- Artificial Analysis: Intelligence Index v4.3
- Artificial Analysis: current Grok 4.6 comparison data
- Snorkel: Grok 4.6 on Senior SWE-Bench
- Senior SWE-Bench leaderboard
Exact identity, access and price
The exact model covered here is Grok 4.6, released August 12. SpaceXAI positions it as its flagship for long-running agents and complex interactive work. The launch page says it is available through the SpaceXAI API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare; later first-party announcements added Amazon Bedrock, Gemini Enterprise Agent Platform, GitHub Copilot and Microsoft Foundry.
SpaceXAI's current first-party launch and cloud-integration pages specify a 500,000-token context window and configurable reasoning effort. Standard API pricing starts at $2 per million input tokens and $6 per million output tokens, with cached input listed at $0.50 per million on the Bedrock and Gemini Enterprise Agent Platform announcements. SpaceXAI also offers a faster variant at twice the standard launch price.
These are current first-party commercial terms, not benchmark-normalized cost estimates. Provider billing, long-context rules and product subscriptions can differ, so a developer should verify the endpoint actually being used before projecting workload cost.
The launch benchmark table was vendor-selected
SpaceXAI's August 12 launch table reported 61 on the then-current Artificial Analysis Intelligence Index, 65.9% on DeepSWE v1.1, 69.9% on CursorBench v3.2, 61.3% on FrontierCode v1.1 Extended, 57.5% on APEX-Agents, 56.4% on APEX-SWE and 26% on Terminal-Bench v3.0 for Grok 4.6 High.
Those numbers are useful, but they are not one homogeneous evaluation. SpaceXAI explicitly says competitor figures came from developers' system cards or public leaderboards. The table therefore mixes benchmark families and provenance rather than representing one identical, independently rerun harness.
This matters especially for Terminal-Bench. A score on Terminal-Bench v3.0 cannot be directly compared with Artificial Analysis's new Terminal-Bench v4.0 result as if the model simply lost accuracy.
Artificial Analysis v4.3 now scores Grok 4.6 High at 44
Artificial Analysis's current v4.3 comparison data reports 44 for Grok 4.6 High. Its component results include 66.7% on AutomationBench-AA and 21% on Terminal-Bench v4.0, alongside 56% SciCode, 43% Humanity's Last Exam and other component measurements.
The important conclusion is not that Grok 4.6 "fell from 61 to 44." The benchmark changed.
Artificial Analysis says v4.3 replaces Terminal-Bench v2.1 with Terminal-Bench v4.0 and replaces τ³-Banking with AutomationBench-AA. It runs all 66 Terminal-Bench v4.0 tasks three times and reports average pass@1. AutomationBench-AA uses a held-out 657-task set based on Zapier AutomationBench v1.0.6, spanning Finance, HR, Marketing, Operations, Sales and Support. A guardrail violation reduces that task's score to zero. Evaluations with private questions or answers now account for 45% of the Index weight, up from 40% in v4.2.
So the v4.3 score of 44 is a fresh independent measurement under a materially different test mixture. It should be compared with other models on v4.3, not treated as a raw time-series regression against the launch-era 61.
Within the same v4.3 comparison, Grok 4.6 High's 66.7% AutomationBench-AA is strong, while its 21% Terminal-Bench v4.0 is much weaker than GPT-5.6 Sol Max's 40% in the same Artificial Analysis table. That split is more informative than a single composite score: Grok looks considerably stronger on simulated business-workflow automation than on the new harder terminal benchmark under this harness.
Senior SWE-Bench: retries help a lot, but tastefulness changes the picture
Snorkel's September 3 analysis gives a second independent coding-agent view. Its Grok 4.6 snapshot covers 95 evaluated tasks out of a 100-task set, with five excluded for reward-hacking concerns.
Grok 4.6 records:
- 65.3% basic solve rate at pass@3
- 38.9% tasteful solve rate at pass@3
- 16.8% pass^3, where all three attempts must succeed
- about 23.9K output tokens per task in Snorkel's reported configuration
The distinction between pass@3 and tasteful solving is critical. Basic solving asks whether at least one of three attempts reaches runtime correctness. Senior SWE-Bench's stricter tasteful bar adds quality constraints around implementation quality, patch bloat, adherence to repository practice and related engineering criteria.
A model that gets 65.3% at the basic bar but 38.9% at the tasteful bar may be capable of finding working solutions while still producing a meaningful number of patches that senior engineers would not want to merge unchanged. The benchmark is intentionally probing that difference.
The current live Senior SWE-Bench leaderboard has continued evolving after the September 3 snapshot, so this article pins the date and evaluated set rather than silently replacing the historical result with a later leaderboard state.
SWE-bench Verified is separate
SWE-bench Verified must not be conflated with DeepSWE, Senior SWE-Bench or Terminal-Bench.
SpaceXAI's August 12 launch article does not publish a SWE-bench Verified score for Grok 4.6. Independent Vals results are widely mirrored by current coding-leaderboard trackers as 95.6% for Grok 4.6 on Vals' full SWE-bench Verified rerun, but I did not recover a stable direct Vals result page exposing that exact Grok row in this bounded review.
Vals describes itself as independently running model evaluations, and the mirrored reports consistently attribute the 95.6% figure to its neutral SWE-bench Verified evaluation. Because the exact primary result row was not directly retrievable here, this article does not upgrade 95.6% to the same evidence tier as the directly reviewed SpaceXAI, Artificial Analysis and Snorkel sources.
The practical lesson is simple: if using the 95.6% figure in procurement or research, archive the exact Vals run metadata—suite revision, 500-task coverage, mini-SWE-agent/harness version, retries, endpoint and run date—rather than copying the number without provenance.
SWE-bench Pro remains an evidence gap
I did not verify an authoritative, version-pinned SWE-bench Pro result for Grok 4.6 in the reviewed first-party xAI release material or the independent benchmark sources accepted for this article.
That absence is not a score of zero and it is not evidence that the model performs poorly on Pro. It means the field should remain unknown until a reproducible run identifies the exact Pro revision, task count, agent scaffold, tool/internet policy, retry settings, endpoint and raw result.
No DeepSWE, Senior SWE-Bench, CursorBench or Terminal-Bench score is substituted for SWE-bench Pro.
Latency: independent API measurements show a long reasoning delay
Artificial Analysis currently measures Grok 4.6 High on SpaceXAI's API at roughly 56–57 output tokens per second and around 45 seconds time to first answer token in its current provider benchmark. It lists the same $2/M input, $0.50/M cache-hit and $6/M output pricing and a blended 7:2:1 rate of about $1.35/M tokens.
Those latency figures are benchmark-workload measurements, not a universal service-level guarantee. Reasoning effort, prompt length, request load, provider path and model updates can all move the result. They are still more useful than saying the model is simply "fast" or "slow" without a measured configuration.
Community reports reinforce that variability but cannot prove its cause. In an August 25 r/cursor thread, several users complained of slow Grok 4.6 coding sessions and one described a UI task taking much longer than expected. Other users elsewhere reported positive Grok Build experiences. These are self-selected anecdotes with uncontrolled prompts, endpoint modes and load conditions, so they should not override measured provider data.
Public feedback is mixed and selection-biased
Accessible public discussion is not a controlled benchmark. An August 13 r/cursor thread highlighted Grok 4.6's apparent calibration/non-hallucination behavior from benchmark data and argued that this could matter for long agent chains. Other r/grok and r/cursor threads contain sharply negative reports about coding frustration, writing quality, product availability or latency, while some users report strong build/research experiences.
The disagreement itself is the useful signal: product mode, reasoning effort, workload and user expectations matter. None of these threads establish a population-level consensus, and their participants self-select into posting.
I did not recover a stable, directly attributable X post adding reproducible Grok 4.6 benchmark evidence during this bounded review. No X quotation or manufactured "community consensus" is inserted.
Community examples:
- r/cursor discussion, August 13, 2026
- r/cursor latency discussion, August 25, 2026
- r/grok hands-on discussion, August 31, 2026
Practical verdict
Grok 4.6 remains a serious agentic model, but its evidence is best understood as a profile rather than one rank.
The launch-era 61 Artificial Analysis score and the current v4.3 score of 44 belong to different index versions, so the delta is not a clean model regression. Under v4.3, Grok 4.6 High is notably strong on AutomationBench-AA at 66.7% but much weaker on Terminal-Bench v4.0 at 21% than several top peers. On Snorkel's September 3 Senior SWE-Bench snapshot, 65.3% pass@3 shows strong ability to reach working solutions across repeated attempts, while 38.9% tasteful pass@3 shows that engineering-quality constraints remove a large part of that apparent success.
SWE-bench Verified and SWE-bench Pro should stay separate. The widely mirrored Vals Verified number needs its exact direct run metadata pinned before being treated as top-tier primary evidence in this article, and no authoritative Grok 4.6 SWE-bench Pro run was accepted at all.
For teams evaluating Grok 4.6, the next useful test is not another composite leaderboard screenshot. It is a workload-matched A/B test that pins reasoning effort, endpoint, tool permissions, context length, retry budget, latency, token use, patch quality and total cost. That is the level of evidence needed to decide whether Grok 4.6's strong automation and retry-assisted coding results translate into a production advantage.
This article is built from the source material below. Open the originals for full context and the latest updates.