Tencent Hy4 Preview Reality Check: 14.7T Weekly Tokens, $0.26 Agent Arena Tasks, 65.7 Pro Is Vendor-Run
Hy4 preview led OpenRouter weekly usage at 14.7T tokens, up 379%, while Agent Arena shows a low $0.26 median task cost. Tencent's 65.7 SWE-bench Pro score remains vendor-run and the benchmark itself has documented quality issues.
Why Hy4 preview is suddenly worth a second look
Tencent released Hy4 preview on August 28, 2026. The model is not new today, but its usage signal is: OpenRouter's weekly ranking through September 6, 2026 places tencent/hy4-preview at #1 by tokens processed with 14.7 trillion tokens, up 379% from the prior weekly comparison shown by OpenRouter.
That is a remarkable adoption signal. It is not a capability benchmark. Token volume can be driven by price, free or bundled access, coding-agent integrations, long prompts, retries, provider routing and a small number of very heavy users. The right question is therefore not "does 14.7T prove Hy4 is best?" It is why developers are sending so much work to it, and what the independent evidence says about the quality/cost tradeoff.
Exact model identity and access
Tencent's official release and repository describe Hy4 preview as a 770B-total / 49B-active Mixture-of-Experts model with a 1M-token context window. Its backbone has 78 layers, 256 routed experts plus one shared expert in each MoE layer, top-8 routed experts activated per token, Gated DeepSeek Sparse Attention with IndexCache, and a native multi-token-prediction layer for speculative decoding.
Tencent publishes the weights under Apache 2.0 and provides vLLM and SGLang deployment recipes, including an FP8 path. That makes this meaningfully more inspectable and self-hostable than an API-only frontier model, but "open weights under Apache 2.0" should not be silently expanded into a claim that the complete pretraining corpus and every production training artifact are public.
Tencent also calls Hy4 an early preview and documents two known behavioral issues: it can spend longer than necessary reasoning through complex tasks and can over-verify its own work.
The freshest independent signal: Agent Arena
Arena's September 5 Agent Arena snapshot covers 2,285,256 sessions across 59 models. Hy4 preview's overall row reports:
- 6.92% ±1.51% net improvement
- 13.56% ±2.96% confirmed success
- 10.20% ±0.79% bash recovery
- 14,405 sessions
- $0.26 median cost per task
- 33.8K median output tokens per task
- $0.83/M input and $2.50/M output list prices in Arena's table
Hy4 is not the highest-capability model on this board. Claude Fable 5.1 Max has a much higher 15.87% ±2.84% net improvement, while Kimi K3 Max is also above Hy4 at 7.96%. But Hy4 sits on Arena's displayed Pareto frontier, because its measured cost per task is far below the leading proprietary models in that workload.
This is the practical story: Hy4 is showing a credible quality-per-dollar position in a real agent environment, not an overall capability victory.
Arena's $0.26 is a workload-specific median, not a guaranteed invoice for your application. Cost per successful task changes with prompt length, reasoning mode, tool loops, retries, cache behavior and whether your own task distribution resembles Arena's.
OpenRouter pricing and live availability are separate from benchmark quality
OpenRouter currently lists Hy4 preview at $0.834 per million input tokens, $2.501 per million output tokens and $0.042 per million cache-read tokens, with a 1,048,576-token context window and 64,000 maximum completion tokens. Its model page showed 99.94% availability over the latest three-day window when checked for this article.
Availability is not latency. I did not find a standardized, current independent Hy4 table that pins time-to-first-token, output tokens per second, hardware/provider route and reasoning mode in the same evaluation used for capability. Provider uptime also cannot be converted into generation speed.
Tencent's coding numbers are strong, but they remain vendor-run
Tencent's benchmark appendix reports 65.7 on SWE-bench Pro (public), 82.9 on SWE-bench Multilingual, 85.4 on Terminal-Bench 2.1, 64.3 on DeepSWE, and 92.3 on GPQA Diamond for Hy4 preview.
Those numbers are useful, but the public release does not expose enough run-level detail to treat them as an independent reproduction. For the coding-agent rows, a careful comparison still needs the exact harness or scaffold, attempted-task count, retry policy, reasoning setting, tool budget, environment version and raw trajectories.
Tencent's own human study is also clearly first-party: 163 internal experts rated 203 engineering tasks. Hy4 averaged 2.99, narrowly ahead of GLM 5.3 at 2.92 and Kimi K3 at 2.94. Tencent itself describes the result as only slightly ahead, which is the appropriate interpretation.
SWE-bench Verified and SWE-bench Pro must stay separate
SWE-bench Verified: I did not find an exact Hy4 preview result in Tencent's current public materials or a sufficiently pinned independent run that should be published as a Hy4 Verified score. It remains unknown here.
SWE-bench Pro: Tencent reports 65.7 on the public split. That is a different suite and must not be relabeled as Verified.
There is another important limitation. SWE-bench Pro's public split contains 731 tasks, but OpenAI published a July 8 audit estimating that about 30% of those tasks are broken and retracted its earlier recommendation to adopt the benchmark. OpenAI's pipeline marked 200 tasks (27.4%) as broken, while a separate five-engineer annotation campaign identified 249 (34.1%).
That does not erase Hy4's 65.7 vendor score. It means the benchmark itself now carries a material validity warning, so small cross-model gaps should not be treated as precise capability differences without task-level analysis.
A useful WebDev signal arrived before the usage surge
Arena publicly announced on August 28 that Hy4 preview had entered Code Arena WebDev at 1,633 ±17 via AutoEval, roughly fifth at that snapshot. Arena explicitly warned that the early value used a reward model trained on human-preference data to cast automatic votes and said live human votes should be allowed to converge before treating the rank as settled.
That caveat is important. A model can look strong in an automatic preference proxy and still move after real votes accumulate. The later Agent Arena evidence is more useful for this article because it reports thousands of real agent sessions, uncertainty, cost and multiple behavioral signals.
Public feedback is enthusiastic but still anecdotal
The public reaction is unusually large for an open-weight model. An August 28 LocalLLaMA release thread drew hundreds of votes and comments. Some users were excited by the jump over Hy3; others immediately focused on the hardware reality of a 770B-total model and the compromises required by very low-bit quantization.
A later commenter reported that a local one-shot run took roughly 3.5 hours on their setup while producing an impressive result. That is exactly the kind of anecdote that should not be converted into a general latency claim: the hardware, quantization, prompt, sampler and serving stack were not controlled.
Arena's own August 28 X post is a better example of measured caution: it publicized the 1,633 AutoEval score while explicitly saying that the early automatic score should converge with live human votes. There is no honest basis for calling the social reaction a consensus about Hy4's production quality.
Multimodal, safety and reproducibility gaps
The current OpenRouter listing presents Hy4 preview as a text/tool-calling model. I did not find a Hy4-specific multimodal benchmark package that should be compared with native vision/audio models in this review, so no multimodal score is inferred.
I also did not find a same-harness independent reproduction of Tencent's 65.7 SWE-bench Pro or 85.4 Terminal-Bench 2.1 result with public raw trajectories. That is the biggest evidence gap remaining.
Practical tradeoff
The best evidence-based reading is:
Hy4 preview has become a real production-usage phenomenon, not merely a launch-day benchmark chart. OpenRouter shows 14.7T weekly tokens and a 379% jump, while Arena's September 5 agent data places it at 6.92% ±1.51% net improvement with a $0.26 median task cost. That combination makes it especially interesting for high-volume coding and tool-use workloads where cost matters.
But three cautions prevent a universal "best model" claim. First, stronger models remain well ahead on Agent Arena's net-improvement metric. Second, Tencent's headline coding scores are vendor-run and do not yet have a like-for-like public reproduction. Third, the 65.7 SWE-bench Pro result inherits the benchmark's own newly documented task-quality concerns.
For teams choosing a model today, test Hy4 on your own agent harness, record cost per successful task rather than price per token alone, and measure latency on the actual provider or self-hosted hardware you intend to use. The next evidence that would materially change this assessment is a version-pinned independent SWE-bench Pro or Terminal-Bench rerun, a clean Hy4 SWE-bench Verified result if anyone still chooses to run that suite, and matched latency/cost measurements under the same agent scaffold.
This article is built from the source material below. Open the originals for full context and the latest updates.