Ai News
Ai News

Qwen3.8-Flash-Next Opens an Early Preview of the Qwen4 Architecture

Published Aug 26, 2026 Sources checked Aug 27, 2026

Qwen has released open weights for Qwen3.8-Flash-Next, a multimodal mixture-of-experts model that previews architectural ideas planned for the Qwen4 family.

Qwen releases the next architecture preview

The Qwen team released Qwen3.8-Flash-Next on August 26, 2026 as an open-weight multimodal mixture-of-experts model and an early preview of architectural ideas intended for the future Qwen4 family. The team compares its role to Qwen3-Next, which exposed architectural changes before they appeared across later Qwen3.x releases.

The released model uses a 125-billion-parameter main network plus 51 billion parameters in an N-gram embedding component. Qwen says roughly 6 billion parameters are activated per token. The model accepts text and images, supports a native context window of 262,144 tokens, and can be extended to one million tokens with YaRN.

Four major architectural changes

Qwen3.8-Flash-Next combines Gated DeltaNet with a new Qwen Sparse Attention mechanism. Three of every four layers use Gated DeltaNet to compress historical information into a fixed-size state, while periodic global-attention layers retrieve information from longer context. Qwen Sparse Attention first compresses the sequence into micro-blocks and uses a lightweight indexer to choose the most relevant regions, reducing both attention computation and indexing overhead on long sequences.

The model also introduces Gated Residual, which widens the residual stream into four branches and applies dynamic gates to reads and writes. A separate N-gram embedding table adds local-pattern capacity with relatively little per-token matrix computation and can be offloaded to host memory with asynchronous prefetching. For optimization, Qwen uses Muon for selected two-dimensional weight matrices while retaining AdamW for embeddings, routers and other components.

Efficiency claims and long-context behavior

Qwen reports that the QSA attention kernel reaches up to 7.6x faster prefill and 4.9x faster decode at one-million-token context in its tests. In an online-serving-style setup with a 90% prefix-cache hit rate, the company reports 8.6x the prefill throughput of Qwen3.7-Plus at one million tokens. These are vendor-reported measurements and should be validated independently on target hardware and workloads.

The team also says training Qwen3.8-Flash-Next required about one ninth of the training cost of Qwen3.7-Plus while improving results on several coding, office-work and tool-use evaluations. Benchmark comparisons in the release include SWE-bench Pro, SWE-bench Multilingual, Toolathlon Verified, OSWorld 2.0 and other agentic or multimodal tasks. As with any model-release benchmark, results depend on prompts, harnesses, evaluation settings and infrastructure.

Open weights now, managed API still rolling out

The model weights are available on Hugging Face and ModelScope. QwenCloud serves the production model under the name qwen3.8-flash, with a one-million-token context target and built-in tools, although the release notes say the public API is being enabled shortly after the announcement rather than treating every managed endpoint as already universally live.

Qwen also documents compatibility patterns for OpenAI-style Responses APIs, Anthropic-compatible tooling, Claude Code, Codex, Qwen Code and other agentic developer environments. The open model repository already contains safetensor weight shards and the model configuration.

Why this release matters

The important part of Qwen3.8-Flash-Next is not only another checkpoint. Qwen is exposing a concrete architectural direction for Qwen4: sparse retrieval layered on compressed recurrent-style memory, gated residual pathways, large lookup-based local memory and a revised optimizer recipe. That makes the release useful to researchers evaluating long-context efficiency and to developers deciding whether newer sparse/hybrid architectures can lower serving cost without giving up agentic capability.

Qwen3.8-Flash-Next is a released open-weight model. Qwen4 itself remains a future model family and should not be described as released.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books