News · 2026-09-22
Alibaba's verified Apsara news is a Qwen4 architecture preview and the Zhenwu M890 chip
Alibaba's verified Apsara development is not a public Qwen4 release or a confirmed 5–10-trillion-parameter plan. It is a concrete preview of the proposed Qwen4 architecture in open Qwen3.8-Flash-Next weights, alongside T-Head's announced Zhenwu M890 accelerator with 144 GB of on-chip memory.
Key facts
- Qwen3.8-Flash-Next calls itself an open-weight early preview of architecture intended for Qwen4.
- Its repository describes a 125B main model, 51B N-gram embedding component, and about 6B parameters active per token.
- Alibaba's M890 announcement names T-Head as the designer.
- Alibaba says M890 offers 144 GB on-chip memory and 800 GB/s inter-chip bandwidth.
The first correction is editorially important. Alibaba's Apsara schedule confirms sessions about Qwen and T-Head computing, but it does not publish a keynote transcript that verifies a released Qwen4 model, a launch date, an open-weight commitment, or a 5–10T target. The official Qwen repository points the other way: Flash-Next is “an early preview of the architecture used in Qwen4,” before the complete family is built. A roadmap preview is not a shipment.
Flash-Next is technically significant on its own. Qwen describes hybrid Gated DeltaNet and Qwen Sparse Attention, four gated residual streams, N-gram embeddings, a revised Muon-based optimizer, and refitted scaling laws. Its 125B main model is joined by a 51B N-gram table, while approximately 6B parameters are active per token. That last number is a compute statement. It does not mean that a user needs storage only for six billion parameters, or that the model behaves like a dense 6B model.
The N-gram component explains part of the serving design. The repository says the table can be offloaded to host memory and prefetched asynchronously. Imagine a library where the main reader keeps the current books at the desk while a large card catalogue remains in the stacks but can be fetched ahead of time. That is a systems trade: it makes a large component available without treating it as identically hot memory, but it introduces offload and bandwidth considerations. Qwen says the design's training cost was about one ninth of Qwen3.7-Plus; that is a vendor-reported comparison, not an independently audited bill.
The hardware announcement is firmer. Alibaba identifies the chip as Zhenwu M890, not “V900,” and says T-Head formally debuted it for high-accuracy training and lower-cost inference. The company says the chip supports FP32 through FP4, uses 144 GB of on-chip memory, and supplies 800 GB/s of inter-chip bandwidth. It also says M890 delivers three times its predecessor Zhenwu 810E's performance. Those are Alibaba specifications and performance claims; they establish what shipped and what the company is promising, not independent benchmark leadership.
Memory is the connective tissue between Flash-Next and M890. Sparse models reduce active arithmetic, but weight banks, embedding tables, attention state, and inter-chip communication continue to determine whether a large model is practical to serve. An accelerator with 144 GB of local memory targets a different deployment envelope from a consumer GPU, especially for models that combine sparse activation with large look-up structures. The release belongs alongside mixture of experts, sparse attention, and offloading.
The strongest counterargument is a simple one: official detail is not independent performance proof. No reviewed Alibaba source confirms the rumored future parameter target, and the M890's three-times figure has not been independently reproduced here. Nor does the M890 announcement prove a Qwen4 release. The so-what is still meaningful. Alibaba has disclosed a concrete model-architecture direction and a memory-rich domestic accelerator in the same window, providing stronger evidence of a hardware-and-systems strategy than the unverified superlative headlines do.
It also explains why parameter-count headlines are increasingly a poor guide to deployment: the economic unit is the whole serving system—model routing, state, memory capacity, interconnect, quantization, and software—not one number in a launch graphic. Readers should expect future vendors to foreground active parameters while infrastructure operators foreground the capacity required to make those claims real.
There is a practical strategic implication too. A lab that controls both the architecture and the accelerator can co-design memory layout, numerical precision, interconnect behavior, and serving software. That can improve an operational system even when a model headline benchmark score does not move. Conversely, a chip specification alone says little about compiler maturity, availability, power use, reliability, or the model stack that customers can actually deploy. Those are the next evidence gaps to watch.
Key questions
Did Alibaba release Qwen4 at Apsara 2026?
What is Alibaba's Zhenwu M890?
Why does Qwen3.8-Flash-Next have 51B extra parameters?
Cite this
APA
Ground Truth. (2026, September 22). Alibaba's verified Apsara news is a Qwen4 architecture preview and the Zhenwu M890 chip. Ground Truth. https://groundtruth.day/news/qwen-flash-next-m890-chip-verified-apsara.html
BibTeX
@misc{groundtruth:qwen-flash-next-m890-chip-verified-apsara,
title = {Alibaba's verified Apsara news is a Qwen4 architecture preview and the Zhenwu M890 chip},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/qwen-flash-next-m890-chip-verified-apsara.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.