On 26 August Z.ai released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series: 320 billion parameters in total, 18 billion active per token, a million-token context, MIT licence, weights on Hugging Face. The same date, Alibaba's Qwen team opened Qwen3.8-Flash-Next, a 125-billion-parameter main model with 6 billion active, plus a 51-billion-parameter n-gram embedding table. Both labs describe a hybrid of linear attention and sparse attention.

GLM-5.3-Flash stacks 45 layers: 34 linear, 11 sparse, a 3:1 ratio. Qwen3.8-Flash-Next stacks 48 layers in a repeating block of three Gated DeltaNet layers plus one Qwen Sparse Attention layer — 36 to 12, the same 3:1. Z.ai uses IndexPool and Manifold-Constrained Hyper-Connections. Qwen uses Gated DeltaNet and Qwen Sparse Attention. Two names. One family of idea.

Z.ai's list price is US$0.15 per million input tokens and US$0.50 per million output. Local serving is named: SGLang, vLLM, and TokenSpeed. The model was tested anonymously as Ox Alpha on OpenRouter the week before the weights dropped. Training corpus is listed at 30 trillion multimodal tokens.

Z.ai and Qwen ship the same hybrid-attention graph
Pokiiri / Wikimedia Commons (CC BY-SA 4.0)Download

The news on 26 August is the coincidence, not a feud. Two labs. Same date. Same hybrid bet. Linear attention and sparse attention, named by both. GLM-5.3-Flash is 320 billion parameters in total, 18 billion active per token, a million-token context, MIT, on Hugging Face, the first natively multimodal model in the GLM-5 series. Qwen3.8-Flash-Next is a 125-billion-parameter main model with 6 billion active, plus a 51-billion-parameter n-gram embedding table. One lab in the GLM-5 line. One lab in the Qwen line. Both Flash. Both open weights on the same Wednesday.

The ratio is the sentence, not the nicknames. GLM-5.3-Flash stacks 45 layers: 34 linear, 11 sparse. Thirty-four to eleven is 3:1. Qwen3.8-Flash-Next stacks 48 layers in a repeating block of three Gated DeltaNet layers plus one Qwen Sparse Attention layer. Thirty-six to twelve is 3:1. Same ratio. Different layer counts. Different names on the sparse slot. Z.ai uses IndexPool and Manifold-Constrained Hyper-Connections. Qwen uses Gated DeltaNet and Qwen Sparse Attention. This file will print those names. It will not invent a mechanism behind them. IndexPool is the name Z.ai put on the card. Manifold-Constrained Hyper-Connections is the other name Z.ai put on the card. Gated DeltaNet is the name Qwen put on the repeating three. Qwen Sparse Attention is the name Qwen put on the fourth layer in the block. Two names. One family of idea. 3:1 on both graphs.

Two desks
Two desks in a dark lab, Qwen on the left, Z.ai on the right, the same graph on both monitors.

The second still is the same-day story as furniture. Two desks. Qwen on the left sign. Z.ai on the right sign. The same graph on both monitors. The still is the picture of a coincidence, not a leak of either lab's training run. The live files are GLM-5.3-Flash and Qwen3.8-Flash-Next, dated 26 August 2026. The still does not replace the parameter counts. 320 billion total, 18 billion active, on the GLM side. 125 billion on the Qwen main model, 6 billion active, plus a 51-billion-parameter n-gram embedding table. Million-token context on GLM-5.3-Flash. Native context 262,144 tokens on Qwen3.8-Flash-Next, extensible to 1 million with YaRN. Those are the numbers in this file. The still is two rooms looking at one family of graph.

Qwen's Flash-Next is an early public preview of the Qwen4 architecture. That is the Qwen sentence. SGLang posted day-0 support on 26 August, the same date as the weights. The FP8 checkpoint is about 173 GiB. That is not an eight-gigabyte afternoon. It is a public graph. MIT on the GLM side is the sentence that makes that file a tool. Vendor coding tables are not the story. The licence and the active-parameter counts are, and the fact that two open-weight Flash models arrived with the same hybrid-attention bet on the same day.

Local serving on the Z.ai side is already named: SGLang, vLLM, and TokenSpeed. Three doors. Day-0 SGLang on the Qwen side is a fourth mention of the same class of server, dated 26 August. Anyone who wants a hosted path can read Z.ai's list price: US$0.15 per million input tokens and US$0.50 per million output. Those two dollar figures are in this body because they were in this body. They stay on the page. It is not converting them into a cost model for the Qwen checkpoint. Qwen's price is not in this file, and none is added here.

Ox Alpha is the anonymous name. Z.ai's GLM-5.3-Flash was tested on OpenRouter the week before the weights dropped, under that name. The week before 26 August is the week on the page. Training corpus, Z.ai's list: 30 trillion multimodal tokens. First natively multimodal model in the GLM-5 series is the other Z.ai claim this file already carries. Multimodal in the series name, 30 trillion in the corpus line, million-token context in the window line. Keep those as Z.ai's lines. The MIT licence is how the weights sit on Hugging Face as a tool rather than as a brochure.

Forty-five and forty-eight are close enough to make a newspaper, and different enough to keep the two graphs from being one file. 45 layers on GLM-5.3-Flash, 34 of them linear, 11 of them sparse. 48 layers on Qwen3.8-Flash-Next, repeating 3 + 1, which is 36 Gated DeltaNet and 12 Qwen Sparse Attention. 34:11 and 36:12 are the same 3:1. The hybrid is linear plus sparse on both posts. That is the entire architecture printed from the 26 August cards. No extra block. No extra indexer. No extra residual trick beyond the names already on the cards: IndexPool, Manifold-Constrained Hyper-Connections, Gated DeltaNet, Qwen Sparse Attention. Naming is filing. Explaining would be inventing.

Two glass panels
Two glass panels, Qwen on one sticker, Z.ai on the other, the same network drawing on both.

The third still is two glass panels in a dark hall. Qwen on one sticker. Z.ai on the other. The same network drawing on both panes. Version labels on the stickers are the still's labels, not a substitute for GLM-5.3-Flash and Qwen3.8-Flash-Next. The news is not that two labs copied a screenshot. The news is that on 26 August both labs shipped a Flash-class open-weight model on a 3:1 linear-to-sparse stack. Hybrid linear and sparse attention. Same ratio. Same Wednesday.

Active parameters are the bill. 18 billion active per token on GLM-5.3-Flash, against 320 billion total. 6 billion active on Qwen3.8-Flash-Next, against a 125-billion-parameter main model and a 51-billion-parameter n-gram embedding table. Total is the warehouse. Active is the token. Both numbers stay on both models. Context is the other half of the warehouse. A million tokens on the GLM card. 262,144 native on the Qwen card, 1 million with YaRN. Two routes to a long window. One date.

173 GiB of FP8 on Flash-Next is why the eight-gigabyte line is a joke printed once and then dropped. A garage can read a public graph without loading 173 GiB. A shop with the iron can serve it. SGLang said day-0. Z.ai named SGLang, vLLM, and TokenSpeed for GLM-5.3-Flash. MIT is the GLM licence. The Qwen licence is not in this file. The absence stays an absence.

The date is 30 August 2026, a Sunday, four days after both drops. Neither lab has pulled the 26 August cards. The 3:1 ratio has not been restated as a different ratio. The layer counts have not moved: 45 on GLM-5.3-Flash, 48 on Qwen3.8-Flash-Next. IndexPool and Manifold-Constrained Hyper-Connections remain the Z.ai names. Gated DeltaNet and Qwen Sparse Attention remain the Qwen names. US$0.15 per million input and US$0.50 per million output remain Z.ai's list price. Ox Alpha remains the OpenRouter name from the week before. 30 trillion multimodal tokens remains the corpus line. First natively multimodal in GLM-5 remains the series line. Early public preview of the Qwen4 architecture remains the Qwen line.

A same-day story is a calendar fact. 26 August 2026. Z.ai and Alibaba's Qwen team. GLM-5.3-Flash and Qwen3.8-Flash-Next. 320B total, 18B active, million-token context, MIT, Hugging Face. 125B main, 6B active, 51B n-gram table. Hybrid linear and sparse. 34:11 and 36:12. 3:1 twice. Two names. One family of idea. Vendor coding tables are not a crown. The date, the ratio, the licences on this page, the active-parameter counts, and the public graphs are the news. That is the weather. August 26 put two Flash models on the same hybrid-attention bet. Sunday is for writing the bet down.