The Gated Residuals (GR) mechanism is in spirit related to mHC, which both replace the residual connection in a transformer block with four parallel ones. In the case of GR, there are two units or modules, GR Read and GR Write. For example, before each MoE sublayer, a GR Read module compresses the 4 streams into one normal-width input. Then, afterwards, a GR Write module then adds the sublayer output back to all four streams (this uses a learned scalar gate for each stream).

This is kind of like mHC because where each sublayer still uses a normal hidden width. Nowever, Qwen 3.8-Flash-Next, which introduced this concept, does not use mHC’s separate constrained matrix for mixing the residual streams.

Qwen3.8-Flash-Next architecture showing GR Read and GR Write around the sequence and MoE sublayers
Figure 1. Qwen3.8-Flash-Next places GR Read and GR Write around both the sequence sublayer and the MoE sublayer in each transformer block. Architecture details come from the model card, released configuration, and implementation.

Residual width

Four parallel streams

Read and write gates

Element-wise read gates and one scalar write gate per stream

Example architectures

Qwen3.8-Flash-Next 125B-A6B

GR Read

Let \(R_1, \ldots, R_4\) denote the four individual residual streams. Here, GR Read normalizes each individual residual stream independently and then predicts an element-wise gate \(g_i\) for each of these streams. Next, it computes the average over these gated streams to compute the normal-width sublayer input:

\[x = \frac{1}{4}\sum_{i=1}^{4} g_i \odot \operatorname{RMSNorm}(R_i).\]

Here, the gate network reads the (concatenated) four-stream state. In Qwen3.8-Flash-Next, it projects the 10,240-dimensional state down to a 320-dim representation and back up to 10,240 element-wise gates. A sigmoid function (“gate”) keeps each Read gate in the range 0 and 1.

GR Write

After an certain module (attention module, or Gated DeltaNet, or MoE sublayer) produces its output \(y\), a GR Write module predicts one scalar gate \(s_i\) for each of the streams and applies the regular residual addition separately:

\[R_i' = R_i + s_i y.\]

The GR Write module computes the write gates as \(2\,\sigma(\cdot)\), so each value lies in the range between 0 and 2. The raw residual streams remain intact, and each receives a differently scaled copy of the same sublayer output.

Qwen applies this read-sublayer-write sequence twice in every transformer block. The first instance surrounds Gated DeltaNet or Qwen Sparse Attention, and the second surrounds the MoE sublayer. A final GR Read collapses the four streams back to one model output after the last layer.

How GR and mHC differ

Both methods widen the residual path while keeping the attention and MoE sublayers at the usual hidden width. Their stream updates differ.

  • mHC includes a residual-stream mixing matrix. The manifold constraint makes this matrix doubly stochastic.
  • GR carries each residual stream forward directly. The streams meet when GR Read forms the sublayer input, and GR Write adds the sublayer output back with one scalar gate per stream.
  • GR uses sigmoid-bounded read and write gates. It does not apply the doubly stochastic projection used by mHC.

Sources