Attention Residuals (AttnRes)
Attention Residuals are a way to improve the residual path, but they work a bit differently from other recent residual-path changes. I.e., mHC made the residual path wider. Attention Residuals (also already part of Kimi Linear) connect the residuals across layers; the connection itself uses an attention score for an importance/contribution weight.
In a regular transformer, residual connections add all earlier updates with a contribution weight of 1. Attention Residuals, or AttnRes for short, replace these fixed weights with learned ones. The Attention Residuals paper reports consistent (but modest) improvements in validation loss and downstream performance, with about 4% in training cost and 2% in inference cost.
From a fixed sum to learned weights
In a standard PreNorm transformer, the input to sublayer \(l\) is the sum of the embedding and all preceding attention and feed-forward updates:
\[h_l = \sum_{i=0}^{l-1} v_i\]The first value, \(v_0\), is the token embedding. Each later \(v_i\) is an output from an earlier sublayer. AttnRes uses the same values but learns how much each one should contribute:
\[h_l = \sum_{i=0}^{l-1} \alpha_{i \rightarrow l} \cdot v_i\]Here, the earlier outputs are the values, and their RMS-normalized versions serve as keys.
As illustrated in the figure above, we compute each weight as the dot product between the normalized key and a learned pseudo-query for the destination sublayer. We normalize these weights by applying the softmax function across model depth, which gives us the normalized weights \(\alpha\).
Finally, the last step is to compute \(h_l\), which is an attention-weighted version of the embedding and earlier sublayer outputs.
The pseudo-query is shared across tokens for a given destination sublayer. The keys remain token-dependent, so the weights can still vary by token. Zero-initialized pseudo-queries give all available outputs the same weight at the beginning of training.
Full and Block AttnRes
For regular self-attention, the weights connect positions in the input sequence. AttnRes instead mixes earlier sublayer outputs for the same token across model depth. The regular sequence layer, such as Kimi Delta Attention or MLA, remains unchanged.
Full AttnRes keeps the embedding and every earlier sublayer output. This list grows with model depth. Block AttnRes keeps regular additions inside each block and stores one representation at the boundary. For \(L\) sublayers grouped into \(N\) blocks, the storage per token goes from \(O(Ld)\) to \(O(Nd)\). The large experiments use about eight blocks.
Experiments
The full-scale comparison uses two 48B models with 3B active parameters. They have the same architecture and training setup except for the residual connections and are trained from scratch on 1.4 trillion tokens, including 1 trillion pretraining tokens and about 400 billion higher-quality mid-training tokens.
After the complete recipe, Block AttnRes matches or outperforms the regular-residual baseline on every reported downstream benchmark.
Sources