IndexShare
IndexShare lets nearby layers reuse the token positions selected by a DeepSeek Sparse Attention (DSA) indexer.
What is shared
Selected token positions
GLM-5.2 pattern
full, shared, shared, shared
Example architectures
The DSA selection being reused
The k in top-k, not to be confused with the k that is used for the keys in the DSA equation, is a hyperparameter that is set to 2048 in the DeepSeek model code. The indexer and token selector result in each token attending to a few past tokens that the model has learned to consider most relevant, rather than all tokens or a fixed local window.
The goal was not to improve performance over the dense base model but to reduce performance degradation from the sparse attention mechanism while benefiting from improved efficiency.
The four-layer reuse pattern
indexer_type = [full, shared, shared, shared]
if indexer_type == full:
selected = top_k(indexer(hidden_state))
else:
selected = previous_selected
output = attention(hidden_state, selected) # New output in every layer
full or shared across its
78-layer stack (Original source
GLM-5.2 and IndexShare for Long-Context Sparse Attention).
Evidence and tradeoff
| Item | Reported detail | What it means |
|---|---|---|
| Adjacent-layer overlap | The IndexCache study measured 70% to 100% overlap between adjacent top-k sets. |
A new indexer often rediscovers positions selected one layer earlier. |
| GLM-5.2 training | The four-layer pattern was introduced during continued mid-training with 128K-token sequences. | Retained indexers can adapt to serving the following layers. |
| GLM-5.2 result | Z.ai reports 2.9× fewer per-token FLOPs at a one-million-token context. | The reported gain is specific to GLM-5.2 and this context length. |
| IndexCache result | Removing 75% of indexer computations on a 30B DSA model produced up to a 1.82× prefill speedup and a 1.48× decode speedup at 200K. | These are separate system measurements and should not be compared directly with the GLM-5.2 FLOP result. |
| Tradeoff | A shared layer cannot select fresh positions from its current hidden states. | It uses the most recent full layer’s positions, while computing its own attention output. |
Sources