Review date: 2026-09-19
Author: Zhongzhu Zhou
Paper reviewed: Attention Residuals
Paper authors: Kimi Team; Guangyu Chen, Yu Zhang, Jianlin Su and collaborators; the complete contribution list is in Appendix A.
arXiv: 2603.15031v1, submitted 2026-03-16
this review uses the 21-page v1, including Appendices A and B.
1. What changes when a residual becomes a retrieval operation?
I read Attention Residuals as a proposal about the interface between layers. A conventional residual stream carries an accumulated vector forward. Each attention or feed-forward sublayer adds a new contribution, and subsequent sublayers receive the sum. AttnRes instead preserves earlier outputs as individually addressable sources and learns how much of each source the next sublayer should receive. The selection varies with the token, even though the query is a learned parameter that stays fixed after training.
That last detail connects the mathematical design to the systems design. A query produced from the current hidden state would be more expressive, but it could not be evaluated before that state existed. A fixed query can score all already completed sources in advance. The paper uses this independence to batch retrieval for several future layers, then merges those results with the evolving local state. The architecture and the execution schedule are therefore closely related.
The strongest result is not simply that another attention operator improves a benchmark. The paper provides a route from full access to every earlier output, through compressed block histories, to a distributed execution plan. The main 48B-total, 3B-activated model uses nine residual blocks across 54 sublayers and an additional embedding source. Its comparison with the baseline follows the same 1.4T-token recipe. Table 3 shows improvements on fourteen of fifteen benchmarks, with MMLU-Pro tied at the displayed precision.
Three distinctions guide this review. First, source selection is token dependent, but it happens over depth rather than over sequence positions. Second, the block architecture approximates the full architecture; the two-phase schedule is an exact rearrangement of the chosen architecture in real arithmetic. Third, lower validation loss at a fixed training-compute budget is not a measured reduction in elapsed training time. Keeping these distinctions explicit makes the paper more useful for architecture decisions.

2. Prerequisites: residual streams, normalization, and two different axes
Let be the hidden width, the number of tokens, and the batch size. A hidden-state tensor has shape . I write most equations for one token, so each source is a vector in . The same mixing rule is independently applied to each token, while the underlying attention sublayer can still mix information across sequence positions.
The paper counts an attention sublayer and an MLP sublayer as two separate layers. I use for this sublayer count and for the number of ordinary Transformer blocks. This matters when comparing the 16-block ablation model, which has 32 sublayers, with a residual block size . That setting has eight AttnRes blocks, not four.
For standard residuals, let enter sublayer , with its transformation denoted by . Expanding the recurrence gives
The expansion reveals a second function of a residual connection beyond its role in optimization: it defines a depth-aggregation rule. Every previous output has coefficient one. A later sublayer can transform the accumulated state, but it cannot directly request one earlier output from the sum unless that information remains recoverable in the representation.
The familiar gradient argument is also worth retaining. Writing , the Jacobian from one depth to another includes
Its algebraic expansion contains an identity term. That is a path, not a guarantee that every singular value of the full Jacobian stays near one: other terms can amplify or cancel. Similarly, replacing the residual sum with attention changes the gradient structure; it does not automatically preserve a unit-weight identity path to every source.
PreNorm places normalization before a sublayer transformation, while the accumulated stream itself can grow. A useful simplified normalization is
where is a learned scale and prevents division by zero. This operator controls the scale presented to a transformation. It does not by itself constrain the norm of an unnormalized sum elsewhere in the network.
I treat the paper’s depth-growth discussion as a motivation rather than a universal growth law. If unit vectors align, their sum has norm ; if they are orthogonal, it has norm . Cancellation can reduce it further. Actual growth depends on learned correlations and magnitudes. The empirical curves in paper Figure 5 are stronger evidence for this model family than asserting linear growth for every possible residual network.

3. Full AttnRes: derive the mixture before interpreting it
Define and for . The input to sublayer is reconstructed from . Each destination has a learned pseudo-query . The score and weight are
followed by
There is no extra factor in the paper’s displayed scoring definition. The learned query can absorb a scale, but that does not justify silently changing the stated rule when explaining it. Keys are normalized; values are the original outputs. Consequently, a large output does not get a larger logit solely because of its norm, yet its magnitude still affects the resulting weighted sum.
Why is the mixture input dependent if the query is fixed? Because each key depends on the input token and all preceding computation. Two tokens can produce different , hence different logits and different mixture weights under the same . What is deliberately removed is the dependence of the query on the current layer’s not-yet-computed state.
Softmax creates competition: increasing one logit reduces the normalized weights of other sources. For the illustrative logits , the weights are approximately . Adding the same constant to all three logits changes nothing because the common exponential factor cancels. Multiplying all logits by a positive constant does change their sharpness. Query norm and key normalization therefore participate in the selectivity of the mechanism.
The initialization is important. The paper initializes every pseudo-query to zero. Initially all logits are zero and
This is a uniform average, not the original unit-weight residual sum. PreNorm can reduce the practical effect of a common scale under idealized assumptions, but normalization includes , learned scales, and downstream nonlinearities. I would not describe zero initialization as exact functional preservation of a pretrained standard-residual checkpoint. The reported experiment trains the modified architecture.
Because the weights are nonnegative and sum to one, the mixed input satisfies the convexity bound
The first inequality is the triangle inequality; the second uses normalized nonnegative weights. This controls the aggregation relative to its sources. It does not bound the sources themselves or every network Jacobian. That narrower claim is sufficient to explain why replacing repeated unweighted accumulation can change the observed magnitude profile.
Algorithm 1: Full depth retrieval for a single token.
- Initialize the source list with the token embedding .
- At sublayer , normalize every available source to form keys; keep original sources as values.
- Score those keys with . Subtract the maximum score before exponentiating.
- Divide exponentials by their sum and form the weighted value sum .
- Evaluate the sublayer transformation and append its output as the next source.
- Repeat in depth order; use the corresponding final aggregation for the model output.
This algorithm exposes the expense: there are source visits. Ignoring batch and token factors, arithmetic is and stored source vectors require . Small parameter overhead should not be confused with small activation traffic.
4. Block AttnRes: choose what information to discard
Partition the sublayers into consecutive blocks of size , assuming divisibility for now. For block , define its completed summary and its partial summary after sublayers as
The embedding remains a separate source . The first sublayer in block attends to . Subsequent sublayers also attend to the one evolving partial sum . The partial sum is absent before the first transformation; adding a zero-valued placeholder to softmax would add probability mass and change the result.
The important algebra is what happens when a block summary receives weight :
Every output in the completed block receives the same effective coefficient. Full AttnRes can emphasize one output and suppress another. Block AttnRes cannot do that after they have been summed. Opposing contributions can cancel, and no later choice of block weight can recover the canceled components. The block approximation is therefore a lossy representation decision, not merely a faster way to calculate the full model.
Algorithm 2: Block retrieval and summary construction.
- Set the completed-history list to .
- At the beginning of each block, initialize the partial sum and the local sublayer counter .
- Use completed history as the source set; include only when .
- Normalize source keys, score them using the destination’s pseudo-query, and softmax over this entire source set.
- Form the weighted input and evaluate .
- Update , then advance the local counter.
- At the boundary, append to history and start a new empty partial sum. If the final block is shorter, store its actual final sum.
- After all sublayers, aggregate completed sources for the final output using the model’s output query.

The architectural history becomes per token, plus local workspace. Each of sublayers consults at most roughly sources, so a useful general arithmetic bound is . The paper sometimes describes this as while its related-work discussion gives . These agree up to a fixed factor only when is held constant. When block size is a variable, retaining both and avoids hiding its cost.
There are two revealing endpoints. If , every completed block contains exactly one output and the full architecture is recovered. If , all previous outputs share one accumulating block, while the embedding remains separate. This is residual-like accumulation inside a block, but the displayed attention formula still mixes the embedding with the partial sum using normalized weights. It is not generally identical to the standard unweighted sum. I interpret the paper’s statement about recovering residuals at as a structural analogy unless extra scaling conditions are supplied.
For a concrete information-loss example, consider two outputs and . Their block sum is . Full retrieval can assign weights and , producing before considering other sources. A single scalar times the block sum cannot produce a nonzero first component. Normalizing keys does not restore the missing direction. This example isolates the approximation without making any claim about how often cancellation occurs in trained models.
5. Exact two-phase attention: keep the denominator
At the start of a block, completed history is known and all pseudo-queries are already known. Phase 1 scores that history for all future destinations in parallel. The partial sum is not yet known, so phase 2 processes it sequentially. To combine the two results correctly, one must preserve softmax normalization statistics.
For any source subset , define a stable representation
Here is an unnormalized weighted numerator and is a scaled denominator. The normalized answer is . I use this notation consistently because a log-sum-exp value is a different quantity: . The paper’s Algorithm 1 uses a division by and exponential rescaling; that algebra requires the scaled-sum interpretation even though its annotation mentions LSE.
For disjoint subsets and , choose . Multiplying each numerator and denominator by its missing scale gives
To derive this, expand the first term: . The second term gives the corresponding sum over . Their union therefore has exactly the numerator and denominator of a single softmax over all sources. No approximation enters this identity. Floating-point reduction order can still cause small numerical differences.
An example shows why averaging normalized partial results fails. Suppose has logits and with scalar values and . Its unscaled denominator is , its numerator is , and its normalized output is . A new source has logit and value , giving denominator and numerator . The union output is . Equal averaging happens to work here only because both groups have the same total exponential mass. If the new logit were , the union would instead be , while an equal average would incorrectly remain .
Algorithm 3: Two-phase execution of one Block AttnRes block.
- Stack its destination pseudo-queries into .
- Normalize completed source keys and calculate all historical attention statistics in one batched operation.
- Initialize the partial sum .
- For the first local sublayer, use ; do not create a fictitious local source.
- For each later sublayer, compute the score of . Its singleton statistics are .
- Merge historical and singleton statistics with the stable formulas above, then divide numerator by denominator.
- Evaluate and add the output into before the next sublayer.
- Store the completed when the block ends.
The completed-history work can be batched, but the partial-sum loop remains sequential. A paper discussion of parallel queries should therefore not be read as parallel execution of all nonlinear layers. The speed benefit comes from reusing memory reads and improving the shape of computation, while respecting the same causal dependency in depth.
6. Distributed training: count transfers, then ask about overlap
Full AttnRes can reuse activations already retained in simple backpropagation. Activation checkpointing changes that comparison: a conventional model may discard those activations and recompute them later, whereas later depth retrieval still needs the sources. Pipeline parallelism adds another constraint because a source created on an earlier stage may be needed on later stages. The paper’s systems work addresses this large-scale regime.
Let be the number of physical pipeline ranks, the number of virtual stages per rank, and the total logical chunks. Use the paper’s simplified average for new block representations per chunk in the transfer model. If transition sends all accumulated blocks, summing transition costs gives
The factor and element size must be restored when calculating actual bytes. Blocks need not align perfectly with pipeline boundaries; is an averaged accounting device rather than a claim that every stage creates an identical integer number of blocks.
With physical-rank caching, a rank retains history from its earlier virtual-stage visit. The first traversal still accumulates history normally. Later traversals only transfer representations created since the receiver’s previous visit. The paper obtains
For , the two simplified totals are and . The saved six units agree with the paper’s illustrative reduction. They imply a reduction of about 21.4% in this transfer model, not an eightfold end-to-end speedup. At , the formulas coincide because there has been no earlier virtual visit whose cached data can be reused.

The paper also argues that peak transition traffic changes from dependence on to dependence on . This can make communication overlap with computation feasible. Whether it actually disappears from the critical path depends on network bandwidth, message latency, microbatch size, scheduling, and the duration of each compute segment. The report gives less than 4% measured training overhead under pipeline parallelism. I treat that as a result for its setup, not as a hardware-independent upper bound.
Cache identity must conceptually include the microbatch and sequence segment; a source from another microbatch is not reusable merely because it has the same block index. Similarly, backward propagation must account for contributions through reused history. These are consequences of the algorithm’s dependency graph, not claims about a particular software implementation. The useful design principle is to remove repeated transfers of the same mathematical object without confusing object reuse with recomputation or cross-example reuse.
7. Inference: distinguish source storage, traffic, and latency
The paper’s Table 1 compares memory access associated with the residual mechanism only. It excludes the internal traffic of attention and MLP transformations. Its representative setting is , with four streams for mHC. Standard residual merging costs elements, Block AttnRes , scheduled Full AttnRes , and mHC approximately , omitting small scalar terms in the last figure.

Appendix B explains the Full AttnRes result. This use of blocks is purely an execution schedule: individual layer outputs remain available. In scheduling block , there are historical sources, and keys plus values cost reads. Summing over blocks gives
Within each block, the local sequential reads sum to
Across blocks, divide by . The amortized reads become . Add two output writes per sublayer to obtain . With , that is .
The resulting expression has an instructive optimum. If the schedule can choose real-valued temporarily, its factor is . Differentiating gives , so the continuous minimum is near . Actual choices must respect integer block sizes, hardware tiles, memory capacity, and sequential latency. This is a scheduling observation, not a prescription for the Block AttnRes architecture, where changing block size also changes the learned function.
Block AttnRes replaces historical individual sources with summaries. The paper’s accounting becomes . At the illustrative setting, . Compared with standard residual traffic, this is about times as much residual I/O, yet the paper reports under 2% end-to-end inference overhead on typical workloads. There is no contradiction: residual traffic is only a fraction of total work, and batching, fusion, and overlap can reduce its wall-clock visibility. Conversely, the comparison must not be presented as a measured sixfold speedup over mHC.
Prefill creates a different memory issue. Keeping summaries for tokens requires elements, or bytes for an element size of . Sequence-sharding those summaries across devices reduces the per-device term to . The paper’s rounded 128K-context example goes from 15 GB to about 1.9 GB with eight-way sharding. Restricting active work to a 16K chunk gives another factor of eight, approximately 0.234 GB for this isolated term.
That calculation does not include weights, sequence-attention KV caches, recurrent states, communication buffers, or peak temporary workspace. It is also not a claim that total prefill memory falls by 64 times. The depth history is conceptually different from the sequence KV cache: the former retrieves representations at one token across layers, while the latter allows an attention layer to consult preceding tokens. A hybrid KDA/MLA model can contain both kinds of state.
8. What the scaling and downstream results establish
Table 2 contains five activated non-embedding model sizes: 194M, 241M, 296M, 436M, and 528M. Their token budgets range from 38.7B to 119.0B. At each size, the compared methods share a training configuration, use an 8192-token context, and follow a cosine learning-rate schedule. The configurations across sizes differ, so the plotted line is a family of matched comparisons rather than a fixed-token parameter sweep.

At 528M, baseline loss is 1.719, Block AttnRes 1.693, and Full AttnRes 1.692. The block variant therefore recovers of the loss reduction at this particular point. At 436M, it recovers . “Most of the gain” is a reasonable local description, but the fraction varies with scale; it is not a single universal compression-efficiency constant.
The fitted curves are reported as
Here is compute in PFLOP/s-days. To compare compute at a target loss , invert a fitted curve:
This is how to interpret the reported approximately 1.25-times advantage. The small exponent makes the inferred ratio sensitive to rounded fit coefficients and target loss. It is an interpolation-based efficiency statement, without a disclosed confidence interval, rather than proof of the same advantage at arbitrary scale or on a different accelerator.
The large-model experiment is separate. It uses Kimi Linear with 27 Transformer blocks, 48B total and 3B activated parameters, eight routed experts selected from 256 plus one shared expert, and a 3:1 interleaving of KDA and MLA. Six sublayers per AttnRes block produce nine summaries plus the embedding. The main recipe uses 4096-token context, Muon, an 8M-token global batch, a 1T-token warmup-stable-decay pretraining phase, and roughly 400B high-quality mid-training tokens. The report also describes extension to 32K context.

The clearest gains include GPQA-Diamond at +7.5 percentage points, Math at +3.6, HumanEval at +3.1, and C-Eval at +2.9. MMLU improves by 1.1 and HellaSwag by 0.2. The pattern is compatible with a benefit to compositional processing, but it does not isolate a causal mechanism: the architecture changes optimization, representation scale, and access to earlier features together. Benchmark improvements alone cannot assign the gains to just one of these effects.
The training curves show lower validation loss, reduced depth-dependent output growth, and more even gradient magnitudes. Those observations reinforce the motivation. They are not substitutes for multiple random seeds, controlled causal interventions, or uncertainty estimates. In particular, a large improvement on a smaller benchmark should prompt attention to sampling and evaluation protocol before being treated as an exact estimate of general reasoning improvement.
9. Design choices: what was bought, and what was given up?
The ablations are unusually useful because the most expressive option is not the default. Full AttnRes with an input-dependent query reaches loss 1.731, compared with 1.737 for the fixed-query version. The paper chooses fixed queries to avoid a projection per sublayer and to retain advance batching over completed history. The choice therefore sacrifices some demonstrated modeling quality for a better execution structure.

Static mixing gives 1.749, compared with 1.737 for Full AttnRes. This supports content-dependent source selection under the tested recipe. It does not establish that every static-mixture design must fail; DenseFormer and static coefficients have their own optimization choices, and only the listed configurations were tested.
Replacing softmax with sigmoid gives 1.741. Competitive normalization is a plausible explanation for the difference: independently activated sigmoid gates need not sum to one, so both scale and selection change. A stronger causal comparison would normalize sigmoid gates or explicitly match mixture norms, separating competition from overall magnitude control.
Removing key RMSNorm changes Full loss to 1.743 and Block loss to 1.750. This is consistent with the need to prevent source norm from determining retrieval scores, especially when complete blocks and partial blocks contain different numbers of outputs. However, the values remain unnormalized, so key normalization cannot remove all magnitude effects from the downstream signal.
Multihead block mixing reaches 1.752 versus 1.746 for the single-head block variant. The paper interprets this as support for whole-vector relevance. I regard that as one hypothesis. Different head widths, optimization sensitivity, or reduced cross-channel coordination could also explain the degradation. One multihead setting is not enough to establish that token features universally share the same preferred depth source.
The sliding-window variant keeps the embedding plus eight recent outputs and achieves 1.764, close to the 1.766 baseline. Block summaries with achieve 1.746. At roughly comparable source budgets, retaining coarse access to distant history appears more valuable than retaining only fine-grained recent history. This is a particularly relevant comparison for memory-limited designs.

The block-size sweep gives losses 1.757, 1.753, 1.748, 1.746, and 1.746 for , while Full at reaches 1.737. The plateau followed by an additional full-access gain is worth investigating: a simple monotone approximation story does not explain why and tie at reported precision. Optimization, interactions with the embedding source, or the particular sequence of layer types may matter.
Finally, the architecture sweep holds about FLOPs and active parameters fixed. The best tested width-to-depth ratio moves from about 60 to 45 under AttnRes. This suggests that the method can make extra depth useful in that regime. Deeper models also lengthen the sequential path at inference. A training-loss optimum is therefore an input to deployment design, not the final deployment objective.
10. A closer mathematical reading of gradients and matrix structure
The attention weights are helpful for visualization, but they are not the complete derivative of a layer input with respect to a source. Treat other sources as fixed and write . For the destination under consideration,
Differentiate with respect to the scalar score :
The source also directly enters the value sum, while its key changes the logit. Applying the chain rule yields
This is a local partial derivative; the whole network additionally contains paths through later source construction. Even this local expression has a value term and a routing term. Thus, observing does not imply a network gradient norm bounded by one, and a low attention weight does not by itself prove causal irrelevance. Normalization derivatives and query magnitude also matter.
For intuition, omit and learned channel scales temporarily. Let . Differentiating gives
The radial direction is removed, while tangential directions scale by . This explains both why normalized keys ignore pure positive rescaling and why very small source norms require care. Actual RMSNorm retains , so the exact derivative differs and remains regularized around zero. The displayed simplified derivative is an explanatory calculation, not a substitute for the full model definition.
The structured-matrix view needs similar care. The unrolled standard residual coefficients form a lower-triangular all-ones matrix with nonzero diagonal. Its ordinary matrix rank is , despite its simple recurrent description. Its off-diagonal blocks across a depth cut have rank one. That latter property is the relevant semiseparable structure: a small recurrent state restricts how information can cross a cut through the computation.
A dense attention coefficient matrix can have richer off-diagonal structure, but saying that it has ordinary rank does not distinguish it from the standard residual matrix. Moreover, the coefficients are input dependent, and the source vectors themselves are constructed recursively. This is not one fixed linear operator describing the entire nonlinear network. I read section 6.2 as a useful representational analogy, while keeping ordinary rank, semiseparable rank, and the paper’s informal effective-rank discussion separate.
11. Limitations and experimental information still needed
The main evaluation establishes a promising result on a specific hybrid MoE family. It does not establish the same gains for dense decoders, every optimizer, post-training objectives, or arbitrarily deep models. The large model uses nine blocks rather than the approximate eight-block default of the scaling study. Hardware, context length, and training recipe belong to the result rather than being incidental details.
The benchmark table does not supply seed variation or confidence intervals. It also follows evaluation procedures from the Kimi Linear report rather than restating every decoding and prompting choice. A careful comparison should preserve those settings, benchmark versions, and scoring procedures. Where they are not specified in this report, the correct conclusion is that this document alone is insufficient to reconstruct the entire experiment.
The report’s latency claims lack a comprehensive workload matrix in the presented tables. Batch size, sequence length, precision, hardware, parallel layout, and the distinction between prefill and decode all affect the fraction of time consumed by residual operations. The under-2% inference figure is encouraging, but insufficient to predict a different production service’s tail latency or throughput.
Block compression still loses individual outputs. A fixed source budget trades representational detail for manageable state. Larger models can also have different distributions of source usefulness, so the success of roughly eight blocks should be tested rather than assumed. Learned boundaries or uneven block sizes could help, but would introduce their own scheduling and optimization costs.
The final limitation concerns interpretation. Token-averaged attention heatmaps hide variability across requests, tasks, and positions. Persistent embedding attention can indicate useful access to original features, a routing preference, or an attention sink. A heatmap alone cannot decide among these explanations. Intervening on a source and measuring the resulting loss is a stronger diagnostic than labeling large weights as the mechanism of improved reasoning.
12. Independent critical analysis and concrete next experiments
Separate normalized averaging from content selection. Zero-query initialization already changes the original sum into an average. A useful controlled study would compare standard residuals, a scale-matched fixed average, learned static mixtures, and content-dependent mixtures, with the same normalization placement and source storage. Matching the number of trained tokens is necessary, but matching optimization opportunities and scale is also relevant. This experiment could identify how much improvement comes from retrieval rather than a better-behaved magnitude trajectory.
Audit mathematical claims through definitions, not slogans. The endpoint, the shorthand, and the use of rank in section 6.2 each require qualifications. These are presentation issues that affect how readers generalize the method. I would state the actual source sets, retain the arithmetic expression, and define an off-diagonal rank measure before making capacity comparisons. This makes the theoretical picture testable without diminishing the empirical results.
Measure the block bottleneck directly. Full-versus-block validation loss is an aggregate outcome. A more informative diagnostic would measure cancellation within blocks, task-conditioned source importance, and the damage from replacing a selected block with zero or a matched-norm alternative. Compare fixed, layer-type-aligned, and learned boundaries under the same number of stored source vectors. If boundaries aligned with attention/MLP specialization recover a substantial fraction of the full-model gap, the result would suggest a practical improvement without increasing history capacity.
Evaluate a latency-quality frontier. The dynamic-query ablation has the lowest listed loss but worse scheduling properties. Rather than treating it as an abandoned variant, measure it alongside fixed-query Full and Block AttnRes under equal memory limits. Report prefill latency, decode latency, throughput at several batch sizes, and training time to a fixed validation target. The relevant winner may differ between offline training, interactive serving, and high-throughput batch inference.
Test whether distant access is causal. The sliding-window result motivates long-range depth retrieval, but does not identify which distant sources matter. Freeze a trained model, mask selected historical sources at evaluation, and compare against masks matched for total removed attention mass and source norm. Then retrain selected architectural variants to distinguish immediate dependence from adaptation. These are proposed experiments; this review has not performed model training or benchmark reproduction.
Use uncertainty where the comparison is close. At 241M, the displayed mHC(-lite) loss is 1.869 and Full AttnRes is 1.874; at other sizes the ordering changes. The general prose claim that Full always outperforms mHC is stronger than this table supports. Repeated runs at the close-comparison settings would be more informative than a single ranking. Likewise, the tie on MMLU-Pro should remain a tie, and rounding should not be interpreted as proof of identical underlying performance.
For readers planning a research comparison, I would first fix the mathematical architecture and source-count convention, then document the recipe and hardware. The most useful small checks are deterministic identities: softmax invariance to a common shift, exact merging of disjoint source sets, the absence of an empty partial source, and the relationship between transfer formulas and the assumed schedule. Those checks establish the explanation’s internal consistency. They do not substitute for training evidence.
13. Conclusion
Attention Residuals makes a familiar operation newly explicit: combining earlier representations is a modeling choice. Full AttnRes grants token-dependent access to individual layer outputs; Block AttnRes keeps a smaller set of summaries; fixed pseudo-queries enable a schedule that reuses completed-history work. The systems contribution is integral to the method because activation traffic and pipeline communication otherwise threaten the appeal of a parameter-light modification.
My main takeaway is the joint design of representation and access. Summaries determine what a later layer can recover, while scheduling determines how expensive that access becomes. The paper offers credible evidence that this trade-off can improve its tested model family. The next step is to characterize where the gains come from, how uncertainty affects close comparisons, and which point on the quality-latency frontier remains attractive under different workloads.
References and provenance
- Kimi Team. Attention Residuals, arXiv:2603.15031v1, 2026-03-16. Equations, experimental values, and the reported systems accounting in this review refer to this version; Appendix A gives the full authorship and contribution order, and Appendix B derives Full AttnRes I/O.
- Official project resource, linked from the paper.
- All nine figures here are original explanatory diagrams or redrawings of explicitly identified numerical facts. Hypothetical vectors, source counts, and analytical memory calculations are labeled as such. The paper is distributed under CC BY-NC-ND 4.0; no original paper figure is reproduced or edited here.
- Mathematical examples and arithmetic in this review are independent explanatory calculations. No model-training run, author experiment, or production latency result is claimed as independently reproduced.