Zhongzhu (Charlie) Zhou
Home
Research
Publication
Experience
Recent News
Blog
CV
↗
Tag
#
Attention & MLA
23 posts tagged with this label. Back to
all tags
or the
main feed
.
2026
08-15
EN
SwiftQK: When Normalization Becomes the Bottleneck of Tensor Parallelism
08-15
中
SwiftQK:当归一化层成为张量并行的瓶颈
08-09
EN
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models
08-09
中
Bole 阅读笔记:为混合注意力大模型重新设计树状投机解码
08-05
EN
AnchorKV: Anchor-Residual KV Cache Compression
08-05
中
AnchorKV:锚点-残差表示下的 KV 缓存压缩阅读笔记
08-01
EN
Back from the Future: KV Cache Management by Counter-Causal Surprise
08-01
中
来自未来的信号:用反因果惊奇度做 KV Cache 淘汰
07-29
EN
LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
07-29
中
LOCKS 阅读笔记:为什么 KV Cache 摘要要按「页」而不是按「序列」或「全局」来建
07-15
EN
COBS: What Block-Sparse Attention Selectors Are Actually Missing (A Second-Order Fix)
07-15
中
COBS 阅读笔记:块稀疏注意力的选择器到底漏掉了什么(一个二阶修正)
07-01
EN
SSV: Sparse Speculative Verification for Efficient LLM Inference
07-01
中
SSV:稀疏投机验证——在动态稀疏注意力中做投机解码
06-24
EN
SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference
06-24
中
SparDA:稀疏解耦注意力,让长上下文推理又快又准
05-27
EN
GQA: Grouped-Query Attention — Bridging Multi-Head Quality and Multi-Query Speed
05-27
中
GQA 分组查询注意力:用分组 KV 头桥接多头质量与多查询速度
05-24
EN
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
05-24
中
FlashAttention-2:更好的并行策略与线程块工作划分
03-23
EN
MiRA: A Subgoal-driven Framework for Improving Long-Horizon LLM Agents — Technical Review
03-14
EN
FlashAttention: The IO-Aware Algorithm That Made Transformers Actually Fast
02-18
EN
DeepSeek-V2: Multi-head Latent Attention and DeepSeekMoE — Technical Review