Page 3 / 30
349 posts in total. Keep on posting.
Showing posts 25–36 of 349. Each entry opens locally on this site; legacy Hexo posts link back to their original article at the bottom for reference.
2026
- EN
Attention Residuals: Learning What to Retrieve Across Network Depth
A technical reading of depth-wise attention, block summaries, exact softmax merging, and the limits of the reported efficiency gains.
- 中
Attention Residuals 阅读笔记:让每一层选择自己需要的历史表示
从残差累加到深度方向的检索,推导分块表示与在线 softmax 合并,并分析训练收益和系统成本的适用边界。
- EN
SVD-LLM V2: Allocating Rank, Preserving Activations, and Understanding Local Optimality
A technical reading of heterogeneous rank allocation and activation-weighted truncation, with derivations, numerical examples, experimental evidence, and limits of local optimality.
- 中
SVD-LLM V2 阅读笔记:秩怎样分,激活怎样保,局部最优意味着什么
从参数预算和激活误差出发,推导两次 SVD 的截断方法,解释层间压缩比例分配,并分析论文实验、数值边界与实际加速条件。
- EN
QeRL: When Quantized Weights Help Reinforcement Learning, and Where the Evidence Stops
A technical review of NVFP4, low-rank policy updates, structured exploration noise, and the accounting needed to evaluate QeRL.
- 中
QeRL 阅读笔记:量化、探索与训练效率
从 NVFP4、LoRA 和 RMSNorm 噪声推导出发,拆解 QeRL 的探索机制、训练效率与可复现性问题。
- EN
SAPO: Smooth Policy Updates, Their Mathematics, and Their Limits
A technical reading of soft policy gates, sequence coherence, and the evidence needed to trust an RL optimizer.
- 中
SAPO 阅读笔记:平滑策略更新的推导、实现与边界
从策略梯度和组内优势出发,理解软门控、序列一致性的条件,以及训练实验尚未回答的问题。
- EN
Strong Drafts Need Compact Memories: A Technical Review of Memory-Augmented Sliding-Window (MASW) Drafting for Long-Context Speculative Decoding
A close technical read of MASW, which equips a strong independent draft model with a compact three-part working memory (sink tokens, exact local window, learned memory slots) so long-context speculative decoding keeps high acceptance length without paying full-KV-cache draft latency.
- 中
强推理草稿也要瘦身:长上下文投机解码的记忆增强滑窗(MASW)阅读笔记
读 MASW 这篇论文:给一个独立的强草稿模型配上三段式紧凑工作记忆(哨兵token、精确局部窗口、可学习记忆槽),让长上下文投机解码既保住高接受率、又不用支付全量KV访存的代价。
- EN
TreeWY: Removing the Memory Wall in Speculative Decoding for Gated DeltaNet Hybrids
A close technical read of TreeWY, which rewrites the gated delta rule as a tree-structured WY transform so that a Gated DeltaNet hybrid model can verify a wide speculative-decoding draft tree with one triangular solve instead of snapshotting a full recurrent state at every draft node.
- 中
TreeWY 阅读笔记:用树结构 WY 变换拆掉 Gated DeltaNet 混合模型投机解码里的内存墙
这篇论文把 Gated DeltaNet 的门控 delta 规则重新写成树结构的 WY 变换,让投机解码在验证一整棵很宽的草稿树时只需要做一次三角方程求解,而不必在每个草稿节点上都完整保存一份循环状态快照。