Page 1 / 30
349 posts in total. Keep on posting.
Showing posts 1–12 of 349. Each entry opens locally on this site; legacy Hexo posts link back to their original article at the bottom for reference.
2026
- EN
FlashInfer-Bench: Closing the Loop Between Kernel Generation and LLM Serving
A technical review of workload contracts, numerical validation, benchmark metrics, and the evidence for deploying AI-generated GPU kernels.
- 中
FlashInfer-Bench 阅读笔记:从生成 GPU 算子到接入真实推理系统
梳理论文的任务契约、数值验证、性能指标与动态替换机制,并重新核对局部算子收益和端到端实验之间的关系。
- EN
Fast-dLLM: Reusing Bidirectional State and Spending Confidence on Parallel Tokens
A technical reading of Fast-dLLM: approximate KV reuse, confidence-aware parallel decoding, the assumptions behind its theorem, and the workload behind its speedups.
- 中
Fast-dLLM 阅读笔记:双向模型怎样复用缓存,又该一次填几个词?
从双向注意力的缓存近似,到置信度并行解码的定理、反例和实验边界,理解 Fast-dLLM 的加速来自哪里,以及 27.6 倍意味着什么。
- EN
ProRL: What Prolonged Reinforcement Learning Changes, and What Its Evidence Can Prove
A technical reading of ProRL, from group-relative updates and moving KL anchors to finite-sample reasoning coverage and the limits of causal attribution.
- 中
ProRL 阅读笔记:长期强化学习与推理边界
从组内相对优势、移动 KL 参照与分阶段训练,理解 ProRL 的收益,并用有限采样和实验控制重新审视推理能力边界。
- EN
ACE: How Agents Accumulate Useful Context Without Rewriting It Away
A technical reading of ACE's incremental playbooks, reflection feedback, evaluation protocols and adaptation costs, with a critical analysis of retention, harmful memories and fair comparisons.
- 中
ACE 阅读笔记:让智能体积累经验,也别把旧经验改没了
从上下文坍缩出发,理解 ACE 如何用增量规则积累经验,并分析反馈可靠性、实验口径、长上下文成本与自我演化的边界。
- EN
FlashInfer: Making Attention Fit the Serving Workload
A technical reading of attention-state composition, sparse KV layouts, dynamic scheduling, and the workload boundaries behind FlashInfer's serving gains.
- 中
FlashInfer 阅读笔记:让注意力计算适应真实推理服务
从注意力状态合并、稀疏 KV 视图到动态调度,理解 FlashInfer 的设计理由、加速来源与失效边界。
- EN
Scalable Kronecker-Fisher Approximation: What Cross-Layer Curvature Can Tell Us About Compression
A technical reading of matrix-free Kronecker approximation, cross-layer compression damage, targeted repair, and the limits of perplexity-based interaction claims.
- 中
Scalable Kronecker-Fisher 阅读笔记:压缩一层之后,另一层还安全吗?
从跨层曲率的矩阵重排推导出发,分析联合压缩与定向修复,并辨析困惑度交互指标、批梯度统计和敏感层排序的边界。