Page 7 / 30
349 posts in total. Keep on posting.
Showing posts 73–84 of 349. Each entry opens locally on this site; legacy Hexo posts link back to their original article at the bottom for reference.
2026
- 中
SwiftQK:当归一化层成为张量并行的瓶颈
SwiftQK 用标量偏和聚合取代张量并行下 Query-Key 归一化所需的整向量 All-Gather 通信,并在一个避免死锁的持久化 kernel 内把剩余的点对点同步延迟与计算重叠,QK-Norm 延迟最多降低 93.9%,端到端 TPOT 平均降低 29.5%。
- EN
Chasing a Moving Target: Calibration and Rank-Allocation Drift in Training-Free Low-Rank LLM Compression
A close read of a COLM 2026 paper showing that training-free low-rank LLM compression pipelines silently drift away from the model they think they are compressing, and how two lightweight corrections close much of that gap.
- 中
追着一个不断移动的目标:训练无关低秩压缩中的校准漂移与秩分配漂移
阅读笔记:COLM 2026 一篇论文指出,主流的训练无关低秩压缩流水线在压缩过程中会悄悄偏离它本应对齐的模型状态,本文提出两个轻量修正来缓解这一问题。
- EN
ZeroLock: Breaking Update Locking in Pipeline-Parallel LLM Fine-Tuning
ZeroLock replaces backpropagation's update-locking pipeline with per-chunk local objectives, cutting per-stage peak memory by up to 55% and lifting throughput by up to 63% over GPipe/1F1B/PipeDream, backed by the first convergence analysis for local-objective pipeline training under general chunk division.
- 中
ZeroLock 阅读笔记:用模块化解耦打破流水线并行训练的更新锁定
ZeroLock 用逐块本地目标替代反向传播的全链路更新依赖,单卡峰值显存最多降低约55%,吞吐最多提升约63%,并首次给出了任意分块数量下本地目标构造类算法的收敛性证明。
- EN
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
OasisKV turns speculative-decoding draft tokens into a free, one-step-ahead KV-cache prefetch signal, letting sparse decoding scale batch size and throughput up to 2.1x over dense vLLM while staying within 0.7 accuracy points of full attention.
- 中
OasisKV 阅读笔记:把投机采样的草稿 token 变成免费的 KV Cache 预取信号
OasisKV 把投机采样早已算出来的草稿 token 直接拿来当 KV cache 预取的『提前一步』信号,让稀疏解码把批大小和吞吐量提升到稠密 vLLM 的 2.1 倍,同时精度损失控制在 0.7 个百分点以内。
- EN
CVPO: Curriculum-Guided Value-Variance Policy Optimization for LLM Reasoning
CVPO turns the variance of a value model's token-level estimates into a signal that both stabilizes correct trajectories and encourages exploration on incorrect ones, then layers a Bayesian curriculum on top to fight difficulty drift during long RL runs.
- 中
CVPO 阅读笔记:用价值方差驱动的课程式策略优化
CVPO 把价值模型对同一条轨迹内部的方差波动,转化成一个能区分对错、动态调节探索强度的优势加权信号,再叠加一套基于贝叶斯后验的难度课程机制来对抗训练中的'难度漂移'。
- EN
GRAFT: Global Optimization and Inference-Time Region Grafting for Agentic Workflows
GRAFT keeps a globally-searched agentic workflow frozen as structural scaffolding, then locally swaps out only the operators that a label-free quality signal flags as weak for the current input — beating supernet- and search-based workflow optimizers without any training or per-query global re-search.
- 中
阅读笔记:GRAFT——给智能体工作流做「全局定骨架、局部动手术」的推理时嫁接
GRAFT 把离线搜好的智能体工作流当作冻结的结构骨架,推理时只用无标签质量信号局部替换表现不佳的算子——不训练、不重搜全局,就在多个基准上超过 supernet 与搜索式工作流优化方法。
- EN
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models
Bole reformulates tree speculative decoding for hybrid full-attention/linear-attention models as one closed-form parallel solve, cutting transient recurrent-state memory by up to 99x and linear-attention verification time by up to 7.7x.