Page 13 / 30
349 posts in total. Keep on posting.
Showing posts 145–156 of 349. Each entry opens locally on this site; legacy Hexo posts link back to their original article at the bottom for reference.
2026
- 中
RSPO:多轮LLM智能体的奖励交换策略优化
RSPO 通过奖励交换机制,让稠密过程奖励训练的探索智能体扩展轨迹多样性,再用结果奖励训练最终策略,同时避免奖励错位与奖励黑客问题。
- EN
The Mirage of Optimizing Training Policies: Monotonic Inference Policy Improvement for LLM RL
MIPU exposes an overlooked objective-level flaw in LLM RL: training-side improvement does not guarantee inference-side improvement under training-inference mismatch, and proposes a two-step framework to fix it.
- 中
训练策略的幻觉:为什么LLM强化学习的真正目标是推理策略单调改进
MIPU揭示了LLM RL训练中被忽视的目标错位问题:在训练-推理不一致的情况下,训练侧的策略改进并不保证推理侧的策略改进,并提出了两步框架来解决这一问题。
- EN
Lynx: Progressive Speculative KV Cache Transfer for Disaggregated LLM Inference
Lynx challenges the assumption that KV caches must be fully received before decoding begins — splitting them into MSB Anchor and LSB Residual streams to overlap transfer with speculative generation, achieving INT4-level TTFT while matching BF16 accuracy in disaggregated LLM serving.
- 中
Lynx 阅读笔记:渐进式推测量化加速解聚合 LLM 推理的 KV 传输
Lynx 打破「KV 缓存必须完整接收才能开始解码」的假设,将 KV 缓存分成高优先级 MSB Anchor 流和低优先级 LSB Residual 流,Anchor 到达后立即推测解码,Residual 到达后无损验证——在解聚合 LLM 服务中同时实现 INT4 的首 token 延迟和 BF16 的推理精度。
- EN
MosaicKV: Dynamic Two-Dimensional KV Cache Compression for Long-Context LLM Serving — Technical Review
MosaicKV solves the long-context KV cache bottleneck by applying dynamic per-vector element selection and segment-adaptive strategies across both sequence and channel dimensions, achieving 16x attention speedup and 7.3x throughput gain at only 1.76pct average accuracy loss.
- 中
MosaicKV:面向超长上下文LLM服务的动态二维KV缓存压缩——阅读笔记
MosaicKV通过逐向量元素选择与分段自适应策略同时压缩序列维度和通道维度,在LongBench和RULER上仅损失1.76pct准确率,实现16倍注意力加速和7.3倍吞吐提升。
- EN
AIR: Activation- and Influence-Aware SVD Compression for LLMs — Technical Review
AIR adds a closed-form element-wise influence ALS sweep on top of SVD-LLM(W) whitening, achieving 18-45pct perplexity gains at 20-60pct parameter retention while cutting peak memory 64pct and per-token latency 53pct on an A100.
- 中
AIR 阅读笔记:激活与影响力双重感知的SVD低秩LLM压缩
AIR 在激活白化基础上引入元素级反向传播影响力矩阵,通过封闭形式ALS迭代实现混合感知低秩近似,在60%参数保留下困惑度降低18%,峰值内存削减64%,推理延迟降低53%。
- EN
Tangram: Hiding GPU Heterogeneity for Efficient LLM Parallelization
Tangram decouples parallelization planning from GPU heterogeneity by abstracting heterogeneous clusters into homogeneous GPU islands, then composing partial plans from existing parallelizers into work-balanced pipelines — achieving up to 2.3× higher throughput than heterogeneous baselines while retaining full support for expert parallelism, ZeRO, and activation recomputation.
- 中
Tangram:为异构GPU集群隐藏硬件差异的高效LLM并行化系统
Tangram将异构GPU集群抽象为同构GPU岛,让现有的同构并行化器生成部分计划,再通过动态规划组合成全局负载均衡的流水线——在保留专家并行、ZeRO、激活重计算等全部特性的同时,比现有异构并行化器吞吐量高出最多2.3倍。
- EN
SSV: Sparse Speculative Verification for Efficient LLM Inference
SSV resolves the structural mismatch between speculative decoding and dynamic sparse attention by grouping overlapping verifier queries, fusing NSA branches across layers, and adaptively orchestrating draft-verify strategies per prompt — achieving up to 3.49x end-to-end throughput on H100 GPUs.