Page 22 / 30

351 posts in total. Keep on posting.

Showing posts 253–264 of 351. Each entry opens locally on this site; legacy Hexo posts link back to their original article at the bottom for reference.

2026

  • EN

    Swift-SVD: Activation-Aware Low-Rank Compression for LLM Weights and KV Cache

    A detailed technical review of Swift-SVD, an activation-aware low-rank compression method for LLM weights and KV cache that uses output covariance eigendecomposition to avoid expensive generalized SVD.

  • EN

    Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism

    A detailed technical review of Piper, a resource-model-driven system for large-scale MoE training with pipelined hybrid parallelism, HALO hierarchical all-to-all, and topology-aware expert placement.

  • EN

    Low-Rank Optimization Trajectories for LLM RLVR Acceleration: A Technical Review of NExt

    A detailed technical review of NExt, a method that models low-rank optimization trajectories to accelerate reinforcement learning with verifiable rewards for large language models.

  • EN

    Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond — Technical Review

    A technical review of agentic world modeling, covering capability levels, governing-law regimes, evaluation, and why decision-centric world models matter for LLM agents.

  • EN

    OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning

    A comprehensive technical review of SAGE, analyzing how to optimize semantic evidence composition for edge-cloud systems under hard uplink budget constraints. The paper challenges importance-only patch selection and proposes a training-free method combining importance filtering with diversity-maximizing sampling.

  • EN

    Generalization at the Edge of Stability: A Random Dynamical Systems Perspective

    1. What This Paper Does Core Problem The edge of stability phenomenon, discovered by Cohen et al. (2021), presents a theoretical puzzle: when training with sufficiently large learning rates η, the largest Hessian eigenvalue λ₁ frequently exceeds the stability threshold 2/η, implying the system should diverge according to classical optimization theory. Yet empirically: Training loss continues to decrease Model generalization often improves in this regime The optimizer doesn't settle at a point but explores a bounded, chaotic set Prior explanations relying on pointwise properties (Hessian trace, spectral norm) fail to capture this phenomenon because they ignore the ensemble behavior of the attractor set. Main Contribution The paper's central insight: characterize generalization through the geometric properties of the random attractor itself, not individual solutions. They prove that: Sharpness Dimension (SD) < ambient dimension d with high probability at EoS Worst-case generalization error depends on SD, not parameter count d The complete Hessian spectrum structure matters, not just the trace or largest eigenvalue The attractor forms a fractal set with intrinsic dimension strictly smaller than the parameter space This explains why overparameterized models generalize: the training dynamics naturally compress into a lower-dimensional manifold despite the high-dimensional parameter space.

  • EN

    SAGE: Training-Free Semantic Evidence Composition for Edge-Cloud Inference Under Hard Uplink Budgets

    Paper: Choi & Park, arXiv:2604.19623 (April 2026) Focus: Efficient inference in edge-cloud hybrid systems through optimal evidence composition Key Contribution: Demonstrates that coverage-aware patch selection outperforms importance-only methods under hard bandwidth constraints What This Paper Does This paper addresses a practical but underexplored problem in edge-cloud inference systems: how should the edge device select which image patches to transmit to the server when the uplink channel strictly limits the number of patches per request? The standard approach—selecting patches by importance (attention score)—turns out to be fundamentally limited. The paper shows that this creates "coverage gaps": high-attention patches cluster in the same semantic region, wasting budget on overlapping information. SAGE proposes a simple but effective alternative that combines importance filtering with diversity-maximizing sampling, achieving 93% of the server's full-transmission accuracy while sending fewer than half the patches. The insight is elegant: under hard budgets, every transmitted patch must count, so we should prioritize information coverage alongside importance.

  • EN

    SpecGuard: Verification-Aware Speculative Decoding for Efficient Multi-Step Reasoning

    A technical review of SpecGuard, a verification-aware speculative decoding method that uses model-internal attention and log-probability signals to improve multi-step reasoning efficiency.

  • 中

    SpecGuard:用于多步推理的验证感知推测解码

    一篇关于 SpecGuard 的阅读笔记:它用模型内部的注意力与对数概率信号改进多步推理场景下的推测解码验证。

  • EN

    GRASP Technical Review: Replacing Redundant LLM Layers with Adaptive Singular Parameters

    A detailed review of GRASP, which replaces redundant transformer layers with gradient-selected adaptive singular parameters instead of simply deleting layers or keeping only the largest singular values.

  • EN

    PipeDream: Turning Pipeline Parallelism into a Practical Training System — Deep Technical Review

    1. Why this paper still matters in 2026 I think PipeDream is one of those papers that is easier to appreciate after the field has moved on. If I explain it in one sentence, I would say: PipeDream turned pipeline parallelism from a vague idea into a system-level recipe: profile the model, partition it automatically, keep multiple minibatches in flight, and repair the optimization semantics enough that training still converges. That sounds modest today because pipeline parallelism is now normal vocabulary in large-model training. But in 2018, this was an important systems step. The paper is historically important for at least four reasons. It clearly shows that data parallelism is not always the right default. When models become large, or when interconnects are weak relative to GPU speed, weight synchronization becomes a real bottleneck. It reframes pipeline parallelism as a joint scheduling and optimization problem, not just a diagram where layers are placed on different GPUs. It identifies the subtle but crucial issue of parameter-version mismatch between forward and backward passes. That is the kind of detail that separates a classroom concept from a production system. It anticipates a lot of the design space that later became standard in large-scale training stacks: stage partitioning, pipeline schedules, weight-version policies, stage replication, and runtime-managed buffer reuse. I also think the paper is still useful for modern readers because it teaches a systems mindset that remains valid: first find the actual bottleneck, then pick the right parallelization dimension, then ask what semantic damage the optimization introduces, then engineer around that damage carefully. That sequence is still exactly how good ML systems work today.

  • 中

    PipeDream:把 Pipeline Parallelism 做成真正可训练系统——深度阅读笔记

    1. 为什么这篇论文到 2026 年仍然值得读 如果让我用一句话概括这篇论文,我会说: PipeDream 的价值,不只是“把模型切成几段在不同 GPU 上跑”,而是把 pipeline parallelism 真正做成了一个完整训练系统:先 profile,后 partition,再 schedule,同时处理参数版本一致性问题,最后用 time-to-accuracy 来衡量系统价值。 今天大家谈大模型训练,已经很习惯使用 pipeline、tensor parallel、ZeRO、FSDP、activation checkpointing 这些术语,所以回头看 PipeDream,好像会觉得它只是早期工作之一。 但如果放回 2018 年的语境,这篇论文做了几件非常关键的事: 它明确说明了:数据并行不是永远正确的默认解。 它把 pipeline parallelism 从“概念图”推进到了可实现、可验证、可比较的系统设计。 它抓住了一个非常本质的问题:同一个 minibatch 的 forward 和 backward 如果看到的不是同一版参数,会不会把训练语义搞坏? 它让后来很多大模型训练系统里的概念变得更容易表达,比如 stage 划分、1F1B 调度、weight version、stage replication 等等。 我觉得它到今天仍然值得认真读,原因不是“它还能直接拿来训练最新 LLM”,而是它教会了我们一个很重要的系统思路: 先找真正的瓶颈; 再决定用哪一种并行方式; 再追问这种并行方式会不会破坏训练语义; 最后才是运行时与实现层面的工程落地。 这个思路今天一点都不过时。