Zhongzhu (Charlie) Zhou
Home
Research
Publication
Experience
Recent News
Blog
CV
↗
Tag
#
Distributed Training
45 posts tagged with this label. Back to
all tags
or the
main feed
.
2026
08-15
EN
SwiftQK: When Normalization Becomes the Bottleneck of Tensor Parallelism
08-15
中
SwiftQK:当归一化层成为张量并行的瓶颈
08-13
EN
ZeroLock: Breaking Update Locking in Pipeline-Parallel LLM Fine-Tuning
08-13
中
ZeroLock 阅读笔记:用模块化解耦打破流水线并行训练的更新锁定
08-06
EN
Fine-Grained Compute-Communication Overlap for MoE: Hiding All-to-All Behind GEMM with Tile-Level Signaling
08-06
中
MoE 计算-通信细粒度重叠:用 Tile 级信号把 All-to-All 藏在 GEMM 背后
08-02
EN
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
08-02
中
FEPLB 阅读笔记:用 Copy Engine 几乎零成本地做 MoE 负载均衡
07-30
EN
Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool
07-30
中
Libra 阅读笔记:用有界序列池驯服长上下文训练中的注意力负载偏斜
07-23
EN
PHOENIX: Recovering LLM Training in 40 Seconds Instead of Restarting the Whole Job
07-23
中
PHOENIX:让大模型训练在故障后 40 秒内复活,而不是整个作业重启
07-16
EN
GIFT: Why the Coordinate System You Quantize In Matters More Than the Quantizer
07-16
中
GIFT 阅读笔记:量化用的坐标系,比量化器本身更重要
07-09
EN
RATrain: Training-State Lifecycle Scheduling for Dense LLM Training on Bandwidth-Constrained Heterogeneous Supercomputers
07-09
中
RATrain:面向带宽受限异构超算的稠密大模型训练状态生命周期调度
07-02
EN
Tangram: Hiding GPU Heterogeneity for Efficient LLM Parallelization
07-02
中
Tangram:为异构GPU集群隐藏硬件差异的高效LLM并行化系统
06-13
EN
ForeMoE: Micro-step-level MoE Load Balancing for RL Post-training via Routing Foresight
06-13
中
ForeMoE:利用路由预见性实现 RL 后训练中 MoE 微步级负载均衡
06-11
EN
MegaScale: Engineering 55% MFU at 12,288 GPUs for LLM Training
06-11
中
MegaScale:ByteDance 如何在 12,288 块 GPU 上实现 55% MFU 的大规模 LLM 训练
05-15
EN
Zero Sum SVD: A Global, Loss-Aware Rank Budget for LLM Compression
05-15
中
Zero Sum SVD:用「损失零和」做全局奇异值预算分配的 LLM 压缩方法
05-14
EN
DisagMoE: Disaggregating Attention and FFN to Beat the MoE All-to-All Bottleneck
05-14
中
DisagMoE:用解耦 Attention 和 FFN 打通 MoE 训练的 all-to-all 瓶颈
05-07
EN
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
04-29
EN
FEPLB Technical Review: Nearly Free MoE Load Balancing with the NVLink Copy Engine
04-24
EN
FEPLB: Zero-Cost MoE Load Balancing via NVLink Copy Engine
04-16
EN
PipeDream: Turning Pipeline Parallelism into a Practical Training System — Deep Technical Review
04-16
中
PipeDream:把 Pipeline Parallelism 做成真正可训练系统——深度阅读笔记
04-04
EN
Switch Transformers: Scaling to Trillion-Parameter Sparse Models — In-Depth Technical Review
04-04
中
Switch Transformers:用简单高效的稀疏性扩展到万亿参数模型 — 深度阅读笔记
04-02
EN
GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism — In-Depth Technical Review
04-02
中
GPipe:微批次流水线并行的大规模模型训练 — 深度阅读笔记
03-29
EN
Ring Attention: Blockwise Transformers for Near-Infinite Context — In-Depth Technical Review
03-26
EN
Alpa: Automating Inter- and Intra-Operator Parallelism — In-Depth Technical Review
03-19
EN
ZeRO: Shattering the Memory Wall — How DeepSpeed Trains Trillion-Parameter Models
03-12
EN
Megatron-LM: NVIDIA's Blueprint for Training Billion-Parameter Models at Scale
03-12
EN
PaRO: Smarter Partitioning for Distributed Training — Beyond ZeRO's One-Size-Fits-All
2020
09-25
EN
Slurm-Day5
09-09
EN
Slurm-Day4
09-05
EN
Slurm-Day2
09-05
EN
Slurm-Day3
09-04
EN
Slurm-Day1