Zhongzhu (Charlie) Zhou
Home
Research
Publication
Experience
Recent News
Blog
CV
↗
Tag
#
Transformer
19 posts tagged with this label. Back to
all tags
or the
main feed
.
2026
09-25
EN
FACTS and CoRS: Which Curvature Should Low-Rank Compression Preserve?
09-25
中
FACTS 与 CoRS 阅读笔记:低秩压缩究竟该保留哪一种曲率?
09-24
EN
Deep Delta Learning: Editing the Residual Stream, and Accounting for the Cost
09-24
中
Deep Delta Learning 阅读笔记:怎样改写残差流,以及这次改写要付出什么代价
09-20
EN
mHC: Stable Residual Mixing, Its Guarantees, and Its Costs
09-20
中
mHC 阅读笔记:残差混合如何稳定,保证又止于哪里
09-19
EN
Attention Residuals: Learning What to Retrieve Across Network Depth
09-19
中
Attention Residuals 阅读笔记:让每一层选择自己需要的历史表示
08-03
EN
BLADE: Boundary-Expanded and Layer-Adaptive Dynamic Exit for Efficient LLM Reasoning
08-03
中
BLADE 阅读笔记:把「早退」这件事从自我怀疑词扩展到普通句子边界
06-12
EN
SliceGPT: Post-Training LLM Compression via Computational Invariance
06-12
中
SliceGPT 阅读笔记:用计算不变性删除 Transformer 的行与列
06-11
EN
MegaScale: Engineering 55% MFU at 12,288 GPUs for LLM Training
06-11
中
MegaScale:ByteDance 如何在 12,288 块 GPU 上实现 55% MFU 的大规模 LLM 训练
04-22
EN
SAGE: Training-Free Semantic Evidence Composition for Edge-Cloud Inference Under Hard Uplink Budgets
03-28
EN
Mamba: Linear-Time Sequence Modeling with Selective State Spaces — In-Depth Technical Review
03-22
EN
Attention Is All You Need: The Transformer — In-Depth Technical Review
02-18
EN
DeepSeek-V2: Multi-head Latent Attention and DeepSeekMoE — Technical Review
02-18
EN
GLM-5 Technical Review: From Vibe Coding to Agentic Engineering