Zhongzhu (Charlie) Zhou
/ZHONG-JOO JOH/
- Senior Research Scientist
Turbo Team
Together AI -
Ph.D.
School of Computer Science, Faculty of Engineering
The University of Sydney -
B.Eng. (Hons)
School of Computer Science and Engineering
Sun Yat-sen University
"Let everything happen to you. Beauty and terror. Just keep going. No feeling is final." — Rainer Maria Rilke
Who am I?
I am a Senior Research Scientist at the Turbo Team, Together AI, supervised by Ben Athiwaratkun.
I am a Ph.D. at the School of Computer Science, Faculty of Engineering, The University of Sydney, supervised by Prof. Shuaiwen Leon Song. I have been fortunate to intern at Dolby, DeepSpeed Microsoft, Weixin Group Tencent, and Microsoft (China), contributing to projects in building machine learning systems. I was also a research associate at the School of Computer Science and Engineering, Sun Yat-sen University from 2019 to 2022, under the supervision of Prof. Dan Huang and Yutong Lu. I received my B.E. degree from the School of Computer Science and Engineering, Sun Yat-sen University in 2019.
Research Highlights
I focus on efficient model architecture, post-training and reinforcement learning, quantization, ML engines, systems, and infrastructure, and agentic AI. For aligned collaborations, feel free to get in touch.
Efficient Post-Training, RL & Quantization
Efficient Loss Design
Efficient ML Engines, Systems & Infrastructure
Training, Serving & Agent Systems
Agents, Coding & Science
Agentic AI & AI for Science
For the complete project portfolio, see the Research page.
Featured Projects
2026
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
Quantization · 2026
Attention-aware offline rotations + clipping that compress the long-history KV cache to 2.28 bits per element on Qwen3-4B/8B/32B and GLM-4.7-FP8 — ~8× KV-memory reduction and up to ~7× higher large-batch throughput, with mean accuracy gap to BF16 of just 1.42 / -0.02 / +0.27 on the 8B / 32B / 358B models. SGLang INT2-serving path with Triton kernels.
Show 6 more from 2026
I-DLM: Introspective Diffusion Language Models
Efficient inference
First diffusion language model to match same-scale AR quality. Introspective Strided Decoding (ISD) verifies prior tokens while advancing new ones in one forward pass. I-DLM-8B beats LLaDA-2.1-mini (16B) by +26 on AIME-24 and +15 on LiveCodeBench-v6 with half the parameters, at 2.9-4.1× throughput. With gated LoRA, R-ISD is bit-for-bit lossless versus the base AR model.
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
Efficient inference
Unified multi-model orchestration for verifier-free evolution — routes generation and aggregation between large and small models based on cross-model confidence. Accepted by COLM 2026.
Aurora: When RL Meets Adaptive Speculative Training — A Unified Training-Serving System
Speculator training system · ICML 2026
Unified training-serving system that closes the loop by continuously learning a speculator from live inference traces. Day-0 deployment, 1.45× speedup on MiniMax M2.1 and Qwen3-Coder-Next, and 1.25× over static speculators.
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
Efficient inference · ICLR 2026
Covariance-aware low-rank decomposition that converts pretrained GQA/MHA into MLA — up to 215× lower one-shot perplexity and 1.70× mean accuracy at matched KV budgets on Llama-3.1-8B/70B and Qwen3-4B/30B.
KITTY: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
Quantization · MLSys 2026
Dynamic Channel-wise Precision Boost keeps a small fraction of K-cache channels at 4-bit while quantizing the rest to 2-bit — cuts KV memory ~8×, enables up to 8× larger batches, and achieves 2.1×–4.1× higher throughput at matched memory.
2025
Ladder-Residual: Parallelism-Aware Architecture for Accelerating Large Model Inference with Communication Overlapping
Efficient inference · ICML 2025
Architectural modification that decouples communication from computation in tensor parallelism — 29% end-to-end wall-clock speedup on a 70B Transformer with TP=8. Pure PyTorch, no custom CUDA kernels.
2024
CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning
Efficient training · NeurIPS 2024
Task-aware adapters built via SVD on covariance-oriented decomposition. Two modes: knowledge-preserved (mitigates forgetting, beats LoRA) and instruction-previewed (matches full fine-tuning, beats LoRA/DoRA/PiSSA).
Show 2 more from 2024
Quant-LLM: Accelerating Large Language Model Serving via FP6-Centric Algorithm-System Co-Design on Modern GPUs
Quantization · USENIX ATC 2024
TC-FPx — first full-stack GPU kernel scheme with unified Tensor-Core support for FP6 and arbitrary bit-widths. Delivers 1.69×–2.65× higher inference throughput over FP16 and enables LLaMA-70B on a single GPU.
Flash-LLM: Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
Quantization · VLDB 2024
Load-as-Sparse and Compute-as-Dense (LSCD) methodology for unstructured-sparse SpMM on GPU tensor cores — 2.9× faster than Sputnik at the kernel level and up to 3.8× over DeepSpeed end-to-end on OPT models.
2023
DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
RL training system · DeepSpeed
End-to-end RLHF training pipeline replicating InstructGPT. Trains a 6.7B model in ~4.1 hours / ≈$132 on 8×A100; a 66B model in 2.1 days / ≈$1,620.
2020
JSidentify: A Hybrid Framework for Detecting Plagiarism Among JavaScript Code in Online Mini Games
Software engineering · ICSE 2020
Hybrid plagiarism-detection framework for JavaScript mini-games. Uses edit-distance + network-flow on V8/Ignition bytecode plus a priority-queue framework to consolidate detection algorithms. Deployed at Tencent.
Recent News
- 2026/08/01 Talk I will give a talk at Cornell RL Seminar Weekly on reinforcement learning for Aurora, summarizing how adaptive speculative training frames speculator training as an online RL problem.
- 2026/07/22 Talk I gave an invited talk at Amazon: Compression-Centric Systems for Large-Model Post Training & Inference. Thank you to the Amazon team for the invitation and engaging discussion!
- 2026/07/09 Paper Our papers I-DLM: Introspective Diffusion Language Models and Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution have been accepted by COLM 2026! I-DLM PaperI-DLM CodeI-DLM ProjectSqueeze Evolve PaperSqueeze Evolve Project
- 2026/06/26 Project Happy to see community interest in bringing OSCAR toward vLLM serving workflows. Grateful to the contributors pushing broader support for deployable rotation-based KV-cache quantization in mainstream LLM serving stacks. vLLM PR #46774ProjectCode
- 2026/06/26 Paper Two new scaling-law papers are now on arXiv: Sketched Linear Contrastive Learning: Approximation, Optimization, and Statistical Scaling, and From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression. Congrats to Andy! Sketched Contrastive LearningMini-Batch Scaling Laws
See the full archive for every announcement.
Experience
Together AI
Hybrid & San Francisco, United States
Research Consultant / Senior Research Scientist May 2024 - Present
Dolby
Sydney, Australia
Research Intern Mar 2024 - Sep 2024
DeepSpeed Team, Microsoft
Sydney, Australia & Remote
Research Intern Mar 2023 - Feb 2024
Weixin Group, Tencent Holdings Ltd.
Champaign, IL, US & Guangzhou, China
Research Intern Jul 2018 - Jul 2020
Microsoft (China) Co., Ltd.
Guangzhou, China
Project Assistant to Senior Cloud Architect Sep 2018 - Feb 2019
Education
The University of Sydney (USYD)
Sydney, Australia
Doctor of Philosophy (Ph.D.) Oct 2022 - Feb 2026
GPA: 4.0/4.0 (High Distinction)
Awards: Excellent progress evaluations; APR Intern Program Scholarship; JD AI Research Scholarship.
The University of Sydney (USYD)
Sydney, Australia
Visiting Scholar Mar 2022 - Oct 2022
Sun Yat-sen University (SYSU)
Guangzhou, China
Research Associate Sep 2019 - Mar 2022
GPA: 3.41/4.0
Awards: Overseas research funding; three Third-Class Scholarships; one Second-Class Scholarship.
University of Illinois Urbana-Champaign (UIUC)
Remote & Champaign, IL, US
Summer Session Student Jun 2018 - Sep 2018
Awards: Illinois Computer Science Summer Research Program.
Sun Yat-sen University (SYSU)
Guangzhou, China
Bachelor of Engineering in Computer Science and Technology Sep 2015 - Jun 2019
GPA: 3.9/4.0
Awards: National Scholarship; Research Honor Degree; two First-Class and one Second-Class Scholarships; selected competition awards.
Professional Service
- IEEE Member ID: 97841404
- ACM Member ID: 6708618
- China Computer Federation (CCF) Member ID: B8293G
- Conference Reviewer: The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), NeurIPS 2025, NeurIPS 2023
- Journal Reviewer: IEEE Transactions on Computers
- Program Committee: 1st International Workshop on Sustainable and Efficient Language, Vision, and Action Models (SELVA 2026) at ACL 2026
- Program Committee: Computer Science Research Methods (CSRM 2023), University of Sydney
- Artifact Evaluation Committee: ASPLOS 2024 (AEC)
- Web Chair: 32nd IEEE International Symposium on High-Performance Computer Architecture (HPCA 2026)
Selected Publications
Selected recent and representative work across efficient machine learning algorithms and systems.
- 2026 OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
- 2026 I-DLM: Introspective Diffusion Language Models
- 2026 Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
- 2026 When RL Meets Adaptive Speculative Training: A Unified Training-Serving System
- 2025 Ladder-Residual: Parallelism-Aware Architecture for Accelerating Large Model Inference with Communication Overlapping
- 2024 Quant-LLM: Accelerating the Serving of Large Language Models via FP6-Centric Algorithm-System Co-Design on Modern GPU
- 2023 DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
- 2021 Binary Neural Network for Automated Visual Surface Defect Detection
- 2020 JSidentify: A Hybrid Framework for Detecting Plagiarism Among JavaScript Code in Online Mini Games
See the complete publication list or my Google Scholar profile.
Selected Talks
- Aurora: Adaptive Speculative Training as Online Reinforcement Learning Cornell RL Seminar Weekly · August 1, 2026
- Compression-Centric Systems for Large-Model Post Training & Inference Amazon · July 22, 2026
- Frontier Research: Scaling Agentic RL & Verified Reasoning AI-tonomy Summit (Models & Agents) — AI Researcher Forum, Plug and Play Tech Center, Silicon Valley · June 5, 2026 · watch
- I-DLM: Introspective Diffusion Language Models Ant Group · April 2026
- CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention ICLR (International Conference on Learning Representations) 2026 · April 2026
- Panel Discussion: AI in Cross-Border Digital Technologies Infrastructure Platforms & Global Launchpads AI in the Cross-Border Digital Technologies Ecosystem: Infrastructure, Platforms & Global Launchpads · November 20, 2025
- JSidentify: A Hybrid Framework for Detecting Plagiarism Among JavaScript Code in Online Mini Games ICSE (International Conference on Software Engineering) · July 11, 2020 · watch
Selected Awards
- [2024] APR Intern Program Scholarship (SC3600), USYD
- [2022] Jingdong Technology (JD) Co Ltd Research Scholarship in Artificial Intelligence, USYD
- [2021] SYSU Overseas Visiting and Collaborative Research Program Funding Plan, SYSU
- [2019] Research Honor Degree, SYSU
- [2018] The Third Prize, Microsoft Hackathon, South China
- [2018] The First Class Scholarship (Top 5% of the major), SYSU × 2 (2015–16 & 2017–18)
- [2017] Meritorious Winner, COMAP's Mathematical Contest in Modeling, United States
- [2017] The Second Prize, Student Innovation Software Development Competition, SYSU
- [2017] The Third Prize, ACM-ICPC, SYSU
- [2016] National Scholarship (Top 1 of the major), China
Teaching
COMP3520: Operating Systems Internals — Tutor, The University of Sydney, Fall 2023.