Zhongzhu (Charlie) Zhou
/ZHONG-JOO JOH/
- Senior Research Scientist
Turbo Team
Together AI -
Ph.D.
School of Computer Science, Faculty of Engineering
The University of Sydney -
B.Eng. (Hons)
School of Computer Science and Engineering
Sun Yat-sen University
"Let everything happen to you. Beauty and terror. Just keep going. No feeling is final." — Rainer Maria Rilke
Who am I?
I am a Senior Research Scientist at the Turbo Team, Together AI, supervised by Ben Athiwaratkun.
I am a Ph.D. at the School of Computer Science, Faculty of Engineering, The University of Sydney, supervised by Prof. Shuaiwen Leon Song. I have been fortunate to intern at Dolby, DeepSpeed Microsoft, Weixin Group Tencent, and Microsoft (China), contributing to projects in building machine learning systems. I was also a research associate at the School of Computer Science and Engineering, Sun Yat-sen University from 2019 to 2022, under the supervision of Prof. Dan Huang and Yutong Lu. I received my B.E. degree from the School of Computer Science and Engineering, Sun Yat-sen University in 2019.
Research Highlights
I focus on efficient model architecture, post-training and reinforcement learning, quantization, systems infrastructure, and agentic AI. For aligned collaborations, feel free to get in touch.
Efficient Post-Training, RL & Quantization
Efficient Loss Design
Efficient ML Engines, Systems & Infrastructure
Training, Serving & Agent Systems
Agents, Coding & Science
Agentic AI & AI for Science
For the complete project portfolio, see the Research page.
Featured Projects
2026
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
Quantization · 2026
Attention-aware offline rotations + clipping that compress the long-history KV cache to 2.28 bits per element on Qwen3-4B/8B/32B and GLM-4.7-FP8 — ~8× KV-memory reduction and up to ~7× higher large-batch throughput, with mean accuracy gap to BF16 of just 1.42 / -0.02 / +0.27 on the 8B / 32B / 358B models. SGLang INT2-serving path with Triton kernels.
Show 6 more from 2026
I-DLM: Introspective Diffusion Language Models
Efficient inference
First diffusion language model to match same-scale AR quality. Introspective Strided Decoding (ISD) verifies prior tokens while advancing new ones in one forward pass. I-DLM-8B beats LLaDA-2.1-mini (16B) by +26 on AIME-24 and +15 on LiveCodeBench-v6 with half the parameters, at 2.9-4.1× throughput. With gated LoRA, R-ISD is bit-for-bit lossless versus the base AR model.
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
Efficient inference
Unified multi-model orchestration for verifier-free evolution — routes generation and aggregation between large and small models based on cross-model confidence. Accepted by COLM 2026.
Aurora: When RL Meets Adaptive Speculative Training — A Unified Training-Serving System
Speculator training system · ICML 2026
Unified training-serving system that closes the loop by continuously learning a speculator from live inference traces. Day-0 deployment, 1.45× speedup on MiniMax M2.1 and Qwen3-Coder-Next, and 1.25× over static speculators.
CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention
Efficient inference · ICLR 2026
Covariance-aware low-rank decomposition that converts pretrained GQA/MHA into MLA — up to 215× lower one-shot perplexity and 1.70× mean accuracy at matched KV budgets on Llama-3.1-8B/70B and Qwen3-4B/30B.
KITTY: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
Quantization · MLSys 2026
Dynamic Channel-wise Precision Boost keeps a small fraction of K-cache channels at 4-bit while quantizing the rest to 2-bit — cuts KV memory ~8×, enables up to 8× larger batches, and achieves 2.1×–4.1× higher throughput at matched memory.
2025
Ladder-Residual: Parallelism-Aware Architecture for Accelerating Large Model Inference with Communication Overlapping
Efficient inference · ICML 2025
Architectural modification that decouples communication from computation in tensor parallelism — 29% end-to-end wall-clock speedup on a 70B Transformer with TP=8. Pure PyTorch, no custom CUDA kernels.
2024
CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuning
Efficient training · NeurIPS 2024
Task-aware adapters built via SVD on covariance-oriented decomposition. Two modes: knowledge-preserved (mitigates forgetting, beats LoRA) and instruction-previewed (matches full fine-tuning, beats LoRA/DoRA/PiSSA).
Show 2 more from 2024
Quant-LLM: Accelerating Large Language Model Serving via FP6-Centric Algorithm-System Co-Design on Modern GPUs
Quantization · USENIX ATC 2024
TC-FPx — first full-stack GPU kernel scheme with unified Tensor-Core support for FP6 and arbitrary bit-widths. Delivers 1.69×–2.65× higher inference throughput over FP16 and enables LLaMA-70B on a single GPU.
Flash-LLM: Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
Quantization · VLDB 2024
Load-as-Sparse and Compute-as-Dense (LSCD) methodology for unstructured-sparse SpMM on GPU tensor cores — 2.9× faster than Sputnik at the kernel level and up to 3.8× over DeepSpeed end-to-end on OPT models.
2023
DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
RL training system · DeepSpeed
End-to-end RLHF training pipeline replicating InstructGPT. Trains a 6.7B model in ~4.1 hours / ≈$132 on 8×A100; a 66B model in 2.1 days / ≈$1,620.
2020
JSidentify: A Hybrid Framework for Detecting Plagiarism Among JavaScript Code in Online Mini Games
Software engineering · ICSE 2020
Hybrid plagiarism-detection framework for JavaScript mini-games. Uses edit-distance + network-flow on V8/Ignition bytecode plus a priority-queue framework to consolidate detection algorithms. Deployed at Tencent.
Recent News
- 2026/08/01 Talk I will give a talk at Cornell RL Seminar Weekly on reinforcement learning for Aurora, summarizing how adaptive speculative training frames speculator training as an online RL problem.
- 2026/07/22 Talk I gave an invited talk at Amazon. Thank you to the Amazon team for the invitation and engaging discussion!
- 2026/07/09 Paper Our papers I-DLM: Introspective Diffusion Language Models and Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution have been accepted by COLM 2026! I-DLM PaperI-DLM CodeI-DLM ProjectSqueeze Evolve PaperSqueeze Evolve Project
- 2026/06/26 Project Happy to see community interest in bringing OSCAR toward vLLM serving workflows. Grateful to the contributors pushing broader support for deployable rotation-based KV-cache quantization in mainstream LLM serving stacks. vLLM PR #46774ProjectCode
- 2026/06/26 Paper Two new scaling-law papers are now on arXiv: Sketched Linear Contrastive Learning: Approximation, Optimization, and Statistical Scaling, and From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression. Congrats to Andy! Sketched Contrastive LearningMini-Batch Scaling Laws
See the full archive for every announcement.
Experience
Together AI
Hybrid & San Francisco, United States
Research Consultant / Senior Research Scientist May 2024 - Present
Dolby
Sydney, Australia
Research Intern Mar 2024 - Sep 2024
DeepSpeed Team, Microsoft
Sydney, Australia & Remote
Research Intern Mar 2023 - Feb 2024
Weixin Group, Tencent Holdings Ltd.
Champaign, IL, US & Guangzhou, China
Research Intern Jul 2018 - Jul 2020
Microsoft (China) Co., Ltd.
Guangzhou, China
Project Assistant to Senior Cloud Architect Sep 2018 - Feb 2019
Education
The University of Sydney (USYD)
Sydney, Australia
Doctor of Philosophy (Ph.D.) Oct 2022 - Feb 2026
GPA: 4.0/4.0 (High Distinction)
Awards: Excellent progress evaluations; APR Intern Program Scholarship; JD AI Research Scholarship.
The University of Sydney (USYD)
Sydney, Australia
Visiting Scholar Mar 2022 - Oct 2022
Sun Yat-sen University (SYSU)
Guangzhou, China
Research Associate Sep 2019 - Mar 2022
GPA: 3.41/4.0
Awards: Overseas research funding; three Third-Class Scholarships; one Second-Class Scholarship.
University of Illinois Urbana-Champaign (UIUC)
Remote & Champaign, IL, US
Summer Session Student Jun 2018 - Sep 2018
Awards: Illinois Computer Science Summer Research Program.
Sun Yat-sen University (SYSU)
Guangzhou, China
Bachelor of Engineering in Computer Science and Technology Sep 2015 - Jun 2019
GPA: 3.9/4.0
Awards: National Scholarship; Research Honor Degree; two First-Class and one Second-Class Scholarships; selected competition awards.
Professional Service
- IEEE Member ID: 97841404
- ACM Member ID: 6708618
- China Computer Federation (CCF) Member ID: B8293G
- Conference Reviewer: The 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), NeurIPS 2025, NeurIPS 2023
- Journal Reviewer: IEEE Transactions on Computers
- Program Committee: 1st International Workshop on Sustainable and Efficient Language, Vision, and Action Models (SELVA 2026) at ACL 2026
- Program Committee: Computer Science Research Methods (CSRM 2023), University of Sydney
- Artifact Evaluation Committee: ASPLOS 2024 (AEC)
- Web Chair: 32nd IEEE International Symposium on High-Performance Computer Architecture (HPCA 2026)
Selected Publications
One representative work for each active publication year from 2020–2026.
2026 OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
2025 Ladder-Residual: Parallelism-Aware Architecture for Accelerating Large Model Inference with Communication Overlapping
2024 Quant-LLM: Accelerating the Serving of Large Language Models via FP6-Centric Algorithm-System Co-Design on Modern GPU
2023 DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
2021 Binary Neural Network for Automated Visual Surface Defect Detection
Special Issue on Intelligent Sensing and Monitoring for Industrial Process.
See the complete publication list or my Google Scholar profile.
Selected Talks
- Aurora: Adaptive Speculative Training as Online Reinforcement Learning Cornell RL Seminar Weekly · August 1, 2026
- Invited Talk Amazon · July 22, 2026
- Frontier Research: Scaling Agentic RL & Verified Reasoning AI-tonomy Summit (Models & Agents) — AI Researcher Forum, Plug and Play Tech Center, Silicon Valley · June 5, 2026 · watch
- I-DLM: Introspective Diffusion Language Models Ant Group · April 2026
- CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent Attention ICLR (International Conference on Learning Representations) 2026 · April 2026
- Panel Discussion: AI in Cross-Border Digital Technologies Infrastructure Platforms & Global Launchpads AI in the Cross-Border Digital Technologies Ecosystem: Infrastructure, Platforms & Global Launchpads · November 20, 2025
- JSidentify: A Hybrid Framework for Detecting Plagiarism Among JavaScript Code in Online Mini Games ICSE (International Conference on Software Engineering) · July 11, 2020 · watch
Selected Awards
- [2024] APR Intern Program Scholarship (SC3600), USYD
- [2022] Jingdong Technology (JD) Co Ltd Research Scholarship in Artificial Intelligence, USYD
- [2021] SYSU Overseas Visiting and Collaborative Research Program Funding Plan, SYSU
- [2019] Research Honor Degree, SYSU
- [2018] The Third Prize, Microsoft Hackathon, South China
- [2018] The First Class Scholarship (Top 5% of the major), SYSU × 2 (2015–16 & 2017–18)
- [2017] Meritorious Winner, COMAP's Mathematical Contest in Modeling, United States
- [2017] The Second Prize, Student Innovation Software Development Competition, SYSU
- [2017] The Third Prize, ACM-ICPC, SYSU
- [2016] National Scholarship (Top 1 of the major), China
Teaching
COMP3520: Operating Systems Internals — Tutor, The University of Sydney, Fall 2023.