I am a PhD student in Computer Science at Berkeley AI Research (BAIR), UC Berkeley, advised by Prof. Kurt Keutzer. My research focuses on efficient machine learning, especially quantization and sparsity, to compress and accelerate language models and video generation models while preserving accuracy and robustness.
I graduated from the Yao Class at Tsinghua University, led by Prof. Andrew Yao. I work closely with Prof. Song Han at MIT. At Tsinghua, I was fortunate to be advised by Prof. Jianfei Chen and Prof. Jun Zhu. I was also a visiting student at the University of Washington, advised by Prof. Sheng Wang.
I received the Citadel Fellowship in 2026, and our StreamDiffusion V2 work received the Best Research Paper Award at MLSys 2026.
May 2026 - Aug 2026
May 2026 - Aug 2026

May 2025 - Apr 2026
San Jose, CA, USA
Research internship at NVIDIA, advised by Prof. Song Han (MIT Han Lab).
May 2025 - Apr 2026

Mar 2024 - Aug 2024
Beijing, China
Research on FP8 training for large language models, advised by Prof. Song Han.
Mar 2024 - Aug 2024
Video-native hybrid attention for livestream video generation. Code · Project website
Compositing spatio-temporal memory for autoregressive video generation. Code · Project website
Efficient linear attention through accurate quantization of recurrent states.
We propose LoSA, a locality-aware sparse attention for block-wise diffusion language models. LoSA reuses cached prefix-attention for stable tokens and applies sparse attention only to active tokens, achieving up to +9 accuracy points at aggressive sparsity, 1.54x lower attention density, and 4.14x attention speedup.
We propose Flash-KMeans, an IO-aware and contention-free k-means for modern GPUs. FlashAssign fuses distance computation with online argmin; Sort-Inverse Update transforms atomic scatters into localized reductions. Achieves up to 17.9x end-to-end speedup, 33x over cuML, and 200x+ over FAISS on H200 GPUs.
We present Quant VideoGen (QVG), a training-free KV cache quantization framework for autoregressive video diffusion with semantic-aware smoothing and progressive residual quantization. Reduces KV cache memory by up to 7× with under 4% end-to-end latency overhead while improving quality over baselines.
We unify FP8 precision across RL training and rollout to avoid off-policy numerical mismatch from BF16-train + FP8-rollout, which can destabilize long-horizon RL. Achieves up to 33% rollout speedup, 41% training speedup, and 16% end-to-end speedup over BF16 with stable convergence.
We present StreamDiffusionV2, a training-free streaming system that adapts video diffusion models for interactive, low-latency live generation. It integrates an SLO-aware batching scheduler, sink-token-guided rolling KV cache, motion-aware noise controller, and scalable pipeline orchestration across denoising steps and network layers. Achieves 0.5s time-to-first-frame and up to 58.28 FPS (14B) / 64.52 FPS (1.3B) on 4xH100 without TensorRT or quantization.
We propose a training-free sparse attention framework that uses semantic-aware permutation - clustering and reordering tokens via k-means based on semantic similarity - to achieve a new pareto-frontier in speed-quality tradeoff.
Oscillation-reduced MXFP4 training for vision transformers.
We propose a self-speculative decoding framework, QuantSpec, to speedup long-context inference. QuantSpec maintains high acceptance rates (>90%) and reliably provides consistent end-to-end speedups upto ∼ 2.5×.
We identify the spatial head and temporal head pattern in attention map and propose to use sparse attention to accelerate. Achieves up to 2.28x and 2.33x end-to-end speedup on CogVideoX-v1.5 and HunyuanVideo.
We propose Dynamic range expansion for FP8 optimizer, and propose FP8 precision flow for FP8 activations. Achieve Lossless performance, end-to-End 1.54x memory reduction and 1.43x training speedup over BF16.
We propose a new method for efficient and accurate transformer pretraining with INT8 data flow and per-block quantization. Demonstrate effectiveness on GPT2-774M model. Achieve End-to-End 1.42x training speedup and 1.49x memory reduction.
Propose Hadamard Quantizer and Leverage Score Sampling to enable INT4 Precision Matmul in training for speedup. Both the forward and backward pass are quantized into INT4 precision for maximized speedup. Outperforms all existing 4-bit training baselines.













