I keep this page as a compact summary of selected research projects rather than a database of publications and manuscripts. For a complete publication list, please see my Google Scholar.
Last manually curated: .
My recent work and near-term research interests focus on:
Designing better optimization algorithms for LLM pre-training.
Understanding how components of the training pipeline interact, and co-designing optimization algorithms with those components.
Understanding pre-training beyond the final validation cross-entropy loss. When two pre-trained models achieve nearly the same validation loss, how can we tell which will perform better during post-training?
Selected projects
Click to enlarge.
NeurIPS · 2025
ASGO: Adaptive Structured Gradient Optimization
Kang An, Yuxing Liu, Rui Pan, Yi Ren, Shiqian Ma, Donald Goldfarb, Tong Zhang
We develop ASGO, a matrix-aware adaptive optimizer with a fine-grained convergence analysis that exploits low-rank gradient structure. Under the analyzed assumptions, ASGO improves convergence rates over Shampoo. Our practical variant reduces optimizer-state memory by over 50% relative to Shampoo, with 8% wall-clock overhead relative to AdamW in our experiments.
Demystifying Manifold Constraints in LLM Pre-training
Kang An, Jiaxiang Li, Donald Goldfarb, Shiqian Ma
We develop MACRO to study how manifold constraints shape LLM pre-training. The work connects manifold constraints with RMSNorm, weight decay, and BF16 precision, from the perspectives of activation control, rotation regulation, and update-to-weight ratio control, respectively. Experiments on Qwen3-based models (240M–1.2B parameters) show competitive performance with respect to other constrained optimizers.
Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence
Yuxing Liu, Yuze Ge, Rui Pan, Kang An, Tong Zhang
We investigate why learning-rate warmup can accelerate convergence through generalized smoothness assumptions that connect local curvature to the loss gap. Our analysis covers gradient descent in deterministic and stochastic settings, establishing when increasing learning rates improves convergence over non-increasing schedules.