Research

I keep this page as a compact summary of selected research projects rather than a database of publications and manuscripts. For a complete publication list, please see my Google Scholar.

Last manually curated: .

My recent work and near-term research interests focus on:

  1. Designing better optimization algorithms for LLM pre-training.
  2. Understanding how components of the training pipeline interact, and co-designing optimization algorithms with those components.
  3. Understanding pre-training beyond the final validation cross-entropy loss. When two pre-trained models achieve nearly the same validation loss, how can we tell which will perform better during post-training?

Selected projects

Relationships between ASGO, Muon, Adam, SignSGD, and Shampoo, including ASGO's one-sided preconditioner
Click to enlarge.

NeurIPS · 2025

ASGO: Adaptive Structured Gradient Optimization

Kang An, Yuxing Liu, Rui Pan, Yi Ren, Shiqian Ma, Donald Goldfarb, Tong Zhang

We develop ASGO, a matrix-aware adaptive optimizer with a fine-grained convergence analysis that exploits low-rank gradient structure. Under the analyzed assumptions, ASGO improves convergence rates over Shampoo. Our practical variant reduces optimizer-state memory by over 50% relative to Shampoo, with 8% wall-clock overhead relative to AdamW in our experiments.

Poster presenting MACRO and the interactions of manifold constraints with RMSNorm, weight decay, and BF16 precision
Click to open poster (PDF).

Preprint · 2026

Demystifying Manifold Constraints in LLM Pre-training

Kang An, Jiaxiang Li, Donald Goldfarb, Shiqian Ma

We develop MACRO to study how manifold constraints shape LLM pre-training. The work connects manifold constraints with RMSNorm, weight decay, and BF16 precision, from the perspectives of activation control, rotation regulation, and update-to-weight ratio control, respectively. Experiments on Qwen3-based models (240M–1.2B parameters) show competitive performance with respect to other constrained optimizers.

ResNet and NanoGPT experiments plotting log smoothness against log loss gap, with points colored by training iteration
Click to enlarge.

Preprint · 2025

Theoretical Analysis on how Learning Rate Warmup Accelerates Convergence

Yuxing Liu, Yuze Ge, Rui Pan, Kang An, Tong Zhang

We investigate why learning-rate warmup can accelerate convergence through generalized smoothness assumptions that connect local curvature to the loss gap. Our analysis covers gradient descent in deterministic and stochastic settings, establishing when increasing learning rates improves convergence over non-increasing schedules.