Research
My research interests are in reliable optimization methods for modern machine learning, especially stochastic methods, distributed systems, robustness, privacy, and variational inequalities.
For my complete filterable list of papers, see Publications; for my presentation materials, see Talks.
Research map
The ultimate goal of my research is to bridge the gap between theory and practice in Optimization in Machine Learning.
Distributed learning
Reliability and trust
Modern assumptions and algorithmic tools
Representative papers
Selected recent publications
From Optimization to Generalization under Heavy-Tailed Data: The Role of Gradient Clipping
Recent work connecting heavy-tailed optimization theory with generalization and the role of gradient clipping in modern training.
General Analysis of LMO-based Optimizers: Beyond Bounded Variance
A general analysis of linear-minimization-oracle optimizers beyond the classical bounded-variance stochastic-gradient setting.
On the Role of Batch Size in Stochastic Conditional Gradient Methods
Theoretically derived scaling laws that help tune batch size and stepsize in large-scale stochastic conditional gradient training.
High-Probability Bounds for the Last Iterate of Clipped SGD
Last-iterate convergence guarantees for clipped stochastic gradient methods under heavy-tailed noise.
Error Feedback under $(L_0,L_1)$-Smoothness: Normalization and Momentum
Error-feedback methods under generalized smoothness, with normalization and momentum for communication-efficient training.
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
High-probability convergence guarantees for differentially private clipped SGD with arbitrary clipping levels.
Clipping Improves Adam-Norm and AdaGrad-Norm When the Noise Is Heavy-Tailed
Shows how clipping stabilizes adaptive normalization methods under heavy-tailed stochastic gradients.
Selected foundational and high-visibility publications
Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity
Clipping, acceleration, and adaptivity for generalized smoothness assumptions beyond the classical Lipschitz-gradient model.
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise
Clipping stochastic gradient differences for composite, distributed, and variational inequality problems with heavy-tailed noise.
Distributed Learning with Compressed Gradient Differences
A compressed-gradient-difference approach for communication-efficient distributed learning.
Stochastic Optimization with Heavy-Tailed Noise via Accelerated Gradient Clipping
The first near-optimal high-probability complexity guarantees under heavy-tailed gradient noise assumptions.
Variance Reduction Is an Antidote to Byzantines
The first variance-reduced Byzantine-robust method with strong convergence guarantees, weaker assumptions, and communication compression.
Extragradient Method: $O(1/K)$ Last-Iterate Convergence
Resolved a long-standing open question on last-iterate convergence of extragradient for monotone variational inequalities.
Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable Devices
Communication-efficient decentralized training for heterogeneous and unreliable devices.
MARINA: Faster Non-Convex Distributed Learning with Compression
A communication-efficient method for non-convex distributed learning with compressed gradient differences.
Linearly Converging Error Compensated SGD
The unified theory of distributed methods with error feedback and the first linearly converging method in this class.
A Unified Theory of SGD: Variance Reduction, Sampling, Quantization and Coordinate Descent
A unifying view of SGD variants through the parametric assumption on the stochastic estimator.