Research
I study optimization, dynamics and learning, with a focus on modern machine learning. I have also worked at the intersection of systems and theory. Below are the themes my group works on now, with our papers on each. The full list is on the publications page.
Optimization for modern deep learning
All papers on this themeWhy optimizers like Adam work, and how to make training faster and more robust: batch sizes, sign and spectral methods, distributed and asynchronous training.
- Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
- Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
- Understanding Adam Requires Better Rotation Dependent Assumptions
- LEAD: Least-Action Dynamics for Min-Max Optimization
- No Wrong Turns: The Simple Geometry Of Neural Networks Optimization Paths
- Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
Generalization, out-of-distribution robustness and compositionality
All papers on this themeModels that keep working when the data changes: domain shift, OOD detection, identifiable and compositional representations, and pretraining objectives.
- DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
- A Derandomization Framework for Structure Discovery: Applications in Neural Networks and Beyond
- Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
- Compositional risk minimization
- An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
- Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
Privacy and machine unlearning
All papers on this themeRemoving the influence of data from trained models, with guarantees, and measuring whether unlearning methods actually work.
- Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
- Are we making progress in unlearning? Findings from the first NeurIPS unlearning competition
Dynamics of games and reinforcement learning
All papers on this themeMin-max optimization, variational inequalities and the learning dynamics of multi-agent and reinforcement learning.
- Solving hidden monotone variational inequalities with surrogate losses
- LEAD: Least-Action Dynamics for Min-Max Optimization
- A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games
- Stochastic Gradient Descent-Ascent and Consensus Optimization for Smooth Games: Convergence Analysis under Expected Co-coercivity
- A Tight and Unified Analysis of Gradient-Based Methods for a Whole Spectrum of Differentiable Games
- Accelerating Smooth Games by Manipulating Spectral Shapes
Earlier work and resources
- Projects archive: representative projects grouped by area, curated in early 2022.
- Summary of research contributions from my group (PDF, not fully up to date).
- Slides summarizing our work on games and min-max optimization.
- Recent manuscripts on Google Scholar.
Funding acknowledgements
Special thanks to Intel and NVIDIA for donating access to hardware, and to SigOpt for access to their platform for some of our work.







