Publications
Peer-reviewed papers, workshop papers and preprints, newest first. Also on Google Scholar and DBLP.
Peer-reviewed
2026
- Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
- A Derandomization Framework for Structure Discovery: Applications in Neural Networks and Beyond
- DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
- Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
- Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
2025
- An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
- Understanding Adam Requires Better Rotation Dependent Assumptions
- Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
- Compositional risk minimization
- Solving hidden monotone variational inequalities with surrogate losses
2024
- Feature learning as alignment: a structural property of gradient descent in non-linear neural networks
- Are we making progress in unlearning? Findings from the first NeurIPS unlearning competition
- Expecting The Unexpected: Towards Broad Out-Of-Distribution Detection
- No Wrong Turns: The Simple Geometry Of Neural Networks Optimization Paths
- Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
- LEAD: Least-Action Dynamics for Min-Max Optimization
2023
- Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation
- CADet: Fully Self-Supervised Anomaly Detection With Contrastive Learning
- A Reproducible and Realistic Evaluation of Partial Domain Adaptation Methods
- Stochastic Mirror Descent: Convergence Analysis and Adaptive Variants via the Mirror Stochastic Polyak Stepsize
- Synergies Between Disentanglement and Sparsity: a Multi-Task Learning Perspective
- Empirical Study on Optimizer Selection for Out-of-Distribution Generalization
- A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games
- Neural Networks Efficiently Learn Low-Dimensional Representations with SGD
- Performative Prediction with Neural Networks
2022
- Gradient Descent Is Optimal Under Lower Restricted Secant Inequality And Upper Error Bound
- Optimal transport meets noisy label robust loss and MixUp regularization for domain adaptation
- Towards efficient representation identification in supervised learning
2021
- Stochastic Gradient Descent-Ascent and Consensus Optimization for Smooth Games: Convergence Analysis under Expected Co-coercivity
- Invariance Principle Meets Information Bottleneck for Out-of-Distribution Generalization
- Adversarial score matching and improved sampling for image generation
- A Study of Condition Numbers for First-Order Optimization
2020
- In search of robust measures of generalization
- Stochastic Hamiltonian Gradient Methods for Smooth Games
- Linear Lower Bounds and Conditioning of Differentiable Games
- Accelerating Smooth Games by Manipulating Spectral Shapes
- A Tight and Unified Analysis of Gradient-Based Methods for a Whole Spectrum of Differentiable Games
2019
- Reducing the variance in online optimization by transporting past gradients
- Multi-objective training of Generative Adversarial Networks
- Manifold Mixup: Better Representations by Interpolating Hidden States
- State-Reification Networks: Improving Generalization by Modeling the Distribution of Hidden Representations
- h-detach: Modifying the LSTM Gradient Towards Better Optimization
- Negative Momentum for Improved Game Dynamics
- YellowFin and the Art of Momentum Tuning
2018
- Learning Representations and Generative Models for 3D Point Clouds
- YellowFin: Adaptive optimization for (A)synchronous systems
- Accelerated stochastic power iteration
2017
- Deep Learning at 15PF: Supervised and Semi-Supervised Classification for Scientific Data
- Improving Gibbs Sampler Scan Quality with DoGS
2016
- Asynchrony begets Momentum, with an Application to Deep Learning
- Scan Order in Gibbs Sampling: Models in Which it Matters and Bounds on How Much
2015
2014
2013
2011
- User Rankings from Comparisons: Learning Permutations in High Dimensions
- Joint Power and Admission Control for Ad-hoc and Cognitive Underlay Networks: Convex Approximation and Distributed Implementation
2010
- Strong Information-Theoretic Limits for Source/Model Recovery
- Distributed Joint Power and Admission Control for Ad-hoc and Cognitive Underlay Networks
2008
- Convex Approximation-based Joint Power and Admission Control for Cognitive Underlay Networks
Workshop papers
2026
2024
- Smoothness-adaptive sharpness-aware minimization for finding flatter minima
- Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
2019
- In Support of Over-Parametrization in Deep Reinforcement Learning: an Empirical Study
- A Modern Take on the Bias-Variance Tradeoff in Neural Networks
2016
Preprints
2026
- Navigating Potholes with Geometry-Aware Sharpness Minimization
- Revisiting Generalization Measures Beyond IID: An Empirical Study under Distributional Shift
2025
2024
2022
2021
- Convergence Analysis and Implicit Regularization of Feedback Alignment for Deep Linear Networks
- Gotta Go Fast When Generating Data with Score-Based Models
2020
- Generalizing to unseen domains via distribution matching
- Gradient penalty from a maximum margin perspective