Ioannis Mitliagkas
I study optimization, dynamics and learning in modern machine learning: why training works, when models generalize, and how to make both reliable.
- Associate Professor, Computer Science (DIRO), Université de Montréal
- Core faculty member, Mila, and Canada CIFAR AI Chair
- Part-time research scientist, Google DeepMind, Montréal
- Affiliated researcher, Archimedes, Athena Research Center, Athens
Prospective students: Fall 2027
I am recruiting PhD and MSc students for Fall 2027. Please go over my recent publications and the research themes below. If you think we have a strong overlap in interests, submit a Mila supervision request between October 15 and December 1, 2026, and list me as one of your faculty of choice. You also need to apply separately to the Université de Montréal MSc or PhD program.
I cannot respond to all emails; the supervision request is how you will be considered. I will consider all good candidates, but pay extra attention to those from unusual backgrounds, underrepresented groups, and candidates coming from regions under threat of war, occupation, political instability, etc. (Ukraine, Palestine, Africa, …).
Research
More on each themeOptimization for modern deep learning
Why optimizers like Adam work, and how to make training faster and more robust: batch sizes, sign and spectral methods, distributed and asynchronous training.
-
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
Adaptive batch-size rules rely on the gradient noise scale, which assumes the Euclidean geometry of SGD. We derive the noise scale for sign and spectral methods (Signum, Muon) from their dual norms, and estimate it cheaply from gradients already on each data-parallel worker. On a 160M-parameter Llama model, this matches the loss of fixed batch sizes with up to 66% fewer training steps.
-
Understanding Adam Requires Better Rotation Dependent Assumptions
Theory often explains Adam with assumptions that do not change when the parameter space is rotated. We show that Adam on transformers gets worse under random rotations, while some structured rotations help, and that existing rotation- dependent assumptions do not explain this. The orthogonality of Adam's update emerges as a promising quantity for a better theory.
-
No Wrong Turns: The Simple Geometry Of Neural Networks Optimization Paths
Why do stochastic first-order methods train non-convex networks so reliably? We measure two quantities along the optimization path, tied to the restricted secant inequality and the error bound, and find that they behave predictably across vision and language tasks, architectures and optimizers. These properties are enough to guarantee linear convergence and to motivate learning-rate schedules like those used in practice.
Generalization, out-of-distribution robustness and compositionality
Models that keep working when the data changes: domain shift, OOD detection, identifiable and compositional representations, and pretraining objectives.
-
Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
Next-token prediction trains a model to look one step ahead, which limits long- horizon reasoning and planning. We add an auxiliary head that predicts a summary of the long-term future, either handcrafted (a bag of words) or learned with a reverse language model. In 3B and 8B pretraining runs, it beats next-token and multi-token prediction on math, reasoning and coding benchmarks.
-
Compositional risk minimization
Classifiers fail when test data contain combinations of attributes never seen together in training. We model each attribute as an additive energy term, train a classifier with that structure, then adjust it for the new combinations. The method, compositional risk minimization, provably extrapolates beyond the training combinations and is more robust than methods built for subpopulation shift.
-
Additive Decoders for Latent Variables Identification and Cartesian-Product Extrapolation
When an image is a sum of object-specific parts, a decoder with the same additive structure recovers each object's latent variables from reconstruction alone, with no independence assumptions on the latents. The same decoder can then generate combinations of objects never seen together in training. This gives a theoretical account of object-centric decoders and of how disentanglement enables extrapolation.
Privacy and machine unlearning
Removing the influence of data from trained models, with guarantees, and measuring whether unlearning methods actually work.
-
Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data
Certified unlearning adds noise, and the noise costs utility, especially for large deletion requests. We show that mixing in public data cuts the cost of unlearning by a factor of O(1/n_pub²), and we account for shifts between the public and private data. This makes it practical to unlearn a constant fraction of the training set while keeping utility.
-
Are we making progress in unlearning? Findings from the first NeurIPS unlearning competition
We ran the first NeurIPS machine unlearning competition, with nearly 1,200 teams, scoring forgetting through a formal definition of unlearning together with model utility. The top entries beat existing algorithms, and their ranking stays stable across variations of the evaluation, so it can be made cheaper. We also analyze where the methods succeed and fail.
Dynamics of games and reinforcement learning
Min-max optimization, variational inequalities and the learning dynamics of multi-agent and reinforcement learning.
-
Solving hidden monotone variational inequalities with surrogate losses
Many problems in deep learning, such as min-max games and projected Bellman error minimization, are variational inequalities, and standard gradient methods cycle or diverge on them. We solve them through a sequence of surrogate losses that work with standard optimizers such as Adam, with convergence guarantees under hidden monotonicity. The approach also yields a more compute- and sample-efficient variant of TD(0).
-
LEAD: Least-Action Dynamics for Min-Max Optimization
Min-max training, as in GANs, stalls because of rotational dynamics. Treating the two players as particles under physical forces, we derive LEAD, a second-order method that damps the rotation. It converges linearly on quadratic games and improves GAN training on CIFAR-10.
-
A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games
We introduce magnetic mirror descent, one algorithm that serves both as an equilibrium solver and as a reinforcement learning method for two-player zero-sum games. It is the first solver of quantal response equilibria with linear convergence in extensive-form games from first-order feedback, and the first standard RL algorithm competitive with counterfactual regret minimization in tabular settings. As a deep self-play method, it does well on Dark Hex and Phantom Tic-Tac-Toe.
News
All news- A Surrogate Perspective on Convergence of Fixed-Target DQN accepted at NeurIPS 2026.
- I am recruiting PhD and MSc students for Fall 2027. See how to apply.
- Hiroki Naganuma has completed his PhD and joins NVIDIA. Congratulations Dr. Naganuma!
- Two papers accepted at ICML 2026, on unlearning with public data and on adaptive batch sizes for sign and spectral descent.
- Three papers accepted at ICLR 2026, including Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries.
- Understanding Adam Requires Better Rotation Dependent Assumptions accepted at NeurIPS 2025.
Group
Members and alumniMy most important responsibility is supervising a group of very talented junior researchers.
-
Ryan D'Orazio
PhD candidate
-
Divyat Mahajan
PhD candidate
- Zichu Liu PhD candidate
-
Ahmed (Mehdi) Inane
PhD candidate
-
Mehrab Hamidi
PhD candidate
- Yukti Makhija PhD student
-
Mahdi Ghaznavi
PhD student (starting January 2027)
Where alumni went next
- Nicolas Loizou: Assistant Professor, Johns Hopkins University
- Manuela Girotti: Assistant Professor, Emory University
- Kartik Ahuja: Research Scientist, FAIR Paris
- Kilian Fatras: Research Scientist, Dreamfold
- Hiroki Naganuma: NVIDIA
- Charles Guille-Escuret: MBZUAI Institute of Foundation Models
- Reyhane Askari Hemmat: Research Scientist, Meta, Montréal
- Adam Ibrahim: Google DeepMind, Mountain View
- Alexia Jolicoeur-Martineau: Research Scientist, Microsoft Research
- Vincent Quirion: Soma Energy
- Mehrnaz Mofakhami: Research assistant, Mila
- Rémi Piché-Taillefer: Microsoft Research, Montréal
- Brady Neal: Senior Research Scientist, Dataiku
About
Researcher in machine learning. Academic, immigrant, amateur musician, runner.
Every fall, I teach Fundamentals of Machine Learning (IFT 3395/6390, in French) to a large class of undergraduate and graduate students. In winter 2027 I am teaching my advanced research class on deep learning theory (IFT 6169) again.
I co-founded and hosted the first two seasons of MTL MLOpt, a bi-weekly meeting of optimization experts from Mila, UdeM, McGill (CS and math), Google DeepMind, SAIL, FAIR and MSR. We share our guest speaker videos.
In the early days of interest in the area, I co-organized the Smooth Games Optimization and ML workshop series at NeurIPS. The opening remarks from NeurIPS 2019 summarize our motivation. In spring 2022 I was invited to the semester on Learning and Games at the Simons Institute, Berkeley. For several summers I taught optimization for ML at the Neuromatch Academy deep learning course.
Before joining the Université de Montréal, I was a postdoc with the Departments of Computer Science and Statistics at Stanford University, and a PhD student at The University of Texas at Austin.