CSE 291C-b, Spring 2026
Topics on Numerical MethodsUniversity of California, San Diego Instructor
Teaching Assistant
- CK Cheng, Office: CSE2130, email: ckcheng+291@ucsd.edu, tel: 858 534-6184
- Office hour: TBA
Discussion Forum
- Ying Yuan, yiy090@ucsd.edu
- Office hours: TBA
Schedule
- TBA
References
- Lectures: 2:00-3:20PM TTH, Room: DIB 122
Prerequisite
- Lectures on Convex Optimization, Y. Nesterov, Springer, 2018.
- Linear and Nonlinear Programming, D.G. Luenberger and Y. Ye, Fifth Edition, Springer, 2021.
- Nonlinear Programming, Third Edition, D.P. Bertsekas, Athena Scientific, 2016.
- Numerical Optimization, J. Nocedal and S.J. Wright, Second Edition, Springer, 2006.
- An Introduction to Optimization on Smooth Manifolds, N. Boumal, Cambridge, 2023.
- Convex Optimization Algorithms, D.P. Bertsekas, Athena Scientific, 2015.
Random Matrix Methods for Machine Learning, R. Couillet and Z. Liao, Cambridge, 2022.- Matrix Mathematics, A Scond Course in Linear Algebra, Second Edition, S.R. Garcia, and R. A. Horn, Cambridge, 2023.
- Matrix Computations, G.H. Golub and C.F. Van Loan, Fouth Edition, Johns Hopkins, 2013
- Numerical Recipes: The Art of Scientific Computing, Third Edition, W.H. Press, S.A. Teukolsky, W.T. Vetterling, and B.P. Flannery, Cambridge University Press, 2007.
- Convex Optimization, S. Boyd and L. Vandenberghe, Cambridge, 2004
- Numerical Optimizatin, J. Nocedal and S.J. Wright, Second Edition, Springer, 2009.
- Electronic Circuit and System Simulation Methods, T.L. Pillage, R.A. Rohrer, C. Visweswariah, McGraw-Hill, 1998
Basic knowledge of linear algebra, numerical methods, convex optimization, or intention of conducting projects related to scientific computation.
ContentWe cover topics on numerical methods for dynamic system analysis and optimization algorithms. Our scope aims at systems in high-dimensional space with temporal behavior. We discuss techniques such as integration methods, first-order methods, random matrix theory, and Gumbel tricks.
Lecture Notes and Reference Papers
- 1. Gumbel Softmax Trick:
- Huijben, I.A., Kool, W., Paulus, M.B. and Van Sloun, R.J., 2022. A review of the gumbel-max trick and its extensions for discrete stochasticity in machine learning. IEEE transactions on pattern analysis and machine intelligence, 45(2), pp.1353-1371. pdf.
- Li, Y., Liu, J., Lin, G., Hou, Y., Mou, M. and Zhang, J., 2021. Gumbel-softmax-based optimization: a simple general framework for optimization problems on graphs. Computational Social Networks, 8(1), p.5. pdf.
- Balog, M., Tripuraneni, N., Ghahramani, Z. and Weller, A., 2017, July. Lost relatives of the Gumbel trick. In International Conference on Machine Learning (pp. 371-379). PMLR. pdf.
- 2. Neural Scaling Laws:
- Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D.D.L., Hendricks, L.A., Welbl, J., Clark, A. and Hennigan, T., 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 10. pdf.
- Tissue, H., Wang, V. and Wang, L., 2024. Scaling law with learning rate annealing. arXiv preprint arXiv:2408.11029. analysis and machine intelligence, 45(2), pp.1353-1371. pdf.
- Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B. and Sutskever, I., 2021. Deep double descent: Where bigger models and more data hurt. Journal of Statistical Mechanics: Theory and Experiment, 2021(12), p.124003. pdf.
- 3. Random Matrix Methods:
- Martin, C.H. and Mahoney, M.W., 2021. Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning. Journal of Machine Learning Research, 22(165), pp.1-73. pdf.
- Seddik, M.E.A., Louart, C., Tamaazousti, M. and Couillet, R., 2020, November. Random matrix theory proves that deep learning representations of gan-data behave as gaussian mixtures. In International Conference on Machine Learning (pp. 8573-8582). PMLR. pdf.
- Random Matrix Methods for Machine Learning, R. Couillet and Z. Liao, Cambridge, 2022.
- Lectures on Randome Matrices, R. Speicher, EMS Press, 2024.
- Topics in Random Matrix Theory, T. Tao, Ameican Mathematical Society, 2012.
- A Dynamical Approach to Random Matrix Theory, L. Erdos, H-T Yau, Ameican Mathematical Society, 2017.
- 4. Gradient Descent:
- Lecture Notes on Optimization for Machine Learning, A. Cutkowskyy. pdf.
- Differential Equation for Modeling Nesterov's Accelerated Gradient Method: Theory and Insights, W. Su et al. Journal of Machine Learning Researech, 2016, pdf.
- Understanding the Accelerated Phenomenon via High-Resolution Differential Equations, B. Shi, et al., Mathematical Programming, 2022, pdf file. pdf.
- Fast Iterative Shrinkage-Thresholding Algorithm for Lineare Inverse Problems, A. Beck, and M. Teboulle, SIAM J. Imageing Sciences, 2009, pdf file, pdf.
- In Search of Adam's Secret Sauce, A. Orvieto and R.M. Gower, NeurIPS 2025, pdf.
- K. Jordan, et al. "Muon: An Optimizer for Hidden Layers in Neural Networks," https://kellerjordan.github.io/posts/muon/.
- Bernstein, J. and Newhouse, L., 2024. Old Optimizer, New Norm: An Anthology. arXiv preprint arXiv:2409.20325, pdf.
- 5. Ordinary Differential Equation Integration:
- From Circuit Theory Simulation to SPICE_Diego: A Matrix Exponential Approach for Time-Domain Analysis of Large-Scale Circuits, H. Zhuang, X. Wang, Q Chen, P. Chen, and C.K. Cheng, IEEE Circuits and Systems Magazine, 2016. pdf.