Yongtao Wu (2026)

 

Research Interests

  • Deep learning
  • Over-parametrized models

Biography

On August 24th, 2026 I successfully defended my PhD thesis. The thesis, entitled “Optimization in Modern Machine Learning: Steepest Descent Theory and Trustworthy Models” was supervised by Professor Volkan Cevher.

I received a Bachelor’s degree in Telecommunication Engineering from Sun Yat-sen University in 2020, and a Master’s degree in Machine Learning from KTH Royal Institute of Technology in 2022. From July to September 2022 I was an intern at LIONS. In September 2022 I started my PhD there.

Diffusion-based Cumulative Adversarial Purification for Vision Language Models

J. Fu; Y. Wu; Y. Chen; K. Peng; X. Zhang et al. 

Transactions on Machine Learning Research. 2026. Vol. 2026-June.

Optimization in Modern Machine Learning: Steepest Descent Theory and Trustworthy Models

Y. Wu / V. Cevher (Dir.)  

Lausanne, EPFL, 2026. 

Robustness in Both Domains: CLIP Needs a Robust Text Encoder

E. Abad Rocamora; C. Schlarmann; N. Singh; Y. Wu; M. Hein et al. 

2025. 39th Conference on Neural Information Processing Systems (NeurIPS 2025), San Diego, USA, 2025-12-02 – 2025-12-07.

Linear Attention for Efficient Bidirectional Sequence Modeling

A. Afzal; E. Abad Rocamora; L. Candogan; P. Puigdemont; F. Tonin et al. 

2025. 39th Conference on Neural Information Processing Systems (NeurIPS 2025), San Diego, USA, 2025-12-02 – 2025-12-07.

Multi-Step Alignment as Markov Games: An Optimistic Online Mirror Descent Approach with Convergence Guarantees

Y. Wu; L. Viano; K. Antonakopoulos; Y. Chen; Z. Zhu et al. 

Transactions on Machine Learning Research. 2025. num. 12/2025.

Machine learning-aided hysteretic response prediction of double skin composite wall under earthquake loads

S. Wang; W. Wang; Y. Wu; Z. Xie; Y. Gao 

Journal of Building Engineering. 2025. Vol. 101, p. 111837. DOI : 10.1016/j.jobe.2025.111837.

Quantum-Peft: Ultra Parameter-Efficient Fine-Tuning

T. Koike-Akino; F. Tonin; Y. Wu; F. Z. Wu; L. Candogan et al. 

2025. The Thirteenth International Conference on Learning Representations, Singapore, 2025-04-24-2025-04-28.

Single-pass Detection of Jailbreaking Input in Large Language Models

L. Candogan; Y. Wu; E. Abad Rocamora; G. Chrysos; V. Cevher 

Transactions on Machine Learning Research. 2025. Vol. 02/2025.

Hadamard product in deep learning: Introduction, Advances and Challenges

G. G. Chrysos; Y. Wu; R. Pascanu; P. Torr; V. Cevher 

IEEE Transactions on Pattern Analysis and Machine Intelligence. 2025. DOI : 10.1109/TPAMI.2025.3560423.

Membership Inference Attacks against Large Vision-Language Models

Zhan Li; Y. Wu; Y. Chen; F. Tonin; E. Abad Rocamora et al. 

2024. 38th Annual Conference on Neural Information Processing Systems, Vancouver Convention Center, 2024-12-10 – 2024-12-15.

Universal Gradient Methods for Stochastic Convex Optimization

A. Rodomanov; A. Kavis; Y. Wu; K. Antonakopoulos; V. Cevher 

2024. 41st International Conference on Machine Learning (ICML 2024), Vienna, Austria, 2024-07-21.

Revisiting Character-level Adversarial Attacks for Language Models

E. Abad Rocamora; Y. Wu; F. Liu; G. Chrysos; V. Cevher 

2024. 41st International Conference on Machine Learning (ICML 2024), Vienna, Austria, July 21-27, 2024.

Robust NAS under adversarial training: benchmark, theory, and beyond

Y. Wu; F. Liu; C-J. Simon-Gabriel; G. Chrysos; V. Cevher 

2024. 12th International Conference on Learning Representations (ICLR 2024), Vienna, Austria, May 7-11, 2024.

On the Convergence of Encoder-only Shallow Transformers

Y. Wu; F. Liu; G. Chrysos; V. Cevher 

2023. 37th Conference on Neural Information Processing Systems (NeurIPS 2023), New Orleans, Louisiana, USA, 2023-12-10 – 2023-12-16.

Extrapolation and Spectral Bias of Neural Nets with Hadamard Product: a Polynomial Net Study

Y. Wu; Z. Zhu; F. Liu; G. Chrysos; V. Cevher 

2022. 36th Conference on Neural Information Processing Systems (NeurIPS 2022), New Orleans, USA, 2022-11-28 – 2022-12-03.