Understanding the Failure of Batch Normalization for Transformers in NLP

Jiaxi Wang, Ji Wu, Lei Huang. Understanding the Failure of Batch Normalization for Transformers in NLP. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, A. Oh, editors, Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022. 2022. [doi]

Abstract

Abstract is missing.