Reducing Transformer Depth on Demand with Structured Dropout

Angela Fan, Edouard Grave, Armand Joulin. Reducing Transformer Depth on Demand with Structured Dropout. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. [doi]

Authors

Angela Fan

This author has not been identified. Look up 'Angela Fan' in Google

Edouard Grave

This author has not been identified. Look up 'Edouard Grave' in Google

Armand Joulin

This author has not been identified. Look up 'Armand Joulin' in Google