Broken Tokens? Your Language Model can Secretly Handle Non-Canonical Tokenizations

Brian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase, Yejin Choi 0001, Noah A. Smith. Broken Tokens? Your Language Model can Secretly Handle Non-Canonical Tokenizations. In Danielle Belgrave, Cheng Zhang 0005, Laura N. Montoya, Hsuan-Tien Lin, Razvan Pascanu, Piotr Koniusz, Marzyeh Ghassemi, Nancy Chen, Iván Vladimir Meza Ruíz, Arturo Loaiza-Bonilla, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diago, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025. 2025. [doi]

Authors

Brian Siyuan Zheng

This author has not been identified. Look up 'Brian Siyuan Zheng' in Google

Alisa Liu

This author has not been identified. Look up 'Alisa Liu' in Google

Orevaoghene Ahia

This author has not been identified. Look up 'Orevaoghene Ahia' in Google

Jonathan Hayase

This author has not been identified. Look up 'Jonathan Hayase' in Google

Yejin Choi 0001

This author has not been identified. Look up 'Yejin Choi 0001' in Google

Noah A. Smith

This author has not been identified. Look up 'Noah A. Smith' in Google