CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs

Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp Koehn. CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs. In Bonnie Webber, Trevor Cohn, Yulan He, Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020. pages 5960-5969, Association for Computational Linguistics, 2020. [doi]

Abstract

Abstract is missing.