Transcending Scaling Laws with 0.1% Extra Compute

Yi Tay, Jason Wei, Hyung Won Chung, Vinh Q. Tran 0002, David R. So, Siamak Shakeri, Xavier Garcia, Huaixiu Steven Zheng, Jinfeng Rao, Aakanksha Chowdhery, Denny Zhou, Donald Metzler, Slav Petrov, Neil Houlsby, Quoc Le, Mostafa Dehghani 0001. Transcending Scaling Laws with 0.1% Extra Compute. In Houda Bouamor, Juan Pino 0001, Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023. pages 1471-1486, Association for Computational Linguistics, 2023. [doi]

@inproceedings{TayWC0SSGZRCZMP23,
  title = {Transcending Scaling Laws with 0.1% Extra Compute},
  author = {Yi Tay and Jason Wei and Hyung Won Chung and Vinh Q. Tran 0002 and David R. So and Siamak Shakeri and Xavier Garcia and Huaixiu Steven Zheng and Jinfeng Rao and Aakanksha Chowdhery and Denny Zhou and Donald Metzler and Slav Petrov and Neil Houlsby and Quoc Le and Mostafa Dehghani 0001},
  year = {2023},
  url = {https://aclanthology.org/2023.emnlp-main.91},
  researchr = {https://researchr.org/publication/TayWC0SSGZRCZMP23},
  cites = {0},
  citedby = {0},
  pages = {1471-1486},
  booktitle = {Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023},
  editor = {Houda Bouamor and Juan Pino 0001 and Kalika Bali},
  publisher = {Association for Computational Linguistics},
  isbn = {979-8-89176-060-8},
}