How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation Under the One-Time-Pad-Based Framework

Zi Liang, Liantong Yu, Shiyu Zhang, Qingqing Ye 0001, Haibo Hu 0001. How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation Under the One-Time-Pad-Based Framework. In Sven Koenig, Chad Jenkins, Matthew E. Taylor, editors, Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026. pages 37636-37644, AAAI Press, 2026. [doi]

Authors

Zi Liang

This author has not been identified. Look up 'Zi Liang' in Google

Liantong Yu

This author has not been identified. Look up 'Liantong Yu' in Google

Shiyu Zhang

This author has not been identified. Look up 'Shiyu Zhang' in Google

Qingqing Ye 0001

This author has not been identified. Look up 'Qingqing Ye 0001' in Google

Haibo Hu 0001

This author has not been identified. Look up 'Haibo Hu 0001' in Google