Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling

Pankayaraj Pathmanathan, Furong Huang. Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling. In Maria Liakata, Viviane P. Moreira, Jiajun Zhang 0001, David Jurgens, editors, Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026. pages 9230-9263, Association for Computational Linguistics, 2026. [doi]

Authors

Pankayaraj Pathmanathan

This author has not been identified. Look up 'Pankayaraj Pathmanathan' in Google

Furong Huang

This author has not been identified. Look up 'Furong Huang' in Google