Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs

Xiaozhe Li, XinYu Fang, Shengyuan Ding, Yang Li 0189, Linyang Li, Haodong Duan, Qingwen Liu 0001, Kai Chen 0026. Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs. In Maria Liakata, Viviane P. Moreira, Jiajun Zhang 0001, David Jurgens, editors, Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026. pages 28351-28368, Association for Computational Linguistics, 2026. [doi]

Authors

Xiaozhe Li

This author has not been identified. Look up 'Xiaozhe Li' in Google

XinYu Fang

This author has not been identified. Look up 'XinYu Fang' in Google

Shengyuan Ding

This author has not been identified. Look up 'Shengyuan Ding' in Google

Yang Li 0189

This author has not been identified. Look up 'Yang Li 0189' in Google

Linyang Li

This author has not been identified. Look up 'Linyang Li' in Google

Haodong Duan

This author has not been identified. Look up 'Haodong Duan' in Google

Qingwen Liu 0001

This author has not been identified. Look up 'Qingwen Liu 0001' in Google

Kai Chen 0026

This author has not been identified. Look up 'Kai Chen 0026' in Google