Xubo Lin, Mingze Wang, Grace Hui Yang, Daniel Chen. Reward-on-the-Line: A Novel Offline Reinforcement Learning Method for Building Legal Conversational Agents. In Emanuelle Burton, Nicholas Mattei, Andrés Páez, editors, Proceedings of the Eighth AAAI/ACM Conference on AI, Ethics, and Society, AIES 2025, Madrid, Spain, October 20-22, 2025, Main Track II. pages 1575-1584, AAAI Press, 2025. [doi]