Tricky²: Towards a Benchmark for Evaluating Human and LLM Error Interactions

Cole Granger, Dipin Khati, Daniel Rodríguez-Cárdenas, Denys Poshyvanyk. Tricky²: Towards a Benchmark for Evaluating Human and LLM Error Interactions. In Proceedings of the 2026 IEEE/ACM Third International Conference on AI Foundation Models and Software Engineering, FORGE 2026, Rio de Janeiro, Brazil, April 12-13, 2026. pages 249-253, ACM, 2026. [doi]

Abstract

Abstract is missing.