OmniCode: A Benchmark for Evaluating Software Development Agents

Atharv Sonwane, Eng-Shen Tu, Wei-Chung Lu, Claas Beger, Carter Larsen, Debjit Dhar, Simon Alford, Rachel Chen, Ronit Pattanayak, Tuan Anh Dang, Guohao Chen, Gloria Geng, Kevin Ellis, Saikat Dutta 0001. OmniCode: A Benchmark for Evaluating Software Development Agents. In Maria Liakata, Viviane P. Moreira, Jiajun Zhang 0001, David Jurgens, editors, Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026. pages 40634-40661, Association for Computational Linguistics, 2026. [doi]

Authors

Atharv Sonwane

This author has not been identified. Look up 'Atharv Sonwane' in Google

Eng-Shen Tu

This author has not been identified. Look up 'Eng-Shen Tu' in Google

Wei-Chung Lu

This author has not been identified. Look up 'Wei-Chung Lu' in Google

Claas Beger

This author has not been identified. Look up 'Claas Beger' in Google

Carter Larsen

This author has not been identified. Look up 'Carter Larsen' in Google

Debjit Dhar

This author has not been identified. Look up 'Debjit Dhar' in Google

Simon Alford

This author has not been identified. Look up 'Simon Alford' in Google

Rachel Chen

This author has not been identified. Look up 'Rachel Chen' in Google

Ronit Pattanayak

This author has not been identified. Look up 'Ronit Pattanayak' in Google

Tuan Anh Dang

This author has not been identified. Look up 'Tuan Anh Dang' in Google

Guohao Chen

This author has not been identified. Look up 'Guohao Chen' in Google

Gloria Geng

This author has not been identified. Look up 'Gloria Geng' in Google

Kevin Ellis

This author has not been identified. Look up 'Kevin Ellis' in Google

Saikat Dutta 0001

This author has not been identified. Look up 'Saikat Dutta 0001' in Google