Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection

Lv Tang, Peng-Tao Jiang, Zhihao Shen, Hao Zhang, Jin-Wei Chen, Bo Li 0115. Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection. In Jianfei Cai 0001, Mohan S. Kankanhalli, Balakrishnan Prabhakaran 0001, Susanne Boll, Ramanathan Subramanian, Liang Zheng 0001, Vivek K. Singh 0001, Pablo César, Lexing Xie, Dong Xu 0001, editors, Proceedings of the 32nd ACM International Conference on Multimedia, MM 2024, Melbourne, VIC, Australia, 28 October 2024 - 1 November 2024. pages 8805-8814, ACM, 2024. [doi]

Authors

Lv Tang

This author has not been identified. Look up 'Lv Tang' in Google

Peng-Tao Jiang

This author has not been identified. Look up 'Peng-Tao Jiang' in Google

Zhihao Shen

This author has not been identified. Look up 'Zhihao Shen' in Google

Hao Zhang

This author has not been identified. Look up 'Hao Zhang' in Google

Jin-Wei Chen

This author has not been identified. Look up 'Jin-Wei Chen' in Google

Bo Li 0115

This author has not been identified. Look up 'Bo Li 0115' in Google