Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection

Lv Tang, Peng-Tao Jiang, Zhihao Shen, Hao Zhang, Jin-Wei Chen, Bo Li 0115. Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection. In Jianfei Cai 0001, Mohan S. Kankanhalli, Balakrishnan Prabhakaran 0001, Susanne Boll, Ramanathan Subramanian, Liang Zheng 0001, Vivek K. Singh 0001, Pablo César, Lexing Xie, Dong Xu 0001, editors, Proceedings of the 32nd ACM International Conference on Multimedia, MM 2024, Melbourne, VIC, Australia, 28 October 2024 - 1 November 2024. pages 8805-8814, ACM, 2024. [doi]

@inproceedings{TangJSZC024,
  title = {Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection},
  author = {Lv Tang and Peng-Tao Jiang and Zhihao Shen and Hao Zhang and Jin-Wei Chen and Bo Li 0115},
  year = {2024},
  doi = {10.1145/3664647.3680730},
  url = {https://doi.org/10.1145/3664647.3680730},
  researchr = {https://researchr.org/publication/TangJSZC024},
  cites = {0},
  citedby = {0},
  pages = {8805-8814},
  booktitle = {Proceedings of the 32nd ACM International Conference on Multimedia, MM 2024, Melbourne, VIC, Australia, 28 October 2024 - 1 November 2024},
  editor = {Jianfei Cai 0001 and Mohan S. Kankanhalli and Balakrishnan Prabhakaran 0001 and Susanne Boll and Ramanathan Subramanian and Liang Zheng 0001 and Vivek K. Singh 0001 and Pablo César and Lexing Xie and Dong Xu 0001},
  publisher = {ACM},
  isbn = {979-8-4007-0686-8},
}