CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Alignment - researchr publication

researchr

You are not signed in
Sign in
Sign up

Hongwei Xue, Yuchong Sun, Bei Liu 0001, Jianlong Fu, Ruihua Song, Houqiang Li, Jiebo Luo. CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Alignment. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. [doi]

Abstract is missing.

runs on WebDSL