Reuben Tan, Bryan A. Plummer, Kate Saenko, Hailin Jin, Bryan Russell. Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos. In Marc'Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, Jennifer Wortman Vaughan, editors, Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual. pages 14476-14487, 2021. [doi]
Abstract is missing.