Space-Time Crop & Attend: improving cross-modal video representation learning

Patrick, M; Huang, P-Y; Misra, I; Metze, F; Vedaldi, A; Asano, YM; Henriques, J

AI Collection

Conference item

Space-Time Crop & Attend: improving cross-modal video representation learning

Abstract:: The quality of the image representations obtained from self-supervised learning depends strongly on the type of data augmentations used in the learning formulation. Recent papers have ported these methods from still images to videos and found that leveraging both audio and video signals yields strong gains; however, they did not find that spatial augmentations such as cropping, which are very important for still images, work as well for videos. In this paper, we improve these formulations in two ways unique to the spatio-temporal aspect of videos. First, for space, we show that spatial augmentations such as cropping do work well for videos too, but that previous implementations, due to the high processing and memory cost, could not do this at a scale sufficient for it to work well. To address this issue, we first introduce Feature Crop, a method to simulate such augmentations much more efficiently directly in feature space. Second, we show that as opposed to naïve average pooling, the use of transformer-based attention improves performance significantly, and is well suited for processing feature crops. Combining both of our discoveries into a new method, Space-Time Crop & Attend (STiCA) we achieve state-of-the-art performance across multiple video-representation learning benchmarks. In particular, we achieve new state-of-the-art accuracies of 67.0% on HMDB-51 and 93.1% on UCF-101 when pre-training on Kinetics-400. Code and pretrained models are available 1 .

Publication status:: Published

Peer review status:: Peer reviewed

Actions

Email

Email this record

Send the bibliographic details of this record to your email address.

Your Email
Please enter the email address that the record information will be sent to.

-
Your message (optional)
Please add any additional information to be included within the email.
Share
Cite

Cite this record

APA Style

Patrick, M., Huang, P.-Y., Misra, I., Metze, F., Vedaldi, A., Asano, Y. M., & Henriques, J. (2022). Space-Time Crop & Attend: improving cross-modal video representation learning. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 10540–10552.

MLA Style

Patrick, M, et al. “Space-Time Crop & Attend: Improving Cross-Modal Video Representation Learning.” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2022, pp. 10540–52.

Chicago Style

Patrick, M, P-Y Huang, I Misra, F Metze, A Vedaldi, YM Asano, and J Henriques. 2022. “Space-Time Crop & Attend: Improving Cross-Modal Video Representation Learning.” In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 10540–52. IEEE.
Print

Access Document

Files:: Patrick_et_al_2021_Space_Time_crop.pdf

(Preview, pdf, 580.2KB, Terms of use)

Publisher copy:: 10.1109/iccv48922.2021.01039

Authors

+ Patrick, M More by this author

Role:: Author

+ Huang, P-Y More by this author

Role:: Author

+ Misra, I More by this author

Role:: Author

+ Metze, F More by this author

Role:: Author

+ Vedaldi, A More by this author

Institution:: University of Oxford
Division:: MPLS
Department:: Engineering Science
Role:: Author

More authors...

Publisher:: IEEE
Host title:: 2021 IEEE/CVF International Conference on Computer Vision (ICCV)
Pages:: 10540-10552
Publication date:: 2022-02-28
Acceptance date:: 2021-07-22
Event title:: 2021 IEEE/CVF International Conference on Computer Vision (ICCV)
Event location:: Virtual event
Event website:: http://iccv2021.thecvf.com/home
Event start date:: 2021-10-11
Event end date:: 2021-10-17
DOI:: 10.1109/iccv48922.2021.01039
EISSN:: 2380-7504
ISSN:: 1550-5499
EISBN:: 9781665428125
ISBN:: 9781665428132

Language:: English
Keywords:: FFR
Pubs id:: 1242667
Local pid:: pubs:1242667
Deposit date:: 2022-03-07
ARK identifier:: ark:/29072/ora_eb0cf1339ad64362ade2917075ebb727

Terms of use

Copyright holder:: IEEE
Notes:: This is the accepted manuscript version of the paper. The final version is available online from IEEE at: https://doi.org/10.1109/ICCV48922.2021.01039

Licence:: Terms and Conditions of Use for Oxford University Research Archive

Views and Downloads

About views and downloads

If you are the owner of this record, you can report an update to it here: Report update to this record

Conference item

Space-Time Crop & Attend: improving cross-modal video representation learning

Actions

Access Document

Authors

Terms of use

Views and Downloads

Altmetrics

Dimensions

Conference item

Space-Time Crop & Attend: improving cross-modal video representation learning

Actions

Access Document

Authors

Bibliographic Details

Item Description

Terms of use

Metrics

Views and Downloads

Altmetrics

Dimensions