Conference item icon

Conference item

Localizing visual sounds the hard way

Abstract:
The objective of this work is to localize sound sources that are visible in a video without using manual annotations. Our key technical contribution is to show that, by training the network to explicitly discriminate challenging image fragments, even for images that do contain the object emitting the sound, we can significantly boost the localization performance. We do so elegantly by introducing a mechanism to mine hard samples and add them to a contrastive learning formulation automatically. We show that our algorithm achieves state-of-the-art performance on the popular Flickr SoundNet dataset. Furthermore, we introduce the VGG-Sound Source (VGG-SS) benchmark, a new set of annotations for the recently-introduced VGG-Sound dataset, where the sound sources visible in each video clip are explicitly marked with bounding box annotations. This dataset is 20 times larger than analogous existing ones, contains 5K videos spanning over 200 categories, and, differently from Flickr SoundNet, is video-based. On VGG-SS, we also show that our algorithm achieves state-of-the-art performance against several baselines. Code and datasets can be found at http://www.robots.ox.ac.uk/˜vgg/research/lvs/.
Publication status:
Published
Peer review status:
Peer reviewed

Actions

Access Document

Files:
Publisher copy:
10.1109/CVPR46437.2021.01659

Authors

More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author
More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author
More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author
More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author
More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Oxford college:
Brasenose College
Role:
Author


More from this funder
Funder identifier:
http://dx.doi.org/10.13039/501100000266
Grant:
EP/M013774/1


Publisher:
IEEE
Host title:
Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR)
Pages:
16862-16871
Publication date:
2021-11-02
Acceptance date:
2021-02-28
Event title:
Conference on Computer Vision and Pattern Recognition (CVPR 2021)
Event location:
Virtual event
Event website:
http://cvpr2021.thecvf.com/
Event start date:
2021-06-19
Event end date:
2021-06-25
DOI:
EISSN:
2575-7075
ISSN:
1063-6919
EISBN:
9781665445092
ISBN:
9781665445108


Language:
English
Pubs id:
1173942
Local pid:
pubs:1173942
Deposit date:
2021-04-28
ARK identifier:

Terms of use


Views and Downloads

Views and downloads will return soon






If you are the owner of this record, you can report an update to it here: Report update to this record

TO TOP