Conference item
Automatic dense annotation of large-vocabulary sign language videos
- Abstract:
- Recently, sign language researchers have turned to sign language interpreted TV broadcasts, comprising (i) a video of continuous signing and (ii) subtitles corresponding to the audio content, as a readily available and large-scale source of training data. One key challenge in the usability of such data is the lack of sign annotations. Previous work exploiting such weakly-aligned data only found sparse correspondences between keywords in the subtitle and individual signs. In this work, we propose a simple, scalable framework to vastly increase the density of automatic annotations. Our contributions are the following: (1) we significantly improve previous annotation methods by making use of synonyms and subtitle-signing alignment; (2) we show the value of pseudo-labelling from a sign recognition model as a way of sign spotting; (3) we propose a novel approach for increasing our annotations of known and unknown classes based on in-domain exemplars; (4) on the BOBSL BSL sign language corpus, we increase the number of confident automatic annotations from 670K to 5M. We make these annotations publicly available to support the sign language research community.
- Publication status:
- Published
- Peer review status:
- Peer reviewed
Actions
Access Document
- Files:
-
-
(Preview, Accepted manuscript, pdf, 28.0MB, Terms of use)
-
- Publication website:
- https://www.ecva.net/papers/eccv_2022/papers_ECCV/html/1028_ECCV_2022_paper.php
Authors
- Publisher:
- European Computer Vision Association
- Host title:
- ECCV Conference Papers
- Publication date:
- 2022-11-22
- Acceptance date:
- 2022-07-03
- Event title:
- European Conference on Computer Vision (ECCV 2022)
- Event location:
- Tel Aviv
- Event website:
- https://eccv2022.ecva.net/
- Event start date:
- 2022-10-23
- Event end date:
- 2022-10-27
- Language:
-
English
- Keywords:
- Pubs id:
-
1277875
- Local pid:
-
pubs:1277875
- Deposit date:
-
2022-09-07
- ARK identifier:
Terms of use
- Copyright holder:
- Momeni et al.
- Copyright date:
- 2022
- Rights statement:
- © The Authors 2022.
- Notes:
- This paper was presented at the European Conference on Computer Vision (ECCV 2022), 23-27 October 2022, Tel Aviv. This is the accepted manuscript version of the paper.
If you are the owner of this record, you can report an update to it here: Report update to this record