Conference item icon

Conference item

Lip reading in the wild

Abstract:

Our aim is to recognise the words being spoken by a talking face, given only the video but not the audio. Existing works in this area have focussed on trying to recognise a small number of utterances in controlled environments (e.g. digits and alphabets), partially due to the shortage of suitable datasets.


We make two novel contributions: first, we develop a pipeline for fully automated large-scale data collection from TV broadcasts. With this we have generated a dataset with over a million word instances, spoken by over a thousand different people; second, we develop CNN architectures that are able to effectively learn and recognize hundreds of words from this large-scale dataset.


We also demonstrate a recognition performance that exceeds the state of the art on a standard public benchmark dataset.

Publication status:
Published
Peer review status:
Peer reviewed

Actions

Access Document

Files:
Publisher copy:
10.1007/978-3-319-54184-6_6

Authors

More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author
More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Engineering Science
Role:
Author


Publisher:
Springer, Cham
Host title:
Asian Conference on Computer Vision: ACCV 2016: Computer Vision – ACCV 2016
Journal:
Asian Conference on Computer Vision 2016 More from this journal
Volume:
10112
Pages:
87-103
Publication date:
2017-01-01
Acceptance date:
2016-05-27
Event location:
Taipei
DOI:
ISSN:
0302-9743
ISBN:
9783319541839


Pubs id:
pubs:656449
UUID:
uuid:c3238375-ec8b-4ecd-9543-8b179a6b74ba
Local pid:
pubs:656449
Source identifiers:
656449
Deposit date:
2016-11-01
ARK identifier:

Terms of use


Views and Downloads

Views and downloads will return soon






If you are the owner of this record, you can report an update to it here: Report update to this record

TO TOP