Inductive visual localisation: factorised training for superior generalisation

Gupta, A; Vedaldi, A; Zisserman, A

Internet publication

Inductive visual localisation: factorised training for superior generalisation

Abstract:: End-to-end trained Recurrent Neural Networks (RNNs) have been successfully applied to numerous problems that require processing sequences, such as image captioning, machine translation, and text recognition. However, RNNs often struggle to generalise to sequences longer than the ones encountered during training. In this work, we propose to optimise neural networks explicitly for induction. The idea is to first decompose the problem in a sequence of inductive steps and then to explicitly train the RNN to reproduce such steps. Generalisation is achieved as the RNN is not allowed to learn an arbitrary internal state; instead, it is tasked with mimicking the evolution of a valid state. In particular, the state is restricted to a spatial memory map that tracks parts of the input image which have been accounted for in previous steps. The RNN is trained for single inductive steps, where it produces updates to the memory in addition to the desired output. We evaluate our method on two different visual recognition problems involving visual sequences: (1) text spotting, i.e. joint localisation and reading of text in images containing multiple lines (or a block) of text, and (2) sequential counting of objects in aerial images. We show that inductive training of recurrent models enhances their generalisation ability on challenging image datasets.

Publication status:: Published

Peer review status:: Not peer reviewed

Actions

Email

Email this record

Send the bibliographic details of this record to your email address.

Your Email
Please enter the email address that the record information will be sent to.

-
Your message (optional)
Please add any additional information to be included within the email.
Cite

Cite this record

APA Style

Gupta, A., Vedaldi, A., & Zisserman, A. (2018). Inductive visual localisation: factorised training for superior generalisation.

MLA Style

Gupta, A., et al. Inductive Visual Localisation: Factorised Training for Superior Generalisation. 2018.

Chicago Style

Gupta, A, A Vedaldi, and A Zisserman. 2018. “Inductive Visual Localisation: Factorised Training for Superior Generalisation.”
Share
Print

Access Document

Files:: Gupta_et_al_2018_Inductive_visual_localisation.pdf

(Preview, Version of record, pdf, 1.9MB, Terms of use)

Publisher copy:: 10.48550/arxiv.1807.08179

Authors

+ Gupta, A More by this author

Institution:: University of Oxford
Division:: MPLS
Department:: Engineering Science
Role:: Author

+ Vedaldi, A More by this author

Institution:: University of Oxford
Division:: MPLS
Department:: Engineering Science
Oxford college:: New College
Role:: Author
ORCID:: 0000-0003-1374-2858

+ Zisserman, A More by this author

Institution:: University of Oxford
Division:: MPLS
Department:: Engineering Science
Oxford college:: Brasenose College
Role:: Author
ORCID:: 0000-0002-8945-8573

+ Engineering and Physical Sciences Research Council More from this funder

Funder identifier:: https://ror.org/0439y7842
Grant:: EP/M013774/1; EP/L015987/2

Host title:: arXiv
Publication date:: 2018-07-21
DOI:: 10.48550/arxiv.1807.08179

Language:: English
Pubs id:: 1771184
Local pid:: pubs:1771184
Deposit date:: 2024-07-11

Terms of use

Copyright holder:: Gupta et al.
Notes:: The final, peer-reviewed version of this paper was published in the Proceedings of the British Machine Vision Conference 2018 and is available in ORA at: https://ora.ox.ac.uk/objects/uuid:c110d333-44c3-4e77-ac0b-01259881dc61

Licence:: Terms and Conditions of Use for Oxford University Research Archive

Views and Downloads

About views and downloads

If you are the owner of this record, you can report an update to it here: Report update to this record

Internet publication

Inductive visual localisation: factorised training for superior generalisation

Actions

Access Document

Authors

Terms of use

Views and Downloads

Altmetrics

Dimensions

Internet publication

Inductive visual localisation: factorised training for superior generalisation

Actions

Access Document

Authors

Funding

Bibliographic Details

Item Description

Terms of use

Metrics

Views and Downloads

Altmetrics

Dimensions