N2F2: hierarchical scene understanding with nested neural feature fields

Bhalgat, YS; Laina, I; Henriques, J; Zisserman, A; Vedaldi, A

AI Collection

Conference item

N2F2: hierarchical scene understanding with nested neural feature fields

Abstract:: Understanding complex scenes at multiple levels of abstraction remains a formidable challenge in computer vision. To address this, we introduce Nested Neural Feature Fields (N2F2), a novel approach that employs hierarchical supervision to learn a single feature field, wherein different dimensions within the same high-dimensional feature encode scene properties at varying granularities. Our method allows for a flexible definition of hierarchies, tailored to either the physical dimensions or semantics or both, thereby enabling a comprehensive and nuanced understanding of scenes. We leverage a 2D class-agnostic segmentation model to provide semantically meaningful pixel groupings at arbitrary scales in the image space, and query the CLIP vision-encoder to obtain language-aligned embeddings for each of these segments. Our proposed hierarchical supervision method then assigns different nested dimensions of the feature field to distill the CLIP embeddings using deferred volumetric rendering at varying physical scales, creating a coarse-to-fine representation. Extensive experiments show that our approach outperforms the state-of-the-art feature field distillation methods on tasks such as open-vocabulary 3D segmentation and localization, demonstrating the effectiveness of the learned nested feature field.

Publication status:: Published

Peer review status:: Peer reviewed

Actions

Email

Email this record

Send the bibliographic details of this record to your email address.

Your Email
Please enter the email address that the record information will be sent to.

-
Your message (optional)
Please add any additional information to be included within the email.
Share
Cite

Cite this record

APA Style

Bhalgat, Y. S., Laina, I., Henriques, J., Zisserman, A., & Vedaldi, A. (2024). N2F2: hierarchical scene understanding with nested neural feature fields. 20th European Conference on Computer Vision (ECCV 2024), 197–214.

MLA Style

Bhalgat, YS, et al. “N2F2: Hierarchical Scene Understanding with Nested Neural Feature Fields.” 20th European Conference on Computer Vision (ECCV 2024), Lecture Notes in Computer Science, 2024, pp. 197–214.

Chicago Style

Bhalgat, YS, I Laina, J Henriques, A Zisserman, and A Vedaldi. 2024. “N2F2: Hierarchical Scene Understanding with Nested Neural Feature Fields.” In 20th European Conference on Computer Vision (ECCV 2024), 197–214. Lecture Notes in Computer Science. Springer.
Print

Access Document

Files:: Bhalgat_et_al_2024_N2F2_hierarchical_scene.pdf

(Preview, Accepted manuscript, pdf, 4.6MB, Terms of use)

Publisher copy:: 10.1007/978-3-031-73202-7_12

Authors

+ Bhalgat, YS More by this author

Institution:: University of Oxford
Division:: MPLS
Department:: Engineering Science
Role:: Author

+ Laina, I More by this author

Institution:: University of Oxford
Division:: MPLS
Department:: Engineering Science
Role:: Author

+ Henriques, J More by this author

Institution:: University of Oxford
Division:: MPLS
Department:: Engineering Science
Role:: Author

+ Zisserman, A More by this author

Institution:: University of Oxford
Division:: MPLS
Department:: Engineering Science
Role:: Author
ORCID:: 0000-0002-8945-8573

+ Vedaldi, A More by this author

Institution:: University of Oxford
Division:: MPLS
Department:: Engineering Science
Oxford college:: New College
Role:: Author
ORCID:: 0000-0003-1374-2858

+ European Research Council More from this funder

Funder identifier:: https://ror.org/0472cxd90
Grant:: 101001212

+ Engineering and Physical Sciences Research Council More from this funder

Funder identifier:: https://ror.org/0439y7842
Grant:: EP/T028572/1

Publisher:: Springer
Host title:: Computer Vision – ECCV 2024 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LIX
Pages:: 197–214
Series:: Lecture Notes in Computer Science
Series number:: 15117
Publication date:: 2024-11-21
Acceptance date:: 2024-07-01
Event title:: 20th European Conference on Computer Vision (ECCV 2024)
Event location:: Milan, Italy
Event website:: https://eccv.ecva.net/
Event start date:: 2024-09-29
Event end date:: 2024-10-04
DOI:: 10.1007/978-3-031-73202-7_12
EISSN:: 1611-3349
ISSN:: 0302-9743
EISBN:: 978-3-031-73202-7
ISBN:: 978-3-031-73201-0

Language:: English
Keywords:: feature field distillation

hierarchical scene understanding

open-vocabulary 3d segmentation
Pubs id:: 2017721
Local pid:: pubs:2017721
Deposit date:: 2024-07-22
ARK identifier:: ark:/29072/ora_cedfd58ad82e46328d7890811757ba74

Terms of use

Copyright holder:: Bhalgat et al.
Notes:: This paper was presented at the 20th European Conference on Computer Vision (ECCV 2024), 29th September - 4th October 2024, Mlian, Italy. This is the accepted manuscript version of the article. The final version is available online from Springer at https://dx.doi.org/10.1007/978-3-031-73202-7_12

Licence:: Terms and Conditions of Use for Oxford University Research Archive

Views and Downloads

About views and downloads

If you are the owner of this record, you can report an update to it here: Report update to this record

Conference item

N2F2: hierarchical scene understanding with nested neural feature fields

Actions

Access Document

Authors

Terms of use

Views and Downloads

Altmetrics

Dimensions

Conference item

N2F2: hierarchical scene understanding with nested neural feature fields

Actions

Access Document

Authors

Funding

Bibliographic Details

Item Description

Terms of use

Metrics

Views and Downloads

Altmetrics

Dimensions