Conference item
Discovering objects and their location in images
- Abstract:
- We seek to discover the object categories depicted in a set of unlabelled images. We achieve this using a model developed in the statistical text literature: probabilistic latent semantic analysis (pLSA). In text analysis, this is used to discover topics in a corpus using the bag-of-words document representation. Here we treat object categories as topics, so that an image containing instances of several categories is modeled as a mixture of topics. The model is applied to images by using a visual analogue of a word, formed by vector quantizing SIFT-like region descriptors. The topic discovery approach successfully translates to the visual domain: for a small set of objects, we show that both the object categories and their approximate spatial layout are found without supervision. Performance of this unsupervised method is compared to the supervised approach of Fergus et al. (2003) on a set of unseen images containing only one object per image. We also extend the bag-of-words vocabulary to include 'doublets' which encode spatially local co-occurring regions. It is demonstrated that this extended vocabulary gives a cleaner image segmentation. Finally, the classification and segmentation methods are applied to a set of images containing multiple objects per image. These results demonstrate that we can successfully build object class models from an unsupervised analysis of images.
- Publication status:
- Published
- Peer review status:
- Peer reviewed
Actions
Access Document
- Files:
-
-
(Preview, Accepted manuscript, pdf, 4.8MB, Terms of use)
-
- Publisher copy:
- 10.1109/ICCV.2005.77
Authors
- Publisher:
- IEEE
- Host title:
- Tenth IEEE International Conference on Computer Vision (ICCV'05) Volume 1
- Volume:
- 1
- Pages:
- 370-377
- Publication date:
- 2005-12-05
- Event title:
- 10th IEEE International Conference on Computer Vision (ICCV 2005)
- Event location:
- Beijing, China
- Event start date:
- 2005-10-17
- Event end date:
- 2005-10-21
- DOI:
- EISSN:
-
2380-7504
- ISSN:
-
1550-5499
- ISBN:
- 076952334X
- Language:
-
English
- Keywords:
- Pubs id:
-
62058
- Local pid:
-
pubs:62058
- Deposit date:
-
2024-07-25
- ARK identifier:
Terms of use
- Copyright holder:
- IEEE
- Copyright date:
- 2005
- Rights statement:
- © 2005 IEEE.
- Notes:
- This paper was presented at the 10th IEEE International Conference on Computer Vision (ICCV 2005), 17th-21st October 2005, Beijing, China. This is the accepted manuscript version of the article. The final version is available online from IEEE at: https://dx.doi.org/10.1109/ICCV.2005.77
If you are the owner of this record, you can report an update to it here: Report update to this record