- Number
- 20180276499
- Published
- 2018-09-27
- Filed
- 2017-03-24
- Assignee
- Disney Enterprises, Inc.
- Inventors
- BAK; Slawomir W.; CARR; G. Peter K.
- CPC
- G06V20/52; G06V10/56
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
CV metric learning for object re-ID.
Abstract
Techniques for detecting objects across images captured by camera devices. Embodiments capture, using first and second camera devices, first and second pluralities of images, respectively. First and second reference images are captured using the first and second camera devices. Color descriptors are extracted from the first plurality of images and the second plurality of images, and texture descriptors are extracted from the first plurality of images and the second plurality of images. Embodiments model a first color subspace and a second color subspace for the first camera device and the second camera device, respectively, based on the first and second pluralities of images and the first and second reference images. A data model for identifying objects appearing in images captured using the first and second camera devices is generated, based on the extracted color descriptors, texture descriptors and the first and second color subspaces.
Background
BACKGROUNDField of the Invention
The present disclosure relates to digital video processing, and more specifically, to techniques for unsupervised learning for recognizing objects appearing across multiple, non-overlapping video streams.Description of the Related Art
Object identification is useful in many different fields. However, many conventional techniques require the individual objects be affirmatively registered with the recognition system, and thus are ill-suited for situations in which many objects (such as people, animals, devices. vehicles and other things) are moving quickly through a location (e.g., an airport). Image processing techniques can be used to recognize objects within frames captured by a camera device. One challenge when detecting objects across multiple camera devices is handling variations in lighting, camera position and object pose across multiple camera devices. For instance, an object may have a certain set of color characteristics in an image captured by a first camera device as a function of the ambient light temperature and direction as well as the first camera device's image sensor and processing electronics, and may be captured with different color characteristics in an image captured by a different camera device. As a result, conventional techniques for determining that objects captured across images from multiple camera devices are indeed the same object may be inaccurate.SUMMARY
Embodiments provide a method, system, and
Claims
1. A method, comprising: capturing, using a first camera device and a second camera device, a first plurality of images and a second plurality of images, respectively; capturing, using the first and second camera devices, a first reference image and a second reference image; extracting color descriptors from the first plurality of images and the second plurality of images; extracting texture descriptors from the first plurality of images and the second plurality of images; modelling a first color subspace and a second color subspace for the first camera device and the second camera device, respectively, based on the first and second pluralities of images and the first and second reference images; and generating a data model for identifying objects appearing in images captured using the first and second camera devices, based on the extracted color descriptors, texture descriptors and the first and second color subspaces.
12. A system, comprising: one or more computer processors; and a memory containing computer program code that, when executed by operation of the one or more computer processors, performs an operation comprising: capturing, using a first camera device and a second camera device, a first plurality of images and a second plurality of images, respectively; capturing, using the first and second camera devices, a first reference image and a second reference image; extracting color descriptors from the first plurality of images and the second plurality of images; extracting texture descriptors from the first plurality of images and the second plurality of images; modelling a first color subspace and a second color subspace for the first camera device and the second camera device, respectively, based on the first and second pluralities of images and the first and second reference images; and generating a data model for identifying objects appearing in images captured using the first and second camera devices, based on the extracted color descriptors, texture descriptors and the first and second color subspaces.
14. The system of claim, 13 wherein generating the data model further comprises: apply a first metric learning algorithm to the color descriptors extracted from the first plurality of images, using the first color subspace.
20. A non-transitory computer-readable medium containing computer program code that, when executed by operation of one or more computer processors, performs an operation comprising: capturing, using a first camera device and a second camera device, a first plurality of images and a second plurality of images, respectively; capturing, using the first and second camera devices, a first reference image and a second reference image; extracting color descriptors from the first plurality of images and the second plurality of images; extracting texture descriptors from the first plurality of images and the second plurality of images; modelling a first color subspace and a second color subspace for the first camera device and the second camera device, respectively, based on the first and second pluralities of images and the first and second reference images; and generating a data model for identifying objects appearing in images captured using the first and second camera devices, based on the extracted color descriptors, texture descriptors and the first and second color subspaces.