Outer Rim Archives
Archives · 2018 · 20180286081

Application (pre-grant publication)

OBJECT RE-IDENTIFICATION WITH TEMPORAL CONTEXT

Number
20180286081
Published
2018-10-04
Filed
2017-03-31
Assignee
Disney Enterprises, Inc.
Inventors
KOPERSKI; Michal; BAK; Slawomir W.; CARR; G. Peter K.
CPC
G06V20/52; G06V10/761; G06T7/292
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

CV object re-identification (temporal).

Abstract

Techniques for object re-identification based on temporal context. Embodiments extract, from a first image corresponding to a first camera device and a second image corresponding to a second camera device, a first plurality of patch descriptors and a second plurality of patch descriptors, respectively. A measure of visual similarity between the first image and the second image is computed, based on the first plurality of patch descriptors and the second plurality of patch descriptors. A temporal cost between the first image and the second image is computed, based on a first timestamp at which the first image was captured and a second timestamp at which the second image was captured. The measure of visual similarity and the temporal cost are combined into a single cost function, and embodiments determine whether the first image and the second image depict a common object, using the single cost function.

Background

BACKGROUNDField of the Invention

The present disclosure relates to digital image processing, and more specifically, to techniques for unsupervised learning for recognizing objects appearing across multiple, non-overlapping images.Description of the Related Art

Object recognition is useful in many different fields. However, many conventional techniques require affirmative interaction with one or more electronic devices, and thus are ill-suited for situations in which objects such as people, animals, vehicles and other things are moving quickly through a location (e.g., an airport). Image processing techniques can be used to recognize objects within frames captured by a camera. One challenge when detecting objects across multiple cameras is handling variations in lighting, camera position and object pose across the cameras. For instance, an object may have a certain set of color characteristics as in an image captured by a first camera, and may be captured with different color characteristics in an image captured by a different camera. As a result, conventional techniques for determining that an object captured by multiple cameras are indeed the same object may be inaccurate.SUMMARY

Embodiments provide a method, system and computer readable medium for re-identifying objects depicted in images captured by two non-overlapping cameras. A first image corresponding to a first camera and a second image corresponding to a second camera, a first plurality of patch des

Claims

1. A method, comprising: extracting, from a first image corresponding to a first camera device and a second image corresponding to a second camera device, a first plurality of patch descriptors and a second plurality of patch descriptors, respectively; computing a measure of visual similarity between the first image and the second image, based on the first plurality of patch descriptors and the second plurality of patch descriptors; computing a temporal cost between the first image and the second image, based on a first timestamp at which the first image was captured and a second timestamp at which the second image was captured; combining the measure of visual similarity and the temporal cost into a single cost function; and determining whether the first image and the second image depict a common object, using the single cost function. 17. A system, comprising: one or more computer processors; and a non-transitory memory containing computer program code that, when executed by operation of the one or more computer processors, performs an operation comprising: extracting, from a first image corresponding to a first camera device and a second image corresponding to a second camera device, a first plurality of patch descriptors and a second plurality of patch descriptors, respectively; computing a measure of visual similarity between the first image and the second image, based on the first plurality of patch descriptors and the second plurality of patch descriptors; computing a temporal cost between the first image and the second image, based on a first timestamp at which the first image was captured and a second timestamp at which the second image was captured; combining the measure of visual similarity and the temporal cost into a single cost function; and determining whether the first image and the second image depict a common object, using the single cost function. 20. A non-transitory computer-readable medium containing computer program code that, when executed by operation of one or more computer processors, performs an operation comprising: extracting, from a first image corresponding to a first camera device and a second image corresponding to a second camera device, a first plurality of patch descriptors and a second plurality of patch descriptors, respectively; computing a measure of visual similarity between the first image and the second image, based on the first plurality of patch descriptors and the second plurality of patch descriptors; computing a temporal cost between the first image and the second image, based on a first timestamp at which the first image was captured and a second timestamp at which the second image was captured; combining the measure of visual similarity and the temporal cost into a single cost function; and determining whether the first image and the second image depict a common object, using the single cost function.