Outer Rim Archives
Archives · 2018 · 20180096192

Application (pre-grant publication)

Systems and Methods for Identifying Objects in Media Contents

Number
20180096192
Published
2018-04-05
Filed
2016-10-04
Assignee
Disney Enterprises, Inc.
Inventors
Sigal; Leonid et al.
CPC
G06V10/75; G06V10/955; G06V20/20; G06V10/454; G06V10/82
Verdict
Set aside media content object identification, content indexing
Source
Google Patents · FreePatentsOnline

Abstract

There is provided a system comprising a memory and a processor configured to receive a plurality of images, train a plurality of independent detectors for identifying a first plurality of objects in the plurality of images based on individual attributes including a first attribute and a second attribute, train a plurality of joint detectors for identifying the first plurality of objects in the plurality of images based on composite attributes including a plurality of composite attributes each including the first attribute and the second attribute, analyze features of the plurality of images to determine a difference between a first training performance of the independent detectors and a second training performance of the joint detectors, and select, based on the analyzing, between using the independent detectors and using the joint detectors for identifying a second plurality of objects in the plurality of images using a first new attribute and a second new attribute in the attribute database.

Background

BACKGROUND

Conventional object recognition schemes enable identification of objects in images based on image attributes and object attributes. Computer vision may be trained to identify parts of an image using independent attributes, such as a noun describing an object or an adjective describing the object, or using composite attributes, such as describing an object using a noun describing the object and an adjective describing the object. However, the conventional schemes do not offer an effective method for identifying some objects. Even more, the conventional schemes are computationally expensive and prohibitively inefficient.SUMMARY

The present disclosure is directed to systems and methods for identifying objects in media contents, such as images and video contents, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.

Claims

1. A system comprising: a non-transitory memory storing an attribute database; and a hardware processor executing an executable code to: receive a plurality of images; train, using the plurality of images, a plurality of independent detectors for identifying a first plurality of objects in the plurality of images based on a first set of individual attributes in the attribute database including a first attribute and a second attribute; train, using the plurality of images, a plurality of joint detectors for identifying the first plurality of objects in the plurality of images based on a first set of composite attributes in the attribute database including a plurality of composite attributes, each of the plurality of composite attributes including the first attribute and the second attribute; analyze a set of features of the plurality of images to determine a difference between a first training performance of the plurality of independent detectors and a second training performance of the plurality of joint detectors; and select, based on the analyzing, between using the plurality of independent detectors and using the plurality of joint detectors for identifying a second plurality of objects in the plurality of images using at least a first new attribute and a second new attribute in the attribute database. 11. A method for use with a system including a non-transitory memory and a hardware processor, the method comprising: receiving, using the hardware processor, a plurality of images; training, using the hardware processor, using the plurality of images, a plurality of independent detectors for identifying a first plurality of objects in the plurality of images based on a first set of individual attributes in the attribute database including a first attribute and a second attribute; training, using the hardware processor, using the plurality of images, a plurality of joint detectors for identifying the first plurality of objects in the plurality of images based on a first set of composite attributes in the attribute database including a plurality of composite attributes, each of the plurality of composite attributes including the first attribute and the second attribute; analyzing, using the hardware processor, a set of features of the plurality of images to determine a difference between a first training performance of the plurality of independent detectors and a second training performance of the plurality of joint detectors; selecting, using the hardware processor and based on the analyzing, between using the plurality of independent detectors and using the plurality of joint detectors for identifying a second plurality of objects in the plurality of images using at least a first new attribute and a second new attribute in the attribute database.