There is provided a system including a memory and a processor to receive images. The processor is further to train independent detectors for identifying first objects in the images based on individual attributes including a first attribute and a second attribute. The processor is also to train joint detectors for identifying first objects in the images based on composite attributes including a composite attributes each including the first attribute and the second attribute. The processor is to analyze features of the images to determine a difference between a first training performance of the independent detectors and a second training performance of the joint detectors. Lastly, the processor is to select, based on the analyzing, between using the independent detectors and using the joint detectors for identifying second objects in the images using a third attribute and a fourth attribute in the attribute database.
BACKGROUND(1) Conventional object recognition schemes enable identification of objects in images based on image attributes and object attributes. Computer vision may be trained to identify parts of an image using independent attributes, such as a noun describing an object or an adjective describing the object, or using composite attributes, such as describing an object using a noun describing the object and an adjective describing the object. However, the conventional schemes do not offer an effective method for identifying some objects. Even more, the conventional schemes are computationally expensive and prohibitively inefficient.SUMMARY(2) The present disclosure is directed to systems and methods for identifying objects in media contents, such as images and video contents, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.
1. A system comprising: a non-transitory memory storing an attribute database; and a hardware processor executing an executable code to: receive a plurality of images; train, using the plurality of images, a plurality of independent detectors for identifying a first plurality of objects in the plurality of images based on a first set of individual attributes in the attribute database including a first attribute and a second attribute; train, using the plurality of images, a plurality of joint detectors for identifying the first plurality of objects in the plurality of images based on a first set of composite attributes in the attribute database including a plurality of composite attributes, each of the plurality of composite attributes including the first attribute and the second attribute; analyze a set of features of the plurality of images to determine a difference between a first training performance of the plurality of independent detectors and a second training performance of the plurality of joint detectors; and select, based on the analyzing, between using the plurality of independent detectors and using the plurality of joint detectors for identifying a second plurality of objects in the plurality of images using at least a first new attribute and a second new attribute in the attribute database.
11. A method for use with a system including a non-transitory memory and a hardware processor, the method comprising: receiving, using the hardware processor, a plurality of images; training, using the hardware processor, using the plurality of images, a plurality of independent detectors for identifying a first plurality of objects in the plurality of images based on a first set of individual attributes in the attribute database including a first attribute and a second attribute; training, using the hardware processor, using the plurality of images, a plurality of joint detectors for identifying the first plurality of objects in the plurality of images based on a first set of composite attributes in the attribute database including a plurality of composite attributes, each of the plurality of composite attributes including the first attribute and the second attribute; analyzing, using the hardware processor, a set of features of the plurality of images to determine a difference between a first training performance of the plurality of independent detectors and a second training performance of the plurality of joint detectors; selecting, using the hardware processor and based on the analyzing, between using the plurality of independent detectors and using the plurality of joint detectors for identifying a second plurality of objects in the plurality of images using at least a first new attribute and a second new attribute in the attribute database.