Outer Rim Archives
Archives · 2018 · 20180189569

Application (pre-grant publication)

Systems and Methods for Identifying Activities and/or Events in Media Contents Based on Object Data and Scene Data

Number
20180189569
Published
2018-07-05
Filed
2018-02-28
Assignee
Disney Enterprises, Inc.
Inventors
Wu; Zuxuan et al.
CPC
G06N3/0442; G06N3/0464; G06N3/09; G06T7/62; G06T7/90; G06V10/454; G06V20/20; G06V20/41; G06V20/46; G06V20/47; G06V40/20; G11B27/102; H04L65/61
Verdict
Set aside media content activity/event identification (PGPUB dup)
Source
Google Patents · FreePatentsOnline

Abstract

There is provided a system including a non-transitory memory storing an executable code and a hardware processor executing the executable code to receive a plurality of training contents depicting a plurality of activities, extract training object data from the plurality of training contents including a first training object data corresponding to a first activity, extract training scene data from the plurality of training contents including a first training scene data corresponding to the first activity, determine that a probability of the first activity is maximized when the first training object data and the first training scene data both exist in a sample media content.

Background

BACKGROUND

Video content has become a part of everyday life with an increasing amount of video content becoming available online, and people spending an increasing amount of time online. Additionally, individuals are able to create and share video content online using video sharing websites and social media. Recognizing visual contents in unconstrained videos has found a new importance in many applications, such as video searches on the Internet, video recommendations, smart advertising, etc. Conventional approaches to content identification rely on manual annotations of video contents, and supervised computer recognition and categorization. However, manual annotations and supervised computer processing are time consuming and expensive.SUMMARY

The present disclosure is directed to systems and methods for identifying activities and/or events in media contents based on object data and scene data, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.

Claims

21. A system comprising: a non-transitory memory storing an executable code including an object data module, a scene data module, an image data module, and a plurality of fusion layers of a neural network; and a hardware processor to: receive a media content having video frames; extract object data, by executing the object data module, from the video frames of the media content; generate an average object data representation of the extracted object data; extract scene data, by executing the scene data module, from the frames of the video media content; generate an average scene data representation of the extracted scene data; extract image data, by executing the image data module, from the frames of the video media content; generate an average image data representation of the extracted image data; feed the average object data representation, the average scene data representation and the average image data representation to the plurality of fusion layers of the neural network; and classify an action in the video frames of the media content by executing the plurality of fusion layers using the average object data representation, the average scene data representation and the average image data representation. 30. A method for use with a system including a hardware processor and a non-transitory memory storing an executable code including an object data module, a scene data module, an image data module, and a plurality of fusion layers of a neural network, the method comprising: receiving, using the hardware processor, a media content having video frames; extracting object data, by executing the object data module using the hardware processor, from the video frames of the media content; generating, using the hardware processor, an average object data representation of the extracted object data; extracting scene data, by executing the scene data module using the hardware processor, from the frames of the video media content; generating, using the hardware processor, an average scene data representation of the extracted scene data; extracting image data, by executing the image data module using the hardware processor, from the frames of the video media content; generating, using the hardware processor, an average image data representation of the extracted image data; feeding, using the hardware processor, the average object data representation, the average scene data representation and the average image data representation to the plurality of fusion layers of the neural network; and classifying, by executing the plurality of fusion layers using the hardware processor, an action in the video frames of the media content using the average object data representation, the average scene data representation and the average image data representation.

Claims truncated at the source; see the full document.