Outer Rim Archives
Archives · 2017 · 20170308754

Application (pre-grant publication)

Systems and Methods for Determining Actions Depicted in Media Contents Based on Attention Weights of Media Content Frames

Number
20170308754
Published
2017-10-26
Filed
2016-07-05
Assignee
Disney Enterprises, Inc.
Inventors
Torabi; Atousa et al.
CPC
G06N3/0442; G06N3/0464; G06N3/09; G06T7/62; G06T7/90; G06V10/454; G06V20/20; G06V20/41; G06V20/46; G06V20/47; G06V40/20; G11B27/102; H04L65/61
Verdict
Set aside media content action determination - content analytics
Source
Google Patents · FreePatentsOnline

Abstract

There isprovided a system comprising a label database including a plurality of label, a non-transitory memory storing an executable code, and a hardware processor executing the executablecode to receive a media content including a plurality of segments, each segment includinga plurality of frames, extract a first plurality of features from a segment, extract a second plurality of features from each frame of the segment, determine an attention weight for each frame of the segment based on the first plurality of features extracted from the segment and the second plurality of features extracted from the segment, and determine thatthe segment depicts one of the plurality of labels in a label database based on the firstplurality of features, the second plurality of features, and the attention weight of eachframe of the plurality of frames of the segment.

Background

BRIEF DESCRIPTION OF THE DRAWINGS

FIG. 1shows a diagram of an exemplary system for determining actions depicted in media contentsbased on attention weights of media content frames, according to one implementation of the present disclosure;

FIG. 2a shows a plurality of exemplary frames of media contents that are analyzed using the system of FIG. 1, according to one implementation of the present disclosure;

FIG. 2b shows another plurality of exemplary frames of media contents that are analyzed using the system of FIG. 1, according to one implementation of the present disclosure;

FIG. 2c shows another plurality of exemplary frames of media contents for that are analyzed using the system of FIG. 1, according to one implementation of the present disclosure;

FIG. 2d shows another plurality of exemplary frames of media contents that are analyzed using the system of FIG. 1, according to one implementation of the present disclosure;

FIG. 2e shows another plurality of exemplary frames of mediacontents that are analyzed using the system of FIG. 1, according to one implementation ofthe present disclosure;

FIG. 2f shows another plurality of exemplary frames of media contents that are analyzed using the system of FIG. 1, according to one implementation of the present disclosure;

FIG. 3 shows a diagram of exemplary frames of a media content analyzed using the system of FIG. 1, according to one implementation of the present disclosure;[0012

Claims

1. A system comprising: a label database including a plurality of labels; a non-transitory memory storing an executablecode; and a hardware processor executing the executable code to: receive a media content including a plurality of segments, each segment of the plurality of segments including a plurality of frames; extract a first plurality of features from a segment of the plurality of segments; extract a second plurality of features from each frame of the plurality of frames of the segment; determine an attention weight for each frame of the plurality of frames of the segment based on the first plurality of features extracted from the segment and the second plurality of features extracted from each frame of the plurality of frames of the segment; and determine that the first segment depicts one of the plurality of labels inthe label database based on the first plurality of features, the second plurality of features, and the attention weight of each frame of the plurality of frames of the segment. 11. A method for use with a system including a non-transitory memory and a hardware processor, the method comprising: receiving, using the hardware processor, a media content including a plurality of segments, each segment of the plurality of segments including a plurality of frames; extracting, using the hardware processor, a first plurality of features from a segment of the plurality of segments; extracting, using the hardware processor, a secondplurality of features from each frame of the plurality of frames of the segment; determining, using the hardware processor, an attention weight for each frame of the plurality of frames of the segment based on the first plurality of features extracted from the segment and the second plurality of features extracted from each frame of the plurality of frames of the segment; and determining, using the hardware processor, that the segment depicts one of the plurality of labels in the label database based on the first plurality of features, the second plurality of features, and the attention weight of each frame of the plurality of frames of the segment.