- Number
- 11055537
- Published
- 2021-07-06
- Filed
- 2016-07-05
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Torabi; Atousa, Sigal; Leonid
- CPC
- G06N3/0442; G06N3/0464; G06N3/09; G06T7/62; G06T7/90; G06V10/454; G06V20/20; G06V20/41; G06V20/46; G06V20/47; G06V40/20; G11B27/102; H04L65/61
- Verdict
- Set aside video action recognition for content indexing, analytics/business
- Source
- Google Patents · FreePatentsOnline
Abstract
There is provided a system comprising a label database including a plurality of label, a non-transitory memory storing an executable code, and a hardware processor executing the executable code to receive a media content including a plurality of segments, each segment including a plurality of frames, extract a first plurality of features from a segment, extract a second plurality of features from each frame of the segment, determine an attention weight for each frame of the segment based on the first plurality of features extracted from the segment and the second plurality of features extracted from the segment, and determine that the segment depicts one of the plurality of labels in a label database based on the first plurality of features, the second plurality of features, and the attention weight of each frame of the plurality of frames of the segment.
Background
BACKGROUND (1) Conventional action recognition in videos relies on manual annotations of the video frames and/or trained motion tracking, and may be combined with object recognition methods. The dramatic increase in video content, such as amateur videos posted on the Internet, has led to a large number of videos that are not annotated for activity recognition, and may include footage and perspective shifts that make conventional action recognition programs ineffective. Additionally, scene changes, commercial interruptions, and other shifts the video content may make identification of activities in a video difficult using conventional methods. SUMMARY (2) The present disclosure is directed to systems and methods for determining actions depicted in media contents based on attention weights of media content frames, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.
Claims
1. A system comprising: a label database including a plurality of labels; a non-transitory memory storing an executable code; and a hardware processor executing the executable code to: receive a media content including a plurality of segments, each segment of the plurality of segments including a plurality of frames; extract segment data from a segment of the plurality of segments; extract frame data from each frame of the plurality of frames of the segment; determine an object attention weight, a scene attention weight, and an action attention weight for each frame of the plurality of frames of the segment based on the segment data extracted from the segment and the frame data extracted from each frame of the plurality of frames of the segment; determine that the segment depicts a label of the plurality of labels in the label database based on the segment data, the frame data, the object attention weight, the scene attention weight, and the action attention weight of each frame of the plurality of frames of the segment; tag the segment with the label; receive a user input from a user device; and perform an act on the media content based on the user input and the label. ||
11. A method for use with a system including a non-transitory memory and a hardware processor, the method comprising: receiving, using the hardware processor, a media content including a plurality of segments, each segment of the plurality of segments including a plurality of frames; extracting, using the hardware processor, segment data from a segment of the plurality of segments; extracting, using the hardware processor, frame data from each frame of the plurality of frames of the segment; determining, using the hardware processor, an object attention weight, a scene attention weight, and an action attention weight for each frame of the plurality of frames of the segment based on the segment data extracted from the segment and the frame data extracted from each frame of the plurality of frames of the segment; determining, using the hardware processor, that the segment depicts a label of the plurality of labels in the label database based on the segment data, the frame data, the object attention weight, a scene attention weight, and an action attention weight of each frame of the plurality of frames of the segment; tagging, using the hardware processor, the segment with the label; receiving, using the hardware processor, a user input from a user device; and performing, using the hardware processor, an act on the media content based on the user input and the label.