Outer Rim Archives
Archives · 2018 · 10061986

Granted patent

Systems and methods for identifying activities in media contents based on prediction confidences

Number
10061986
Published
2018-08-28
Filed
2016-07-14
Assignee
Disney Enterprises, Inc.
Inventors
Sigal; Leonid et al.
CPC
G06N3/0442; G06N3/0464; G06N3/09; G06T7/62; G06T7/90; G06V10/454; G06V20/20; G06V20/41; G06V20/46; G06V20/47; G06V40/20; G11B27/102; H04L65/61
Verdict
Set aside media content activity identification, content indexing
Source
Google Patents · FreePatentsOnline

Abstract

There is provided a system comprising a memory and a processor configured to receive a media content depicting an activity, extract a first plurality of features from a first segment of the media content, make a first prediction that the media content depicts a first activity based on the first plurality of features, wherein the first prediction has a first confidence level, extract a second plurality of features from a second segment of the media content, the second segment temporally following the first segment in the media content, make a second prediction that the media content depicts the first activity based on the second plurality of features, wherein the second prediction has a second confidence level, determine that the media content depicts the first activity based on the first prediction and the second prediction, wherein the second confidence level is at least as high as the first confidence level.

Background

BACKGROUND(1) Video content has become a part of everyday life with an increasing amount of video content becoming available online, and people spending an increasing amount of time online. Additionally, individuals are able to create and share video content online using video sharing websites and social media. Recognizing visual contents in unconstrained videos has found a new importance in many applications, such as video searches on the Internet, video recommendations, smart advertising, etc. However, conventional approaches for recognizing activities in videos do not properly and efficiently make predictions of the activities depicted in the videos.SUMMARY(2) The present disclosure is directed to systems and methods for identifying activities in media contents based on prediction confidences, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.

Claims

1. A system comprising: a non-transitory memory storing an executable code and an activity database; and a hardware processor executing the executable code to: receive a media content including a plurality of segments depicting an activity; extract a first plurality of features from a first segment of the plurality of segments; make a first prediction that the media content depicts a first activity from the activity database based on the first plurality of features, wherein the first prediction has a first confidence level; extract a second plurality of features from a second segment of the plurality of segments, the second segment temporally following the first segment in the media content; make a second prediction that the media content depicts the first activity based on the second plurality of features, wherein the second prediction has a second confidence level; make a third prediction that the media content depicts a second activity, the third prediction having a third confidence level based on the first plurality of features and a fourth confidence level based on the second plurality of features; compare a first difference between the first confidence level and the third confidence level with a second difference between the second confidence level and the fourth confidence level; and determine that the media content depicts the first activity based on when the comparing indicates that the second difference is at least as much as the first difference based on, wherein the second confidence level is at least as high as the first confidence level. 11. A method for use with a system including a non-transitory memory and a hardware processor, the method comprising: receiving, using the hardware processor, a media content including a plurality of segments depicting an activity; extracting, using the hardware processor, a first plurality of features from a first segment of the plurality of segments; making, using the hardware processor, a first prediction that the media content depicts a first activity from the activity database based on the first plurality of features, wherein the first prediction has a first confidence level; extracting, using the hardware processor, a second plurality of features from a second segment of the plurality of segments, the second segment temporally following the first segment in the media content; making, using the hardware processor, a second prediction that the media content depicts the first activity based on the second plurality of features, wherein the second prediction has a second confidence level; making, using the hardware processor, a third prediction that the media content depicts a second activity, the third prediction having a third confidence level based on the first plurality of features and a fourth confidence level based on the second plurality of features; comparing, using the hardware processor, a first difference between the first confidence level and the third confidence level with a second difference between the second confidence level and the fourth confidence level; and determining, using the hardware processor, that the media content depicts the first activity based on when the comparing indicates that the second difference is at least as much as the first difference based on, wherein the second confidence level is at least as high as the first confidence level.