There is provided a method that includes receiving a video having video shots, and creating video shot groups based on similarities between the video shots, where each video shot group of the video shot groups includes one or more of the video shots and has different ones of the video shots than other video shot groups. The method further includes creating at least one video supergroup including at least one video shot group of the video shot groups based on interactions among the one or more of the video shots in each of the video shot groups, and divide the at least one video supergroup into connected video supergroups, each connected video supergroup of the connected video supergroups including one or more of the video shot groups based on the interactions among the one or more of video shots in each of the video shot groups.
BACKGROUND(1) Typical videos, such as television (TV) shows, include a number of different video shots shown in a sequence, the content of which may be processed using video content analysis. Conventional video content analysis may be used to identify motion in a video, recognize objects and/or shapes in a video, and sometimes to track an object or person in a video. However, conventional video content analysis merely provides information about objects or individuals in the video. Identifying different parts of a show requires manual annotation of the show, which is a costly and time-consuming process.SUMMARY(2) The present disclosure is directed to systems and methods for contextual video shot aggregation, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.
1. A system comprising: a memory storing an executable code; and a hardware processor executing the executable code to: receive a video having a plurality of video shots; create a plurality of video shot groups based on feature distances between the plurality of video shots, wherein each video shot group of the plurality of video shot groups includes one or more of the plurality of video shots and has different ones of the plurality of video shots than other video shot groups; create at least one video supergroup including at least one video shot group of the plurality of video shot groups by using a cluster algorithm on the plurality of video shot groups; divide the at least one video supergroup into a plurality of connected video supergroups, each connected video supergroup of the plurality of connected video supergroups including one or more of the plurality of video shot groups based on interactions among the one or more of plurality of video shots in each of the plurality of video shot groups; identify one of the plurality of connected video supergroups as an anchor video subgroup based on (a) a screen time of the one of the plurality of connected video supergroups, and (b) an amount of time between a first appearance and a last appearance of the video shots of the one of the plurality of connected video supergroups, wherein the anchor video subgroup includes video shots that are not temporally adjacent; and assign a category to the video based on the anchor video subgroup of the plurality of connected video supergroups.
13. A method for use by a system having a memory and a hardware processor, the method comprising: receiving, using the hardware processor, a video having a plurality of video shots; creating, using the hardware processor, a plurality of video shot groups based on feature distances between the plurality of video shots, wherein each video shot group of the plurality of video shot groups includes one or more of the plurality of video shots and has different ones of the plurality of video shots than other video shot groups; creating, using the hardware processor, at least one video supergroup including at least one video shot group of the plurality of video shot groups by using a cluster algorithm on the plurality of video shot groups; dividing, using the hardware processor, the at least one video supergroup into a plurality of connected video supergroups, each connected video supergroup of the plurality of connected video supergroups including one or more of the plurality of video shot groups based on interactions among the one or more of plurality of video shots in each of the plurality of video shot groups; identifying, using the hardware processor, one of the plurality of connected video supergroups as an anchor video subgroup based on (a) a screen time of the one of the plurality of connected video supergroups, and (b) an amount of time between a first appearance and a last appearance of the video shots of the one of the plurality of connected video supergroups, wherein the anchor video subgroup includes video shots that are not temporally adjacent; and assigning, using the hardware processor, a category to the video based on the anchor video subgroup of the plurality of connected video supergroups.