Outer Rim Archives
Archives · 2020 · 20200175232

Application (pre-grant publication)

ALIGNMENT OF VIDEO AND TEXTUAL SEQUENCES FOR METADATA ANALYSIS

Number
20200175232
Published
2020-06-04
Filed
2020-02-10
Assignee
Disney Enterprises, Inc.
Inventors
LI; Boyang, SIGAL; Leonid, DOGAN; Pelin
CPC
G06F18/214; G06N3/0442; G06V10/811; G06F18/256; G06V30/2276; G06N3/045; G06V10/82; G06N3/0464; G06V10/454; G06N5/046; G06N3/044; G06F40/45; G06V30/19173; G06V20/41; G06N3/09
Verdict
Set aside video/text metadata alignment analytics, business
Source
Google Patents · FreePatentsOnline

Abstract

Systems, methods and computer program products related to aligning heterogeneous sequential data are disclosed. Video data in a media presentation and textual data corresponding to content of the media presentation are received. An action related to aligning the video data and the textual data is determined using an alignment neural network, such that the video data and the textual data are at least partially aligned following the action. The alignment neural network includes a first fully connected layer that receives as input the video data, the textual data, and data relating to a previously determined action by the alignment neural network related to aligning the video data and the textual data. The determined action related to aligning the video data and the textual data is performed.

Background

BACKGROUNDField of the Invention

The present invention relates to computerized neural networks, and more specifically, to a neural network for aligning heterogeneous sequential data.Description of the Related Art

Alignment of sequential data is a common problem in many different fields, including molecular biology, natural language processing, historic linguistics, and computer vision, among other fields. Aligning heterogeneous sequences of data, with complex correspondences, can be particularly complex. Heterogeneity refers to the lack of a readily apparent surface matching. For example, alignment of visual and textual content can be very complex. This is particularly true where one-to-many and one-to-none correspondences are possible, as in alignment of video from a film or television show with a script relating to the film or television show. One or more embodiments herein describe use of a computerized neural network to align sequential heterogeneous data, for example visual and textual data.SUMMARY

Embodiments described herein include a method for aligning heterogeneous sequential data. The method includes receiving video data in a media presentation and textual data corresponding to content of the media presentation. The method further includes determining an action related to aligning the video data and the textual data using an alignment neural network, such that the video data and the textual data are at least partially aligned following the action.

Claims

1. A method of aligning heterogeneous sequential data, comprising: receiving video data in a media presentation and textual data corresponding to content of the media presentation; determining an action related to aligning the video data and the textual data using an alignment neural network, such that the video data and the textual data are at least partially aligned following the action, the alignment neural network comprising: a first fully connected layer that receives as input: the video data, the textual data, and data relating to a previously determined action by the alignment neural network related to aligning the video data and the textual data; and performing the determined action related to aligning the video data and the textual data. 10. A computer program product for aligning heterogeneous sequential data, the computer program product comprising: a computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code comprising computer-readable program code configured to perform an operation, the operation comprising: receiving video data in a media presentation and textual data corresponding to content of the media presentation; determining an action related to aligning the video data and the textual data using an alignment neural network, such that the video data and the textual data are at least partially aligned following the action, the alignment neural network comprising: a first fully connected layer that receives as input: the video data, the textual data, and data relating to a previously determined action by the alignment neural network related to aligning the video data and the textual data; and performing the determined action related to aligning the video data and the textual data. 16. A system, comprising: a processor; and a memory containing a program that, when executed on the processor, performs an operation, the operation comprising: receiving video data in a media presentation and textual data corresponding to content of the media presentation; determining an action related to aligning the video data and the textual data using an alignment neural network, such that the video data and the textual data are at least partially aligned following the action, the alignment neural network comprising: a first fully connected layer that receives as input: the video data, the textual data, and data relating to a previously determined action by the alignment neural network related to aligning the video data and the textual data; and performing the determined action related to aligning the video data and the textual data.