Outer Rim Archives
Archives · 2021 · 10956685

Granted patent

Alignment of video and textual sequences for metadata analysis

Number
10956685
Published
2021-03-23
Filed
2020-02-10
Assignee
Disney Enterprises, Inc.
Inventors
Li; Boyang, Sigal; Leonid, Dogan; Pelin
CPC
G06V10/82; G06F18/214; G06V30/2276; G06N3/09; G06N5/046; G06N3/0464; G06V30/19173; G06N3/0442; G06V10/454; G06F40/45; G06V20/41; G06V10/811; G06F18/256; G06N3/048; G06N3/045; G06N3/044
Verdict
Set aside video/text metadata alignment analytics, business
Source
Google Patents · FreePatentsOnline

Abstract

Systems, methods and computer program products related to aligning heterogeneous sequential data are disclosed. Video data in a media presentation and textual data corresponding to content of the media presentation are received. An action related to aligning the video data and the textual data is determined using an alignment neural network, such that the video data and the textual data are at least partially aligned following the action. The alignment neural network includes a first fully connected layer that receives as input the video data, the textual data, and data relating to a previously determined action by the alignment neural network related to aligning the video data and the textual data. The determined action related to aligning the video data and the textual data is performed.

Background

BACKGROUND Field of the Invention (1) The present invention relates to computerized neural networks, and more specifically, to a neural network for aligning heterogeneous sequential data. Description of the Related Art (2) Alignment of sequential data is a common problem in many different fields, including molecular biology, natural language processing, historic linguistics, and computer vision, among other fields. Aligning heterogeneous sequences of data, with complex correspondences, can be particularly complex. Heterogeneity refers to the lack of a readily apparent surface matching. For example, alignment of visual and textual content can be very complex. This is particularly true where one-to-many and one-to-none correspondences are possible, as in alignment of video from a film or television show with a script relating to the film or television show. One or more embodiments herein describe use of a computerized neural network to align sequential heterogeneous data, for example visual and textual data. SUMMARY (3) Embodiments described herein include a method for aligning heterogeneous sequential data. The method includes receiving video data in a media presentation and textual data corresponding to content of the media presentation. The method further includes determining an action related to aligning the video data and the textual data using an alignment neural network, such that the video data and the textual data are at least partially aligned following the action. Th

Claims

1. A method comprising: receiving video data in a media presentation and textual data corresponding to content of the media presentation at an alignment neural network previously trained to align training video data and training textual data; determining a first action related to aligning the video data and the textual data using the trained alignment neural network, such that the video data and the textual data are at least partially aligned following the first action, the trained alignment neural network comprising: a first fully connected layer that receives as input: the video data, the textual data, and data relating to a previously determined action by the trained alignment neural network related to aligning the video data and the textual data; and performing the first action related to aligning the video data and the textual data. || 10. A computer program product, comprising: a non-transitory computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code configured to perform one or more operations comprising: receiving video data in a media presentation and textual data corresponding to content of the media presentation at an alignment neural network previously trained to align training video data and training textual data; determining a first action related to aligning the video data and the textual data using the trained alignment neural network, such that the video data and the textual data are at least partially aligned following the first action, the trained alignment neural network comprising: a first fully connected layer that receives as input: the video data, the textual data, and data relating to a previously determined action by the trained alignment neural network related to aligning the video data and the textual data; and performing the first action related to aligning the video data and the textual data. || 16. A system, comprising: a processor; and a memory containing a program that, when executed on the processor, performs one or more operations comprising: receiving video data in a media presentation and textual data corresponding to content of the media presentation at an alignment neural network previously trained to align video data and textual data; determining a first action related to aligning the video data and the textual data using the trained alignment neural network, such that the video data and the textual data are at least partially aligned following the first action, the trained alignment neural network comprising: a first fully connected layer that receives as input: the video data, the textual data, and data relating to a previously determined action by the trained alignment neural network related to aligning the video data and the textual data; and performing the first action related to aligning the video data and the textual data.