Granted patent
Alignment of video and textual sequences for metadata analysis
- Number
- 10956685
- Published
- 2021-03-23
- Filed
- 2020-02-10
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Li; Boyang, Sigal; Leonid, Dogan; Pelin
- CPC
- G06V10/82; G06F18/214; G06V30/2276; G06N3/09; G06N5/046; G06N3/0464; G06V30/19173; G06N3/0442; G06V10/454; G06F40/45; G06V20/41; G06V10/811; G06F18/256; G06N3/048; G06N3/045; G06N3/044
- Verdict
- Set aside video/text metadata alignment analytics, business
- Source
- Google Patents · FreePatentsOnline
Abstract
Systems, methods and computer program products related to aligning heterogeneous sequential data are disclosed. Video data in a media presentation and textual data corresponding to content of the media presentation are received. An action related to aligning the video data and the textual data is determined using an alignment neural network, such that the video data and the textual data are at least partially aligned following the action. The alignment neural network includes a first fully connected layer that receives as input the video data, the textual data, and data relating to a previously determined action by the alignment neural network related to aligning the video data and the textual data. The determined action related to aligning the video data and the textual data is performed.
Background
BACKGROUND Field of the Invention (1) The present invention relates to computerized neural networks, and more specifically, to a neural network for aligning heterogeneous sequential data. Description of the Related Art (2) Alignment of sequential data is a common problem in many different fields, including molecular biology, natural language processing, historic linguistics, and computer vision, among other fields. Aligning heterogeneous sequences of data, with complex correspondences, can be particularly complex. Heterogeneity refers to the lack of a readily apparent surface matching. For example, alignment of visual and textual content can be very complex. This is particularly true where one-to-many and one-to-none correspondences are possible, as in alignment of video from a film or television show with a script relating to the film or television show. One or more embodiments herein describe use of a computerized neural network to align sequential heterogeneous data, for example visual and textual data. SUMMARY (3) Embodiments described herein include a method for aligning heterogeneous sequential data. The method includes receiving video data in a media presentation and textual data corresponding to content of the media presentation. The method further includes determining an action related to aligning the video data and the textual data using an alignment neural network, such that the video data and the textual data are at least partially aligned following the action. Th