Application (pre-grant publication)
JOINT HETEROGENEOUS LANGUAGE-VISION EMBEDDINGS FOR VIDEO TAGGING AND SEARCH
- Number
- 20170357720
- Published
- 2017-12-14
- Filed
- 2017-06-12
- Assignee
- Disney Enterprises, Inc.
- Inventors
- TORABI; Atousa et al.
- CPC
- G06F16/638; G06F16/7844; G06N3/044; G06N3/0442; G06N3/045; G06N3/0464; G06N3/08; G06N3/084; G06N3/09; G06V10/82; G06V20/41; H04N21/8405
- Verdict
- Set aside video tagging/search infra, generic ML
- Source
- Google Patents · FreePatentsOnline
Abstract
Systems, methods and articles of manufacture for modeling a joint language-visual space. A textual query to be evaluated relative to a video library is received from a requesting entity. The video library contains a plurality of instances of video content. One or more instances of video content from the video library that correspond to the textual query are determined, by analyzing the textual query using a data model that includes a soft-attention neural network module that is jointly trained with a language Long Short-term Memory (LSTM) neural network module and a video LSTM neural network module. At least an indication of the one or more instances of video content is returned to the requesting entity.
Background
BRIEF DESCRIPTION OF THE DRAWINGS
So that the manner in which the above recited aspects are attained and can be understood in detail, a more particular description of embodiments of the invention, briefly summarized above, may be had by reference to the appended drawings.
It is to be noted, however, that the appended drawings illustrate only typical embodiments ofthis invention and are therefore not to be considered limiting of its scope, for the invention may admit to other equally effective embodiments.
FIG. 1 is a system diagram illustrating a computing environment in which embodiments described herein can be implemented.
FIG. 2 is a block diagram illustrating a workflow for identifying instances of video content relating to a textual query, according to one embodiment described herein.
FIG. 3 is a block diagram illustrating a workflow for training a data model to identify instances of video content relating to a textual query, according to one embodiment described herein.
FIGS. 4A-C illustrate neural network data models configured to identify instances of video content relating to a textual query, according to embodiments described herein.
FIG. 5 is a flow diagram illustrating a method for determining instances of video content relating to a textual query, according to one embodiment described herein.
FIG. 6 is a flow diagram illustrating a method for training and using a data modelfor identifying instances of video conten