Outer Rim Archives
Archives · 2017 · 20170357720

Application (pre-grant publication)

JOINT HETEROGENEOUS LANGUAGE-VISION EMBEDDINGS FOR VIDEO TAGGING AND SEARCH

Number
20170357720
Published
2017-12-14
Filed
2017-06-12
Assignee
Disney Enterprises, Inc.
Inventors
TORABI; Atousa et al.
CPC
G06F16/638; G06F16/7844; G06N3/044; G06N3/0442; G06N3/045; G06N3/0464; G06N3/08; G06N3/084; G06N3/09; G06V10/82; G06V20/41; H04N21/8405
Verdict
Set aside video tagging/search infra, generic ML
Source
Google Patents · FreePatentsOnline

Abstract

Systems, methods and articles of manufacture for modeling a joint language-visual space. A textual query to be evaluated relative to a video library is received from a requesting entity. The video library contains a plurality of instances of video content. One or more instances of video content from the video library that correspond to the textual query are determined, by analyzing the textual query using a data model that includes a soft-attention neural network module that is jointly trained with a language Long Short-term Memory (LSTM) neural network module and a video LSTM neural network module. At least an indication of the one or more instances of video content is returned to the requesting entity.

Background

BRIEF DESCRIPTION OF THE DRAWINGS

So that the manner in which the above recited aspects are attained and can be understood in detail, a more particular description of embodiments of the invention, briefly summarized above, may be had by reference to the appended drawings.

It is to be noted, however, that the appended drawings illustrate only typical embodiments ofthis invention and are therefore not to be considered limiting of its scope, for the invention may admit to other equally effective embodiments.

FIG. 1 is a system diagram illustrating a computing environment in which embodiments described herein can be implemented.

FIG. 2 is a block diagram illustrating a workflow for identifying instances of video content relating to a textual query, according to one embodiment described herein.

FIG. 3 is a block diagram illustrating a workflow for training a data model to identify instances of video content relating to a textual query, according to one embodiment described herein.

FIGS. 4A-C illustrate neural network data models configured to identify instances of video content relating to a textual query, according to embodiments described herein.

FIG. 5 is a flow diagram illustrating a method for determining instances of video content relating to a textual query, according to one embodiment described herein.

FIG. 6 is a flow diagram illustrating a method for training and using a data modelfor identifying instances of video conten

Claims

1. A method, comprising: receiving, from a requesting entity, a textual query to be evaluated relative to a video library, the video library containing a plurality of instances of video content; determining one or more instances of video content from the video library that correspond to the textual query, by analyzingthe textual query using a data model that includes a soft-attention neural network modulethat is jointly trained with a language Long Short-term Memory (LSTM) neural network module and a video LSTM neural network module; and returning at least an indication of the oneor more instances of video content to the requesting entity. 10. A method, comprising: receiving, from a requesting entity, a textual query to be evaluated relative to a video library, the video library containing a plurality of instances of video content; training a data model based on a plurality of training samples, wherein each of the plurality of training samples includes (i) a respective instance of video content and (ii) a respective plurality of phrases describing the instance of video content, wherein training the data modelfurther comprises, for each of the plurality of training samples: encoding each of the plurality of phrases for the training sample as a matrix, wherein each word within the one or more phrases is encoded as a vector; determining a weighted ranking between the plurality of phrases; encoding the instance of video content for the training sample as a sequenceof frames; extracting frame features from the sequence of frames; performing an object classification analysis on the extracted frame features; and generating a matrix representing the instance of video content, based on the extracted frame features and the object classification analysis; processing the textual query using the trained data model to identifyone or more instances of video content from the plurality of instances of video content; and returning at least an indication of the one or more instances of video content to the requesting entity. 17. A method, comprising: training a data model based in part on a plurality of training samples, wherein each of the plurality of training samples includes (i) a respective instance of video content and (ii) a respective plurality of phrases describing the instance of video content, wherein a first one of the plurality of training samplescomprises a single frame instance of video content generated from an image file, and wherein training the data model further comprises, for each of the plurality of training samples: determining a weighted ranking between the plurality of phrases, based on a respectivelength of each phrase, such that more lengthy phrases are ranked above less lengthy phrases; and generating a matrix representing the instance of video content, based at least in part on an object classification analysis performed on frame features extracted from the instance of video content; and using the trained data model to identify one or more instances of video content that are related to a textual query.