Outer Rim Archives
Archives · 2024 · 11983183

Granted patent

Techniques for training machine learning models using actor data

Number
11983183
Published
2024-05-14
Filed
2018-08-07
Assignee
Disney Enterprises, Inc.
Inventors
Li; Boyang et al.
CPC
G06F16/248; G06N7/01; G06N20/00; G06F16/9535; G06F16/55; G06F16/24578; G06N3/047; G06F16/3344
Verdict
Set aside generic ML training-data management, business
Source
Google Patents · FreePatentsOnline

Abstract

Systems, methods, and articles of manufacture are disclosed for learning models of movies, keywords, actors, and roles, and querying the same. In one embodiment, a recommendation application optimizes a model based on training data by initializing the mean and co-variance matrices of Gaussian distributions representing movies, keywords, and actors to random values, and then performing an optimization to minimize a margin loss function using symmetrical or asymmetrical measures of similarity between entities. Such training produces an optimized model with the Gaussian distributions representing movies, keywords, and actors, as well as shift vectors that change the means of movie Gaussian distributions and model archetypical roles. Subsequent to training, the same similarity measures used to train the model are used to query the model and obtain rankings of entities based on similarity to terms in the query, and a representation of the rankings may be displayed via, e.g., a display device.

Background

BACKGROUND Field of the Invention (1) Embodiments presented in this disclosure generally relate to recommendation and search engines. More specifically, embodiments presented herein relate to techniques for learning models of movies, keywords, actors, and roles, and querying the same. Description of the Related Art (2) The motion picture industry has been extremely risky. Despite the best efforts of directors, casting directors, screenwriters, marketing teams, and experienced executives, it remains difficult to guarantee a return on investment from any movie production. (3) Recently, the computational understanding of narrative content, in textual and visual formats, has received renewed attention. However, in the context of movies in particular, little attempt has been made to understand movie actors in relation to characters they play and movies they appear in. SUMMARY (4) One embodiment of this disclosure provides a computer-implemented method that generally includes training, based at least in part on received training data, a model which includes Gaussian distributions representing actors, movies, and keywords. The method further includes receiving a query including one or more terms, and ranking, using the trained model, one or more of the actors, movies, or keywords, based at least in part on similarity to the one or more terms in the query. (5) Another embodiment provides a computer-implemented method that generally includes receiving information specifying at least m

Claims

1. A method comprising: receiving, at a computing device including one or more processors, a plurality of scripts, wherein the plurality of scripts textually describe at least characters; generating vector representations of the characters based on words associated with actions performed by the characters, actions received by the characters, and descriptions of the characters; generating, by the one or more processors, archetypes of the characters described in the plurality of scripts based on clustering of the vector representations of the characters; initializing, by the one or more processors, a machine learning model using the archetypes, the machine learning model including distributed representations in a dimensional space of entities including actors, content, and keywords, wherein the machine learning model is configured to compare the distributed representations in the dimensional space, wherein during initialization of the machine learning model, the distributed representations are assigned random values for positions in the dimensional space; training the machine learning model based on similarities between the vector representations by: implementing a random dropout for at least one of the distributed representations before the distributed representations are compared; and modifying the positions of the distributed representations in the machine learning model based on relatedness between the distributed representations, such that related distributed representations are pushed together and unrelated distributed representations are pushed apart; wherein the machine learning model is configured to represent actor versatility across the archetypes; configuring, in the machine learning model, shift vectors or changes to co-variance matrices to modify the distributed representations based on the archetypes; receiving, at the computing device, a query including one or more terms and a queried archetype; and ranking, by the one or more processors and using the machine learning model and either the shift vectors or the changes to the co-variance matrices, one or more of the actors based at least in part on the actor versatility across the archetypes. || 15. A computer implemented method of modelling actor versatility based on a plurality of scripts, the computer implemented method comprising: receiving, at a computing device including one or more processors, a plurality of scripts, the plurality of scripts textually specifying at least keywords and actors; generating vector representations of characters in the scripts based on words associated with actions performed by the characters, actions received by the characters, and descriptions of the characters; determining, by the one or more processors, archetypes of the characters based on clustering of the vector representations of the characters; initializing, by the one or more processors, a machine learning model including distributed representations in a dimensional space of entities including the actors, content, and the keywords, wherein the machine learning model is configured to compare the distributed representations in the dimensional space, wherein during initialization of the machine learning model, the distributed representations are assigned random values for positions in the dimensional space; training the machine learning model based on similarities between the vector representations by: implementing a random dropout for at least, one of the distributed representations before the distributed representations are compared; and modifying the positions of the distributed representations in the machine learning model based on relatedness between the distributed representations, such that related distributed representations are pushed together and unrelated distributed representations are pushed apart; wherein the machine learning model is configured to represent actor versatility across the archetypes; configuring, in the machine learning model, shift vectors or changes to co-variance matrices to modify the distributed representations based on the archetypes; receiving, at the computing device, a query including one or more terms and a queried archetype; and ranking, by the one or more processors and using the machine learning model and either the shift vectors or the changes to the co-variance matrices, one or more actors of the actors based at least in part on the actor versatility across the archetypes. || 20. A computer implemented method to model actor versatility based on a plurality of scripts, the computer implemented method comprising: receiving, at a computing device including one or more processors, a plurality of scripts, wherein the plurality of scripts textually describe at least characters; generating, by the one or more processors, archetypes of characters played by actors in the plurality of scripts by: linking pronouns with the characters in the plurality of scripts, identifying words in the plurality of scripts associated with actions performed by the characters, actions received by the characters, and descriptions of the characters, generating vector representations of the characters based on the actions performed by the characters, the actions received by the characters, and the descriptions of the characters, and identifying the archetypes based on clustering of the vector representations of the characters; initializing, by the one or more processors, a machine learning model including distributed representations in a dimensional space of entities including the actors, content, and keywords, wherein the machine learning model is configured to compare the distributed representations in the dimensional space, wherein during initialization of the machine learning model, the distributed representations are assigned random values for positions in the dimensional space; training the machine learning model based on similarities between the vector representations by: implementing a random dropout for at least one of the c distributed representations before the distributed representations are compared; and modifying the positions of the distributed representations in the machine learning model based on relatedness between the distributed representations, such that related distributed representations are pushed together and unrelated distributed representations are pushed apart; wherein the machine learning model is configured to represent actor versatility across the archetypes; configuring, in the machine learning model, shift vectors or changes to co-variance matrices to modify the distributed representations based on the archetypes; and ranking, by the one or more processors and using the machine learning model and either the shift vectors or the changes to the co-variance matrices, one or more actors of the actors based at least in part on the actor versatility across the archetypes.