- Number
- 11403531
- Published
- 2022-08-02
- Filed
- 2017-07-19
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Carr; G. Peter K., Deng; Zhiwei, Navarathna; Rajitha D. B, Yue; Yisong, Mandt; Stephan Marcel
- CPC
- G06F17/16; G06N3/045; G06N3/0455; G06N3/047; G06N3/0475; G06N3/0499; G06N3/084; G06N3/088; G06N3/0895
- Verdict
- Set aside generic ML representation-learning research
- Source
- Google Patents · FreePatentsOnline
Abstract
The disclosure provides an approach for learning latent representations of data using factorized variational autoencoders (FVAEs). The FVAE framework builds a hierarchical Bayesian matrix factorization model on top of a variational autoencoder (VAE) by learning a VAE that has a factorized representation so as to compress the embedding space and enhance generalization and interpretability. In one embodiment, an FVAE application takes as input training data comprising observations of objects, and the FVAE application learns a latent representation of such data. In order to learn the latent representation, the FVAE application is configured to use a probabilistic VAE to jointly learn a latent representation of each of the objects and a corresponding factorization across time and identity.
Background
BACKGROUND Field of the Invention (1) Embodiments of the disclosure presented herein relate to unsupervised machine learning and, more specifically, to factorized variational autoencoders. Description of the Related Art (2) In unsupervised machine learning, inferences are made from input data that are not labeled. Matrix and tensor factorization techniques have been used in unsupervised machine learning to find underlying patterns from noisy data. In particular, such factorization techniques can be applied to identify an underlying low-dimensional representation of latent factors in raw data. However, when the relationship between the latent representation to be learned and the raw data is complex, such as the behavior of a forest of trees in response to wind in a physical simulation or the reactions of an audience viewing a movie, the data may not linearly decompose into a set of underlying factors as required by conventional matrix and tensor factorization techniques. As a result, such techniques may not achieve good reconstructions of training data or produce interpretable latent factors in complex domains. SUMMARY (3) One embodiment provides a computer-implemented method for determining latent representations from data. The method generally includes receiving data associated with observations of one or more objects. The method further includes learning, from the received data, a variational autoencoder, the learning being subject to at least a constraint that outputs of a
Claims
1. A computer-implemented method comprising: receiving data indicating observations of a plurality of persons reacting to a first portion of a plurality of portions of a video, wherein the received data is unfactorizable; learning, from the received data in an unsupervised manner, a variational autoencoder that compresses the received data into a lower-dimensional vector of latent factors by: generating, for the plurality of persons, a first vector representing latent stimuli; generating, for a first person of the plurality of persons, a second vector representing how the first person reacts to the latent stimuli; generating the lower-dimensional vector of latent factors based on the first and second vectors; and factorizing the lower-dimensional vector of latent factors; and predicting, using the learned variational autoencoder, data separate from the received data, the predicted data indicating an unobserved reaction of the first person to at least a second portion of the plurality of portions of the video. ||
12. A non-transitory computer-readable storage medium storing a program, which, when executed by a processor performs operations comprising: receiving data indicating observations of a plurality of persons reacting to a first portion of a plurality of portions of a video, wherein the received data is unfactorizable; learning, from the received data in an unsupervised manner, a variational autoencoder that compresses the received data into a lower-dimensional vector of latent factors by: generating, for the plurality of persons, a first vector representing latent stimuli; generating, for a first person of the plurality of persons, a second vector representing how the first person reacts to the latent stimuli; generating the lower-dimensional vector of latent factors based on the first and second vectors; and factorizing the lower-dimensional vector of latent factors; and predicting, using the learned variational autoencoder, data separate from the received data, the predicted data indicating an unobserved reaction of the first person to at least a second portion of the plurality of portions of the video. ||
22. A system, comprising: a processor; and a memory, wherein the memory includes a program configured to perform operations comprising: receiving data indicating observations of a plurality of persons reacting to a first portion of a plurality of portions of a video, wherein the received data is unfactorizable; learning, from the received data in an unsupervised manner, a variational autoencoder that compresses the received data into a lower-dimensional vector of latent factors by: generating, for the plurality of persons, a first vector representing latent stimuli; generating, for a first person of the plurality of persons, a second vector representing how the first person reacts to the latent stimuli; generating the lower-dimensional vector of latent factors based on the first and second vectors; and factorizing the lower-dimensional vector of latent factors; and predicting, using the learned variational autoencoder, data separate from the received data, the predicted data indicating an unobserved reaction of the first person to at least a second portion of the plurality of portions of the video.