- Number
- 12488524
- Published
- 2025-12-02
- Filed
- 2021-11-15
- Assignee
- Disney Enterprises, INC.
- Inventors
- Bradley; Derek Edward et al.
- CPC
- G06T9/002; G06T13/40; G06N3/044; G06N3/045; G06N3/047; G06N3/08; G06N3/084; G06N3/088; G06T3/00
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
3D-geometry sequence synthesis for character movement-based performance.
Abstract
A technique for generating a sequence of geometries includes converting, via an encoder neural network, one or more input geometries corresponding to one or more frames within an animation into one or more latent vectors. The technique also includes generating the sequence of geometries corresponding to a sequence of frames within the animation based on the one or more latent vectors. The technique further includes causing output related to the animation to be generated based on the sequence of geometries.
Background
BACKGROUND Field of the Various Embodiments (1) Embodiments of the present disclosure relate generally to machine learning and animation and, more specifically, to synthesizing sequences of three-dimensional (3D) geometries for movement-based performance. Description of the Related Art (2) Realistic digital faces are required for various computer graphics and computer vision applications. For example, digital faces are oftentimes used in virtual scenes of film or television productions and in video games. (3) To capture photorealistic faces, a typical facial capture system employs a specialized light stage and hundreds of lights that are used to capture numerous images of an individual face under multiple illumination conditions. The facial capture system additionally employs multiple calibrated camera views, uniform or controlled patterned lighting, and a controlled setting in which the face can be guided into different expressions to capture images of individual faces. These images can then be used to determine three-dimensional (3D) geometry and appearance maps that are needed to synthesize digital versions of the faces. (4) Machine learning models have also been developed to synthesize digital faces. These machine learning models can include a large number of tunable parameters and thus require a large amount and variety of data to train. However, collecting training data for these machine learning models can be time- and resource-intensive. For example, a deep neural net
Claims
1. A computer-implemented method for generating a sequence of three-dimensional (3D) geometries, the computer-implemented method comprising: converting, via an encoder neural network, one or more input 3D geometries corresponding to one or more frames within an animation into one or more latent vectors; combining (i) a capture code that represents one or more attributes of the sequence of 3D geometries with (ii) a plurality of position encodings that represent a plurality of time steps within the animation to produce a plurality of position-encoded representations of the capture code; generating, via a decoder neural network, the sequence of 3D geometries based on input that includes (i) the one or more latent vectors and (ii) the plurality of position-encoded representations of the capture code, wherein each 3D geometry included in the sequence of 3D geometries corresponds to (i) a different time step included in the plurality of time steps and (ii) a different frame included in a sequence of frames within the animation; and causing output related to the animation to be generated based on the sequence of 3D geometries. ||
11. One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: converting, via an encoder neural network, one or more input three-dimensional (3D) geometries corresponding to one or more frames within an animation into one or more latent vectors; combining (i) a capture code that represents one or more attributes of a sequence of 3D geometries with (ii) a plurality of position encodings that represent a plurality of time steps within the animation to produce a plurality of position-encoded representations of the capture code; generating, via a decoder neural network, sequence of 3D geometries based on input that includes (i) the one or more latent vectors and (ii) h plurality of position-encoded representations of the capture code, wherein each 3D geometry included in the sequence of 3D geometries corresponds to (i) a different time step included in the plurality of time steps and (ii) a different frame included in a sequence of frames within the animation; and causing output related to the animation to be generated based on the sequence of 3D geometries. ||
20. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to: convert, via an encoder neural network, one or more input three-dimensional (3D) geometries corresponding to one or more frames within an animation into one or more latent vectors; combine (i) a capture code that represents one or more attributes of a sequence of 3D geometries with (ii) a plurality of position encodings that represent a plurality of time steps within the animation to produce a plurality of position-encoded representations of the capture code; generate, via a decoder neural network, sequence of 3D geometries based on input that includes (i) the one or more latent vectors and (ii) the plurality of position-encoded representations of the capture code, wherein each 3D geometry included in the sequence of 3D geometries corresponds to (i) a different time step included in the plurality of time steps and (ii) a different frame included in a sequence of frames within the animation; and cause output related to the animation to be generated based on the sequence of 3D geometries.