Outer Rim Archives
Archives · 2025 · 12243349

Granted patent

Face reconstruction using a mesh convolution network

Number
12243349
Published
2025-03-04
Filed
2022-03-17
Assignee
Disney Enterprises, INC.
Inventors
Bradley; Derek Edward et al.
CPC
G06V40/176; G06T17/20; G06V40/166; G06V40/161; G06V40/174; G06V40/172
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Mesh-convolution-network face-reconstruction VFX technique.

Abstract

Embodiment of the present invention sets forth techniques for performing face reconstruction. The techniques include generating an identity mesh based on an identity encoding that represents an identity associated with a face in one or more images. The techniques also include generating an expression mesh based on an expression encoding that represents an expression associated with the face in the one or more images. The techniques also include generating, by a machine learning model, an output mesh of the face based on the identity mesh and the expression mesh.

Background

BACKGROUND Field of the Various Embodiments (1) The present invention relates generally to computer science and computer-generated graphics and, more specifically, to face reconstruction using a mesh convolution network, including all aspects of the related hardware, software, graphical user interfaces, and algorithms with implementing the contemplated systems, techniques, functions, and operations set forth herein. Description of the Related Art (2) Realistic digital faces are required for various computer graphics and computer vision applications. For example, digital faces are oftentimes used in virtual scenes of film or television productions and in video games. (3) To capture photorealistic faces, a typical facial capture system employs a specialized light stage and hundreds of lights that are used to capture numerous images of an individual face under multiple illumination conditions. The facial capture system additionally employs multiple calibrated camera views, uniform or controlled patterned lighting, and a controlled setting. Further, a given face is typically scanned during a scheduled block of time, in which the corresponding individual can be guided into different expressions to capture images of individual faces. The resulting images can then be used to determine three-dimensional (3D) geometry and appearance maps that are needed to synthesize digital versions of the face. (4) One drawback that exists with many existing facial capture systems is the dependency

Claims

1. A computer-implemented method for performing reconstruction of a face, the computer-implemented method comprising: generating an identity mesh based on an identity encoding that represents an identity associated with a face in one or more images; generating an expression mesh based on an expression encoding that represents an expression associated with the face in the one or more images; and generating, by a machine learning model, an output mesh of the face via an upsampling operation associated with at least one of the identity mesh or the expression mesh. || 10. One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: generating an identity mesh based on an identity encoding that represents an identity associated with a face in one or more images; generating an expression mesh based on an expression encoding that represents an expression associated with the face in the one or more images; and generating, by a machine learning model, an output mesh of the face via an upsampling operation associated with at least one of the identity mesh or the expression mesh. || 19. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to: generate an identity mesh based on an identity encoding that represents an identity of a face in one or more images; generate an expression mesh based on an expression encoding that represents an expression of the face in the one or more images of the face; and generate, by a machine learning model, an output mesh of the face via an upsampling operation associated with at least one of the identity mesh or the expression mesh.