Outer Rim Archives
Archives · 2024 · 20240303983

Application (pre-grant publication)

TECHNIQUES FOR MONOCULAR FACE CAPTURE USING A PERCEPTUAL SHAPE LOSS

Number
20240303983
Published
2024-09-12
Filed
2024-03-07
Assignee
DISNEY ENTERPRISES, INC.
Inventors
Bradley; Derek Edward et al.
CPC
G06T17/205; G06V10/993; G06V40/174; G06V10/82
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Monocular facial-capture technique using perceptual shape loss.

Abstract

One embodiment of the present invention sets forth a technique for evaluating three-dimensional (3D) reconstructions. The technique includes generating a 3D reconstruction of an object based on one or more mesh parameters. The technique also includes generating, based on the 3D reconstruction, a 3D rendering of the object. The technique further includes generating, using a machine learning model, a perceptual score associated with the 3D rendering and an input image of the object. The generated score represents how closely the 3D rendering matches the input image.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to machine learning and computer vision and, more specifically, to techniques for generating and evaluating three-dimensional (3D) renderings of two-dimensional (2D) representations of faces. Description of the Related Art

Generating a 3D reconstruction of a face from one or more 2D representations of the face is a common task in the fields of computer vision and computer graphics. Such reconstructions may be used in visual effects for, e.g., movies or television shows, computer games, social media, telepresence applications and virtual reality (VR) applications.

Existing techniques for generating 3D reconstructions of faces from 2D representations often require capturing multiple 2D representations of a single face using a predefined capture protocol. For example, a capture protocol may require multiple calibrated cameras capturing 2D representations of a face from several specified distances or viewpoints, under uniform lighting, and in a specified capture sequence.

One drawback of the above techniques is that they are not suitable for generating a 3D reconstruction from a single 2D representation (i.e., monocular face capture), or from multiple 2D representations that are not captured under controlled conditions using calibrated equipment and specified viewpoints as described above. As an example, such techniques may not be suitable for generating

Claims

1. A computer-implemented method for evaluating three-dimensional (3D) reconstructions, the computer-implemented method comprising: generating, based on one or more mesh parameters, a 3D reconstruction of an object; generating, based on the 3D reconstruction, a 3D rendering of the object; and generating, using a machine learning model, a perceptual score associated with the 3D rendering and an input image of the object, wherein the perceptual score represents how closely the 3D rendering matches the input image. || 10. A computer-implemented method for training a machine learning model to evaluate three-dimensional (3D) renderings, the computer-implemented method comprising: generating, using a machine learning model, a perceptual score based on a 3D rendering of an object and an input image of the object, wherein the perceptual score indicates a degree to which the 3D rendering does not match the input image; generating a critic loss based on the perceptual score, and modifying one or more parameters of the machine learning model based on the critic loss. || 18. A system comprising: one or more memories storing instructions; and one or more processors for executing the instructions to: generate, based on one or more mesh parameters, a 3D reconstruction of an object; generate, based on the 3D reconstruction, a 3D rendering of the object; and generate, using a machine learning model, a perceptual score associated with the 3D rendering and an input image of the object, wherein the perceptual score represents how closely the 3D rendering matches the input image.