Outer Rim Archives
Archives · 2018 · 10147023

Granted patent

Markerless face tracking with synthetic priors

Number
10147023
Published
2018-12-04
Filed
2016-10-20
Assignee
DISNEY ENTERPRISES, INC.
Inventors
Klaudiny; Martin; McDonagh; Steven; Bradley; Derek; Beeler; Thabo; Matthews; Iain; Mitchell; Kenneth
CPC
G06T13/40; G06V10/774; G06V40/167; G06V40/176; G06T7/251; G06F18/214; G06T15/50
Verdict
Medium Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Markerless facial CV tracking.

Abstract

Provided are methods, systems, and computer-readable medium for synthetically generating training data to be used to train a learning algorithm that is capable of generating computer-generated images of a subject from real images that include the subject. The training data can be generated using a facial rig by changing expressions, camera viewpoints, and illumination in the training data. The training data can then be used for tracking faces in a real-time video stream. In such examples, the training data can be tuned to expected environmental conditions and camera properties of the real-time video stream. Provided herein are also strategies to improve training set construction by analyzing which attributes of a computer-generated image (e.g., expression, viewpoint, and illumination) require denser sampling.

Background

BACKGROUND OF THE INVENTION(1) Animation is an ever-growing technological field. As computer systems and technology evolve, the methods and systems used to animate are also evolving. For example, conventionally, an animation was created by an artist creating each frame of an animated sequence. The frames were then combined to create an impression of movement. However, manually creating each frame was time consuming. In addition, inadvertent mistakes were introduced by artists simply because of the vast amount of detail required to create accurate animations.(2) Recently, methods and systems have been created to capture movements of subjects to create computer-generated subjects. For example, facial movements can be electronically converted to produce computer-generated animation. Currently, the main technique for facial motion capture is to have reference points associated with different places on a face (sometimes referred to as markers). These markers are then used to detect movement of the face. Unfortunately, capture of small facial movements can be missed when markers are not included where the small movements occur. Therefore, there is a need in the art for improved methods and systems for facial tracking.BRIEF SUMMARY OF THE INVENTION(3) Provided are methods, systems, and computer-readable media for synthetically generating training data to be used to train a learning algorithm that is capable of generating computer-generated images of a subject. In some examples, the

Claims

1. A method for markerless face tracking, the method comprising: obtaining a facial rig associated with a subject, wherein the facial rig includes a plurality of expression shapes, wherein an expression shape defines at least a portion of an expression of the subject and includes one or more values for one or more facial attributes; generating a plurality of model states for the facial rig, wherein a model state describes a combination of expression shapes defining an expression of the subject and a set of camera setup location coordinates in relation to the subject; determining a lighting characteristic to use for rendering a computer-generated image of a model state of the plurality of model states; rendering a plurality of computer-generated images of a face of the subject, wherein a computer-generated image is rendered using the lighting characteristic and a corresponding model state of the facial rig; generating a plurality of training samples, wherein a training sample includes a computer generated image and a corresponding model state; and training a regressor using the plurality of training samples, wherein the trained regressor is configured to infer a model state that corresponds to the face of the subject captured in a frame. 10. A system for face tracking, the system comprising: a memory storing a plurality of instructions; and one or more processors configurable to: obtain a facial rig associated with a subject, wherein the facial rig includes a plurality of expression shapes, wherein an expression shape defines at least a portion of an expression of the subject and includes one or more values for one or more facial attributes; generate a plurality of model states for the facial rig, wherein a model state describes a combination of expression shapes defining an expression of the subject and a set of camera setup location coordinates in relation to the subject; determine a lighting characteristic to use for rendering a computer-generated image of a model state of the plurality of model states; render a plurality of computer-generated images of a face of the subject, wherein a computer-generated image is rendered using the lighting characteristic and a corresponding model state of the facial rig; generate a plurality of training samples, wherein a training sample includes a computer-generated image and a corresponding model state; and train a regressor using the plurality of training samples, wherein the trained regressor is configured to infer a model state that corresponds to the face of the subject captured in a frame. 16. A computer-readable memory storing a plurality of instructions executable by one or more processors, the plurality of instructions comprising instructions that cause the one or more processors to: obtain a facial rig associated with a subject, wherein the facial rig includes a plurality of expression shapes, wherein an expression shape defines at least a portion of an expression of the subject and includes one or more values for one or more facial attributes; generate a plurality of model states for the facial rig, wherein a model state describes a combination of expression shapes defining an expression of the subject and a set of camera setup location coordinates in relation to the subject; determine a lighting characteristic to use for rendering a computer-generated image of a model state of the plurality of model states; render a plurality of computer-generated images of a face of the subject, wherein a computer-generated image is rendered using the lighting characteristic and a corresponding model state of the facial rig; generate a plurality of training samples, wherein a training sample includes a computer-generated image and a corresponding model state; and train a regressor using the plurality of training samples, wherein the trained regressor is configured to infer a model state that corresponds to the face of the subject captured in a frame.