Outer Rim Archives
Archives · 2022 · 11393107

Granted patent

Jaw tracking without markers for facial performance capture

Number
11393107
Published
2022-07-19
Filed
2019-07-12
Assignee
Disney Enterprises, Inc.
Inventors
Beeler; Dominik Thabo, Bradley; Derek Edward, Zoss; Gaspard
CPC
G06T7/75; G06T7/33; G06T7/251; G06T17/00; G06T13/40; G06T7/246
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Markerless jaw-tracking facial performance capture technique (granted).

Abstract

Some implementations of the disclosure are directed to capturing facial training data for one or more subjects, the captured facial training data including each of the one or more subject's facial skin geometry tracked over a plurality of times and the subject's corresponding jaw poses for each of those plurality of times; and using the captured facial training data to create a model that provides a mapping from skin motion to jaw motion. Additional implementations of the disclosure are directed to determining a facial skin geometry of a subject; using a model that provides a mapping from skin motion to jaw motion to predict a motion of the subject's jaw from a rest pose given the facial skin geometry; and determining a jaw pose of the subject using the predicted motion of the subject's jaw.

Background

BRIEF SUMMARY OF THE DISCLOSURE (1) Implementations of the disclosure describe systems and methods for training and using a model to accurately track the jaw of a subject during facial performance capture based on the facial skin motion of the subject. (2) In one embodiment, a method comprises: capturing facial training data for one or more subjects, the captured facial training data including each of the one or more subject's facial skin geometry tracked over a plurality of times and the subject's corresponding jaw poses for each of those plurality of times; and using the captured facial training data to create a model that provides a mapping from skin motion to jaw motion. The mapping may be from a set of skin features that define the skin motion to a set of jaw features that define the jaw motion. In particular implementations, the jaw features are displacements of jaw points. (3) In some implementations, the facial training data is captured for a plurality of subjects over a plurality of facial expressions. In some implementations, using the captured facial training data to create a model that provides a mapping from the set of skin features that define the skin motion to the set of jaw features that define the jaw motion, comprises: for a first of the plurality of subjects, using the first subject's facial skin geometry captured over a plurality of times and the first subject's corresponding jaw poses for each of those plurality of times to learn a first mapping from a s

Claims

1. A non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, cause a system to perform operations comprising: capturing facial training data for one or more subjects, the captured facial training data including each subject's facial skin geometry and corresponding jaw pose tracked over a plurality of times; and creating, using the captured facial training data, a model that predicts, based on multiple spatially distributed facial skin features that define a skin motion, multiple jaw features that define a jaw motion corresponding to the skin motion, wherein creating the model comprises: selecting points on a jaw of a first subject of the one or more subjects while the first subject is in a rest pose; for each of the selected points, finding a corresponding point on a face of the first subject to determine skin feature points in the rest pose; and for each of a plurality of frames: tracking a position of the skin feature points to derive skin features; and tracking a displacement between a current position of the selected points on the first subject's jaw and a position of the selected points on the first subject's jaw in the rest pose to derive target jaw displacements. || 6. The non-transitory computer-readable medium of 1, wherein the operations further comprise: predicting, using the model, a jaw pose of a first subject for a first time using as an input a first facial skin geometry of the first subject for the first time. || 9. A method, comprising: capturing facial training data for one or more subjects, the captured facial training data including each subject's facial skin geometry and corresponding jaw pose tracked over a plurality of times; and creating, using the captured facial training, a model that predicts, based on multiple spatially distributed facial skin features that define a skin motion, multiple jaw features that define a jaw motion corresponding to the skin motion, wherein creating the model comprises: selecting points on a jaw of a first subject of the one or more subjects while the first subject is in a rest pose; for each of the selected points, finding a corresponding point on a face of the first subject to determine skin feature points in the rest pose; and for each of a plurality of frames: tracking a position of the skin feature points to derive skin features; and tracking a displacement between a current position of the selected points on the first subject's jaw and a position of the selected points on the first subject's jaw in the rest pose to derive target jaw displacements. || 13. A non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, cause a system to perform operations comprising: capturing facial training data for a plurality of subjects, the captured facial training data including each subject's facial skin geometry and corresponding jaw pose tracked over a plurality of times; and creating, using the captured facial training data, a model that predicts, based on multiple spatially distributed facial skin features that define a skin motion, multiple jaw features that define a jaw motion corresponding to the skin motion, wherein creating the model comprises: learning, for a first subject of the plurality of subjects, using the first subject's facial skin geometry and corresponding jaw pose tracked over the plurality of times, a first mapping from a set of skin features to a set of jaw features; aligning, using the first mapping, the facial skin geometry of another subject of the plurality of subjects with the facial skin geometry of the first subject; and learning, using the first subject's facial skin geometry, the other subject's aligned facial skin geometry, and each subject's corresponding jaw pose tracked over the plurality of times, a second mapping from a set of skin features to a set of jaw features. || 14. A non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, cause a system to perform operations comprising: capturing facial training data for one or more subjects, the captured facial training data including each subject's facial skin geometry and corresponding jaw pose tracked over a plurality of times; creating, using the captured facial training data, a model that provides a mapping from skin motion to jaw motion; and predicting, using the model, a jaw pose of a first subject for a first time using as an input a first facial skin geometry of the first subject for the first time, the first subject being different from the one or more subjects for which the facial training data is captured. || 15. The non-transitory computer-readable medium of 14, wherein the operations further comprise: capturing a plurality of calibration poses for the first subject; extracting calibration skin features and corresponding calibration bone features from each of the plurality of captured calibration poses; and determining, using at least the extracted calibration skin features and corresponding calibration bone features, a transformation of skin features of the first subject to align with a feature space of the model. || 16. A non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor, cause a system to perform operations comprising: determining a facial skin geometry of a subject; extracting skin features from the facial skin geometry of the subject; transforming the skin features to align with a feature space of a model, wherein the model is created using facial skin features of another subject; after transforming the skin features, predicting, using the model, a motion of the subject's jaw from a rest pose given the facial skin geometry; and determining, based on the motion of the subject's jaw from the rest pose, a jaw pose of the subject.