Outer Rim Archives
Archives · 2021 · 10902343

Granted patent

Deep-learning motion priors for full-body performance capture in real-time

Number
10902343
Published
2021-01-26
Filed
2016-09-30
Assignee
DISNEY ENTERPRISES, INC.
Inventors
Andrews; Sheldon, Huerta Casado; Ivan, Mitchell; Kenneth J., Sigal; Leonid
CPC
G06N20/00; G06N3/044; G06N3/045; G06N3/0464; G06N3/0475; G06N3/08; G06N3/09; G06T7/251; G06V10/764; G06V10/82; G06V40/23
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Real-time full-body motion-capture ML technique.

Abstract

Training data from multiple types of sensors and captured in previous capture sessions can be fused within a physics-based tracking framework to train motion priors using different deep learning techniques, such as convolutional neural networks (CNN) and Recurrent Temporal Restricted Boltzmann Machines (RTRBMs). In embodiments employing one or more CNNs, two streams of filters can be used. In those embodiments, one stream of the filters can be used to learn the temporal information and the other stream of the filters can be used to learn spatial information. In embodiments employing one or more RTRBMs, all visible nodes of the RTRBMs can be clamped with values obtained from the training data or data synthesized from the training data. In cases where sensor data is unavailable, the input nodes may be unclamped and the one or more RTRBMs can generate the missing sensor data.

Background

BACKGROUND OF THE INVENTION (1) Embodiments relate to computer capture of object motion. More specifically, embodiments relate to capturing of movement or performance of an actor. (2) Animation has been based more directly upon physical movement of actors. This animation is known as motion capture or performance capture. For capturing motion of an actor, the actor is typically equipped with a suit with a number of markers, and as the actor moves, a number of cameras or optical sensors track the positions of the markers in space. This technique allows the actor's movements and expressions to be captured, and the movements and expressions can then be manipulated in a digital environment to produce whatever animation is desired. Modern motion capture systems may include various other types of sensors, such as inertial sensors. The inertial sensors typically comprise an inertial measurement unit (IMU) having a combination of gyroscope, magnetometer, and accelerometer for measuring rotational rates, such as orientation, linear acceleration and gyro rate. Markerless Motion Capturing (Mocap) is another active field of research in motion capture. The goal of Mocap is to determine the 3D positions and orientations as well as the joint angles of the actor from image data. In such a tracking scenario, it is common to assume as input a sequence of multiview images of the performed motion as well as a surface mesh of the actor's body. (3) In motion capture sessions, movements of the actor

Claims

1. A method for motion capture, the method being implemented by a processor configured to execute machine-readable instructions, the method comprising: obtaining a deep learning model, wherein the deep learning model includes a convolutional neural network (CNN); obtaining a set of training data and training the deep learning model using the training data to generate a motion model, the training data including temporal and spatial information regarding one or more actors' motion captured previously, wherein the training of the deep learning model includes: receiving the spatial information in a first stream to learn a spatial relationship of the one or more actors' motion captured previously using a first set of one or more filters; receiving the temporal information in a second stream to learn a temporal relationship of the deep learning model using a second set of one or more filters, wherein the second stream is received independent of the first stream; mapping the set of training data to a ground truth skeleton pose using a fully connected feedforward neural network; receiving, from one or more motion capture sensors, real-time motion data for the one or more actor's motion; identifying missing motion information in the received real-time motion data introduced by the one or more motion capture sensors; and estimating the missing motion information regarding the actor's motion based on the received real-time motion data using the deep learning model with inverse dynamics to estimate a reference pose for a frame and combining the received real-time motion data with the estimated reference pose to solve for orientation and position constraints. || 6. A system for motion capture, the system comprising one or more of a processor configured to execute machine-readable instructions such that when the machine-readable instructions are executed, the process is caused to perform: obtaining a deep learning model, wherein the deep learning model includes a convolutional neural network (CNN); obtaining a set of training data and training the deep learning model using the training data to generate a motion model, the training data including temporal and spatial information regarding one or more actors' motion captured previously, wherein the training of the deep learning model includes: receiving the spatial information in a first stream to learn a spatial relationship of the one or more actors' motion captured previously using a first set of one or more filters; receiving the temporal information in a second stream to learn a temporal relationship of the deep learning model using a second set of one or more filters, wherein the second stream is received independent of the first stream; mapping the set of training data to a ground truth skeleton pose using a fully connected feedforward neural network; receiving, from one or more motion capture sensors, real-time motion data for the one or more actor's motion; identifying missing motion information in the received real-time motion data introduced by the one or more motion capture sensors; and estimating the missing motion information regarding the actor's motion based on the received real-time motion data using the deep learning model with inverse dynamics to estimate a reference pose for a frame and combining the received real-time motion data with the estimated reference pose to solve for orientation and position constraints. || 11. A computer-readable medium storing a plurality of instructions that, when executed by one or more processors of a computing device, cause the one or more processors to perform operations comprising: obtaining a deep learning model, wherein the deep learning model includes a convolutional neural network (CNN); obtaining a set of training data and training the deep learning model using the training data to generate a motion model, the training data including temporal and spatial information regarding one or more actors' motion captured previously, wherein the training of the deep learning model includes: receiving the spatial information in a first stream to learn a spatial relationship of the one or more actors' motion captured previously using a first set of one or more filters; receiving the temporal information in a second stream to learn a temporal relationship of the deep learning model using a second set of one or more filters, wherein the second stream is received independent of the first stream; mapping the set of training data to a ground truth skeleton pose using a fully connected feedforward neural network; receiving, from one or more motion capture sensors, real-time motion data for the one or more actor's motion; identifying missing motion information in the received real-time motion data introduced by the one or more motion capture sensors; and estimating the missing motion information regarding the actor's motion based on the received real-time motion data using the deep learning model with inverse dynamics to estimate a reference pose for a frame and combining the received real-time motion data with the estimated reference pose to solve for orientation and position constraints.