- Number
- 10970849
- Published
- 2021-04-06
- Filed
- 2019-04-16
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Öztireli; Ahmet Cengiz, Chandran; Prashanth, Gross; Markus
- CPC
- G06F3/011; G06F3/017; G06N20/00; G06N3/045; G06N3/0455; G06N3/0464; G06N3/08; G06N3/088; G06N3/0895; G06N3/09; G06T12/00; G06T7/20; G06T7/251; G06T7/70; G06T7/75
- Verdict
- Medium Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Neural-network pose estimation and body tracking (granted).
Abstract
According to one implementation, a pose estimation and body tracking system includes a computing platform having a hardware processor and a system memory storing a software code including a tracking module trained to track motions. The software code receives a series of images of motion by a subject, and for each image, uses the tracking module to determine locations corresponding respectively to two-dimensional (2D) skeletal landmarks of the subject based on constraints imposed by features of a hierarchical skeleton model intersecting at each 2D skeletal landmark. The software code further uses the tracking module to infer joint angles of the subject based on the locations and determine a three-dimensional (3D) pose of the subject based on the locations and the joint angles, resulting in a series of 3D poses. The software code outputs a tracking image corresponding to the motion by the subject based on the series of 3D poses.
Background
BACKGROUND (1) Augmented Reality (AR) and virtual reality (VR) experiences merge virtual objects or characters with real-world features in a way that can, in principle, provide a deeply immersive and powerfully interactive experience. Nevertheless, despite the relative success of digital enhancement techniques in augmenting many inanimate objects, digital augmentation of the human body continues to present substantial technical obstacles. For example, due to the ambiguities associated with depth projection, as well as the variations in human body shapes, three-dimensional (3D) human pose estimation remains a significant challenge. (2) In addition to AR and VR applications, accurate body tracking, in particular hand tracking, is important for effective use of the human hand as a Human Computer Interface (HCl). Applications for which use of the human hand as an HCl may be advantageous or desirable include hand tracking based character animation, for example. However, the challenges associated with pose estimation present significant problems for hand tracking as well. Consequently, there is a need in the art for a fast and accurate pose estimation and body tracking solution. SUMMARY (3) There are provided systems and methods for performing pose estimation and body tracking using an artificial neural network, substantially as shown in and/or described in connection with at least one of the figures, and as set forth more completely in the claims.
Claims
1. A pose estimation and body tracking system comprising: a computing platform including a hardware processor and a system memory; a software code stored in the system memory, the software code including a tracking module trained to track motions; the hardware processor configured to execute the software code to: receive a series of images of a motion by a subject; for each image of the series of images, determine, using the tracking module, a plurality of locations each corresponding respectively to a two-dimensional (2D) skeletal landmark of the subject based on constraints imposed by features of a hierarchical skeleton model intersecting at each 2D skeletal landmark; for each image of the series of images, infer, using the tracking module, a plurality of joint angles of the subject based on the plurality of locations; for each image of the series of images, reconstruct, using the tracking module, a three-dimensional (3D) pose of the subject based on the plurality of locations and the plurality of joint angles, resulting in a series of 3D poses by the subject; and output a tracking image corresponding to the motion by the subject based on the series of 3D poses by the subject. ||
9. A method for use by a pose estimation and body tracking system including a computing platform having a hardware processor and a system memory storing a software code including a tracking module trained to track motions, the method comprising: receiving, by the software code executed by the hardware processor, a series of images of a motion by a subject; for each image of the series of images, determining, by the software code executed by the hardware processor and using the tracking module, a plurality of locations each corresponding respectively to a two-dimensional (2D) skeletal landmark of the subject based on constraints imposed by features of a hierarchical skeleton model intersecting at each 2D skeletal landmark; for each image of the series of images, inferring, by the software code executed by the hardware processor and using the tracking module, a plurality of joint angles of the subject based on the plurality of 2D locations; for each image of the series of images, reconstructing, by the software code executed by the hardware processor and using the tracking module, a three-dimensional (3D) pose of the subject based on the plurality of locations and the plurality of joint angles, resulting in a series of 3D poses of the subject; and outputting, by the software code executed by the hardware processor, a tracking image corresponding to the motion by the subject based on the series of 3D poses of the subject. ||
17. A method comprising: training an hourglass network of a landmark detector of a tracking module to determine a plurality of locations corresponding respectively to a plurality of two-dimensional (2D) skeletal landmarks of a body image based on constraints imposed by features of a hierarchical skeleton model intersecting at each of the plurality of 2D skeletal landmarks; training a joint angle encoder of the tracking module to map a plurality of joint angles of a rigged body part corresponding to the body image to a low dimensional latent space from which a plurality of three-dimensional (3D) poses of the body image can be reconstructed; and training an inverse kinematics artificial neural network (ANN) of the tracking module to map the plurality of locations corresponding respectively to the 2D skeletal landmarks of the body image to the low dimensional latent space of the joint angle encoder for accurate reconstruction of the pose of the body image in 3D.