Application (pre-grant publication)
STABLE POSE ESTIMATION WITH ANALYSIS BY SYNTHESIS
- Number
- 20220392099
- Published
- 2022-12-08
- Filed
- 2022-05-19
- Assignee
- DISNEY ENTERPRISES, INC.
- Inventors
- GUAY; Martin, BORER; Dominik Tobias, BUHMANN; Jakob Joachim
- CPC
- G06N20/00; G06N3/045; G06N3/0475; G06N3/084; G06N3/088; G06N3/094; G06T7/70; G06T9/002; G06V10/774
- Verdict
- Medium Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Pose-estimation technique via analysis-by-synthesis.
Abstract
One embodiment of the present invention sets forth a technique for generating a pose estimation model. The technique includes generating one or more trained components included in the pose estimation model based on a first set of training images and a first set of labeled poses associated with the first set of training images, wherein each labeled pose includes a first set of positions on a left side of an object and a second set of positions on a right side of the object. The technique also includes training the pose estimation model based on a set of reconstructions of a second set of training images, wherein the set of reconstructions is generated by the pose estimation model from a set of predicted poses outputted by the one or more trained components.
Background
BACKGROUND Field of the Various Embodiments
Embodiments of the present disclosure relate generally to machine learning and pose estimation and, more specifically, to stable pose estimation with analysis by synthesis. Description of the Related Art
Pose estimation techniques are commonly used to detect and track humans, animals, robots, mechanical assemblies, and other articulated objects that can be represented by rigid parts connected by joints. For example, a pose estimation technique could be used to determine and track two-dimensional (2D) and/or three-dimensional (3D) locations of wrist, elbow, shoulder, hip, knee, ankle, head, and/or other joints of a person in an image or a video.
Recently, machine learning models have been developed to perform pose estimation. These machine learning models typically include deep neural networks with a large number of tunable parameters and thus require a large amount and variety of data to train. However, collecting training data for these machine learning models can be time- and resource-intensive. Continuing with the above example, a deep neural network could be trained to estimate the 2D or 3D locations of various joints for a person in an image or a video. To adequately train the deep neural network for the pose estimation task, the training dataset for the deep neural network would need to capture as many variations as possible on human appearances, human poses, and environments in which humans appear. Each t