Outer Rim Archives
Archives · 2022 · 20220392099

Application (pre-grant publication)

STABLE POSE ESTIMATION WITH ANALYSIS BY SYNTHESIS

Number
20220392099
Published
2022-12-08
Filed
2022-05-19
Assignee
DISNEY ENTERPRISES, INC.
Inventors
GUAY; Martin, BORER; Dominik Tobias, BUHMANN; Jakob Joachim
CPC
G06N20/00; G06N3/045; G06N3/0475; G06N3/084; G06N3/088; G06N3/094; G06T7/70; G06T9/002; G06V10/774
Verdict
Medium Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Pose-estimation technique via analysis-by-synthesis.

Abstract

One embodiment of the present invention sets forth a technique for generating a pose estimation model. The technique includes generating one or more trained components included in the pose estimation model based on a first set of training images and a first set of labeled poses associated with the first set of training images, wherein each labeled pose includes a first set of positions on a left side of an object and a second set of positions on a right side of the object. The technique also includes training the pose estimation model based on a set of reconstructions of a second set of training images, wherein the set of reconstructions is generated by the pose estimation model from a set of predicted poses outputted by the one or more trained components.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to machine learning and pose estimation and, more specifically, to stable pose estimation with analysis by synthesis. Description of the Related Art

Pose estimation techniques are commonly used to detect and track humans, animals, robots, mechanical assemblies, and other articulated objects that can be represented by rigid parts connected by joints. For example, a pose estimation technique could be used to determine and track two-dimensional (2D) and/or three-dimensional (3D) locations of wrist, elbow, shoulder, hip, knee, ankle, head, and/or other joints of a person in an image or a video.

Recently, machine learning models have been developed to perform pose estimation. These machine learning models typically include deep neural networks with a large number of tunable parameters and thus require a large amount and variety of data to train. However, collecting training data for these machine learning models can be time- and resource-intensive. Continuing with the above example, a deep neural network could be trained to estimate the 2D or 3D locations of various joints for a person in an image or a video. To adequately train the deep neural network for the pose estimation task, the training dataset for the deep neural network would need to capture as many variations as possible on human appearances, human poses, and environments in which humans appear. Each t

Claims

1. A computer-implemented method for generating a pose estimation model, the computer-implemented method comprising: generating one or more trained components included in the pose estimation model based on a first set of training images and a first set of labeled poses associated with the first set of training images, wherein each labeled pose included in the first set of labeled poses comprises a first set of positions on a left side of an object and a second set of positions on a right side of the object; and training the pose estimation model based on a set of reconstructions of a second set of training images, wherein the set of reconstructions is generated by the pose estimation model from a set of predicted poses outputted by the one or more trained components. || 11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: generating one or more trained components included in a pose estimation model based on a first set of training images and a first set of labeled poses associated with the first set of training images; and training the pose estimation model based on one or more losses associated with a second set of training images and a set of reconstructions of the second set of training images, wherein the set of reconstructions is generated by the pose estimation model from a set of predicted poses outputted by the one or more trained components. || 20. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to: execute one or more trained components included in a pose estimation model based on an input image; and receive, as output of the one or more trained components, one or more poses associated with an object depicted in the input image, wherein the one or more poses comprise a first set of positions on a left side of the object and a second set of positions on a right side of the object.