Outer Rim Archives
Archives · 2026 · 12725289

Granted patent

Stable pose estimation with analysis by synthesis

Number
12725289
Published
2026-09-01
Filed
2022-05-19
Assignee
Disney Enterprises, Inc.
Inventors
Guay; Martin, Borer; Dominik Tobias, Buhmann; Jakob Joachim
CPC
G06T7/70; G06N3/045; G06N3/0475; G06N3/084; G06N3/088; G06N3/094; G06N20/00; G06T9/002; G06V10/774; G06T2207/20081; G06T2207/20084; G06T2207/30196
Verdict
High Notable software
First reported
2026-W36 (2026-09-05)
Source
Google Patents · FreePatentsOnline

The keeper's note

Analysis-by-synthesis pose estimation that stays stable frame to frame — tracking tech for guests, performers, or robots.

Abstract

One embodiment of the present invention sets forth a technique for generating a pose estimation model. The technique includes generating one or more trained components included in the pose estimation model based on a first set of training images and a first set of labeled poses associated with the first set of training images, wherein each labeled pose includes a first set of positions on a left side of an object and a second set of positions on a right side of the object. The technique also includes training the pose estimation model based on a set of reconstructions of a second set of training images, wherein the set of reconstructions is generated by the pose estimation model from a set of predicted poses outputted by the one or more trained components.

Background

BACKGROUND Field of the Various Embodiments (1) Embodiments of the present disclosure relate generally to machine learning and pose estimation and, more specifically, to stable pose estimation with analysis by synthesis. Description of the Related Art (2) Pose estimation techniques are commonly used to detect and track humans, animals, robots, mechanical assemblies, and other articulated objects that can be represented by rigid parts connected by joints. For example, a pose estimation technique could be used to determine and track two-dimensional (2D) and/or three-dimensional (3D) locations of wrist, elbow, shoulder, hip, knee, ankle, head, and/or other joints of a person in an image or a video. (3) Recently, machine learning models have been developed to perform pose estimation. These machine learning models typically include deep neural networks with a large number of tunable parameters and thus require a large amount and variety of data to train. However, collecting training data for these machine learning models can be time- and resource-intensive. Continuing with the above example, a deep neural network could be trained to estimate the 2D or 3D locations of various joints for a person in an image or a video. To adequately train the deep neural network for the pose estimation task, the training dataset for the deep neural network would need to capture as many variations as possible on human appearances, human poses, and environments in which humans appear. Each training s

Claims

1. A computer-implemented method for generating a pose estimation model, the computer-implemented method comprising: generating one or more trained components included in the pose estimation model based on one or more supervised losses computed from (i) a first set of predicted poses associated with a first set of training images and (ii) a first set of labeled poses associated with the first set of training images, wherein the first set of training images depict a first set of articulated objects against a first set of backgrounds; generating, via execution of the one or more trained components, a set of predicted poses based on input that includes a second set of training images that depict a second set of articulated objects in a first set of poses against a second set of backgrounds; generating, via execution of an image renderer included in the pose estimation model, a set of output images based on input that includes (i) the set of predicted poses and (ii) a set of reference images that depict the second set of articulated objects in a second set of poses against the second set of backgrounds; and training the one or more trained components and the image renderer based on one or more unsupervised losses computed between (i) the set of output images and (ii) the second set of training images to generate a trained pose estimation model. || 11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: generating one or more trained components included in a pose estimation model based on one or more supervised losses computed from (i) a first set of predicted poses associated with a first set of training images that depict a first set of articulated objects against a first set of backgrounds and (ii) a first set of labeled poses associated with the first set of training images; and training the one or more trained components and an image renderer included in the pose estimation model based on one or more unsupervised losses computed between (i) a second set of training images that depict a second set of articulated objects against a second set of backgrounds and (ii) a set of reconstructions of the second set of training images to generate a trained pose estimation model, wherein the set of reconstructions is generated by the image renderer from a set of predicted poses outputted by the one or more trained components. || 20. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to: execute a trained pose estimation model based on an input image, wherein the trained pose estimation model is generated by: generating one or more trained components included in a pose estimation model based on one or more supervised losses computed from (i) a first set of predicted poses associated with a first set of training images that depict a first set of articulated objects against a first set of backgrounds and (ii) a first set of labeled poses associated with the first set of training images; and training the one or more trained components and an image renderer included in the pose estimation model based on one or more unsupervised losses computed between (i) a second set of training images that depict a second set of articulated objects against a second set of backgrounds and (ii) a set of reconstructions of the second set of training images, wherein the set of reconstructions is generated by the image renderer from a set of predicted poses outputted by the one or more trained components; and receive, as output of the one or more trained components, one or more poses associated with an object depicted in the input image.