Outer Rim Archives
Archives · 2023 · 11669999

Granted patent

Techniques for inferring three-dimensional poses from two-dimensional images

Number
11669999
Published
2023-06-06
Filed
2020-05-26
Assignee
DISNEY ENTERPRISES, INC.
Inventors
Guay; Martin et al.
CPC
G06T7/75; G06N3/09; G06T15/02; G06T7/74; G06N3/0464; G06N20/00; G06N3/08; G06N3/045
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

3D pose-inference technique from 2D images.

Abstract

In various embodiments, a training application generates training items for three-dimensional (3D) pose estimation. The training application generates multiple posed 3D models based on multiple 3D poses and a 3D model of a person wearing a costume that is associated with multiple visual attributes. For each posed 3D model, the training application performs rendering operation(s) to generate synthetic image(s). For each synthetic image, the training application generates a training item based on the synthetic image and the 3D pose associated with the posed 3D model from which the synthetic image was rendered. The synthetic images are included in a synthetic training dataset that is tailored for training a machine-learning model to compute estimated 3D poses of persons from two-dimensional (2D) input images. Advantageously, the synthetic training dataset can be used to train the machine-learning model to accurately infer the orientations of persons across a wide range of environments.

Background

BACKGROUND Field of the Various Embodiments (1) The various embodiments relate generally to computer science and computer vision and, more specifically, to techniques for inferring three-dimensional poses from two-dimensional images. Description of the Related Art (2) Computer vision techniques are oftentimes used to understand and track the spatial positions and movements of people for a wide variety of tasks. Some examples of such tasks include, without limitation, creating animated content, replacing actors with digital doubles, controlling games, analyzing gait abnormalities in medical practices, analyzing player movements in team sports, interacting with users in an augmented reality environment, etc. In one particularly useful computer vision technique known as three-dimensional (“3D”) human pose estimation, the 3D pose of a person is automatically estimated based on a two-dimensional (“2D”) image or video of the person acquired from a single camera. As used herein, the 3D pose of a person is defined to be a set of 3D positions associated with a set of joints that delineates a skeleton of the person, where the positions and orientations of bones included in the skeleton indicate the positions and orientations of the body parts of the person. (3) In one approach to 3D human pose estimation, a deep neural network is trained to estimate the 3D pose of a person depicted in a 2D image using training images that depict people oriented in various 3D poses. When an input image

Claims

1. A computer-implemented method for generating training items for three-dimensional (3D) pose estimation, the method comprising: generating a plurality of posed 3D models based on a plurality of 3D poses and a 3D model of a first person wearing a first costume associated with a plurality of visual attributes, wherein generating the plurality of posed 3D models comprises, for a first 3D pose included in the plurality of 3D poses: modifying at least one shape associated with the 3D model of the first person wearing the first costume to generate a modified 3D model, and fitting the modified 3D model to the first 3D pose to generate a first posed 3D model; for each posed 3D model, performing at least one rendering operation to generate at least one synthetic image; and for each synthetic image, generating a training item to include in a synthetic training dataset, wherein the training item includes the synthetic image and a 3D pose associated with a posed 3D model included in the plurality of posed 3D models from which the synthetic image was rendered, and wherein the synthetic training dataset is tailored for training a machine-learning model to compute estimated 3D poses of persons from two-dimensional (2D) input images. || 10. One or more non-transitory computer readable media including instructions that, when executed by one or more processors, cause the one or more processors to generate training items for three-dimensional (3D) pose estimation by performing the steps of: generating a plurality of posed 3D models based on a plurality of 3D poses and a 3D model of a first person wearing a first costume associated with a plurality of visual attributes, wherein generating the plurality of posed 3D models comprises, for a first 3D pose included in the plurality of 3D poses: modifying at least one shape associated with the 3D model of the first person wearing the first costume to generate a modified 3D model, and fitting the modified 3D model to the first 3D pose to generate a first posed 3D model; for each posed 3D model, performing at least one rendering operation to generate at least one synthetic image; and for each synthetic image, generating a training item to include in a synthetic training dataset, wherein the training item includes the synthetic image and a 3D pose associated with a posed 3D model included in the plurality of posed 3D models from which the synthetic image was rendered, and wherein the synthetic training dataset is tailored for training a machine-learning model to compute estimated 3D poses of persons from two-dimensional (2D) input images. || 18. A system comprising: one or more memories storing instructions; and one or more processors coupled to the one or more memories that, when executing the instructions, perform the steps of: generating a plurality of posed 3D models based on a plurality of 3D poses and a 3D model of a first person wearing a first costume associated with a plurality of visual attributes, wherein generating the plurality of posed 3D models comprises, for a first 3D pose included in the plurality of 3D poses: modifying at least one shape associated with the 3D model of the first person wearing the first costume to generate a modified 3D model; and fitting the modified 3D model to the first 3D pose to generate a first posed 3D model, for each posed 3D model, performing at least one rendering operation to generate at least one synthetic image, and for each synthetic image, generating a training item to include in a synthetic training dataset, wherein the training item includes the synthetic image and a 3D pose associated with a posed 3D model included in the plurality of posed 3D models from which the synthetic image was rendered, and wherein the synthetic training dataset is tailored for training a machine-learning model to compute estimated 3D poses of persons from two-dimensional (2D) input images.