Outer Rim Archives
Archives · 2021 · 20210374993

Application (pre-grant publication)

TECHNIQUES FOR INFERRING THREE-DIMENSIONAL POSES FROM TWO-DIMENSIONAL IMAGES

Number
20210374993
Published
2021-12-02
Filed
2020-05-26
Assignee
DISNEY ENTERPRISES, INC.
Inventors
GUAY; Martin, NITTI; Maurizio, BUHMANN; Jakob Joachim, BORER; Dominik Tobias
CPC
G06T7/74; G06N3/045; G06T7/75; G06N3/09; G06N3/0464; G06T15/02; G06N3/08; G06N20/00
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

ML 3D pose inference from 2D images.

Abstract

In various embodiments, a training application generates training items for three-dimensional (3D) pose estimation. The training application generates multiple posed 3D models based on multiple 3D poses and a 3D model of a person wearing a costume that is associated with multiple visual attributes. For each posed 3D model, the training application performs rendering operation(s) to generate synthetic image(s). For each synthetic image, the training application generates a training item based on the synthetic image and the 3D pose associated with the posed 3D model from which the synthetic image was rendered. The synthetic images are included in a synthetic training dataset that is tailored for training a machine-learning model to compute estimated 3D poses of persons from two-dimensional (2D) input images. Advantageously, the synthetic training dataset can be used to train the machine-learning model to accurately infer the orientations of persons across a wide range of environments.

Background

BACKGROUND Field of the Various Embodiments

The various embodiments relate generally to computer science and computer vision and, more specifically, to techniques for inferring three-dimensional poses from two-dimensional images. Description of the Related Art

Computer vision techniques are oftentimes used to understand and track the spatial positions and movements of people for a wide variety of tasks. Some examples of such tasks include, without limitation, creating animated content, replacing actors with digital doubles, controlling games, analyzing gait abnormalities in medical practices, analyzing player movements in team sports, interacting with users in an augmented reality environment, etc. In one particularly useful computer vision technique known as three-dimensional (“3D”) human pose estimation, the 3D pose of a person is automatically estimated based on a two-dimensional (“2D”) image or video of the person acquired from a single camera. As used herein, the 3D pose of a person is defined to be a set of 3D positions associated with a set of joints that delineates a skeleton of the person, where the positions and orientations of bones included in the skeleton indicate the positions and orientations of the body parts of the person.

In one approach to 3D human pose estimation, a deep neural network is trained to estimate the 3D pose of a person depicted in a 2D image using training images that depict people oriented in various 3D poses. When an inp

Claims

1. A computer-implemented method for generating training items for three-dimensional (3D) pose estimation, the method comprising: generating a plurality of posed 3D models based on a plurality of 3D poses and a 3D model of a first person wearing a first costume associated with a plurality of visual attributes; for each posed 3D model, performing at least one rendering operation to generate at least one synthetic image; and for each synthetic image, generating a training item to include in a synthetic training dataset based on the synthetic image and a 3D pose associated with the posed 3D model from which the synthetic image was rendered, wherein the synthetic training dataset is tailored for training a machine-learning model to compute estimated 3D poses of persons from two-dimensional (2D) input images. || 11. One or more non-transitory computer readable media including instructions that, when executed by one or more processors, cause the one or more processors to generate training items for three-dimensional (3D) pose estimation by performing the steps of: generating a plurality of posed 3D models based on a plurality of 3D poses and a 3D model of a first person wearing a first costume associated with a plurality of visual attributes; for each posed 3D model, performing at least one rendering operation to generate at least one synthetic image; and for each synthetic image, generating a training item to include in a synthetic training dataset based on the synthetic image and a 3D pose associated with the posed 3D model from which the synthetic image was rendered, wherein the synthetic training dataset is tailored for training a machine-learning model to compute estimated 3D poses of persons from two-dimensional (2D) input images. || 20. A system comprising: one or more memories storing instructions; and one or more processors coupled to the one or more memories that, when executing the instructions, perform the steps of: generating a plurality of posed 3D models based on a plurality of 3D poses and a 3D model of a first person wearing a first costume associated with a plurality of visual attributes; for each posed 3D model, performing at least one rendering operation to generate at least one synthetic image; and for each synthetic image, generating a training item to include in a synthetic training dataset based on the synthetic image and a 3D pose associated with the posed 3D model from which the synthetic image was rendered, wherein the synthetic training dataset is tailored for training a machine-learning model to compute estimated 3D poses of persons from two-dimensional (2D) input images.