Application (pre-grant publication)
TECHNIQUES FOR INFERRING THREE-DIMENSIONAL POSES FROM TWO-DIMENSIONAL IMAGES
- Number
- 20210374993
- Published
- 2021-12-02
- Filed
- 2020-05-26
- Assignee
- DISNEY ENTERPRISES, INC.
- Inventors
- GUAY; Martin, NITTI; Maurizio, BUHMANN; Jakob Joachim, BORER; Dominik Tobias
- CPC
- G06T7/74; G06N3/045; G06T7/75; G06N3/09; G06N3/0464; G06T15/02; G06N3/08; G06N20/00
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
ML 3D pose inference from 2D images.
Abstract
In various embodiments, a training application generates training items for three-dimensional (3D) pose estimation. The training application generates multiple posed 3D models based on multiple 3D poses and a 3D model of a person wearing a costume that is associated with multiple visual attributes. For each posed 3D model, the training application performs rendering operation(s) to generate synthetic image(s). For each synthetic image, the training application generates a training item based on the synthetic image and the 3D pose associated with the posed 3D model from which the synthetic image was rendered. The synthetic images are included in a synthetic training dataset that is tailored for training a machine-learning model to compute estimated 3D poses of persons from two-dimensional (2D) input images. Advantageously, the synthetic training dataset can be used to train the machine-learning model to accurately infer the orientations of persons across a wide range of environments.
Background
BACKGROUND Field of the Various Embodiments
The various embodiments relate generally to computer science and computer vision and, more specifically, to techniques for inferring three-dimensional poses from two-dimensional images. Description of the Related Art
Computer vision techniques are oftentimes used to understand and track the spatial positions and movements of people for a wide variety of tasks. Some examples of such tasks include, without limitation, creating animated content, replacing actors with digital doubles, controlling games, analyzing gait abnormalities in medical practices, analyzing player movements in team sports, interacting with users in an augmented reality environment, etc. In one particularly useful computer vision technique known as three-dimensional (“3D”) human pose estimation, the 3D pose of a person is automatically estimated based on a two-dimensional (“2D”) image or video of the person acquired from a single camera. As used herein, the 3D pose of a person is defined to be a set of 3D positions associated with a set of joints that delineates a skeleton of the person, where the positions and orientations of bones included in the skeleton indicate the positions and orientations of the body parts of the person.
In one approach to 3D human pose estimation, a deep neural network is trained to estimate the 3D pose of a person depicted in a 2D image using training images that depict people oriented in various 3D poses. When an inp