Outer Rim Archives
Archives · 2021 · 10916046

Granted patent

Joint estimation from images

Number
10916046
Published
2021-02-09
Filed
2019-02-28
Assignee
Disney Enterprises, Inc.
Inventors
Guay; Martin, Borer; Dominik Tobias, Öztireli; Ahmet Cengiz, Sumner; Robert W., Buhmann; Jakob Joachim
CPC
G06T11/10; G06T13/40; G06T7/73
Verdict
High Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Joint/pose estimation technique from images for animation/robotics.

Abstract

Techniques are disclosed for estimating poses from images. In one embodiment, a machine learning model, referred to herein as the 'detector,' is trained to estimate animal poses from images in a bottom-up fashion. In particular, the detector may be trained using rendered images depicting animal body parts scattered over realistic backgrounds, as opposed to renderings of full animal bodies. In order to make appearances of the rendered body parts more realistic so that the detector can be trained to estimate poses from images of real animals, the body parts may be rendered using textures that are determined from a translation of rendered images of the animal into corresponding images with more realistic textures via adversarial learning. Three-dimensional poses may also be inferred from estimated joint locations using, e.g., inverse kinematics.

Background

BACKGROUND Field (1) This disclosure provides techniques for estimating joints of animals and other articulated figures in images. Description of the Related Art (2) Three-dimensional (3D) animal motions can be used to animate 3D virtual models of animals in movie production, digital puppeteering, and other applications. However, unlike humans whose motions may be captured via marker-based tracking, animals do not comply well and are difficult to transport to confined areas. As a result, marker-based tracking of animals can be infeasible. Instead, animal motions are typically created manually via key-framing. SUMMARY (3) One embodiment disclosed herein provides a computer-implemented method for identifying poses in images. The method generally includes rendering a plurality of images, where each of the plurality of images depicts distinct body parts of at least one figure, and each of the distinct body parts is associated with at least one joint location. The method further includes training a machine learning model using, at least in part, the plurality of images and the joint locations associated with the distinct body parts in the plurality of images. In addition, the method includes processing a received image using, at least in part, the trained machine learning model which outputs indications of joint locations in the received image. (4) Another embodiment provides a computer-implemented method for determining texture maps. The method generally includes converting, usin

Claims

1. A computer-implemented method for identifying poses in images, the method comprising: rendering a plurality of training images, wherein each of the plurality of training images depicts distinct body parts of at least one figure, wherein each of the distinct body parts is associated with at least one joint location, and wherein each of the distinct body parts is a rendering of a portion of a first virtual model, wherein the portion is textured using a texture map determined via adversarial learning; training a machine learning model using, at least in part, the plurality of training images and the at least one joint location associated with each of the distinct body parts in the plurality of training images; processing an image using, at least in part, the machine learning model to determine joint locations in the image; and inferring a skeleton based, at least in part, on the joint locations in the image, wherein the skeleton is used to animate the first virtual model or a second virtual model. || 9. A computer-implemented method for determining texture maps, comprising: converting, using adversarial learning, a plurality of rendered images to corresponding images that include different textures than textures of the plurality of rendered images, each of the plurality of rendered images depicting a respective figure; extracting one or more texture maps based, at least in part, on (a) the textures included in the corresponding images, and (b) pose and camera parameters used to render the plurality of rendered images; and training a machine learning model using, at least in part, a plurality of training images, each training image depicting distinct body parts textured using at least one of the one or more texture maps, wherein the machine learning model is configured to determine joint locations in an image, wherein a skeleton is inferred based, at least in part, on the joint locations, wherein the skeleton is used to animate a first virtual model. || 17. A computer-implemented method for extracting poses from images, the computer-implemented method comprising: receiving one or more images, each of the one or more images depicting a respective figure; processing the one or more images using, at least in part, a machine learning model configured to determine joint locations in each of the one or more images, wherein the machine learning model is trained using a plurality of training images, each training image of the plurality of training images depicting a plurality of body parts, each body part of the plurality of body parts comprising a rendering of a portion of a first virtual model, wherein the portion is textured using a texture map determined via adversarial learning; and determining a motion by inferring a respective skeleton for each image of the one or more images based, at least in part, on the joint locations in each image, wherein the motion is used to animate the first virtual model or a second virtual model.