Outer Rim Archives
Archives · 2020 · 20200279428

Application (pre-grant publication)

JOINT ESTIMATION FROM IMAGES

Number
20200279428
Published
2020-09-03
Filed
2019-02-28
Assignee
Disney Enterprises, Inc.
Inventors
GUAY; Martin, BORER; Dominik Tobias, ÖZTIRELI; Ahmet Cengiz, SUMNER; Robert W., BUHMANN; Jakob Joachim
CPC
G06T11/10; G06T13/40; G06T7/73
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Joint/pose estimation from images for animation rigging.

Abstract

Techniques are disclosed for estimating poses from images. In one embodiment, a machine learning model, referred to herein as the “detector,” is trained to estimate animal poses from images in a bottom-up fashion. In particular, the detector may be trained using rendered images depicting animal body parts scattered over realistic backgrounds, as opposed to renderings of full animal bodies. In order to make appearances of the rendered body parts more realistic so that the detector can be trained to estimate poses from images of real animals, the body parts may be rendered using textures that are determined from a translation of rendered images of the animal into corresponding images with more realistic textures via adversarial learning. Three-dimensional poses may also be inferred from estimated joint locations using, e.g., inverse kinematics.

Background

BACKGROUNDField

This disclosure provides techniques for estimating joints of animals and other articulated figures in images.Description of the Related Art

Three-dimensional (3D) animal motions can be used to animate 3D virtual models of animals in movie production, digital puppeteering, and other applications. However, unlike humans whose motions may be captured via marker-based tracking, animals do not comply well and are difficult to transport to confined areas. As a result, marker-based tracking of animals can be infeasible. Instead, animal motions are typically created manually via key-framing.SUMMARY

One embodiment disclosed herein provides a computer-implemented method for identifying poses in images. The method generally includes rendering a plurality of images, where each of the plurality of images depicts distinct body parts of at least one figure, and each of the distinct body parts is associated with at least one joint location. The method further includes training a machine learning model using, at least in part, the plurality of images and the joint locations associated with the distinct body parts in the plurality of images. In addition, the method includes processing a received image using, at least in part, the trained machine learning model which outputs indications of joint locations in the received image.

Another embodiment provides a computer-implemented method for determining texture maps. The method generally includes converting,

Claims

1. A computer-implemented method for identifying poses in images, the method comprising: rendering a plurality of images, wherein: each of the plurality of images depicts distinct body parts of at least one figure, and each of the distinct body parts is associated with at least one joint location; training a machine learning model using, at least in part, the plurality of images and the joint locations associated with the distinct body parts in the plurality of images; and processing a received image using, at least in part, the trained machine learning model which outputs indications of joint locations in the received image. 11. A computer-implemented method for determining texture maps, comprising: converting, using adversarial learning, a plurality of rendered images that each depicts a respective figure to corresponding images that include different textures than the rendered images; and extracting one or more texture maps based, at least in part, on (a) textures of the respective figures as depicted in the corresponding images, and (b) pose and camera parameters used to render the rendered images. 19. A computer-implemented method for extracting poses from images, the method comprising: receiving one or more images, each of the one or more images depicting a respective figure; processing the one or more images using, at least in part, a trained machine learning model which outputs indications of joint locations in the one or more images; and inferring a respective skeleton for each image of the one or more images based, at least in part, on the joint locations in the image.