Virtual-character-based automated image-augmentation technique.
An image processing system includes a computing platform having a hardware processor and a system memory storing an image augmentation software code, a three-dimensional (3D) shapes library, and/or a 3D poses library. The image processing system also includes a two-dimensional (2D) pose estimation module communicatively coupled to the image augmentation software code. The hardware processor executes the image augmentation software code to provide an image to the 2D pose estimation module and to receive a 2D pose data generated by the 2D pose estimation module based on the image. The image augmentation software code identifies a 3D shape and/or a 3D pose corresponding to the image using an optimization algorithm applied to the 2D pose data and one or both of the 3D poses library and the 3D shapes library, and may output the 3D shape and/or 3D pose to render an augmented image on a display.
BACKGROUND (1) Augmented reality (AR), in which real world objects and/or environments are digitally augmented with virtual imagery, offers more immersive and enjoyable educational or entertainment experiences. SUMMARY (2) There are provided systems and methods for performing automated image augmentation using a virtual character, substantially as shown in and/or described in connection with at least one of the figures, and as set forth more completely in the claims.
1. An image processing system comprising: a computing platform including a hardware processor and a system memory; the system memory storing an image augmentation software code; and a two-dimensional (2D) pose estimation software having computing instructions; wherein the hardware processor is configured to execute the image augmentation software code to: provide an image as an input to the 2D pose estimation software, the pose estimation software including a deep neural network trained over a data set of manually annotated images to receive the image as input and output 2D pose data including a plurality of joint positions and a respective confidence value for each of the plurality of joint positions; receive from the 2D pose estimation software, the 2D pose data generated based on the image; generate a virtual character; estimate a three-dimensional (3D) body shape corresponding to the image based on a rigid transformation of one of a plurality of predetermined 3D body shape exemplars that brings the one of the plurality of predetermined 3D body shape exemplars closest, in terms of joint positions and bone direction similarity, to a 2D pose described by the 2D pose data; and utilize the 3D body shape to generate a shadow cast from the image to the virtual character. ||
10. A computer-implemented method for generating a virtual character, the method comprising: providing an image as an input to a two-dimensional (2D) pose estimation software having computing instructions, the pose estimation software including a deep neural network trained over a data set of manually annotated images to receive the image as input and output 2D pose data including a plurality of joint positions and a respective confidence value for each of the plurality of joint positions; receiving from the 2D pose estimation software, a 2D pose data generated based on the image; estimating a three-dimensional (3D) body shape corresponding to the image based on a rigid transformation of one of a plurality of predetermined 3D body shape exemplars that brings the one of the plurality of predetermined 3D body shape exemplars closest, in terms of joint positions and bone direction similarity, to a 2D pose described by the 2D pose data; and utilizing the 3D body shape to generate a shadow cast from the image to the virtual character. ||
19. A computer-implemented method for generating a virtual character, the method comprising: providing an image of a human as an input to a two-dimensional (2D) pose estimation software having computing instructions, the pose estimation software including a deep neural network trained over a data set of manually annotated images to receive the image as input and output 2D pose data including a plurality of joint positions and a respective confidence value for each of the plurality of joint positions; receiving from the 2D pose estimation software, a 2D pose data generated based on the image of the human; identifying a three-dimensional (3D) shape corresponding to the image of the human using the 2D pose data and a 3D shapes library including a plurality of predetermined 3D body shape exemplars, based on a rigid transformation of one of the plurality of predetermined 3D body shape exemplars that brings the one of the plurality of predetermined 3D body shape exemplars closest, in terms of joint positions and bone direction similarity, to a 2D human pose described by the 2D pose data; generating the virtual character having the identified 3D shape; and utilizing the 3D shape to generate a partial occlusion of the virtual character by the image of the human.