Outer Rim Archives
Archives · 2023 · 20230377214

Application (pre-grant publication)

IDENTITY-PRESERVING IMAGE GENERATION USING DIFFUSION MODELS

Number
20230377214
Published
2023-11-23
Filed
2023-05-19
Assignee
DISNEY ENTERPRISES, INC.
Inventors
Kansy; Manuel Jakob et al.
CPC
G06T11/00; G06T5/70; G06V10/82; G06V40/169; G06V40/171
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Diffusion-model identity-preserving image generation (creative ML).

Abstract

One embodiment of the present invention sets forth a technique for performing identity-preserving image generation. The technique includes converting an identity image depicting a facial identity into an identity embedding. The technique further includes generating a combined embedding based on the identity embedding and a diffusion iteration identifier. The technique further includes converting, using a neural network and based on the combined embedding, a first input image that includes first noise into a first predicted image depicting one or more facial features that include one or more first facial identity features, wherein the one or more first facial identity features correspond to one or more respective second facial identity features of the identity image and are based at least on the identity embedding.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to machine learning and computer vision and, more specifically, to techniques for generating images of individuals. Description of the Related Art

As used herein, a facial identity of an individual refers to an appearance of a face that is considered distinct from facial appearances of other entities due to differences in personal identity, age, or the like. Two facial identities may be of different individuals or the same individual under different conditions, such as the same individual at different ages.

A face of an individual depicted in an image, such as a digital photograph or a frame of digital video, can be converted to a value referred to as an identity embedding in a latent space. The identity embedding includes a representation of identity features of the individual and can, in practice, also include other features that do not represent the individual. The identity embedding can be generated by a machine-learning model such as a face recognition model. The identity embedding does not have a readily-discernible correspondence to the image, since the identity embedding depends on the particular configuration of the network in addition to the content of the image. Although the identity embedding does not necessarily encode sufficient information to reconstruct the exact original image from the embedding, reconstruction of an approximation of the ori

Claims

1. A computer-implemented method for performing identity-preserving image generation, the computer-implemented method comprising: converting a first identity image depicting a first facial identity into a first identity embedding; generating a combined embedding based on the first identity embedding and a diffusion iteration identifier; and converting, using a neural network and based on the combined embedding, an input image that includes noise into a first predicted image depicting one or more facial features that include one or more first facial identity features, wherein the one or more first facial identity features correspond to one or more respective second facial identity features of the first identity image and are based at least on the first identity embedding. || 17. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the steps of: converting a clean image depicting a facial identity into an identity embedding; converting, using a neural network and based on the identity embedding and a diffusion iteration identifier, a first input image that includes first noise into predicted noise; and training the neural network based on a loss associated with the predicted noise and an image of pure noise. || 19. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: converting an identity image depicting a facial identity into an identity embedding; generating a combined embedding based on the identity embedding and a diffusion iteration identifier; and converting, using a neural network and based on the combined embedding, a first input image that includes first noise into a first predicted image depicting one or more facial features that include one or more first facial identity features, wherein the one or more first facial identity features correspond to one or more respective second facial identity features of the identity image and are based at least on the identity embedding.