Latent-space facial-feature-editing face-swap VFX technique.
A computer-implemented method of changing a face within an output image or video frame that includes: receiving an input image that includes a face presenting a facial expression in a pose; processing the image with a neural network encoder to generate a latent space point that is an encoded representation of the image; decoding the latent space point to generate an initial output image in accordance with a desired facial identity but with the facial expression and pose of the face in the input image; identifying a feature of the facial expression in the initial output image to edit; applying an adjustment vector to a latent space point corresponding to the initial output image to generate an adjusted latent space point; and decoding the adjusted latent space point to generate an adjusted output image in accordance with the desired facial identity but with the facial expression and pose of the face in the input image altered in accordance with the adjustment vector
BACKGROUND OF THE INVENTION
Face swapping is the process of replacing an actor's face in a plate with another person's face. In visual effects, face swapping is desirable for many creative goals, like replacing the face of a stunt double with that of the main actor or to achieve de-aging by swapping the face of a present-day actor with a younger looking face learned on archival footage. In the recent past, face-swapping techniques based on deep learning have become popular and are starting to see adoption for high quality visual effects production. These techniques typically employ encoder-decoder neural networks where the encoder ingests images of the actor to be replaced (e.g. the stunt double) and outputs a “latent space point” (a lower-dimensional abstract representation of that input data). An identity-specific decoder can then transform this latent space point back into an image in which the stunt double's face is replaced with the main actor's face.
While some currently available deep-learning face swapping techniques can do a good job at translating the facial expression of a source actor (e.g., stunt double) to target character in many instances, improvements in this regard are continuously being sought. In particular, one challenge with many deep learning techniques is the lack of control over the swapped image. For example, the eye gaze in the replaced face might be slightly off and there is no clear, easy way to correct the eye gaze direction. BRIEF
1. A computer-implemented method of changing a face within an image or video frame, the method comprising: receiving an input image that includes a face presenting a facial expression in a pose; processing the image with a neural network encoder to generate a latent space point that is an encoded representation of the image; decoding the latent space point to generate an initial output image in accordance with a desired facial identity but with the facial expression and pose of the face in the input image; identifying a feature of the facial expression in the initial output image to edit; applying an adjustment vector to a latent space point corresponding to the initial output image to generate an adjusted latent space point; and decoding the adjusted latent space point to generate an adjusted output image in accordance with the desired facial identity but with the facial expression and pose of the face in the input image altered in accordance with the adjustment vector. ||
2. The computer-implemented method of changing a face within an image or video frame set forth in claim 1 further comprising repeating the steps of applying an adjustment vector to the latent space point corresponding to the initial output image to generate an adjusted latent space point and decoding the adjusted space point to generate an adjusted output image until the adjusted output image has the desired facial expression. ||
3. The computer-implemented method of changing a face within an image or video frame set forth in claim 1 wherein the adjustment vector is generated from a plurality of key poses from selected images having a facial expression with a selected trait. ||
4. The computer-implemented method of changing a face within an image or video frame set forth in claim 1 wherein the adjustment vector is generated from a plurality of key poses from selected images having a facial expression with a selected trait, calculating latent space points for the selected images, and generating the adjustment vectors by computing differences between an average of latent space points for the selected images and a neutral latent space point. ||
5. The computer-implemented method of changing a face within an image or video frame set forth in claim 1 wherein the neural network is trained to be identity agnostic. ||
6. The computer-implemented method of changing a face within an image or video frame set forth in claim 1 wherein the input image is normalized prior to the receiving step. ||
7. The computer-implemented method of changing a face within an image or video frame set forth in claim 6 wherein the input image is resized to a predetermined size prior to the receiving step. ||
8. The computer-implemented method of changing a face within an image or video frame set forth in claim 1 further comprising allowing a user to select one or more features in the initial output image to adjust via a user interface. ||
9. The computer-implemented method of changing a face within an image or video frame set forth in claim 8 wherein the user interface comprises a slider that allows the user to control a weighting of the adjustment vector that is applied to the latent space point corresponding to the initial output image. ||
10. The computer-implemented method of changing a face within an image or video frame set forth in claim 1 further comprising incorporating the output image into one or more of a movie, a video, a video game or virtual or augmented reality content. ||
11. The computer-implemented method of changing a face within an image or video frame set forth in claim 1 wherein processing the input image with the neural network encoder to generate a latent space point that is an encoded representation of the image comprises: separately encoding different portions of the image by, for each separately encoded portion, generating a latent space point of the portion, thereby generating a plurality of multi-dimensional vectors where each multi-dimensional vector is an encoded representation of a different portion of the input image; and concatenating the plurality of multi-dimensional vectors into a combined vector that is the latent space point which, in turn, is an encoded representation of the image. ||
12. The computer-implemented method of changing a face within an image or video frame set forth in claim 11 wherein: identifying a feature of the facial expression in the initial output image to edit corresponds to identifying at least one of the separately encoded image portions, and wherein applying an adjustment vector comprises selecting an adjustment vector that corresponds to the at least one identified separately encoded image portion. ||
13. The computer-implemented method of changing a face within an image or video frame set forth in claim 12 wherein decoding the adjusted latent space point to generate an adjusted output image alters only a portion of the output image that corresponds to the identified feature. ||
14. A system for changing a face within an output image or video frame, the system comprising: a memory storing a plurality of computer-readable instructions; and one or more processors operable to execute the computer-readable instructions and cause the system to: receive an input image that includes a face presenting a facial expression in a pose; process the image with a neural network encoder to generate a latent space point that is an encoded representation of the image; decode the latent space point to generate an initial output image in accordance with a desired facial identity but with the facial expression and pose of the face in the input image; identify a feature of the facial expression in the initial output image to edit; apply an adjustment vector to a latent space point corresponding to the initial output image to generate an adjusted latent space point; and decode the adjusted latent space point to generate an adjusted output image in accordance with the desired facial identity but with the facial expression and pose of the face in the input image altered in accordance with the adjustment vector. ||
15. The system set forth in claim 14 wherein the plurality of computer readable instructions further comprise instructions to cause the system to repeat the steps of: (i) applying an adjustment vector to the latent space point corresponding to the initial output image to generate an adjusted latent space point and (ii) decoding the adjusted space point to generate an adjusted output image until the adjusted output image has the desired facial expression. ||
16. The system set forth in claim 15 wherein the adjustment vector is generated from a plurality of key poses from selected images having a facial expression with a selected trait. ||
17. The system set forth in claim 14 wherein the neural network is trained to be identity agnostic. ||
18. The system set forth in claim 14 wherein the input image is normalized and resized prior to the receiving step. ||
19. A non-transitory computer-readable memory comprising a plurality of computer-readable instructions that, when executed by one or more processors, cause the one or more processors to: receive an input image that includes a face presenting a facial expression in a pose; process the image with a neural network encoder to generate a latent space point that is an encoded representation of the image; decode the latent space point to generate an initial output image in accordance with a desired facial identity but with the facial expression and pose of the face in the input image; identify a feature of the facial expression in the initial output image to edit; apply an adjustment vector to a latent space point corresponding to the initial output image to generate an adjusted latent space point; and decode the adjusted latent space point to generate an adjusted output image in accordance with the desired facial identity but with the facial expression and pose of the face in the input image altered in accordance with the adjustment vector. ||
20. The non-transitory computer-readable memory set forth in claim 19 comprising additional computer-readable instructions that, when executed by one or more processors, cause the one or more processors to repeat the steps of: (i) applying an adjustment vector to the latent space point corresponding to the initial output image to generate an adjusted latent space point, and (ii) decoding the adjusted space point to generate an adjusted output image until the adjusted output image has the desired facial expression.