Frame-interpolation rendering technique.
One embodiment of the present invention sets forth a technique for performing frame interpolation. The technique includes generating (i) a first set of feature maps based on a first set of rendering features associated with a first key frame, (ii) a second set of feature maps based on a second set of rendering features associated with a second key frame, and (iii) a third set of feature maps based on a third set of rendering features associated with a target frame. The technique also includes applying one or more neural networks to the first, second, and third set of feature maps to generate a set of mappings from a first set of pixels in the first key frame to a second set of pixels in the target frame. The technique further includes generating the target frame based on the set of mappings.
BACKGROUND Field of the Various Embodiments (1) Embodiments of the present disclosure relate generally to frame interpolation and, more specifically, to frame interpolation for rendered content. Description of the Related Art (2) Computer-animated video content is commonly created via a multi-stage process that involves generating models of characters, objects, and settings; applying textures, lighting, shading, and other effects to the models; animating the models; and rendering individual frames based on the animated models. During this process, the rendering stage typically incurs significantly more computational overhead than earlier stages. This computational overhead has also increased over time, as more complex and sophisticated rendering techniques are developed to improve the realism and detail in the rendered frames. For example, a single frame of computer-animated video content could require multiple hours to multiple days to render on a single processor. A feature-length film with over 100,000 frames would thus consume hundreds of years of processor hours. (3) To reduce rendering overhead associated with computer-animated video content, a subset of “key frames” in the computer-animated video content can be rendered, and remaining frames between pairs of consecutive key frames can be generated via less computationally expensive frame interpolation techniques. However, an interpolated frame is not created using the same amount of information (e.g., models, texture,
1. A computer-implemented method for performing frame interpolation, the method comprising: converting a first set of rendering features used to render a first key frame from one or more models of one or more objects in the first key frame into a first set of feature maps; converting a second set of rendering features of a target frame into a second set of feature maps, wherein the target frame is to be rendered based on at least the first key frame, wherein the second set of rendering features comprises features of the target frame that are available prior to the target frame being rendered and that are independent of features associated with key frames that have been rendered, and wherein the second set of rendering features is associated with one or more additional models of one or more additional objects to be rendered in the target frame that are different from the one or more objects in the first key frame; applying one or more neural networks to the first set of feature maps and the second set of feature maps to generate a set of mappings from a first set of pixels in the first key frame to a second set of pixels in the target frame; and generating the target frame based on the set of mappings. ||
13. One or more non-transitory computer readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: converting a first set of rendering features used to render a first key frame from one or more models of one or more objects in the first key frame into a first set of feature maps; converting a second set of rendering features of a target frame into a second set of feature maps, wherein the target frame is to be rendered based on at least the first key frame, wherein the second set of rendering features comprises features of the target frame that are available prior to the target frame being rendered and that are independent of features associated with key frames that have been rendered, and wherein the second set of rendering features is associated with one or more additional models of one or more additional objects to be rendered in the target frame that are different from the one or more objects in the first key frame; applying one or more neural networks to the first set of feature maps and the second set of feature maps to generate a set of mappings from a first set of pixels in the first key frame to a second set of pixels in the target frame; and generating the target frame based on the set of mappings. ||
20. A system, comprising: a memory that stores instructions, and a processor that is coupled to the memory and, when executing the instructions, is configured to perform the steps of: converting a first set of rendering features used to render a first key frame from one or more models of one or more objects in the first key frame into a first set of feature maps; converting a second set of rendering features of a target frame into a second set of feature maps, wherein the target frame is to be rendered based on at least the first key frame, wherein the second set of rendering features comprises features of the target frame that are available prior to the target frame being rendered and that are independent of features associated with key frames that have been rendered, and wherein the second set of rendering features is associated with one or more additional models of one or more additional objects to be rendered in the target frame that are different from the one or more objects in the first key frame; applying one or more neural networks to the first set of feature maps and the second set of feature maps to generate a set of mappings from a first set of pixels in the first key frame to a second set of pixels in the target frame; and generating the target frame based on the set of mappings.