Controllable 3D style-transfer technique for radiance-field (NeRF) rendering.
The present invention sets forth a technique for performing style transfer. The technique includes converting a first style sample into a first set of features and determining one or more style masks associated with the style sample. For each content sample included in a set of content samples, the technique also includes converting the content sample into an additional set of features, determining one or more two-dimensional content masks associated with the content sample, and determining a set of matches between (i) one or more subsets of the additional set of features corresponding to the one or more content masks and (ii) one or more subsets of the first set of features corresponding to the one or more style masks. The technique further includes generating a style transfer result, wherein the style transfer result comprises structural elements of the 3D scene and stylistic elements of the first style sample.
BACKGROUND Field of the Various Embodiments
Embodiments of the present disclosure relate generally to machine learning and image processing and, more specifically, to techniques for transferring styles in three-dimensional (3D) scenes. Description of the Related Art
Style transfer is a technique for generating stylized output by combining one or more structural elements included in one or more content samples with stylistic elements included in one or more style samples. Structural elements may include features such as objects, lines, edges, outlines, or surfaces. Stylistic elements may include one or more of colors, textures, patterns, or lighting characteristics included in the style samples. Style transfer is applicable to two-dimensional (2D) content samples, such as still images, or to 3D representations of the contents of a scene, such as neural radiance fields (NeRFs).
Existing techniques for performing style transfer in 3D representations of scenes are typically limited to transferring a style from a single style sample to the entirety of a content sample. Consequently, these techniques tend to lack fine-grained controllability, such as the ability to transfer a style to a specified element or object included in the content sample and/or transfer different styles to different regions within the content sample.
Other existing techniques may operate on 3D inputs, such as 3D point clouds or 3D mesh representations of content and/or style sampl
1. A computer-implemented method for performing artistic style transfer, the method comprising: converting a first style sample into a first set of features; determining one or more two-dimensional (2D) style masks associated with the style sample; determining a set of content samples corresponding to a plurality of views of a 3D scene; for each content sample included in the set of content samples: converting the content sample into an additional set of features; determining one or more two-dimensional content masks associated with the content sample; and determining a set of matches between (i) one or more subsets of the additional set of features corresponding to the one or more 2D content masks and (ii) one or more subsets of the first set of features corresponding to the one or more 2D style masks; and generating a style transfer result that includes a representation of the 3D scene based on one or more losses associated with the sets of matches determined for the set of content samples, wherein the style transfer result comprises one or more structural elements of the 3D scene and one or more stylistic elements of the first style sample at one or more locations corresponding to the one or more 2D content masks. ||
11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: converting a first style sample into a first set of features; determining one or more two-dimensional (2D) style masks associated with the style sample; determining a set of content samples corresponding to a plurality of views of a 3D scene; for each content sample included in the set of content samples: converting the content sample into an additional set of features; determining one or more two-dimensional content masks associated with the content sample; and determining a set of matches between (i) one or more subsets of the additional set of features corresponding to the one or more 2D content masks and (ii) one or more subsets of the first set of features corresponding to the one or more 2D style masks; and generating a style transfer result that includes a representation of the 3D scene based on one or more losses associated with the sets of matches determined for the set of content samples, wherein the style transfer result comprises one or more structural elements of the 3D scene and one or more stylistic elements of the first style sample at one or more locations corresponding to the one or more 2D content masks. ||
18. A system comprising: one or more memories storing instructions; and one or more processors for executing the instructions to: convert a first style sample into a first set of features; determine one or more two-dimensional (2D) style masks associated with the style sample; determine a set of content samples corresponding to a plurality of views of a 3D scene; for each content sample included in the set of content samples: convert the content sample into an additional set of features; determine one or more two-dimensional content masks associated with the content sample; and determine a set of matches between (i) one or more subsets of the additional set of features corresponding to the one or more 2D content masks and (ii) one or more subsets of the first set of features corresponding to the one or more 2D style masks; and generate a style transfer result that includes a representation of the 3D scene based on one or more losses associated with the sets of matches determined for the set of content samples, wherein the style transfer result comprises one or more structural elements of the 3D scene and one or more stylistic elements of the first style sample at one or more locations corresponding to the one or more 2D content masks.