Outer Rim Archives
Archives · 2025 · 20250356540

Application (pre-grant publication)

STYLE TRANSFER USING GENERATIVE DIFFUSION FEATURES

Number
20250356540
Published
2025-11-20
Filed
2025-05-12
Assignee
DISNEY ENTERPRISES, INC.
Inventors
DJELOUAH; Abdelaziz et al.
CPC
G06T11/00; G06V10/82; G06T5/50; G06T5/60; G06T5/70; G06V10/40; G06V10/762; G06V10/7715; G06V10/774
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Generative-diffusion style-transfer creative-ML technique.

Abstract

The present invention sets forth techniques for performing style transfer from multiple supplied style images to a supplied content image to generate novel images that include style elements from the multiple supplied style images and content elements from the supplied content image. The techniques include guiding one or more self-attention and cross-attention layers included in a machine learning model based on the multiple supplied style images, such that content elements and style elements included in the style images are not entangled when generating the novel images. The techniques also distill a small subset of representative attention map values from multiple style images, improving performance while reducing computational costs compared to processing all attention map values from the multiple style images.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to computer vision and image processing and, more specifically, to techniques for performing style transfer using generative diffusion features, including all aspects of the related hardware, software, graphical user interfaces, and algorithms associated with implementing the contemplated systems, techniques, functions, and operations set forth herein, Description of the Related Art

In the fields of machine learning and computer vision, domain adaptation or style transfer refers to the generation of novel images that exhibit content features inherited from a supplied content image and stylistic features inherited from one or more supplied style images. For example, a supplied content image may include a photograph of a building against a background, and one or more supplied style images may collectively exhibit one or more style elements, such as an impressionist or cubist artistic style, brush strokes, drawn lines, and/or colors. In this example, style transfer techniques may generate one or more novel images depicting the building, background, and/or other content elements included in the supplied content image, such that the generated image(s) also exhibit one or more style elements included in the supplied style images. Content elements may include features such as objects, lines, edges, outlines, or surfaces. Style elements may further include, but are not lim

Claims

1. A computer-implemented method for performing style transfer, the computer-implemented method comprising: receiving a content image including one or more content elements, and multiple style images each including one or more style elements; generating an average embedding and an average style image based on the multiple style images; generating, via a clustering technique, a representative set of attention map keys and values associated with the multiple style images; and generating, via a trained machine learning model and based at least on the average embedding, the average style image, and the representative set of attention map keys and values, a stylized output image including at least one of the one or more content elements and at least one of the one or more style elements. || 10. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: receiving a content image including one or more content elements, and multiple style images each including one or more style elements; generating an average embedding and an average style image based on the multiple style images; generating, via a clustering technique, a representative set of attention map keys and values associated with the multiple style images; and generating, via a trained machine learning model and based at least on the average embedding, the average style image, and the representative set of attention map keys and values, a stylized output image including at least one of the one or more content elements and at least one of the one or more style elements. || 19. A system comprising: one or more memories storing instructions; and one or more processors for executing the instructions to: receive a content image including one or more content elements, and multiple style images each including one or more style elements; generate an average embedding and an average style image based on the multiple style images; generate, via a clustering technique, a representative set of attention map keys and values associated with the multiple style images; and generate, via a trained machine learning model and based at least on the average embedding, the average style image, and the representative set of attention map keys and values, a stylized output image including at least one of the one or more content elements and at least one of the one or more style elements.