Outer Rim Archives
Archives · 2025 · 20250111485

Application (pre-grant publication)

SPATIALLY CORRELATED NOISE WARPING FOR DIFFUSION MODELS

Number
20250111485
Published
2025-04-03
Filed
2024-09-27
Assignee
DISNEY ENTERPRISES, INC.
Inventors
DA COSTA DE AZEVEDO; Vinicius et al.
CPC
G06T3/18; G06T3/40; G06T5/50; G06T5/60; G06T5/70; G06T7/248
Verdict
Set aside generic diffusion-model research
Source
Google Patents · FreePatentsOnline

Abstract

One embodiment of the present invention sets forth a technique for generating data. The technique includes determining a plurality of flow vectors between a plurality of regions within a canonical space and a plurality of target spaces and generating, based on the plurality of flow vectors and a first noise sample associated with the canonical space, a plurality of noise samples associated with the plurality of target spaces. The technique also includes generating, via execution of a diffusion model based on the plurality of noise samples, a plurality of denoised intermediate samples associated with the plurality of target spaces and blending the plurality of denoised intermediate samples based on the plurality of flow vectors to generate a plurality of blended denoised intermediate samples associated with the plurality of target spaces. The technique further includes generating an output frame based on the plurality of blended denoised intermediate samples.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to machine learning and generative models and, more specifically, to spatially correlated noise warping for diffusion models. Description of the Related Art

Generative models refer to deep neural networks and/or other types of machine learning models that are trained to generate new instances of data and/or augment existing data. For example, a generative model may be trained on a training dataset of images of cats. During the training process, the generative model “learns” the visual attributes of various cats depicted in the images. These learned visual attributes may then be used by the generative model to produce new images of cats that are not found in the training dataset. In another example, a generative model may be used to perform denoising, sharpening, blurring, colorization, compositing, super-resolution, inpainting, outpainting, and/or other types of image editing that involves altering the appearance, structure, and/or content of an image.

A diffusion model is one type of generative model. A diffusion model typically includes a forward diffusion process that gradually perturbs input data (e.g., an image) into noise that follows a certain noise distribution over a series of time steps. The diffusion model also includes a reverse denoising process that generates new data by iteratively converting random noise from the noise distribution into the n

Claims

1. A computer-implemented method for generating data, the method comprising: determining a plurality of flow vectors between a plurality of regions within a canonical space and a plurality of target spaces; generating, based on the plurality of flow vectors and a first noise sample associated with the canonical space, a plurality of noise samples associated with the plurality of target spaces; generating, via execution of a diffusion model based on the plurality of noise samples, a plurality of denoised intermediate samples associated with the plurality of target spaces; blending the plurality of denoised intermediate samples based on the plurality of flow vectors to generate a plurality of blended denoised intermediate samples associated with the plurality of target spaces; and generating an output frame based on the plurality of blended denoised intermediate samples, wherein the output frame comprises a projection of a plurality of diffusion outputs that correspond to the plurality of blended denoised intermediate samples from the plurality of target spaces onto the plurality of regions within the canonical space. || 11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: determining a plurality of flow vectors between a plurality of regions within a canonical space and a plurality of target spaces; generating, based on the plurality of flow vectors and a first noise sample associated with the canonical space, a plurality of noise samples associated with the plurality of target spaces; generating, via execution of a diffusion model based on the plurality of noise samples, a plurality of denoised intermediate samples associated with the plurality of target spaces; blending the plurality of denoised intermediate samples based on the plurality of flow vectors to generate a plurality of blended denoised intermediate samples associated with the plurality of target spaces; and generating an output frame based on the plurality of blended denoised intermediate samples, wherein the output frame comprises a projection of a plurality of diffusion outputs that correspond to the plurality of blended denoised intermediate samples from the plurality of target spaces onto the plurality of regions within the canonical space. || 20. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the steps of: determining a plurality of flow vectors between a plurality of regions within a canonical space and a plurality of target spaces; generating, based on the plurality of flow vectors and a first noise sample associated with the canonical space, a plurality of noise samples associated with the plurality of target spaces; generating, via execution of a diffusion model based on the plurality of noise samples, a plurality of denoised intermediate samples associated with the plurality of target spaces; blending the plurality of denoised intermediate samples based on the plurality of flow vectors to generate a plurality of blended denoised intermediate samples associated with the plurality of target spaces; and generating an output frame based on the plurality of blended denoised intermediate samples, wherein the output frame comprises a projection of a plurality of diffusion outputs that correspond to the plurality of blended denoised intermediate samples from the plurality of target spaces onto the plurality of regions within the canonical space.