Outer Rim Archives
Archives · 2025 · 20250117892

Application (pre-grant publication)

TEMPORALLY CORRELATED NOISE WARPING FOR DIFFUSION MODELS

Number
20250117892
Published
2025-04-10
Filed
2024-09-27
Assignee
DISNEY ENTERPRISES, INC.
Inventors
DA COSTA DE AZEVEDO; Vinicius et al.
CPC
G06T3/18; G06T3/40; G06T5/50; G06T5/60; G06T5/70; G06T7/248
Verdict
Set aside generic diffusion-model research
Source
Google Patents · FreePatentsOnline

Abstract

One embodiment of the present invention sets forth a technique for generating data. The technique includes determining a first set of flow vectors between a first input frame and a second input frame. The technique also includes generating, based on the first set of flow vectors and a first noise sample associated with the first input frame, a second noise sample associated with the second frame. The technique further includes converting, via execution of a diffusion model, the first input frame into a first output frame based on the first noise sample and converting, via execution of the diffusion model, the second input frame into a second output frame based on the second noise sample.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to machine learning and generative models and, more specifically, to temporally correlated noise warping for diffusion models. Description of the Related Art

Generative models refer to deep neural networks and/or other types of machine learning models that are trained to generate new instances of data and/or augment existing data. For example, a generative model may be trained on a training dataset of images of cats. During the training process, the generative model “learns” the visual attributes of various cats depicted in the images. These learned visual attributes may then be used by the generative model to produce new images of cats that are not found in the training dataset. In another example, a generative model may be used to perform denoising, sharpening, blurring, colorization, compositing, super-resolution, inpainting, outpainting, and/or other types of image editing that involves altering the appearance, structure, and/or content of an image.

A diffusion model is one type of generative model. A diffusion model typically includes a forward diffusion process that gradually perturbs input data (e.g., an image) into noise that follows a certain noise distribution over a series of time steps. The diffusion model also includes a reverse denoising process that generates new data by iteratively converting random noise from the noise distribution into the

Claims

1. A computer-implemented method for generating data, the method comprising: determining a first set of flow vectors between a first input frame and a second input frame; generating, based on the first set of flow vectors and a first noise sample associated with the first input frame, a second noise sample associated with the second input frame; converting, via execution of a diffusion model, the first input frame into a first output frame based on the first noise sample; and converting, via execution of the diffusion model, the second input frame into a second output frame based on the second noise sample. || 11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: determining a first set of flow vectors between a first input frame and a second input frame; generating, based on the first set of flow vectors and a first noise sample associated with the first input frame, a second noise sample associated with the second input frame; converting, via execution of a diffusion model, the first input frame into a first output frame based on the first noise sample; and converting, via execution of the diffusion model, the second input frame into a second output frame based on the second noise sample. || 20. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the steps of: determining a first set of flow vectors between a first input frame and a second input frame; generating, based on the first set of flow vectors and a first noise sample associated with the first input frame, a second noise sample associated with the second input frame; converting, via execution of a diffusion model, the first input frame into a first output frame based on the first noise sample; and converting, via execution of the diffusion model, the second input frame into a second output frame based on the second noise sample.