Outer Rim Archives
Archives · 2025 · 20250356467

Application (pre-grant publication)

TECHNIQUES FOR TEMPORALLY CONSISTENT VIDEO RESTORATION USING LATENT DIFFUSION MODELS

Number
20250356467
Published
2025-11-20
Filed
2025-05-19
Assignee
DISNEY ENTERPRISES, INC.
Inventors
ZHANG; Yang et al.
CPC
G06T5/50; G06T5/70; G06T5/60
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Latent-diffusion video-restoration rendering technique.

Abstract

Embodiments of the present disclosure provide techniques for restoring video content. An example method generally includes receiving a set of input video frames that include artifacts, generating one or more conditioning features based on the set of video frames, wherein the conditioning features represent content information included in the set of video frames while reducing representation of the artifacts, denoising, using a latent diffusion model and based on the conditioning features, a representation of the set of input video frames that includes noise, and generating a set of output frames based on the denoised representation, wherein the set of output video frames include fewer artifacts relative to the set of input video frames.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to video processing and, more specifically, to techniques for temporally consistent video restoration using latent diffusion models. Description of the Related Art

Video quality enhancement aims to improve visual details from

low-quality (LQ) videos while removing distorted artifacts, such as noise, blur, and compression artifacts etc. Compared to the synthetic data with specialized degradation, the real-world LQ videos are more challenging where the underlying degradation process is often more complicated and stochastic. To improve perceptual realism, recent research attempts to leverage the pretrained generative vision models, including generative adversarial network (GAN) and latent diffusion models. With the aid of richer prior knowledge of texture and semantics from large-scale datasets and models, these methods elevate the perceptual quality to a higher standard. However, the generative capability of these methods is deficient for video restoration tasks in at least two ways. First, the excessive visual details compromise the fidelity of the corresponding high-quality videos and, second, maintaining pixel-level temporal consistency becomes more demanding.

Thus, what is needed in the art are more effective techniques for video restoration using generative models. SUMMARY

One embodiment of the present disclosure sets forth techniques for re

Claims

1. A processor-implemented method, comprising: receiving a set of input video frames that include artifacts; generating one or more conditioning features based on the set of video frames, wherein the conditioning features represent content information included in the set of video frames while reducing representation of the artifacts; denoising, using a latent diffusion model and based on the conditioning features, a representation of the set of input video frames that includes noise; and generating a set of output frames based on the denoised representation, wherein the set of output video frames include fewer artifacts relative to the set of input video frames. || 8. One or more non-transitory computer readable media that, when executed by one or more computing devices, cause the one or more computing devices to perform the steps of: receiving a set of input video frames that include artifacts; generating one or more conditioning features based on the set of video frames, wherein the conditioning features represent content information included in the set of video frames while reducing representation of the artifacts; denoising, using a latent diffusion model and based on the conditioning features, a representation of the set of input video frames that includes noise; and generating a set of output frames based on the denoised representation, wherein the set of output video frames include fewer artifacts relative to the set of input video frames. || 15. A processing system, comprising: at least one memory having executable instructions stored thereon; and one or more processors configured to execute the executable instructions to cause the processing system to: receive a set of input video frames that include artifacts; generate one or more conditioning features based on the set of video frames, wherein the conditioning features represent content information included in the set of video frames while reducing representation of the artifacts; denoise, using a latent diffusion model and based on the conditioning features, a representation of the set of input video frames that includes noise; and generate a set of output frames based on the denoised representation, wherein the set of output video frames include fewer artifacts relative to the set of input video frames.